Please mention DailyRemote when applying
This is a remote position.
What You'll Do
Design, implement, and manage complex DAGs and SubDAGs in Apache Airflow.
Utilize Operators, macros, connections, variables, and XCom for efficient task execution.
Conduct unit testing to ensure the robustness of Airflow workflows.
Perform file operations and data movement between HDFS and S3 (and vice versa).
Manage Hive external tables, update table DDL, and handle partition operations.
Execute basic Hadoop commands and navigate/traverse HDFS.
Work with Hive on Tez for optimized query execution.
Utilize AWS command line tools for tasks such as listing and copying files from S3.
Implement and manage secrets using tools like Ansible Vault and AWS Secrets Manager.
Develop and optimize data processing tasks using PySpark.
Understand and leverage Livy for efficient Spark job execution.
Import and export data between Microsoft SQL Server and HDFS using Sqoop.
Apply dimension modeling principles, define facts and dimensions, and manage surrogate keys.
Implement effective branching strategies in Git.
Participate in the process of raising and reviewing Merge Requests in GitLab.
Utilize SSIS for migrating data workflows, ensuring a smooth transition.
Model data in Snowflake, differentiating between external and internal tables.
Understand the use of views and materialized views in Snowflake.
What You'll Need
Bachelor's degree in Computer Science, Information Technology, or a related field.
Proficient in Python (both Python 2.7 and 3.6+).
Experience in unit testing and test-driven development.
Familiarity with Secrets Management tools and techniques.
Knowledge of data modeling principles and practices.
Strong understanding of Git version control and GitLab workflows.
Excellent problem-solving and communication skills.
Awareness or knowledge of IT security best practices as defined by ISO / SOC or similar standards.
Experience with SSIS for data migration workflows.
Experience with Snowflake (external/internal tables, views, materialized views).
Familiarity with Hive on Tez for optimized query execution.
Experience with AWS command line tools and S3 operations.
Knowledge of Livy for efficient Spark job execution.
Be part of a remote-first organization where flexibility is embraced.
Work and learn alongside talented engineers and technology leaders.
Explore opportunities to learn and grow through technical and non-technical training programs.
Gain global exposure by working on products with international teams and clients.
Attend virtual and in-person international technology conferences to expand your knowledge and network.
Nursery reimbursement benefit and Aspire Wellness Program.
Exposure to working in an IT environment that adheres to rigorous security and compliance standards defined by ISO / SOC.
Stop the endless job search. Our AI finds and applies to the best jobs for you.
Discover remote opportunities in Data Engineer
Answer easy questions
200,000+ jobs across 15+ categories
Get your best job matches
Only hand-screened, legit jobs
Find a remote job faster
No ads, scams, or junk
“ I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!