For Employers
Apply Now

Please mention DailyRemote when applying

?
Resume Match Score

See how much of this job your resume covers, and what’s missing.

Want a recruiter to go through it line by line?

Get professional review

Create a cover letter for this job

Upload your resume and we draft a letter for this exact role, tailored to what it asks for.

  • Tailored to this role
  • Based on your resume
  • Fully editable
AI Summary

You will build and operate scalable ingestion, ELT/ETL, and orchestration pipelines to transform raw data into clean, AI-ready datasets. Additionally, you will implement real-time data ingestion flows and ensure high data quality through rigorous testing and observability.

This is a hands-on building role: you turn raw, messy fabrication data into the clean, well-modeled, AI-ready datasets that our AI/ML and analytics workloads run on ๐Ÿš€

๐Ÿง‘๐Ÿปโ€๐Ÿ’ป Responsibilities: 

  • Pipeline Development & Operation: Build and operate scalable ingestion, ELT/ETL, and orchestration pipelines (batch and real-time streaming) within Microsoft Fabric and cloud lakehouse environments.

  • Real-Time Data Ingestion: Design and implement low-latency, real-time data ingestion flows to support live operational analytics and streaming workloads.

  • Data Modeling & Layering: Implement layered (medallion-style: Bronze/Silver/Gold) architectures using PySpark/SQL with idempotent, backfillable, and incrementally loaded jobs.

  • Data Quality & Governance: Apply deduplication, normalization, schema validation, and lineage tracking to ensure downstream data is high-quality, trustworthy, and audit-ready.

  • AI & Analytics Readiness: Deliver feature-ready, curated datasets to support business intelligence, analytics, vector search, and AI/ML agentic workloads.

  • Observability & Reliability: Establish testing, monitoring, and pipeline observability (freshness, volume, schema drift) with clear alerting to resolve failures proactively.

  • Tooling & AI Development: Utilize AI-assisted development tools (Claude Code, Copilot, Cursor) as a force multiplier for writing pipelines, query tuning, and data transformation scripts.

๐Ÿค If you have:

  • Experience: 5+ years of hands-on data engineering experience building and operating production data pipelines at scale.

  • Core Technical Stack: Strong proficiency in Python, SQL, and PySpark / Apache Spark, backed by solid software engineering fundamentals (Git, CI/CD, unit/integration testing).

  • Real-Time Data Processing: Demonstrated hands-on experience implementing real-time data ingestion and streaming pipelines (not limited to batch processing).

  • Data Architecture & Modeling: Proven experience in end-to-end data modeling, schema design, and layered lakehouse architectures (Medallion architecture).

  • Platform Experience: Experience with cloud-native lakehouse platforms; hands-on experience or familiarity with Microsoft Fabric is highly preferred.

  • Data Quality & Observability: Strong grasp of data testing frameworks, pipeline monitoring, and data quality enforcement.

  • AI Tooling: Active experience leveraging AI-assisted development tools (Cursor, Copilot, Claude) to accelerate engineering velocity.

๐Ÿฆพ Itโ€™s a plus:

  • Hands-on experience with Microsoft Fabric (Fabric Lakehouse, Data Factory, Synapse Analytics).

  • Experience extracting data from document stores / NoSQL databases (specifically MongoDB / MongoDB Atlas and Change Streams / CDC).

  • Streaming frameworks experience (Event Hubs, Kafka, Spark Structured Streaming).

  • Exposure to vector embeddings, RAG-ready datasets, or feature stores for AI/ML workloads.

  • AEC / Construction / MEP domain experience.


This call is made within the framework of Law 19.691 on the Promotion of Employment for Persons with Disabilities, including individuals registered in the National Registry of Persons with Disabilities of the Ministry of Social Development

Automatically Apply to the Best Remote Jobs

Stop the endless job search. Our AI finds and applies to the best jobs for you.

Try it Now
Keep looking

Similar Jobs

See all Remote Software Development jobs →

Batch Developer (UNIX/C/SQL)

Full Time United States $60000 - $80000 per year Software Development

Site Reliability Engineer

Full Time Argentina, Brazil, Costa Rica +1 more Software Development

Senior Data Engineer (all genders)

Full Time, Part Time Germany Software Development

Developer

Full Time Portugal Software Development

Outside Sales Engineer - OEM (South)

Full Time United States $83500 - $130K per year Software Development

Outside Sales Engineer - OEM (Midwest US)

Full Time United States $83500 - $130K per year Software Development
Apply Now

Personalize your Remote Job Search in 3 Easy Steps!

Featuring 216,738+ Jobs in Data Engineer

Answer easy questions

Answer easy questions

216,738+ jobs across 15+ categories

Get your best job matches

Get your best job matches

Only hand-screened, legit jobs

Find a remote job faster

Find a remote job faster

No ads, scams, or junk

“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”

Sarah J. — Sarah J. · Marketing Manager ★★★★★ Verified