Apply Now

Please mention DailyRemote when applying

AI Summary

The Data Engineer will develop, maintain, and monitor robust data pipelines to process AST outputs, knowledge graph structures, and vector embeddings. They will also collaborate with AI/ML engineers to normalize data and ensure consistency across pipeline runs.

Ciklum is looking for a Data Engineer to join our team full-time in Poland.

We are a custom product engineering company that supports both multinational organizations and scaling startups to solve their most complex business challenges. With a global team of over 4,000 highly skilled developers, consultants, analysts and product owners, we engineer technology that redefines industries and shapes the way people live.

About the role:

As a Data Engineer, become a part of a cross-functional development team engineering experiences of tomorrow.

The Legacy Code Semantic Documentation Project is a 26-week enterprise initiative for a global industrial automation leader. The primary objective is to engineer an automated, AI-assisted pipeline to generate structured, system-level Markdown documentation directly from an undocumented ~400K LOC codebase (spanning IEC 61131-3 languages and ANSI C/C++).

Operating within a dedicated, zero-data-egress secure tenant, the technical architecture combines Tree-sitter AST parsing, Neo4j knowledge graphs, Qdrant vector search, and self-hosted open-weight LLMs (Llama 3.1 / Mixtral family) with NLI-based validation. Delivery is structured across 7 work packages executing over a 6-month period, governed by strict contractual KPIs: ≥95% code coverage, ≥92% NLI-verified factual precision, and ≥90% SME validation acceptance. All outputs align with EU Cyber Resilience Act (EU CRA) requirements for SBOM and source-level traceability.

Responsibilities:

  • Data Pipeline Execution: Develop, test, and maintain robust data pipelines that process AST outputs, knowledge graph structures, and vector embeddings
  • Knowledge Base Ingestion: Write data processing scripts to ingest, normalize, and update code dependency graphs in Neo4j and vector indexes in Qdrant
  • Data Cleansing & Transformation: Parse raw source code structures and raw metadata into structured Markdown assets and standardized JSON/RAG inputs
  • Pipeline Monitoring & Debugging: Monitor execution throughput, resolve batch processing errors, and ensure data state consistency across pipeline runs
  • Collaboration: Work closely with Senior Data Engineers and AI/ML Engineers to optimize data retrieval speeds and pipeline efficiency

Requirements:

  • Professional Experience: 3+ years of commercial Data Engineering experience building and maintaining production data pipelines
  • Core Technical Proficiency: Solid proficiency in Python, SQL, and standard data manipulation frameworks
  • Hands-on Stack Exposure: Direct working knowledge of graph databases (Neo4j), vector search engines (Qdrant), or static code parsing frameworks (Tree-sitter)
  • Engineering Best Practices: Proficiency with Git version control, Docker containerization, unit/integration testing for data pipelines, and CI/CD workflows
  • Problem-Solving & Detail Orientation: Strong analytical skills with a focus on data accuracy, schema consistency, and output validation
  • Language: Professional proficiency in English (B2+/C1)

What’s in it for you?

  • Strong community: Work alongside top professionals in a friendly, open-door environment
  • Growth focus: Take on large-scale projects with a global impact and expand your expertise
  • Tailored learning: Boost your skills with internal events (meetups, conferences, workshops), Udemy access, language courses, and company-paid certifications
  • Endless opportunities: Explore diverse domains through internal mobility, finding the best fit to gain hands-on experience with cutting-edge technologies
  • Flexibility: Enjoy flexibility – full remote working possibilities
  • Care: We've got you covered with company-paid premium medical package through Luxmed
  • Benefits: Access the MyBenefit cafeteria platform, allowing you to choose perks that best suit your lifestyle and needs

About us:

At Ciklum, we are always exploring innovations, empowering each other to achieve more, and engineering solutions that matter. With us, you’ll work with cutting-edge technologies, contribute to impactful projects, and be part of a One Team culture that values collaboration and progress.

With delivery centers in Wrocław and Gdańsk, our 300+ professionals in Poland drive forward-thinking solutions for global clients. Join a community where collaboration sparks innovation—and your impact reaches millions.

Explore, empower, engineer with Ciklum!

Interested already? We would love to get to know you! Submit your application. We can’t wait to see you at Ciklum.

#LI-RS1

Similar Jobs

See all Remote Software Development jobs →

Personalize your Remote Job Search in 3 Easy Steps!

Discover remote opportunities in Data Engineer

Answer easy questions

Answer easy questions

200,000+ jobs across 15+ categories

Get your best job matches

Get your best job matches

Only hand-screened, legit jobs

Find a remote job faster

Find a remote job faster

No ads, scams, or junk

I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!

Sarah J. — Sarah J. · Marketing Manager ★★★★★ Verified