Match your resume skills with our AI powered skill match!
Questions interviewers often ask for this role, with sample answers.
Upload your resume and we draft a letter for this exact role, tailored to what it asks for.
You will own the end-to-end data lifecycle, from scraping and ingestion to modeling and exposing data through the platform's AI and API features. You will also manage the data platform's reliability, build analytical layers, and collaborate with product and AI squads to improve data-driven features.
Remote (Spain) or hybrid from Barcelona · Full-time
Tendios is the tender intelligence platform for the Spanish public procurement market. Companies use Bid to find, qualify and win public tenders. Public institutions use Create to draft and manage their own. Behind both products is one of the most complete datasets on Spanish public procurement: tenders, lots, CPVs, contracting bodies, resolutions and awards. It’s collected continuously from hundreds of public sources and enriched with AI.
That data is our product. We’re looking for the person who will own it end to end.
This is not a classic data architect role. You won’t just design schemas and hand them off. You’ll own Tendios data from the moment it’s scraped to the moment a customer sees it in Bid, asks Vera about it, or queries it through our MCP server. You’ll write the code, propose the features, and make the calls on how data is modelled, stored and exposed, including how it feeds our AI features.
You’ll report directly to the Head of Technology and work closely with product, the AI & Data squad and the backend squads.
Propose and build data-driven features. Examples: market and competitor intelligence from award data, pricing benchmarks, contracting-body profiles, better tender matching and alerting, and data-quality signals shown to customers.
Own the data side of our AI features. Design and improve the retrieval layer behind Vera and our Dynamic RAG service: document ingestion and conversion (Docvert), chunking, embeddings, Qdrant indexing, hybrid search with Elasticsearch, and reranking.
Use LLMs where they add real value in the pipeline: structured extraction from tender documents, classification (e.g. CPVs), entity resolution and summarisation. Build evaluation sets to measure quality instead of guessing.
Work with the AI team so Vera and the Tendios MCP server get clean, well-structured, well-documented data and tools.
Define and measure data and retrieval quality: coverage per source, freshness, parsing accuracy, deduplication and RAG answer quality. Treat it as a product metric, not an afterthought.
Take ownership of the collection pipeline (Crawl Manager, Tenders Discovery, Collector, Tenders and Resolution Parsers, RabbitMQ consumers) and make it more reliable, observable and cheaper to run..
Keep our read models consistent with the system of record: Elasticsearch for search and filtering, and Qdrant for semantic search and RAG.
Build out the analytics layer on ClickHouse, orchestrated with Prefect and exposed through Metabase. This includes replacing our manual, Excel-based SaaS metrics reporting (MRR, churn, NRR) with a proper warehouse.
Write production code in Python for the pipeline and data services and in TypeScript/Node.js in our Turborepo backend (NestJS) where data features touch the API.
Review code, set standards for data modelling and migrations, and document decisions in Confluence.
Contribute to data governance and security as part of our compliance work (ENS, ISMS), including data ownership, retention, access and auditability.
6+ years in software engineering, with at least 3 years focused on data-intensive systems.
Strong Python and solid SQL. Deep, hands-on PostgreSQL experience: modelling, performance and migrations at scale.
Comfortable working across the stack when needed, including reading and writing TypeScript/NestJS.
Experience building and running production data pipelines, including event-driven or queue-based architectures (RabbitMQ, Kafka or similar).
Experience with search or analytical stores such as Elasticsearch/OpenSearch or ClickHouse.
A solid, hands-on understanding of how LLM applications work: RAG, embeddings and vector search, chunking strategies, prompt design, tool calling, and how to evaluate retrieval and answer quality. You’ve shipped at least one LLM or RAG feature to production.
A product mindset. You look at a dataset and see features, and you can write a clear proposal and defend it with product and business stakeholders.
Fluent Spanish and English
Web scraping and document parsing at scale, including PDFs and messy semi-structured sources.
Qdrant or other vector databases at scale, hybrid search and reranking.
LLM observability and evaluation tooling (e.g. Langfuse), or running self-hosted open models (e.g. Qwen) on GPU infrastructure.
Orchestration and ELT tooling such as Prefect, Airflow, dlt or dbt.
Experience with large or legacy data migrations (MongoDB to PostgreSQL is a big bonus).
Knowledge of public procurement, open data or regulated environments.
Docker, and experience with cloud and hybrid infrastructure (Hetzner, DigitalOcean, AWS).
Your work is the product. Better data means better tender matching, better answers from Vera and better decisions for our customers, and you’ll see that directly.
Real ownership. You’ll help define the data strategy of a growing SaaS company, not execute someone else’s.
Interesting problems. You’ll work on scraping at scale, entity resolution across public bodies and suppliers, polyglot persistence, and AI on top of a unique domain dataset.
An engineering team organised in autonomous squads, with a modern stack and a pragmatic culture. Work fully remote from anywhere in Spain, or hybrid from our Barcelona office.
Stop the endless job search. Our AI finds and applies to the best jobs for you.
Featuring 216,786+ Jobs in Full Stack Developer
Answer easy questions
216,786+ jobs across 15+ categories
Get your best job matches
Only hand-screened, legit jobs
Find a remote job faster
No ads, scams, or junk
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”