You will develop and maintain scalable data processing pipelines and machine learning models for large-scale NLP and LLM applications. Additionally, you will collaborate with cross-functional teams to improve dataset quality and implement data-driven solutions for AI training.
Maintain and optimize database queries and data systems to ensure efficient data access and reliability. Assist in developing and improving data pipelines for collecting, processing, and delivering large-scale datasets.
You will build and maintain backend infrastructure using TypeScript while optimizing production systems through metrics, monitoring, and automated deployments. Additionally, you will troubleshoot production issues and improve the existing codebase through rigorous testing and code enhancements.
You will write, test, and refine code to extract data from online sources while ensuring reliability and efficiency. Additionally, you will manage database storage, monitor scraping processes, and clean extracted data to meet quality standards.
Review applications before release to ensure they meet requirements and identify potential compliance issues. Provide actionable feedback to engineering and operations teams while maintaining review documentation.
Design and optimize scalable data pipeline infrastructure for real-time and batch processing. Develop reliable backend features and maintain high engineering standards through technical design and code reviews.
Design and maintain large-scale web crawlers and high-throughput data collection systems for research and model development. Develop pipelines for data cleaning, normalization, and quality monitoring while collaborating with research teams to meet modeling needs.