Careers
Data Acquisition & Pipeline Engineer
Job Description
High-quality pre-training data is a critical foundation for building capable foundation models.
Our team develops the systems that collect, clean, process, and manage large-scale, multimodal data for foundation-model training.
You will work on web-scale data acquisition and processing infrastructure that improves data quality, coverage, freshness, and reliability.
Job Responsibility
Design and develop web-scale data selection and scheduling systems based on data-quality feedback, including URL quality scoring, metadata aggregation, crawl quotas, and data prioritization.
For time-sensitive use cases such as RAG, identify, schedule, and validate up-to-date data sources and seed URLs to maintain the freshness of critical information.
Design and develop the end-to-end data acquisition pipeline, including download services, task scheduling, stream processing, DNS resolution, and network connectivity, to build a high-throughput and fault-tolerant distributed system.
Job Requirement
Bachelor’s degree or higher in computer science, mathematics, or a related field.
Proficiency in at least one programming language, such as Python, C++, or Rust.
Preferred Qualifications
Strong performance in mathematics or informatics competitions.
Experience building large-scale data collection, web-scraping, or data-pipeline systems.