Careers


Data Acquisition & Pipeline Engineer

Job Description

High-quality pre-training data is a critical foundation for building capable foundation models.

Our team develops the systems that collect, clean, process, and manage large-scale, multimodal data for foundation-model training.

You will work on web-scale data acquisition and processing infrastructure that improves data quality, coverage, freshness, and reliability.

Job Responsibility

Design and develop web-scale data selection and scheduling systems based on data-quality feedback, including URL quality scoring, metadata aggregation, crawl quotas, and data prioritization.

For time-sensitive use cases such as RAG, identify, schedule, and validate up-to-date data sources and seed URLs to maintain the freshness of critical information.

Design and develop the end-to-end data acquisition pipeline, including download services, task scheduling, stream processing, DNS resolution, and network connectivity, to build a high-throughput and fault-tolerant distributed system.

Job Requirement

Bachelor’s degree or higher in computer science, mathematics, or a related field.

Proficiency in at least one programming language, such as Python, C++, or Rust.

Preferred Qualifications

Strong performance in mathematics or informatics competitions.

Experience building large-scale data collection, web-scraping, or data-pipeline systems.