Data Scientist (Foundation Models & Data Curation)
San Francisco · Hybrid
About the Role
Overview
A well-funded AI infrastructure company is building next-generation multimodal foundation models alongside a high-efficiency serving platform. Backed by major industry partners and deep technical collaboration, the team is scaling rapidly to deliver the full stack powering frontier AI models and real-time applications.
About the Role
We are seeking a highly technical and forward-thinking Data Scientist to lead the strategy, creation, and curation of massive datasets that power large-scale foundation models. In this role, data is treated as a first-class product and the primary competitive advantage.
You will own the entire data lifecycle—from web-scale raw data ingestion to high-quality human-aligned datasets that shape model behavior. This role is ideal for someone who sees data as both a large-scale engineering problem and a scientific challenge, with direct impact on model reasoning, safety, and multimodal performance.
Equal Opportunity
We are an equal opportunity employer and value diversity at our company. We do not discriminate on the basis of race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, disability, veteran status, or any other legally protected characteristic.
Responsibilities
- Own end-to-end creation of large-scale pre-training datasets for foundation models
- Design data mixing strategies across web, code, books, and technical sources
- Build scalable pipelines for data cleaning, deduplication, and quality filtering
- Curate and maintain petabyte-scale unstructured datasets
- Develop post-training datasets including SFT, multi-turn dialogue, and preference data (RLHF/DPO)
- Manage human-labeling workflows and define quality gold-standard datasets
- Drive acquisition and processing of image and video datasets for multimodal models
- Develop high-performance Python data pipelines using multiprocessing and multithreading
- Perform statistical analysis to identify bias, gaps, and quality regressions in datasets
- Design synthetic data pipelines to augment reasoning and knowledge gaps
Required Qualifications
- 5+ years of industry experience in Data Science or Machine Learning
- Expert-level Python, with strong performance optimization skills
- Hands-on experience working with petabyte-scale datasets
- Proven experience building LLM or large vision model datasets from scratch
- Deep experience with post-training datasets (RLHF, DPO, instruction tuning)
- Strong familiarity with large-scale data frameworks (Spark, Ray) and formats (Parquet, WebDataset)
Nice to Have
- Experience curating large-scale image or video datasets (e.g., LAION-style pipelines)
- Familiarity with multimodal crawling and video processing pipelines
- Experience designing taxonomies for reasoning, coding, or math benchmarks
- Master’s or PhD in a quantitative field (CS, ML, Data Science, IR, etc.)
Benefits & Perks
- Medical, dental, and vision insurance
- 401(k)
- Daily meals and snacks
- Flexible time off
- Competitive compensation and equity
Interested in this role?
Apply now and hear back within 48 hours
Join & ApplyAlready have an account? Sign in
Posted 8 months ago
San Francisco
