Back to Jobs

Data Scientist (Foundation Models & Data Curation)

San Francisco · Hybrid

San FranciscoHybrid$190,000 - $275,000Posted 8 months ago
Data EngineersSenior (5-8 years)

About the Role

Overview

A well-funded AI infrastructure company is building next-generation multimodal foundation models alongside a high-efficiency serving platform. Backed by major industry partners and deep technical collaboration, the team is scaling rapidly to deliver the full stack powering frontier AI models and real-time applications.

About the Role

We are seeking a highly technical and forward-thinking Data Scientist to lead the strategy, creation, and curation of massive datasets that power large-scale foundation models. In this role, data is treated as a first-class product and the primary competitive advantage.

You will own the entire data lifecycle—from web-scale raw data ingestion to high-quality human-aligned datasets that shape model behavior. This role is ideal for someone who sees data as both a large-scale engineering problem and a scientific challenge, with direct impact on model reasoning, safety, and multimodal performance.

Equal Opportunity

We are an equal opportunity employer and value diversity at our company. We do not discriminate on the basis of race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, disability, veteran status, or any other legally protected characteristic.

Responsibilities

  • Own end-to-end creation of large-scale pre-training datasets for foundation models
  • Design data mixing strategies across web, code, books, and technical sources
  • Build scalable pipelines for data cleaning, deduplication, and quality filtering
  • Curate and maintain petabyte-scale unstructured datasets
  • Develop post-training datasets including SFT, multi-turn dialogue, and preference data (RLHF/DPO)
  • Manage human-labeling workflows and define quality gold-standard datasets
  • Drive acquisition and processing of image and video datasets for multimodal models
  • Develop high-performance Python data pipelines using multiprocessing and multithreading
  • Perform statistical analysis to identify bias, gaps, and quality regressions in datasets
  • Design synthetic data pipelines to augment reasoning and knowledge gaps

Required Qualifications

  • 5+ years of industry experience in Data Science or Machine Learning
  • Expert-level Python, with strong performance optimization skills
  • Hands-on experience working with petabyte-scale datasets
  • Proven experience building LLM or large vision model datasets from scratch
  • Deep experience with post-training datasets (RLHF, DPO, instruction tuning)
  • Strong familiarity with large-scale data frameworks (Spark, Ray) and formats (Parquet, WebDataset)

Nice to Have

  • Experience curating large-scale image or video datasets (e.g., LAION-style pipelines)
  • Familiarity with multimodal crawling and video processing pipelines
  • Experience designing taxonomies for reasoning, coding, or math benchmarks
  • Master’s or PhD in a quantitative field (CS, ML, Data Science, IR, etc.)

Benefits & Perks

  • Medical, dental, and vision insurance
  • 401(k)
  • Daily meals and snacks
  • Flexible time off
  • Competitive compensation and equity

Interested in this role?

Apply now and hear back within 48 hours

Join & Apply

Already have an account? Sign in

Posted 8 months ago

San Francisco

Hiring

First candidates by day three.

One form, one call, a written search plan inside 24 hours. You pay nothing until someone signs.

Start a search

Looking

Roles that never hit a job board.

One profile covers every search we run. We only reach out when the role clears your bar — always free, never mass mail.

👋Hey — hiring engineers? Click me for jokes.