This role is with a well-funded AI infrastructure company building next-generation multimodal AI models and a high-performance training and serving platform. The team is pushing the boundaries of large-scale AI systems and translating cutting-edge research into real-world, production-grade applications.
Machine Learning Engineer
San Francisco · Hybrid
About the Role
Overview
A well-funded AI infrastructure company is building next-generation multimodal foundation models alongside a high-efficiency serving platform. With deep industry backing and close collaboration with hardware partners, the team is scaling rapidly to deliver the full stack powering frontier AI models and real-world, high-performance applications.
About the Role
As a Research Engineer / Machine Learning Engineer, you will work across the entire foundation model lifecycle: large-scale pretraining, post-training and reinforcement learning, sandbox environments for evaluation and agentic learning, and deployment plus inference optimization.
This role is ideal for someone who enjoys moving quickly between research and production—turning ideas into scalable systems, contributing to core infrastructure, and shipping models that operate reliably at scale.
Equal Opportunity
We are an equal opportunity employer and value diversity at our company. All qualified applicants will be considered without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, disability, veteran status, or any other protected characteristic.
Responsibilities
- Train large-scale foundation models across massive, heterogeneous datasets
- Design stable training recipes and scaling strategies for novel architectures
- Build and maintain distributed training infrastructure and fault-tolerant pipelines
- Develop post-training pipelines including SFT, preference optimization, and RL
- Curate and generate targeted datasets to improve specific model capabilities
- Build reward models and evaluation frameworks for iterative model improvement
- Create scalable sandbox environments for agent evaluation and learning
- Design high-signal automated evaluations for reasoning, tool use, and safety
- Optimize inference throughput and latency for large-scale model serving
- Build high-performance serving systems with batching, caching, and quantization
- Profile and optimize GPU utilization, runtime bottlenecks, and memory behavior
- Own end-to-end ML systems from research prototypes to production deployment
Required Qualifications
- Strong software engineering fundamentals with experience building performant systems
- Experience training or serving large neural networks (LLMs or similar models)
- Solid understanding of deep learning fundamentals and modern ML literature
- Experience working in GPU-based and distributed computing environments
- Proficiency in Python and familiarity with high-performance ML frameworks
- Hands-on experience with distributed training or large-scale inference systems
Nice to Have
- Experience with large-scale distributed training frameworks (FSDP, ZeRO, Megatron)
- Experience with post-training methods such as RLHF, RLAIF, or preference optimization
- Experience building reinforcement learning environments or agent frameworks
- Experience with inference optimization, quantization, or kernel-level profiling
- Experience building internet-scale data ingestion or ETL pipelines
- Master’s or PhD in Computer Science, Machine Learning, AI, or a related field
Benefits & Perks
- Medical, dental, and vision insurance
- 401(k)
- Daily meals and snacks
- Flexible time off
- Competitive compensation and equity
Interested in this role?
Apply now and hear back within 48 hours
Join & ApplyAlready have an account? Sign in
Posted 8 months ago
San Francisco
Similar Jobs
Overview This role is with a rapidly growing AI infrastructure company building next-generation multimodal AI systems and a high-performance training and serving platform. The team is well-funded and works closely with leading accelerator partners to push the limits of performance for large-scale AI workloads. About the Role We are looking for a GPU Kernel Engineer who is passionate about extracting maximum performance from modern accelerators. In this role, you will design, implement, and optimize custom GPU kernels that power large-scale AI training and inference systems. You will work across the full hardware–software stack, from low-level kernel development to integrating optimized operations into high-level machine learning frameworks. This role is ideal for engineers who thrive at the intersection of GPU programming, systems engineering, and cutting-edge AI workloads. Compensation: $190K–$250K base + equity Equal Opportunity This employer is an equal opportunity organization. All qualified applicants will be considered without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran status, or disability.
Overview This role is with a fast-growing AI infrastructure company building next-generation multimodal models and a high-performance model training and serving platform. The team is backed by significant funding and works closely with leading accelerator partners to build the full software stack powering frontier AI systems. About the Role We are seeking a highly skilled Distributed Training Engineer to design, optimize, and maintain the software stack that enables large-scale AI training workloads. You will work across the entire machine learning infrastructure—from low-level CUDA/ROCm runtimes to high-level frameworks like JAX and PyTorch—ensuring systems are fast, stable, and scalable. This role is ideal for engineers who enjoy deep systems work, debugging complex hardware–software interactions, and optimizing performance across the full ML stack. You will play a critical role in enabling the training and deployment of large language models and generative AI systems. Compensation: $190K–$250K base + equity
