A well-funded AI infrastructure company is building next-generation multimodal foundation models alongside a high-efficiency serving platform. With deep industry backing and close collaboration with hardware partners, the team is scaling rapidly to deliver the full stack powering frontier AI models and real-world, high-performance applications.
Distributed Training Engineer
San Francisco · Hybrid
About the Role
Overview
This role is with a fast-growing AI infrastructure company building next-generation multimodal models and a high-performance model training and serving platform. The team is backed by significant funding and works closely with leading accelerator partners to build the full software stack powering frontier AI systems.
About the Role
We are seeking a highly skilled Distributed Training Engineer to design, optimize, and maintain the software stack that enables large-scale AI training workloads. You will work across the entire machine learning infrastructure—from low-level CUDA/ROCm runtimes to high-level frameworks like JAX and PyTorch—ensuring systems are fast, stable, and scalable.
This role is ideal for engineers who enjoy deep systems work, debugging complex hardware–software interactions, and optimizing performance across the full ML stack. You will play a critical role in enabling the training and deployment of large language models and generative AI systems.
Compensation: $190K–$250K base + equity
Responsibilities
- ML Stack Ownership: Build, maintain, and continuously improve the full ML training stack, from GPU drivers and runtimes to high-level framework tooling.
- Framework Optimization: Maintain and optimize core libraries such as PyTorch, JAX, CUDA, and ROCm across multiple environments and hardware configurations.
- Distributed Training: Ensure models are efficiently sharded and configured for large-scale, multi-node distributed training.
- System Integration: Validate correctness, memory efficiency, and scalability across accelerator clusters.
- Performance Profiling: Profile compilation graphs and training workloads to identify and eliminate bottlenecks.
- Debugging & Reliability: Troubleshoot complex issues such as runtime failures, memory leaks, kernel inconsistencies, and distributed system instability.
- Cross-Team Collaboration: Work closely with research, infrastructure, and kernel engineering teams to improve throughput, stability, and developer experience.
Required Qualifications
- 5+ years of experience in ML systems, distributed training, or related fields
- Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or similar
- Strong programming skills in Python and C++
- Experience with profiling tools (e.g., Nsight, ROCm Profiler, XLA profiler)
- Deep understanding of distributed training and partitioning in frameworks like PyTorch and JAX
- Hands-on experience with CUDA, ROCm, NCCL, XLA, or similar technologies
- Experience working with multi-node distributed training systems and orchestration frameworks
Nice to Have
- Deep experience with XLA/JAX internals and compilation paths
- Familiarity with large-scale inference or serving frameworks (e.g., vLLM, TensorRT)
- Background in GPU kernel optimization or accelerator-aware model partitioning
- Strong understanding of low-level C++ components used in ML frameworks
Benefits & Perks
- Medical, dental, and vision insurance
- 401(k) plan
- Daily lunch, snacks, and beverages
- Flexible time off
- Competitive salary and equity
Interested in this role?
Apply now and hear back within 48 hours
Join & ApplyAlready have an account? Sign in
San Francisco
Similar Jobs
This role is with a well-funded AI infrastructure company building next-generation multimodal AI models and a high-performance training and serving platform. The team is pushing the boundaries of large-scale AI systems and translating cutting-edge research into real-world, production-grade applications.
Overview This role is with a rapidly growing AI infrastructure company building next-generation multimodal AI systems and a high-performance training and serving platform. The team is well-funded and works closely with leading accelerator partners to push the limits of performance for large-scale AI workloads. About the Role We are looking for a GPU Kernel Engineer who is passionate about extracting maximum performance from modern accelerators. In this role, you will design, implement, and optimize custom GPU kernels that power large-scale AI training and inference systems. You will work across the full hardware–software stack, from low-level kernel development to integrating optimized operations into high-level machine learning frameworks. This role is ideal for engineers who thrive at the intersection of GPU programming, systems engineering, and cutting-edge AI workloads. Compensation: $190K–$250K base + equity Equal Opportunity This employer is an equal opportunity organization. All qualified applicants will be considered without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran status, or disability.
