Back to Jobs

Distributed Training Engineer

San Francisco · Hybrid

San FranciscoHybrid$190,000 - $250,000
ML/AI EngineersSenior (5-8 years)

About the Role

Overview

This role is with a fast-growing AI infrastructure company building next-generation multimodal models and a high-performance model training and serving platform. The team is backed by significant funding and works closely with leading accelerator partners to build the full software stack powering frontier AI systems.

About the Role

We are seeking a highly skilled Distributed Training Engineer to design, optimize, and maintain the software stack that enables large-scale AI training workloads. You will work across the entire machine learning infrastructure—from low-level CUDA/ROCm runtimes to high-level frameworks like JAX and PyTorch—ensuring systems are fast, stable, and scalable.

This role is ideal for engineers who enjoy deep systems work, debugging complex hardware–software interactions, and optimizing performance across the full ML stack. You will play a critical role in enabling the training and deployment of large language models and generative AI systems.

Compensation: $190K–$250K base + equity

Responsibilities

  • ML Stack Ownership: Build, maintain, and continuously improve the full ML training stack, from GPU drivers and runtimes to high-level framework tooling.
  • Framework Optimization: Maintain and optimize core libraries such as PyTorch, JAX, CUDA, and ROCm across multiple environments and hardware configurations.
  • Distributed Training: Ensure models are efficiently sharded and configured for large-scale, multi-node distributed training.
  • System Integration: Validate correctness, memory efficiency, and scalability across accelerator clusters.
  • Performance Profiling: Profile compilation graphs and training workloads to identify and eliminate bottlenecks.
  • Debugging & Reliability: Troubleshoot complex issues such as runtime failures, memory leaks, kernel inconsistencies, and distributed system instability.
  • Cross-Team Collaboration: Work closely with research, infrastructure, and kernel engineering teams to improve throughput, stability, and developer experience.

Required Qualifications

  • 5+ years of experience in ML systems, distributed training, or related fields
  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or similar
  • Strong programming skills in Python and C++
  • Experience with profiling tools (e.g., Nsight, ROCm Profiler, XLA profiler)
  • Deep understanding of distributed training and partitioning in frameworks like PyTorch and JAX
  • Hands-on experience with CUDA, ROCm, NCCL, XLA, or similar technologies
  • Experience working with multi-node distributed training systems and orchestration frameworks

Nice to Have

  • Deep experience with XLA/JAX internals and compilation paths
  • Familiarity with large-scale inference or serving frameworks (e.g., vLLM, TensorRT)
  • Background in GPU kernel optimization or accelerator-aware model partitioning
  • Strong understanding of low-level C++ components used in ML frameworks

Benefits & Perks

  • Medical, dental, and vision insurance
  • 401(k) plan
  • Daily lunch, snacks, and beverages
  • Flexible time off
  • Competitive salary and equity

Interested in this role?

Apply now and hear back within 48 hours

Join & Apply

Already have an account? Sign in

San Francisco

Similar Jobs

San Francisco · Hybrid

A well-funded AI infrastructure company is building next-generation multimodal foundation models alongside a high-efficiency serving platform. With deep industry backing and close collaboration with hardware partners, the team is scaling rapidly to deliver the full stack powering frontier AI models and real-world, high-performance applications.

$200,000 - $275,000Apply
San Francisco · Hybrid

This role is with a well-funded AI infrastructure company building next-generation multimodal AI models and a high-performance training and serving platform. The team is pushing the boundaries of large-scale AI systems and translating cutting-edge research into real-world, production-grade applications.

$190,000 - $250,000Apply
San Francisco · Hybrid

Overview This role is with a rapidly growing AI infrastructure company building next-generation multimodal AI systems and a high-performance training and serving platform. The team is well-funded and works closely with leading accelerator partners to push the limits of performance for large-scale AI workloads. About the Role We are looking for a GPU Kernel Engineer who is passionate about extracting maximum performance from modern accelerators. In this role, you will design, implement, and optimize custom GPU kernels that power large-scale AI training and inference systems. You will work across the full hardware–software stack, from low-level kernel development to integrating optimized operations into high-level machine learning frameworks. This role is ideal for engineers who thrive at the intersection of GPU programming, systems engineering, and cutting-edge AI workloads. Compensation: $190K–$250K base + equity Equal Opportunity This employer is an equal opportunity organization. All qualified applicants will be considered without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran status, or disability.

$190,000 - $250,000Apply

Hiring

First candidates by day three.

One form, one call, a written search plan inside 24 hours. You pay nothing until someone signs.

Start a search

Looking

Roles that never hit a job board.

One profile covers every search we run. We only reach out when the role clears your bar — always free, never mass mail.

👋Hey — hiring engineers? Click me for jokes.