A well-funded AI infrastructure company is building next-generation multimodal foundation models alongside a high-efficiency serving platform. With deep industry backing and close collaboration with hardware partners, the team is scaling rapidly to deliver the full stack powering frontier AI models and real-world, high-performance applications.
GPU Kernel Engineer
San Francisco · Hybrid
About the Role
Overview
This role is with a rapidly growing AI infrastructure company building next-generation multimodal AI systems and a high-performance training and serving platform. The team is well-funded and works closely with leading accelerator partners to push the limits of performance for large-scale AI workloads.
About the Role
We are looking for a GPU Kernel Engineer who is passionate about extracting maximum performance from modern accelerators. In this role, you will design, implement, and optimize custom GPU kernels that power large-scale AI training and inference systems.
You will work across the full hardware–software stack, from low-level kernel development to integrating optimized operations into high-level machine learning frameworks. This role is ideal for engineers who thrive at the intersection of GPU programming, systems engineering, and cutting-edge AI workloads.
Compensation: $190K–$250K base + equity
Equal Opportunity
This employer is an equal opportunity organization. All qualified applicants will be considered without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran status, or disability.
Responsibilities
- Design, implement, and optimize custom GPU kernels using C++, PTX, CUDA, ROCm, Triton, and/or JAX Pallas
- Profile and optimize end-to-end ML performance, focusing on large-scale LLM training and inference
- Integrate low-level GPU kernels into frameworks such as PyTorch, JAX, and internal runtimes
- Develop performance models, identify bottlenecks, and deliver kernel-level optimizations
- Collaborate with ML researchers, distributed systems engineers, and serving teams to improve system-wide efficiency
- Work closely with hardware vendors and stay current with evolving GPU architectures and toolchains
- Contribute to benchmarking, testing, tooling, and documentation to ensure correctness and reproducibility
Required Qualifications
- 5+ years of experience in GPU kernel development, high-performance computing, or related fields
- Bachelor’s, Master’s, or PhD in Computer Science, Computer Engineering, Electrical Engineering, Applied Mathematics, or similar
- Strong programming skills in C++ and Python
- Deep expertise in CUDA and/or ROCm, GPU memory models, and performance optimization
- Hands-on experience with Triton and/or JAX Pallas
- Strong understanding of PTX, GPU assembly, and low-level execution models
- Proven experience integrating custom GPU kernels into PyTorch, JAX, or similar frameworks
- Experience working with large-scale LLM training or inference workloads
Nice to Have
- Experience optimizing for AMD GPUs and ROCm
- Familiarity with JAX FFI and custom ML operator development
- Experience with high-performance inference frameworks (e.g., large-scale serving runtimes)
- Exposure to TPUs, XLA, or other accelerator programming environments
- Contributions to open-source ML systems, compilers, or GPU kernel projects
Benefits & Perks
- Medical, dental, and vision insurance
- 401(k) plan
- Daily lunch, snacks, and beverages
- Flexible time off
- Competitive salary and equity
Interested in this role?
Apply now and hear back within 48 hours
Join & ApplyAlready have an account? Sign in
San Francisco
Similar Jobs
This role is with a well-funded AI infrastructure company building next-generation multimodal AI models and a high-performance training and serving platform. The team is pushing the boundaries of large-scale AI systems and translating cutting-edge research into real-world, production-grade applications.
Overview This role is with a fast-growing AI infrastructure company building next-generation multimodal models and a high-performance model training and serving platform. The team is backed by significant funding and works closely with leading accelerator partners to build the full software stack powering frontier AI systems. About the Role We are seeking a highly skilled Distributed Training Engineer to design, optimize, and maintain the software stack that enables large-scale AI training workloads. You will work across the entire machine learning infrastructure—from low-level CUDA/ROCm runtimes to high-level frameworks like JAX and PyTorch—ensuring systems are fast, stable, and scalable. This role is ideal for engineers who enjoy deep systems work, debugging complex hardware–software interactions, and optimizing performance across the full ML stack. You will play a critical role in enabling the training and deployment of large language models and generative AI systems. Compensation: $190K–$250K base + equity
