Back to Jobs

Lead Software Engineer, Model Serving Platform

San Francisco · Hybrid

San FranciscoHybrid$230,000 - $300,000
Software EngineersSenior (5-8 years)

About the Role

Overview

This role is with a well-funded AI infrastructure company building next-generation multimodal models and a high-performance model serving platform. The team is scaling rapidly to deliver production-grade systems that power real-time AI applications, working closely across research, systems, and infrastructure.

About the Role

This is a rare opportunity to help architect and lead the development of a next-generation model serving platform—the core engine that brings highly efficient multimodal foundation models into production.

As a senior technical leader, you will both build critical components yourself and guide other engineers, shaping architectural decisions, engineering standards, and execution quality. You’ll work across the full AI stack, from GPU execution and optimized runtimes to distributed serving, scheduling, and APIs that power low-latency, real-time inference.

This role is ideal for engineers who enjoy deep systems work, thrive on ownership, and want to lead the development of foundational AI infrastructure.

Equal Opportunity

This employer is an equal opportunity organization. All qualified applicants will be considered without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran status, or disability.

Compensation: $230K–$300K base + equity

Responsibilities

  • Lead the technical direction of the model serving platform, owning architecture and execution decisions
  • Design and build core serving components, including execution runtimes, batching, scheduling, and distributed inference systems
  • Develop high-performance C++ and GPU-accelerated modules, including custom kernels and memory-efficient runtimes
  • Collaborate closely with ML researchers to productionize new multimodal models with low-latency and high throughput
  • Build Python APIs and services that expose model capabilities to downstream applications
  • Mentor engineers through design reviews, code reviews, and hands-on technical guidance
  • Drive performance profiling, benchmarking, and observability across the inference stack
  • Ensure system reliability and maintainability through testing, monitoring, and strong engineering practices

Required Qualifications

  • Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience
  • 5+ years of experience building scalable backend systems or distributed infrastructure
  • Strong understanding of LLM inference mechanics (prefill vs. decode, batching strategies, KV cache management)
  • Experience with Kubernetes, Ray, and containerized production environments
  • Strong proficiency in C++ and Python
  • Excellent debugging, profiling, and system-level performance optimization skills
  • Ability to collaborate closely with ML researchers and translate research requirements into production systems
  • Strong communication skills and experience leading technical discussions and mentoring engineers
  • Comfortable working in a fast-paced, high-ownership, in-office environment

Nice to Have

  • Experience with ML systems engineering and distributed GPU scheduling
  • Familiarity with high-performance inference engines (e.g., open-source or custom serving runtimes)
  • Experience building large-scale ML or MLOps infrastructure
  • Hands-on experience with CUDA or ROCm and GPU profiling tools
  • Background in AI startups, research labs, or large-scale infrastructure teams
  • Exposure to multimodal model architectures or advanced inference optimization techniques
  • Contributions to open-source ML, systems, or HPC projects

Benefits & Perks

  • Medical, dental, and vision insurance
  • 401(k) plan
  • Daily lunch, snacks, and beverages
  • Flexible time off
  • Competitive salary and equity

Interested in this role?

Apply now and hear back within 48 hours

Join & Apply

Already have an account? Sign in

San Francisco

Similar Jobs

San Francisco · Hybrid

A well-funded AI infrastructure company is building next-generation multimodal foundation models and a high-efficiency serving platform. With deep industry backing and close collaboration with hardware partners, the team is scaling rapidly to deliver the full stack powering real-time AI applications and developer-facing platforms.

$190,000 - $250,000Apply
San Francisco · Hybrid

A well-funded AI infrastructure company is building next-generation multimodal foundation models and a high-efficiency serving platform. With deep industry backing and close collaboration with hardware partners, the team is scaling rapidly to deliver the full stack powering real-time, production-grade AI applications.

$190,000 - $250,000Apply
New York City, San Francisco · Remote

This role is with a rapidly growing, venture-backed AI software company building an intelligent operating layer for complex, integration-heavy enterprises. The platform helps organizations plan, deploy, and optimize their teams through advanced analytics, automation, and AI-driven workflows. The company works in real-world, security-constrained environments—including hybrid, on-prem, and restricted systems—and is focused on shipping durable, production-grade enterprise software.

$180,000 - $210,000Apply

Hiring

First candidates by day three.

One form, one call, a written search plan inside 24 hours. You pay nothing until someone signs.

Start a search

Looking

Roles that never hit a job board.

One profile covers every search we run. We only reach out when the role clears your bar — always free, never mass mail.

👋Hey — hiring engineers? Click me for jokes.