Back to Jobs

Software Engineer, Backend (AI Infrastructure & Model Serving)

San Francisco · Hybrid

San FranciscoHybrid$190,000 - $250,000Posted 8 months ago
Software EngineersMid-Level (2-5 years)

About the Role

Overview

A well-funded AI infrastructure company is building next-generation multimodal foundation models and a high-efficiency serving platform. With deep industry backing and close collaboration with hardware partners, the team is scaling rapidly to deliver the full stack powering real-time, production-grade AI applications.

About the Role

This role focuses on building the core backend systems that power large-scale AI model serving. You will work across C++, Python, runtime execution, and distributed infrastructure to create a fast, reliable platform for deploying and operating multimodal AI models in production.

You’ll collaborate closely with ML researchers and systems engineers, gain hands-on exposure to performance engineering, and help optimize how large AI models are deployed and served at scale. This role is ideal for engineers who care deeply about performance, reliability, and low-level systems work while wanting exposure to the full AI stack.

Responsibilities

  • Build and maintain the backend model serving platform, including APIs and distributed inference systems
  • Develop control-plane services such as billing, monitoring, and service orchestration
  • Collaborate with ML researchers to integrate new multimodal models into production workflows
  • Write reliable, maintainable, and well-tested backend code with strong documentation practices
  • Provide operational support to ensure production services are performant, available, and reliable
  • Troubleshoot complex issues across runtime, service, and GPU layers
  • Work closely with systems and ML engineers to improve performance and scalability

Required Qualifications

  • Bachelor’s degree in Computer Science, Engineering, or equivalent practical experience
  • 3+ years of software engineering experience, preferably in infrastructure or ML systems
  • Strong proficiency in one or more of: C++, Python, Go, or Rust
  • Experience with Kubernetes and containerized environments
  • Experience building or operating large-scale ML or MLOps infrastructure
  • Strong collaboration and communication skills across engineering and ML teams
  • Comfortable working in a fast-moving, high-ownership, in-office environment

Nice to Have

  • Experience with ML systems engineering or inference engines (e.g., vLLM, SGLang, TRT-LLM)
  • Experience with CUDA or ROCm and GPU profiling/debugging tools
  • Contributions to open-source ML, systems, or HPC infrastructure
  • Experience optimizing distributed systems for performance and reliability

Benefits & Perks

  • Medical, dental, and vision insurance
  • 401(k)
  • Daily meals and snacks
  • Flexible time off
  • Competitive compensation and equity

Interested in this role?

Apply now and hear back within 48 hours

Join & Apply

Already have an account? Sign in

Posted 8 months ago

San Francisco

Similar Jobs

San Francisco · Hybrid

A well-funded AI infrastructure company is building next-generation multimodal foundation models and a high-efficiency serving platform. With deep industry backing and close collaboration with hardware partners, the team is scaling rapidly to deliver the full stack powering real-time AI applications and developer-facing platforms.

$190,000 - $250,000Apply
New York City, San Francisco · Remote

This role is with a rapidly growing, venture-backed AI software company building an intelligent operating layer for complex, integration-heavy enterprises. The platform helps organizations plan, deploy, and optimize their teams through advanced analytics, automation, and AI-driven workflows. The company works in real-world, security-constrained environments—including hybrid, on-prem, and restricted systems—and is focused on shipping durable, production-grade enterprise software.

$180,000 - $210,000Apply
San Francisco · Hybrid

Overview This role is with a well-funded AI infrastructure company building next-generation multimodal models and a high-performance model serving platform. The team is scaling rapidly to deliver production-grade systems that power real-time AI applications, working closely across research, systems, and infrastructure. About the Role This is a rare opportunity to help architect and lead the development of a next-generation model serving platform—the core engine that brings highly efficient multimodal foundation models into production. As a senior technical leader, you will both build critical components yourself and guide other engineers, shaping architectural decisions, engineering standards, and execution quality. You’ll work across the full AI stack, from GPU execution and optimized runtimes to distributed serving, scheduling, and APIs that power low-latency, real-time inference. This role is ideal for engineers who enjoy deep systems work, thrive on ownership, and want to lead the development of foundational AI infrastructure. Equal Opportunity This employer is an equal opportunity organization. All qualified applicants will be considered without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran status, or disability. Compensation: $230K–$300K base + equity

$230,000 - $300,000Apply

Hiring

First candidates by day three.

One form, one call, a written search plan inside 24 hours. You pay nothing until someone signs.

Start a search

Looking

Roles that never hit a job board.

One profile covers every search we run. We only reach out when the role clears your bar — always free, never mass mail.

👋Hey — hiring engineers? Click me for jokes.