Skip to main content
Share:
OpenAI26 days ago

Software Engineer, Model Runtime

hybrid🇺🇸USSan Francisco, USsenior

Recruiter Fit Breakdown & Candid Summary

1

This is a high-stakes systems engineering role focused on the intersection of hardware and software for custom AI accelerators.

2

Ideal candidates have deep experience in low-level systems programming (C++/Rust) and a strong grasp of LLM inference internals like KV-cache management and model parallelism.

3

This is not a standard application-layer software role; it requires expertise in compilers, kernels, and distributed systems.

4

Candidates without experience in performance-critical systems or hardware-software co-design will likely struggle.

Role Responsibilities

  • 1Design and implement LLM inference runtimes for custom silicon.
  • 2Build scheduling, continuous batching, and memory management systems.
  • 3Develop distributed execution strategies across chips, hosts, and racks.
  • 4Optimize end-to-end latency, throughput, and hardware utilization.
  • 5Partner with kernel, compiler, and silicon teams to co-design interfaces.
  • 6Create profiling and observability tools for runtime performance.

Skills Matrix

Must-Have Skills

7 required
C++
Rust
Python
LLM Inference(large language model inference)
Distributed Systems
Systems Programming
Performance Optimization

Nice-to-Have Skills

5 preferred
Compiler Design
Kernel Development
Hardware Architecture
vLLM
SGLang
Free Instant Match Check

Will an ATS filter reject your resume for this role?

Check your keyword match score, missing must-have skills, and formatting flags — before you apply.

Checking saved resume…

No signups. Free check in 3 seconds.

Ready to submit your application?

Apply directly on OpenAI's official job portal.