OpenAI•26 days ago
Software Engineer, Model Runtime
Recruiter Fit Breakdown & Candid Summary
1
This is a high-stakes systems engineering role focused on the intersection of hardware and software for custom AI accelerators.
2
Ideal candidates have deep experience in low-level systems programming (C++/Rust) and a strong grasp of LLM inference internals like KV-cache management and model parallelism.
3
This is not a standard application-layer software role; it requires expertise in compilers, kernels, and distributed systems.
4
Candidates without experience in performance-critical systems or hardware-software co-design will likely struggle.
Role Responsibilities
- 1Design and implement LLM inference runtimes for custom silicon.
- 2Build scheduling, continuous batching, and memory management systems.
- 3Develop distributed execution strategies across chips, hosts, and racks.
- 4Optimize end-to-end latency, throughput, and hardware utilization.
- 5Partner with kernel, compiler, and silicon teams to co-design interfaces.
- 6Create profiling and observability tools for runtime performance.
Skills Matrix
Must-Have Skills
C++
Rust
Python
LLM Inference(large language model inference)
Distributed Systems
Systems Programming
Performance Optimization
Nice-to-Have Skills
Compiler Design
Kernel Development
Hardware Architecture
vLLM
SGLang
Free Instant Match Check
Will an ATS filter reject your resume for this role?
Check your keyword match score, missing must-have skills, and formatting flags — before you apply.
Checking saved resume…
No signups. Free check in 3 seconds.
Ready to submit your application?
Apply directly on OpenAI's official job portal.