Skip to main content
Share:
ElevenLabs23 days ago

Research Engineer - Inference

remote🇬🇧GBUnited Kingdomsenior

Recruiter Fit Breakdown & Candid Summary

1

This role is for high-impact engineers who specialize in the intersection of deep learning research and low-level systems performance.

2

Ideal candidates have a track record of shipping production-grade ML systems where latency and throughput are critical.

3

This is not a pure research role; it requires a strong systems engineering mindset to bridge the gap between model checkpoints and real-time user experiences.

4

Candidates who lack experience in GPU-level optimization or production-scale inference infrastructure will likely struggle.

Role Responsibilities

  • 1Deploy and optimize state-of-the-art AI models for production environments.
  • 2Implement performance optimizations including quantization, distillation, and custom kernel development.
  • 3Build and maintain high-performance serving infrastructure for real-time, streaming workloads.
  • 4Develop tooling to enable researchers to ship models to production with high confidence.
  • 5Profile and eliminate bottlenecks across the entire serving stack, from model architecture to orchestration.

Skills Matrix

Must-Have Skills

6 required
Machine Learning Engineering(ML Engineering)
GPU Programming(CUDA)
Inference Optimization(quantization)
Triton
TensorRT
Serving Frameworks(vLLM)

Nice-to-Have Skills

0 preferred
No optional skills specified for this role.
Free Instant Match Check

Will an ATS filter reject your resume for this role?

Check your keyword match score, missing must-have skills, and formatting flags — before you apply.

Checking saved resume…

No signups. Free check in 3 seconds.

Grounded Observations to Note

Ambiguous organizational structure

"We don’t have job titles. Instead, it’s about the impact you have."

Ready to submit your application?

Apply directly on ElevenLabs's official job portal.