Engineering Manager, GPU Infrastructure
Recruiter Fit Breakdown & Candid Summary
This role is for an experienced engineering leader who thrives at the intersection of hardware, distributed systems, and large-scale infrastructure.
The ideal candidate has a strong background in managing SRE or infrastructure teams and deep technical expertise in Kubernetes at scale.
This is not a pure people-management role; you must be willing to get hands-on with GPU networking, hardware, and cluster operations.
Candidates without significant experience in large-scale compute fleet management or those uncomfortable with cross-functional stakeholder management should not apply.
Role Responsibilities
- 1Hire, mentor, and grow a team of GPU infrastructure engineers.
- 2Own the technical roadmap for deploying, operating, and scaling Kubernetes clusters.
- 3Partner with ML researchers to optimize training and inference stacks for new GPU architectures.
- 4Collaborate with Finance, Legal, and Security teams on capacity planning, cost optimization, and compliance.
- 5Drive operational excellence through observability, automation, and vendor management.
Skills Matrix
Must-Have Skills
Nice-to-Have Skills
Will an ATS filter reject your resume for this role?
Check your keyword match score, missing must-have skills, and formatting flags — before you apply.
No signups. Free check in 3 seconds.
Ready to submit your application?
Apply directly on Cohere's official job portal.
Similar Jobs
Software Engineer Intern
Together AI • San Francisco, US
Software Engineer Intern
Together AI • San Francisco, US
Software Development In Test Intern
Together AI • San Francisco, US
Associate, Real Estate and Office Services
Gemini • Tempe, United States