Skip to main content
Share:
Cohere19 days ago

Engineering Manager, GPU Infrastructure

onsite🇺🇸USUnited Statesmanager$240K - $380K / year

Recruiter Fit Breakdown & Candid Summary

1

This role is for an experienced engineering leader who thrives at the intersection of hardware, distributed systems, and large-scale infrastructure.

2

The ideal candidate has a strong background in managing SRE or infrastructure teams and deep technical expertise in Kubernetes at scale.

3

This is not a pure people-management role; you must be willing to get hands-on with GPU networking, hardware, and cluster operations.

4

Candidates without significant experience in large-scale compute fleet management or those uncomfortable with cross-functional stakeholder management should not apply.

Role Responsibilities

  • 1Hire, mentor, and grow a team of GPU infrastructure engineers.
  • 2Own the technical roadmap for deploying, operating, and scaling Kubernetes clusters.
  • 3Partner with ML researchers to optimize training and inference stacks for new GPU architectures.
  • 4Collaborate with Finance, Legal, and Security teams on capacity planning, cost optimization, and compliance.
  • 5Drive operational excellence through observability, automation, and vendor management.

Skills Matrix

Must-Have Skills

6 required
Engineering Management(people management)
Kubernetes(k8s)
GPU Infrastructure(GPU clusters)
Capacity Planning
Infrastructure as Code(IaC)
Distributed Systems

Nice-to-Have Skills

3 preferred
Multi-cloud environments
GPU Networking
Cost Optimization
Free Instant Match Check

Will an ATS filter reject your resume for this role?

Check your keyword match score, missing must-have skills, and formatting flags — before you apply.

Checking saved resume…

No signups. Free check in 3 seconds.

Ready to submit your application?

Apply directly on Cohere's official job portal.