Talent.com
NVIDIA
Director, Engineering Operations and Site Reliability Engineering - Datacenter Server SystemsNVIDIA • Santa Clara, CA, United States
Director, Engineering Operations and Site Reliability Engineering - Datacenter Server Systems

Director, Engineering Operations and Site Reliability Engineering - Datacenter Server Systems

NVIDIA • Santa Clara, CA, United States
30+ days ago
Job type
  • Full-time
Job description

NVIDIA Engineering Operations And Site Reliability Engineering

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It's a unique legacy of innovation that's fueled by great technologyand amazing people. Today, we're tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what's never been done before takes vision, innovation, and the world's best talent. As an NVIDIAN, you'll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.

NVIDIA is seeking a strong technology leader for our Engineering Operations and Site Reliability Engineering for our next-generation datacenter server systems. This role sits at the intersection of execution, reliability, automation, and large-scale system operations, where we keep NVIDIA's rack-scale systems healthy, observable, and highly available for internal engineering users. These systems bring together the full power of NVIDIA CPUs, GPUs, NVLink, InfiniBand/Spectrum-X networking, cluster management technologies, and our optimized AI/HPC software stack. We enable fast product development by ensuring large internal racks, clusters, and lab infrastructure are reliable, well-instrumented, and operated with scalable engineering practices. This is a technical leadership role focused on execution excellence for large-scale internal datacenter systems. The ideal candidate has strong engineering judgment, experience operating complex distributed infrastructure, and the ability to build teams that combine focused operations with automation-first software engineering.

What you will be doing:

  • Lead teams that help us ensure NVIDIA's internal rack-scale server systems, clusters, and lab facilities remain available, healthy, and reliable.
  • Drive execution across fleet operations, incident response, roadmap planning, change management, operational readiness, and reliability metrics.
  • Build automation, telemetry, alerting, and dashboards that improve visibility and help teams resolve issues faster.
  • Partner with hardware, firmware, software, networking, validation, and infrastructure teams to deploy, sustain, and debug complex systems.
  • Create feedback loops into NPI and sustaining teams to improve product quality, serviceability, and development velocity.
  • Grow and mentor a high-performing technical team with a culture of ownership, learning, and automation-first execution.

What we need to see:

  • BS or MS in Computer Science, Electrical Engineering, Computer Engineering, or related field (or equivalent experience).
  • 12+ overall years of experience in infrastructure, systems engineering, reliability, datacenter operations, distributed systems, or related areas, including 7+ years of people management experience.
  • Strong understanding of server systems, Linux, cluster operations, high-speed networking, and large-scale infrastructure.
  • Experience operating complex systems with high availability expectations, including monitoring, incident management, automation, and fleet-health practices.
  • Proven track record of driving execution across multiple teams, priorities, and technical domains, including close partnership with hardware, firmware, software, networking, validation, and infrastructure organizations.
  • Clear written and verbal communication skills, including executive-level reporting on operational health, risks, and priorities.
  • Track record of building cohesive teams and developing technical leaders who improve reliability and execution.

Ways to stand out from the crowd:

  • Prior Director or Senior Manager experience leading infrastructure, reliability, platform engineering, or large-scale lab operations teams.
  • Experience operating GPU, AI, HPC, cloud, or hyperscale datacenter infrastructure.
  • Broad knowledge of rack-scale systems, including server management, networking, storage, power, thermal, and RAS concepts.
  • Experience building automation, telemetry, fleet health, or dashboarding systems that improve product quality, serviceability, or engineering velocity.

Do you enjoy making complex AI infrastructure reliable at scale while enabling engineering teams to move faster? Come join our datacenter server systems team and help build the reliable, token-efficient computing platforms driving NVIDIA's success in this exciting and rapidly growing field.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 292,000 USD - 442,750 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until July 8, 2026.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Create a job alert for this search

Director, Engineering Operations and Site Reliability Engineering - Datacenter Server Systems • Santa Clara, CA, United States

Similar jobs

Sr. Engineering Manager - Storage Engineering

ClouderaSan Jose, CA, United States
Full-time

At Cloudera, we empower people to transform complex data into clear and actionable insights.With as much data under management as the hyperscalers, we're the preferred data partner for the top comp... Show more

 • Promoted

Director of Engineering, Core Database

YugabyteSunnyvale, CA, United States
Full-time

Director of Engineering, Core Database.Yugabyte is the company behind YugabyteDB, the AI-ready, multi-modal, distributed PostgreSQL database for cloud-native apps.Trusted by industry leaders includ... Show more

 • Promoted

Director, System Engineering & Integration

CadenceSan Jose, CA, United States
Full-time

Director of System Engineering & Integration.At Cadence, we hire and develop leaders and innovators who want to make an impact on the world of technology.The Director of System Engineering & Integr... Show more

 • Promoted

Director System Application Engineering

LumilensSan Jose, CA, United States
Full-time

Director / Senior Director, System Applications Engineering.We are seeking a Director / Senior Director, System Applications Engineering to lead customer-facing system engineering for next-generati... Show more

 • Promoted

Director of Engineering -Exchange

Extreme NetworksSan Jose, CA, United States
Full-time

Over 50,000 customers globally trust our end-to-end, cloud-driven networking solutions.They rely on our top-rated services and support to accelerate their digital transformation efforts and deliver... Show more

 • Promoted

Director, Systems Engineering, CCDI

LittelfuseMilpitas, CA, United States
Full-time

Director, Systems Engineering, CCDI.Littelfuse is a diversified industrial technology manufacturing company shaping solutions for the safe and efficient transfer of electrical energy.Headquartered ... Show more

 • Promoted

Sr. Engineering Manager, Distributed Cloud

F5San Jose, CA, United States
Full-time

At F5, we strive to bring a better digital world to life.Our teams empower organizations across the globe to create, secure, and run applications that enhance how we experience our evolving digital... Show more

 • Promoted

Director, Automation Systems Engineering

QuantumScapeSan Jose, CA, United States
Full-time

Director, Automation Systems Engineering.QuantumScape is on a mission to transform energy storage with solid-state lithium-metal battery technology.The company's next-generation batteries are desig... Show more

 • Promoted

Director, Advanced Manufacturing and Data Center Engineering

GoogleSunnyvale, CA, United States
Full-time

Director, Advanced Manufacturing and Data Center Engineering.In accordance with Washington state law, we are highlighting our comprehensive benefits package, which is available to all eligible US b... Show more

 • Promoted

Director, Engineering Operations and Site Reliability Engineering - Datacenter Server Systems

NVIDIASanta Clara, CA, United States
Full-time

NVIDIA Engineering Operations And Site Reliability Engineering.NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years.It's a unique legacy of in... Show more

 • Promoted

Sr Director, Operations Engineering

FlextronicsSan Jose, CA, United States
Full-time

Sr Director, Operations Engineering.To support our extraordinary team who build great products and contribute to our growth, we're looking to add a Sr Director, Operations Engineering located in Sa... Show more

 • Promoted

Director, Engineering Redshift, Analytics Engines

AmazonPalo Alto, CA, United States
Full-time

Amazon Redshift is seeking a Director of Engineering to lead the Redshift Engine Team, a global organization spanning the United States and Europe.Amazon Redshift is the most widely used cloud data... Show more

 • Promoted

Director of Software Engineering, Core Platform

Pure StorageSanta Clara, CA, United States
Full-time

Director of Software Engineering, Core Platform.We're in an unbelievably exciting area of tech and are fundamentally reshaping the data storage industry.Here, you lead with innovative thinking, gro... Show more

 • Promoted

AI Data Center Debug Director (Remote)

AMDSanta Clara, CA, United States
Remote
Full-time

A leading technology company is looking for a Technical Director to lead the Instinct Customer Engineering Platform Debug and Qualification team.This strategic role focuses on AI datacenter hardwar... Show more

 • Promoted

Director of Engineering

Ensemble Investments, LLCSanta Cruz, CA, United States
Full-time

Nestled along the Pacific Coast, La Bahia Hotel & Spa celebrates its dramatic setting where the tip of Monterey Bay touches Sana Cruz's coveted Main Beach.Steeped in the romantic beauty of Spanish-... Show more

 • Promoted

Director of Operations - IT Infrastructure & Server Rack Manufacturing

AMAXFremont, CA, United States
Full-time

The Director of Operations is responsible for leading and optimizing all operational functions within a fast-paced IT infrastructure and server rack manufacturing environment.This role drives opera... Show more

 • Promoted

Earn up to $400 per Day by Playing Games and TakingSurveys

Freecash.comSanta Cruz, California, United States
Full-time

Since its launch 6 years ago, over 60MN users have earned and withdrawn over $250MN! The platform is rated 4.TrustPilot with over 230k+ reviews, establishing Freecash as one of the highest-rated op... Show more

 • Promoted

Director, Hardware Systems Engineering

Advanced Micro DevicesSanta Clara, CA, United States
Full-time

What You Do At AMD Changes Everything.At AMD, our mission is to build great products that accelerate next-generation computing experiencesfrom AI and data centers, to PCs, gaming and embedded syste... Show more

 • Promoted

Director of Storage Operations

Lawson Drayage IncHayward, CA, United States
Full-time

Director of Storage Operations.Lawson Drayage is seeking a Director of Storage Operations to oversee and grow the company's warehouse and storage business across multiple locations, including Haywa... Show more

 • Promoted

Sr. Director Engineering

Vistance NetworksSunnyvale, CA, United States
Full-time

In our 'always on' world, we believe it's essential to have a genuine connection with the work you do.RUCKUS Networks builds and delivers purpose-driven networks that perform in the tough, unique e... Show more