HPC Operations Manager - Hardware Engineering

NVIDIA Austin, TX

Company

NVIDIA

Location

Austin, TX

Type

Full Time

Job Description

Widely considered to be one of the technology world's most desirable employers, NVIDIA is an industry leader with groundbreaking developments in High-Performance Computing, Artificial Intelligence and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables outstanding creativity and discovery and powers what were once science fiction inventions from artificial intelligence to autonomous cars. We are now looking for a highly motivated HPC Operations Manager to join this multifaceted and innovative infrastructure team to craft global and dynamic HPC clusters used by Nvidia's hardware design teams. We are looking for leaders to help us grow and evolve a reliable computing environment to enable our hardware designers to build the next generation of GPUs and SOCs.

Want more jobs like this?

Get Software Engineering jobs delivered to your inbox every week.

Select a location
By signing up, you agree to our Terms of Service & Privacy Policy.


What You'll be Doing:

  • A huge part of the day-to-day job is collaborating with partners to develop programs driving around storage, networking, and compute in our growing fleet of data centers.
  • Lead, cultivate, and mentor a multi-national team of sysadmins and devops engineers, in support of the chip design teams
  • Ensure the highest reliability of HPC clusters. Develop critical metrics, program schedules to measure program health, predictability, and achievements
  • Identify failures, lead retrospective analysis, and help to develop improvement action plans. Build standard methodologies that cut through complexity and can be used across Nvidia and influence other partners for continuous improvement
  • Evaluate the latest technologies (hardware and cloud computing) and recommend future evolution of the infrastructure. Plan deployments and refresh of hardware (compute, storage, network equipment), and associated software stack (e.g. OS)
  • Work multi-functionally with hardware engineering leaders to support their future chip design needs, understand their workflow characteristics, and engineer an efficient HPC environment. Work with IT and engineering infrastructure teams on the different subsystems that comprise the computing environment.
  • Lead all aspects of the HPC scheduler (LSF), set/adjust policy, ensure delivery of forecasted compute demand to each hardware division, and drive high utilization.
  • Track software licensing servers and drive efficient license utilization
  • Develop and manage program schedules, milestones and deliverables. Adjust in the face of a highly fluid customer product roadmap.
  • Regularly communicate program status and key issues to senior management at NVIDIA's headquarters. Accurately represent the importance of issues and call out issues appropriately. Be the evangelist of data driven project management

What We Need to See:

  • B.S. or M.S. in Computer Science, Computer Engineering, Information Science (or equivalent experience)
  • 15+ years overall
  • 5+ years managing IT infrastructure teams of 10+ people
  • 10+ years experience running Linux servers, NFS storage, and Ethernet networks
  • Knowledge of HPC schedulers (IBM LSF preferred)
  • Knowledge of hardware design workflows (EDA tools and methodology)
  • Experience using project management and capacity planning software
  • Datacenter operations (rack and stack, maintenance)

Ways to stand out from the crowd:

  • HPC storage (e.g. Netapp, Pure Storage, Lustre, ZFS, Isilon)
  • Infiniband (operations, debugging, performance tuning)
  • Software development, especially in a devops context
  • Knowledge of relational databases, data lakes, metrics/visualization/analytics platforms
  • Deploying and maintaining FlexLM-based software license servers
  • Established relationships with enterprise-level equipment suppliers

The base salary range is 272,000 USD - 419,750 USD. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions.

You will also be eligible for equity and benefits. NVIDIA accepts applications on an ongoing basis.

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Apply Now

Date Posted

12/04/2024

Views

0

Back to Job Listings ❤️Add To Job List Company Info View Company Reviews
Positive
Subjectivity Score: 0.9

Similar Jobs

Internal Audit Intern (Summer 2025) - Cloudflare

Views in the last 30 days - 0

The Internal Audit IA organization is offering an internship opportunity in Austin TX The role involves supporting the team in executing operational a...

View Details

Cybersecurity Audit Intern (Summer 2025) - Cloudflare

Views in the last 30 days - 0

The Internal Audit IA organization is offering an internship opportunity for students majoring in Management Information Systems Computer Science Data...

View Details

Investment Research Senior Associate - Austin - CAIS

Views in the last 30 days - 0

CAIS a leading platform for alternative investments is seeking an experienced Associate to join their Investments team The role involves sourcing revi...

View Details

Home Office Services, Broker Dealer – Austin - CAIS

Views in the last 30 days - 0

CAIS is hiring a Vice President for Home Office Services to manage highlevel client relationships provide toptier whiteglove service and drive client ...

View Details

Field CTO (US Remote) - Anomali

Views in the last 30 days - 0

Anomali a Silicon Valleybased company is seeking a Field CTO to drive the adoption of their AIPowered Security Operations Platform The role involves t...

View Details

Principal Machine Learning Engineer- AI Platform - Visa Inc,

Views in the last 30 days - 0

Visa a global leader in payments and technology is seeking a Principal Machine Learning Scientist with extensive experience in machine learning system...

View Details