Senior Deep Learning Systems Software Engineer - AI Infrastructure
Company
NVIDIA
Location
Bangalore, India
Type
Full Time
Job Description
NVIDIA is an industry leader with groundbreaking developments in High-Performance Computing, Artificial Intelligence and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery and powers what were once science fiction inventions from artificial intelligence to autonomous cars. NVIDIA is seeking senior engineers who are mindful of performance analysis and optimization to help us squeeze every last clock cycle out of all facets of Deep Learning such as training and inferencing, one of today's most important workloads in the world. If you are unafraid to work across all layers of the hardware/software stack from GPU architecture to Deep Learning Framework to achieve peak performance, we want to hear from you! This role offers an opportunity to directly impact the hardware and software roadmap in a fast-growing technology company that leads the AI revolution while helping deep learning users around the globe enjoy ever-higher training speeds.
Want more jobs like this?
Get Computer and IT jobs delivered to your inbox every week.
What you'll be doing:
- Understand, analyze, profile, and optimize deep learning workloads on state-of-the-art hardware and software platforms.
- Build tools to automate workload analysis, workload optimization, and other critical workflows.
- Collaborate with cross-functional teams to analyze and optimize cloud application performance on diverse GPU architectures.
- Identify bottlenecks and inefficiencies in application code and propose optimizations to enhance GPU utilization.
- Drive end-to-end platform optimization from a hardware level to the application and service levels
- Design and implement performance benchmarks and testing methodologies to evaluate application performance.
- Provide guidance and recommendations on optimizing cloud-native applications for speed, scalability, and resource efficiency.
- Share knowledge and best practices with domain expert teams as they transition applications to distributed environments.
What we need to see:
- Masters in CS, EE or CSEE or equivalent experience
- 5+ years of experience in application performance engineering
- Experience using large scale multi node GPU infrastructure on premise or in CSPs
- Background in deep learning model architectures and experience with Pytorch and large scale distributed training
- Experience with application profiling tools such as NVIDIA NSight, Intel VTune etc.
- Deep understanding of computer architecture, and familiarity with the fundamentals of GPU architecture. Experience with NVIDIA's Infrastructure and software stacks.
- Proven experience analyzing, modeling and tuning DL application performance.
- Proficiency in Python and C/C++ for analyzing and optimizing application code
Ways to stand out from the crowd:
- Strong fundamentals in algorithms and GPU programming experience (CUDA or OpenCL)
- Understanding of NVIDIA's server and software ecosystem
- Hands-on experience in performance optimization and benchmarking on large-scale distributed systems
- Hands-on experience with NVIDIA GPUs, HPC storage, networking, and cloud computing.
- In-depth understanding storage systems, Linux file systems, RDMA networking
NVIDIA is widely considered to be one of the technology world's most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you're creative and autonomous, we want to hear from you.
Date Posted
12/21/2024
Views
0
Similar Jobs
Senior Solution Consultant - Coursera
Views in the last 30 days - 0
This role involves supporting various Coursera Business teams through Salesforce Solution Architecture and administration skills Key responsibilities ...
View DetailsSenior Product Manager - Mobile - G-P
Views in the last 30 days - 0
The company is seeking a Senior Product Manager with extensive experience in mobile app development to lead the launch and growth of Gias AI Advisor f...
View DetailsTalent Guide - Twilio
Views in the last 30 days - 0
Twilio is seeking a Talent Guide to ensure a seamless global interview experience The role involves providing global interview scheduling coverage del...
View DetailsManager - ML Practice - Databricks
Views in the last 30 days - 0
Databricks is seeking a worldclass Manager to lead its Machine Learning Practice in India The role involves managing hiring and team growth developing...
View DetailsEnglish Physics content creator - Khan Academy
Views in the last 30 days - 0
Khan Academy is a nonprofit organization offering free worldclass education to millions of students globally They aim to provide locally relevant cont...
View DetailsSoftware Engineer (P3) - Twilio
Views in the last 30 days - 0
Twilio is seeking a Software Engineer with 5 years of experience in designing building and deploying largescale distributed systems and microservices ...
View Details