Senior Kubernetes Operations Engineer

Lambda • Remote

Company

Lambda

Location

Remote

Type

Full Time

Job Description

Lambda's GPU cloud is used by deep learning engineers at Stanford, Berkeley, and Carnegie Mellon. Lambda's on-prem systems power research and engineering at Intel, Microsoft, Kaiser Permanente, major universities, and the Department of Defense.

If you'd like to build the world's best deep learning cloud, join us. 

What You’ll Do

  • Remotely install, upgrade, operate and maintain bare-metal Kubernetes clusters (up to thousands of nodes each)
  • Handle cluster degradation, recovery and resizing using our fleet management tooling
  • Perform out-of-hours on-call response for critical incidents as part of a well-balanced on-call rotation
  • Work on improving our tooling, automation, and processes, for both daily operations, alerting, and incident response
  • Dive into systems at a low level to solve unique cluster problems and write up your findings
  • Assist customers with high-level Kubernetes questions and integration with applications, storage and authentication
  • Assist with initial cluster build-outs and validation to help identify failed hardware before customer delivery
  • Work closely with our HPC Ops and Datacenter Ops teams on issues that require lower-level expertise or cross-functional solutions
  • Mentor and assist less-experienced team members
  • Have a voice in our product direction and help us think about how to minimize operational costs and complexity

You

  • Are an experienced operations engineer, SRE, sysadmin or similar with a deep knowledge of running Linux clusters and systems
  • Are very familiar with running on bare-metal (including knowledge of BMCs, kernel drivers, PXE, RAID, VLANs, hypervisors)
  • Have a good understanding of containers, virtualisation, and the mechanisms underpinning them
  • Have a good understanding of daily operation, bug-fixing and maintenance of Kubernetes
  • Have experience in an on-call environment and with incident response
  • Can perform incident post-mortems and develop procedures and tooling to prevent root causes from reoccurring
  • Have an excellent ability to learn on-the-fly and adapt to solve problems
  • Are able to work either independently with limited direction, or as part of a team
  • Are able to work with customers during incidents either via tickets, live messaging, or as part of a larger call.

Nice to Have

  • Deep Kubernetes experience
  • Experience with user-level restrictions and hardening (e.g. AppArmor)
  • Experience with network engineering
  • Experience with HPC clusters, environments & tooling
  • Experience with large-scale AI/ML training clusters
  • Experience with machine learning/AI frameworks
  • A passion for running your own bare-metal lab

About Lambda

  • We offer generous cash & equity compensation
  • Investors include Gradient Ventures, Google’s AI-focused venture fund
  • We are experiencing extremely high demand for our systems, with quarter over quarter, year over year profitability
  • Our research papers have been accepted into top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG
  • We have a wildly talented team of 200, and growing fast
  • Health, dental, and vision coverage for you and your dependents
  • Commuter/Work from home stipends
  • 401k Plan with 2% company match
  • Flexible Paid Time Off Plan that we all actually use

Salary Range Information 

Based on market data and other factors, the salary range for this position is $169,000 - $243,000. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description. 

A Final Note:

You do not need to match all of the listed expectations to apply for this position. We are committed to building a team with a variety of backgrounds, experiences, and skills.

Equal Opportunity Employer

Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.

Apply Now

Date Posted

03/15/2024

Views

5

Back to Job Listings ❤️Add To Job List Company Info View Company Reviews
Positive
Subjectivity Score: 0.8

Similar Jobs

Account Manager, Care Partnerships - Headway

Views in the last 30 days - 0

Headway a mental health care company founded in 2019 aims to revolutionize mental healthcare by building a national network of providers accepting ins...

View Details

Director of Pricing - Garner Health

Views in the last 30 days - 0

Garner Health is a rapidly growing company backed by toptier venture capital firms Their mission is to transform the healthcare economy by delivering ...

View Details

Linux Support Engineer - Voltage Park

Views in the last 30 days - 0

Voltage Park is seeking a Linux Support Engineer for a fulltime remote position The ideal candidate will have command line level Linux sys administrat...

View Details

Data Analyst - Agero

Views in the last 30 days - 0

Agero a leading B2B whitelabel provider of digital driver assistance services is revolutionizing the vehicle ownership experience through datadriven t...

View Details

Director, Product (Remote) - Dscout

Views in the last 30 days - 0

Dscout is a leading company in experience research technology offering a platform for major companies to gain insights into user needs and behaviors T...

View Details

Technical Architect - CDW

Views in the last 30 days - 0

CDW offers a rewarding career opportunity for a Technical Architect with expertise in ServiceNow The role involves delighting customers by collaborati...

View Details