Remote

Senior Engineer, Inference Data Plane

Remote Posted Jul 24, 2026 0 views

Compensation

Compensation not disclosed

This employer didn't list pay. Model a likely range with the calculator.

Model this offer in the calculator

Job description

DigitalOceanJobs
Senior Engineer Inference Data Plane

Senior Engineer Inference Data Plane

Posted Yesterday
Be an Early Applicant
Denver CO USA
In-Office
139K-174K Annually
Senior level
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
DigitalOcean is the Inference Cloud built for production AI.
The Role
Lead design and delivery of high-scale resilient distributed inference data plane services. Architect and optimize inference hosting (tensor/data parallelism KV-cache routing prefill/decode disaggregation) contribute upstream open-source projects mentor engineers and maintain operational excellence and SLOs for a multi-tenant AI inference cloud.
Summary Generated by Built In

Dive in and do the best work of your career at DigitalOcean. Journey alongside a strong community of top talent who are relentless in their drive to build the simplest scalable cloud. If you have a growth mindset naturally like to think big and bold and are energized by the fast-paced environment of a true industry disruptor you’ll find your place here.  We value winning together—while learning having fun and making a profound difference for the dreamers and builders in the world. 

DigitalOcean is expanding its AI Infrastructure layer to support the next generation of AI-driven applications. We are seeking a Senior Engineer 2 to join our AI Inference Data Plane team. In this role you will be a key technical leader responsible for designing developing and delivering high-scale resilient data plane services that power our "Inference as a Service" offering. You will work at the intersection of distributed systems and specialized AI hardware to ensure our customers can deploy and scale their models with industry-leading performance and reliability. This is a hands-on role requiring you to be able to develop high quality software while availing of all the productivity boosts granted by the latest AI coding agents. 

What You’ll Do:
  • Technical Leadership: Act as a technical leader on the team driving the end-to-end design development and delivery of critical data plane components hosting large generative AI models.
  • System Design: Architect and refine system design proposals for our high-scale multi-tenant AI inference cloud ecosystem ensuring they meet rigorous availability and resiliency standards.
  • Performance Optimization: Implement and optimize distributed inference hosting using techniques like tensor/data parallelism KV cache optimizations and smart routing.
  • Collaboration: Work cross-functionally with Product Managers customer-facing teams and other engineering teams to align technical roadmaps with customer needs.
  • Distributed Serving at Scale: Build on Kubernetes-native distributed inference frameworks like llm-d (or alternatives such as NVIDIA Dynamo Ray Serve KServe) to deliver prefill/decode disaggregation KV-cache-aware routing tiered prefix caching and wide expert parallelism for MoE models.
  • Flow Control & Load Balancing: Solve the distributed-systems problems unique to LLM serving — inference-aware load balancing on queue depth cache locality and predicted latency; flow control and fairness across tenants; autoscaling inference pools; and moving gigabytes of KV-cache between prefill and decode instances with negligible overhead.
  • Open Source Contributions: Contribute upstream to llm-d vLLM and the inference gateway ecosystem and represent DigitalOcean in these communities.
  • Mentorship: Coach and mentor junior engineers fostering a culture of technical excellence and continuous improvement.
  • Operational Excellence: Maintain and operate critical high-scale services utilizing observability tools and defining SLOs to ensure superior platform health.
What You’ll Bring to DigitalOcean:
  • AI/ML Domain Knowledge: Hands-on experience hosting large language or multimodal models using inference engines like vLLM SGLang or TensorRT.
  • Inference Frameworks: Familiarity with distributed inference serving frameworks such as llm-d NVIDIA Dynamo or Ray Serve.
  • Inference Engine Depth: Hands-on experience with vLLM or alternatives (SGLang TensorRT-LLM TGI Modular MAX) including internals like continuous batching paged attention and prefix caching.
  • Distributed Inference Fluency: Understanding of why cluster-scale serving is hard: KV-cache locality is partitioned across workers naive round-robin routing destroys cache hit rates and tail latency and disaggregated prefill/decode requires fast cross-pod KV transfer (e.g. NIXL).
  • Upstream Track Record: Merged contributions to vLLM llm-d SGLang or similar projects strongly preferred.
  • Architecture Proficiency: Knowledge of common LLM architectures and optimization techniques (e.g. continuous batching quantization).
  • Software Engineering: Expert-level proficiency in GoLang or Python and familiarity with gRPC.
  • Cloud Operations: Proven experience shipping customer-facing software products and running critical services in a high-scale environment similar to DigitalOcean.
  • Open Source Mindset: Experience integrating and building with open-source software.
Compensation Range: 
  • $139200 - $174000

*This is a remote role

JR: 2026-7624

#LI-Remote

Why You’ll Like Working for DigitalOcean
  • We innovate with purpose. You’ll be a part of a cutting-edge technology company with an upward trajectory who are proud to simplify cloud and AI so builders can spend more time creating software that changes the world. As a member of the team you will be a Shark who thinks big bold and scrappy like an owner with a bias for action and a powerful sense of responsibility for customers products employees and decisions.
  • We prioritize career development. At DO you’ll do the best work of your career. You will work with some of the smartest and most interesting people in the industry. We are a high-performance organization that will always challenge you to think big. Our organizational development team will provide you with resources to ensure you keep growing. We provide employees with reimbursement for relevant conferences training and education. All employees have access to LinkedIn Learning's 10000+ courses to support their continued growth and development.
  • We care about your well-being. Regardless of your location we will provide you with a competitive array of benefits to support you from our Employee Assistance Program to Local Employee Meetups to flexible time off policy to name a few. While the philosophy around our benefits is the same worldwide specific benefits may vary based on local regulations and preferences.
  • We reward our employees. The salary range for this position is based on market data relevant years of experience and skills. You may qualify for a bonus in addition to base salary; bonus amounts are determined based on company and individual performance. We also provide equity compensation to eligible employees including equity grants upon hire and the option to participate in our Employee Stock Purchase Program.
  • DigitalOcean is an equal-opportunity employer. We do not discriminate on the basis of race religion color ancestry national origin caste sex sexual orientation gender gender identity or expression age disability medical condition pregnancy genetic makeup marital status or military service.

Application Limit: You may apply to a maximum of 3 positions within any 180-day period. This policy promotes better role-candidate matching and encourages thoughtful applications where your qualifications align most strongly.

Skills Required

  • Expert-level proficiency in Go or Python
  • Familiarity with gRPC
  • Hands-on experience hosting large language or multimodal models using inference engines like vLLM SGLang or TensorRT
  • Familiarity with distributed inference serving frameworks such as llm-d NVIDIA Dynamo Ray Serve or KServe
  • Experience building and operating Kubernetes-native distributed inference systems
  • Knowledge of LLM architectures and optimization techniques (continuous batching quantization prefix caching)
  • Proven experience shipping customer-facing software products and running critical services at scale
  • Technical leadership and mentorship experience
  • Merged upstream contributions to vLLM llm-d SGLang or similar projects
  • Open source mindset and experience integrating/building with open-source software

What the Team is Saying

DigitalOcean Compensation & Benefits Highlights

  • Healthcare StrengthHealth coverage is described as market‑leading across medical dental vision and mental‑health with company‑paid life and income‑protection. This positions core healthcare and protection benefits as a standout element of the package.
  • Parental & Family SupportParental leave is characterized as above‑average frequently highlighted alongside a phased return‑to‑work program. This indicates strong support for new parents beyond baseline policies.
  • Equity Value & AccessibilityThe package includes recurring equity grants and access to an Employee Stock Purchase Plan. This provides meaningful ownership opportunities alongside cash compensation.

DigitalOcean Insights

Am I A Good Fit?
beta
Expert contributor network
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Broomfield CO
1400 Employees
Year Founded: 2012

What We Do

DigitalOcean is the Inference Cloud — a full-stack production-ready cloud platform built to run AI applications with predictable performance sustainable economics and radically simpler operations at scale. We are built for teams turning AI into real products — not just training models. Our advantage is not fewer features but fewer failure modes when operating AI at scale — combining minimal operational overhead predictable cost efficiency and a full-stack cloud that works as a system. Hyperscalers are broad by design. Neoclouds are infrastructure-first. DigitalOcean is inference-first — with a real cloud underneath. It combines inference-optimized compute managed inference software and integrated cloud capabilities that reduce operational burden for teams running real workloads. Inference is the foundation—not the boundary. Everything else builds on top of it.

Why Work With Us

At DO we do career-defining work. We innovate with AI and build cutting-edge tech. Our rewards to match that intensity - to motivate you recognize your impact and give you what you need to thrive. If you have a growth mindset like to think big and bold and are energized by the fast-paced environment you'll find your place here.

Gallery

DigitalOcean Offices

Remote Workspace

Employees work remotely.

We commit to both remote work and in-person collaboration. These ways of working are dependent on specific roles and are mutually agreed upon by employees. In the US we are mainly remote. In our APAC locations we have a hybrid in-office approach.

Typical time on-site:
Company Office Image
HQBroomfield CO
Company Office Image
Seattle WA
Company Office Image
Hyderabad Telangana
Learn more

Similar Jobs

DigitalOcean

Staff Engineer

Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
In-Office
Denver CO USA
1400 Employees
191K-239K Annually

DigitalOcean

Revenue Enablement Manager

Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
In-Office
Denver CO USA
1400 Employees
160K-195K Annually

DigitalOcean

Financial Reporting Manager

Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
In-Office
Denver CO USA
1400 Employees
113K-141K Annually

Related roles

Similar jobs