Senior Site Reliability Engineer

US Posted Jul 15, 2026 0 views

Compensation

Compensation not disclosed

This employer didn't list pay. Model a likely range with the calculator.

Model this offer in the calculator

Job description

Team: IT

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability Engineer based in United States.

This role offers the opportunity to shape the reliability, scalability, and security of modern cloud infrastructure supporting mission-critical applications.
You will work at the intersection of Kubernetes, automation, AI infrastructure, and developer productivity to build resilient systems at scale.
The position combines hands-on engineering with technical leadership, enabling teams to deliver software faster while maintaining high operational standards.
You will design and optimize infrastructure platforms, CI/CD workflows, observability solutions, and production environments.
The ideal candidate is passionate about automation, reliability engineering, and emerging AI-driven operational workflows.
You will collaborate with engineering teams across the organization to improve delivery processes, system performance, and infrastructure maturity.
This is an opportunity to make a direct impact on secure, high-availability platforms in a fast-moving technology environment.

Accountabilities:

  • Design, build, and scale Kubernetes-based infrastructure supporting secure, multi-tenant, and highly available applications.
  • Develop and operate AI tooling infrastructure, including secure AI access patterns, MCP servers, and governance frameworks for production environments.
  • Optimize CI/CD pipelines to improve deployment speed, reliability, automation, and rollback safety.
  • Implement progressive delivery practices such as blue/green deployments and canary releases.
  • Advance Infrastructure as Code practices using tools such as Terraform, Helm, and GitOps workflows to create reusable infrastructure patterns.
  • Operate and improve streaming and analytics infrastructure, including Kafka, Flink, and ClickHouse environments.
  • Establish and enhance observability practices through monitoring, SLOs, alerting systems, and operational dashboards.
  • Lead incident response activities, perform root cause analysis, and drive long-term reliability improvements.
  • Build automated testing practices into the software delivery lifecycle.
  • Mentor engineers and promote best practices across Kubernetes, cloud infrastructure, automation, and reliability engineering.
  • Requirements:

    • 6+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or related roles with significant production Kubernetes experience.
    • Hands-on experience integrating AI/LLM tools into engineering or operational workflows, including understanding security, governance, and access control considerations.
    • Proven experience designing and maintaining CI/CD pipelines using tools such as GitHub Actions, Jenkins, GitLab CI, or similar technologies.
    • Strong knowledge of Kubernetes internals and managed cloud Kubernetes services such as EKS, GKE, or AKS.
    • Experience with Infrastructure as Code tools including Terraform, Helm, Pulumi, or equivalent solutions.
    • Proficiency in scripting or programming languages such as Python, Bash, or Go.
    • Experience with observability platforms such as Prometheus, Grafana, Datadog, or OpenTelemetry.
    • Production experience working with distributed systems, streaming technologies, and analytics platforms such as Kafka, Flink, and ClickHouse.
    • Strong understanding of cloud infrastructure, automation, system reliability, and operational excellence.
    • Excellent communication and collaboration skills with the ability to work effectively across engineering teams.
    • Experience with multi-region Kubernetes environments, chaos engineering, security automation, policy-as-code, or MLOps workflows is a plus.
    • Benefits:

      • Competitive compensation package ranging from BRL 422,500 – BRL 485,000 total compensation (base salary plus bonus).
      • Stock options and equity opportunities.
      • Health benefits and country-specific employee support programs.
      • Unlimited paid time off and flexible leave policies.
      • Paid parental leave.
      • Tuition reimbursement and learning and development opportunities.
      • Flexible remote working environment.
      • Additional employee benefits designed to support professional growth and well-being.

Related roles

Similar jobs