Site Reliability Engineer

IBM · IN Bangalore

Company

IBM

Location

IN Bangalore

Type

Full Time

Job Description

Introduction
At IBM work is more than a job – it’s a calling: To build. To design. To code. To consult. To think along with clients and sell. To make markets. To invent. To collaborate. Not just to do something better but to attempt things you’ve never thought possible. Are you ready to lead in this new era of technology and solve some of the world’s most challenging problems? If so lets talk.

Your Role and Responsibilities
We are seeking an experienced Site Reliability Engineer (SRE) to join our team. The ideal candidate will have 5 to 8 years of experience in ensuring the reliability availability and performance of critical services and systems. This role involves building and maintaining infrastructure automating processes and responding to incidents to ensure our systems run smoothly and efficiently.
Key Responsibilities:
Infrastructure Management:
  • Design build and maintain scalable resilient infrastructure using cloud platforms (AWS and Azure).
  • Manage and optimize Kubernetes clusters containers and microservices.
  • Implement Infrastructure as Code (IaC) using tools like Terraform (Must) Ansible (Good to have) or CloudFormation(Good to have).
Automation & CI/CD:
  • Maintain automated CI/CD pipelines to ensure rapid safe and reliable delivery of software.
  • Automate repetitive tasks processes and workflows to increase efficiency and reduce human error.
  • Implement and maintain monitoring logging and alerting systems to ensure visibility into system performance.
Cost Optimization:
  • Set up monitoring and reporting tools to track cloud spending in real-time.
  • Regularly review the architecture and operations to identify areas where costs can be reduced. This includes evaluating new tools services or practices that could lead to further cost savings.
  • Collaborate with development teams to ensure that cost-efficient practices are followed in software design and deployment.
  • Recommend and manage the purchase of reserved instances savings plans or other discounts offered by cloud providers to reduce costs for long-term workloads.
Incident Response & Troubleshooting:
  • Respond to and resolve incidents in a timely manner ensuring minimal downtime and impact on customers.
  • Perform root cause analysis and post-mortem reviews to prevent recurrence of issues.
  • Collaborate with development teams to improve system reliability through proactive issue identification and resolution.
Performance Optimization:
  • Monitor system performance and capacity and implement improvements to optimize efficiency and scalability.
  • Analyze and improve application performance ensuring high availability and low latency.
Security & Compliance:
  • Ensure security best practices are followed across the infrastructure.
  • Implement security controls and monitoring to protect against vulnerabilities and threats.
  • Work with compliance teams to ensure systems adhere to regulatory requirements.
Collaboration & Communication:
  • Work closely with software engineers product managers platform team Global Support and other stakeholders to ensure system reliability aligns with business goals.
  • Provide guidance and mentorship to junior SREs and other team members.
  • Document processes procedures and best practices for the broader team.


Required Technical and Professional Expertise

  • 5-8 years of experience in Site Reliability Engineering DevOps or a similar role.
  • Strong experience with cloud platforms (AWS and Azure) and cloud-native technologies.
  • Proficiency in scripting languages (e.g. Python Bash) and automation tools.
  • Experience with containerization (Docker Kubernetes) and orchestration.
  • Knowledge of networking security and infrastructure best practices.
  • Familiarity with monitoring and logging tools (e.g. Prometheus Grafana ELK Stack).
  • Strong problem-solving skills and ability to work under pressure.
  • Excellent communication and collaboration skills.


Preferred Technical and Professional Expertise

  • Experience with database management (SQL).
  • Knowledge of software development practices and experience with Agile methodologies.
  • Certifications in cloud platforms (AWS Certified Azure Certified K8s certification etc.)
Apply Now

Date Posted

10/18/2024

Views

0

Back to Job Listings Add To Job List Company Profile View Company Reviews
Positive
Subjectivity Score: 0.8

Similar Jobs

Quality Engineer: Automation - IBM

Views in the last 30 days - 0

In this role youll work in one of IBMs Consulting Client Innovation Centers delivering deep technical and industry expertise to clients worldwide As a...

View Details

DevOps Engineer - IBM

Views in the last 30 days - 0

The text is an invitation to join IBM where work is more than just a job Its a calling to build design code consult think along with clients sell make...

View Details

Logic Design Engineer - IBM

Views in the last 30 days - 0

This job posting is for a Hardware Developer position at IBM where you will work on systems driving the quantum revolution and AI era The role involve...

View Details

Quality Engineer: Middleware - IBM

Views in the last 30 days - 0

The role of a Test Specialist at IBM involves working in a delivery center using analytical and technical skills to ensure software quality The Middle...

View Details

Infrastructure Engineer - IBM

Views in the last 30 days - 0

IBM Research is seeking a candidate with experience in implementing innovative solutions for resilient and robust computing environments focusing on I...

View Details

SRE Engineer - IBM

Views in the last 30 days - 0

The IBM Cloud Networking Tribe is seeking a Software Engineering professional to build the next generation IAAS The role involves running the producti...

View Details