IBM

Site Reliability Engineer (SRE)

Bengaluru, India
Kubernetes Terraform Ansible
Description
 
  • Infrastructure Management:
  • Design, build, and maintain scalable, resilient infrastructure using cloud platforms (AWS and Azure).
  • Manage and optimize Kubernetes clusters, containers, and microservices.
  • Implement Infrastructure as Code (IaC) using tools like Terraform (Must), Ansible (Good to have), or CloudFormation(Good to have).
  • Automation & CI/CD:
  • Maintain automated CI/CD pipelines to ensure rapid, safe, and reliable delivery of software.
  • Automate repetitive tasks, processes, and workflows to increase efficiency and reduce human error.
  • Implement and maintain monitoring, logging, and alerting systems to ensure visibility into system performance.
  • Cost Optimization:
  • Set up monitoring and reporting tools to track cloud spending in real-time.
  • Regularly review the architecture and operations to identify areas where costs can be reduced. This includes evaluating new tools, services, or practices that could lead to further cost savings.
  • Collaborate with development teams to ensure that cost-efficient practices are followed in software design and deployment.
  • Recommend and manage the purchase of reserved instances, savings plans, or other discounts offered by cloud providers to reduce costs for long-term workloads.
  • Incident Response & Troubleshooting:
  • Respond to and resolve incidents in a timely manner, ensuring minimal downtime and impact on customers.
  • Perform root cause analysis and post-mortem reviews to prevent recurrence of issues.
  • Collaborate with development teams to improve system reliability through proactive issue identification and resolution.
  • Performance Optimization:
  • Monitor system performance and capacity, and implement improvements to optimize efficiency and scalability.
  • Analyze and improve application performance, ensuring high availability and low latency.
  • Security & Compliance:
  • Ensure security best practices are followed across the infrastructure.
  • Implement security controls and monitoring to protect against vulnerabilities and threats.
  • Work with compliance teams to ensure systems adhere to regulatory requirements.
  • Collaboration & Communication:
  • Work closely with software engineers, product managers, platform team, Global Support and other stakeholders to ensure system reliability aligns with business goals.
  • Provide guidance and mentorship to junior SREs and other team members.
  • Document processes, procedures, and best practices for the broader team."    
Glsab24
 










 
IBM
IBM
Business Development Business Information Systems CRM Data Management Software

0 applies

0 views

There are more than 50,000 engineering jobs:

Subscribe to membership and unlock all jobs

Engineering Jobs

60,000+ jobs from 4,500+ well-funded companies

Updated Daily

New jobs are added every day as companies post them

Refined Search

Use filters like skill, location, etc to narrow results

Become a member

🥳🥳🥳 401 happy customers and counting...

Overall, over 80% of customers chose to renew their subscriptions after the initial sign-up.

To try it out

For active job seekers

For those who are passive looking

Cancel anytime

Frequently Asked Questions

  • We prioritize job seekers as our customers, unlike bigger job sites, by charging a small fee to provide them with curated access to the best companies and up-to-date jobs. This focus allows us to deliver a more personalized and effective job search experience.
  • We've got about 70,000 jobs from 5,000 vetted companies. No fake or sleazy jobs here!
  • We aggregate jobs from 5,000+ companies' career pages, so you can be sure that you're getting the most up-to-date and relevant jobs.
  • We're the only job board *for* software engineers, *by* software engineers… in case you needed a reminder! We add thousands of new jobs daily and offer powerful search filters just for you. 🛠️
  • Every single hour! We add 2,000-3,000 new jobs daily, so you'll always have fresh opportunities. 🚀
  • Typically, job searches take 3-6 months. EchoJobs helps you spend more time applying and less time hunting. 🎯
  • Check daily! We're always updating with new jobs. Set up job alerts for even quicker access. 📅

What Fellow Engineers Say