Specialist - Cloud & Infra Management
Role description
Reliability Operations
Maintain high availability performance and reliability of businesscritical applications
Monitor production systems and proactively identify reliability risks
Respond to incidents perform root cause analysis RCA and implement preventive measures
Participate in oncall support and incident management activities
Drive adherence to SLAs SLOs and operational best practices
Monitoring Observability
Implement and manage monitoring logging and ing solutions
Build operational dashboards and health checks for applications and infrastructure
Analyze trends and system behaviors to improve reliability
Work with observability tools such as Grafana Prometheus OpenSearch Instana Splunk or similar platforms
Cloud Container Platform Management
Manage and support Kubernetesbased environments
Deploy configure and maintain cloudnative infrastructure on AWS andor Azure
Support containerized applications and platform services
Optimize resource utilization scalability and performance
Automation DevOps
Automate operational tasks to eliminate manual effort toil
Develop scripts and tools using Python Shell Scripting or similar technologies
Build and maintain CICD pipelines using Jenkins or equivalent platforms
Implement Infrastructure as Code IaC using Terraform and Ansible
Collaboration Continuous Improvement
Work closely with developers architects product owners and infrastructure teams
Drive reliability reviews capacity planning and performance engineering initiatives
Promote SRE best practices and operational excellence across teams
Support release management and deployment activities
Required Skills
57 years of experience in Site Reliability Engineering DevOps or Production Support
Strong knowledge of LinuxUnix administration
Handson experience with Kubernetes
Expertise in Prometheus Grafana OpenSearch Instana or equivalent monitoring tools
Strong troubleshooting and incident management skills
Knowledge of ITIL processes and service management practices
Experience with AWS andor Azure cloud platforms
Proficiency in scripting using Python Shell or Bash
Good understanding of networking databases and distributed systems
Preferred Skills
Terraform Ansible and Infrastructure as Code IaC
Jenkins and CICD pipeline management
Kafka and eventdriven architectures
SQL and NoSQL databases
Site Reliability Engineering concepts including SLI SLO SLA and Toil reduction
Experience in largescale enterprise environments