Specialist - Cloud & Infra Management

LTM
Greater Kolkata Area

Role description

Reliability Operations

Maintain high availability performance and reliability of businesscritical applications

Monitor production systems and proactively identify reliability risks

Respond to incidents perform root cause analysis RCA and implement preventive measures

Participate in oncall support and incident management activities

Drive adherence to SLAs SLOs and operational best practices

Monitoring Observability

Implement and manage monitoring logging and ing solutions

Build operational dashboards and health checks for applications and infrastructure

Analyze trends and system behaviors to improve reliability

Work with observability tools such as Grafana Prometheus OpenSearch Instana Splunk or similar platforms

Cloud Container Platform Management

Manage and support Kubernetesbased environments

Deploy configure and maintain cloudnative infrastructure on AWS andor Azure

Support containerized applications and platform services

Optimize resource utilization scalability and performance

Automation DevOps

Automate operational tasks to eliminate manual effort toil

Develop scripts and tools using Python Shell Scripting or similar technologies

Build and maintain CICD pipelines using Jenkins or equivalent platforms

Implement Infrastructure as Code IaC using Terraform and Ansible

Collaboration Continuous Improvement

Work closely with developers architects product owners and infrastructure teams

Drive reliability reviews capacity planning and performance engineering initiatives

Promote SRE best practices and operational excellence across teams

Support release management and deployment activities

Required Skills

57 years of experience in Site Reliability Engineering DevOps or Production Support

Strong knowledge of LinuxUnix administration

Handson experience with Kubernetes

Expertise in Prometheus Grafana OpenSearch Instana or equivalent monitoring tools

Strong troubleshooting and incident management skills

Knowledge of ITIL processes and service management practices

Experience with AWS andor Azure cloud platforms

Proficiency in scripting using Python Shell or Bash

Good understanding of networking databases and distributed systems

Preferred Skills

Terraform Ansible and Infrastructure as Code IaC

Jenkins and CICD pipeline management

Kafka and eventdriven architectures

SQL and NoSQL databases

Site Reliability Engineering concepts including SLI SLO SLA and Toil reduction

Experience in largescale enterprise environments

Score my resume against this job, free →

Get your ATS score for this role — free. Score my resume free →