SRE Cloud Engineer
SRE – Cloud Engineer
Job ID: NIM2963
Location: Chennai
Experience: 8+ Years
Job Type: Permanent Hire
Domain: Capital Markets / Financial Services
Job Summary
We are looking for an experienced SRE / Cloud Engineer with 8+ years of experience supporting highly available, business-critical production systems, preferably within the Capital Markets or Financial Services domain.
The role focuses on production reliability, cloud infrastructure, observability, incident management, automation, Kubernetes, performance, and resilience for critical trading and post-trade platforms.
Key Responsibilities
- Own reliability, availability, monitoring, and operational stability of production services.
- Design and maintain monitoring, alerting, dashboards, SLIs, and SLOs.
- Handle production incidents, escalations, RCA, problem management, and post-incident reviews.
- Monitor critical batch, EOD, settlement, and time-sensitive processing.
- Automate operational activities using Python, Bash, or PowerShell to reduce manual effort and improve MTTR.
- Develop runbooks, automated remediation, and self-healing solutions.
- Support AWS and/or Azure cloud infrastructure and Linux production environments.
- Manage and troubleshoot Kubernetes and Docker-based workloads.
- Implement Infrastructure as Code using Terraform and support CI/CD pipelines.
- Work with application and engineering teams to improve production readiness and deployment reliability.
- Perform performance monitoring, capacity planning, and reliability improvements.
- Support HA, DR, failover, BCP, and resilience engineering activities.
- Support Kafka/MQ and event-driven platforms where required.Mandatory Skills
- 8+ years of experience in SRE, Production Engineering, DevOps, or Cloud Engineering.
- Strong experience supporting high-availability, business-critical production systems.
- Hands-on experience with AWS and/or Azure.
- Strong Linux administration and troubleshooting.
- Hands-on Kubernetes and Docker experience.
- Terraform / Infrastructure as Code.
- CI/CD using Jenkins, GitHub Actions, Azure DevOps, or similar tools.
- Python, Bash, and/or PowerShell scripting.
- Strong monitoring and observability experience with tools such as Prometheus, Grafana, ELK, Splunk, CloudWatch, or Datadog.
- Incident management, on-call support, escalation, and RCA.
- Experience with Autosys or similar enterprise scheduling tools.
- Experience with PagerDuty or similar alerting/incident-management tools.
- Strong understanding of HA, DR, failover, resilience, performance, and capacity engineering.Domain Experience
Preferred candidates should have experience in Capital Markets / Financial Services, particularly in production environments supporting:
- Front-office trading systems
- FX / Foreign Exchange
- Interest Rate Derivatives / IRS
- Trade lifecycle processing
- Post-trade processing
- Settlement systems
- EOD / batch processing
- Market data or event-driven platforms
- Regulatory or financial reporting
- Trading platform modernizationExperience with banking or financial-services production environments will also be considered where the candidate has strong SRE and production-engineering experience.
Preferred Skills
- Helm and Argo CD.
- Kafka and/or MQ.
- SLIs, SLOs, and error-budget management.
- Databricks or data-platform environments.
- Cloud security, IAM, secrets management, and encryption.
- CloudFormation, ARM, or Bicep.
- Automated remediation and self-healing.
- Performance and capacity engineering.Candidate Profile
The ideal candidate will be an SRE / Cloud Engineer with strong hands-on experience across cloud infrastructure, Kubernetes, observability, automation, Infrastructure as Code, incident management, and resilience engineering, along with experience supporting critical Financial Services or Capital Markets applications.
Strong troubleshooting, analytical, communication, and cross-functional collaboration skills are required.
Education
Bachelor’s or Master’s degree in Computer Science, Engineering, Information Technology, or a related discipline.