Site Reliability Engineer

Qualtrix Consulting Inc.
India

JOB TITLE: Site Reliability Engineer (SRE) – Performance & Chaos Engineering

Location: India (Remote)

Experience: 3–5 Years

Job Description:

Scope of work:

Strong experience in Site Reliability Engineering or production operations role supporting large-scale, customer-facing systems.

  • Own service reliability end to end: define and maintain SLIs, SLOs and error budgets, establish steady-state behaviour for critical user journeys, and drive the reliability backlog that comes out of it.
  • Performance and load engineering with k6 – author JavaScript/TypeScript test scripts, build scenarios and executors (constant-VUs, ramping-VUs, ramping-arrival rate), define custom metrics, checks and thresholds, parameterise test data, and cover HTTP, WebSocket and gRPC protocols.
  • Integrate k6 into CI/CD pipelines (Harness, GitHub Actions, Jenkins) as automated performance regression gates; run distributed and containerised tests via the k6 Operator on Kubernetes or Grafana Cloud k6; stream results to Prometheus, InfluxDB, Grafana or Dynatrace for trend analysis.
  • Chaos engineering – design and run hypothesis-driven experiments with an explicit steady state, controlled blast radius, abort conditions and rollback plan; plan and facilitate GameDays with application, infrastructure and client teams; convert every finding into a tracked remediation item.
  • Fault injection using AWS Fault Injection Service (FIS), Gremlin, Chaos Mesh, LitmusChaos or Chaos Toolkit – CPU and memory pressure, pod and node termination, network latency and packet loss, dependency and third-party API failure, AZ evacuation, and database failover.
  • Dynatrace administration and engineering – OneAgent deployment and lifecycle, management zones, entity model and tagging strategy, alerting profiles, SLO definitions, dashboards and notebooks, and day-to-day administration of tooling in the APM space.
  • Dynatrace Davis AI and the Davis / Dynatrace API v2 – programmatic access to the problems, metrics, events, entities and SLO endpoints; ingest custom metrics and deployment / chaos events; tune Davis anomaly detection, alerting sensitivity and root-cause behavior; correlate load tests and chaos experiments with Davis-detected problems to validate detection and MTTR.
  • Automate release validation and quality gates using Dynatrace Site Reliability Guardian / Cloud Automation (or equivalent) so that performance and resilience evidence is evaluated automatically as part of every deployment.
  • Run production workloads on AWS – EKS (node groups, clusters, autoscaling), EC2, ALB/NLB, Route 53, RDS, Lambda and other AWS native services, with a strong grasp of container monitoring best practices.
  • Infrastructure and observability as code with Terraform; scripting in Python, Go, Bash and JavaScript/TypeScript for automation, API integration and internal tooling.
  • Participate in the on-call rotation; drive incident triage, root cause analysis and corrective actions, and work with cross-functional teams and Problem Management on escalations.
  • Manage uptime and availability reporting, capacity planning and performance tuning, using both load-test evidence and production telemetry.

REQUIRED QUALIFICATIONS - KNOWLEDGE/SKILLS

  • Demonstrable hands-on experience building and maintaining a k6 performance testing suite, not just running someone else’s scripts.
  • Practical chaos engineering experience in a production or production-like environment, with evidence of the reliability defects it surfaced.
  • Working knowledge of the Dynatrace API v2 and Davis AI, including authentication and token scopes, rate limits, and consuming problem and metric data from automation.
  • Ability to define alert standards for production environments and implement them, including tuning to reduce false positives and alert fatigue.
  • Strong understanding of distributed systems, networking and troubleshooting techniques – latency, timeouts, retries, connection pooling, DNS, TLS and load balancing.
  • Experience with automated build pipelines and continuous integration. Source control, branching and merging: git/svn/etc (Repository Management).
  • Familiarity with configuration management software and observability standards such as OpenTelemetry.
  • Provide support to teams for alarms and outages on an as needed basis, and work with development teams and management to ensure high availability.
  • Communication Skills- The ability to communicate verbally and in writing with all levels of employees and management, speaks and writes clearly and understandably at the right level.
  • Integrity and Trust- Involves being widely trusted, being seen as a direct, truthful individual, can present the unvarnished truth in an appropriate and helpful manner, keeps confidences, admits mistakes, and doesn’t misrepresent him/herself for personal gain.
  • Teamwork- Works well in a collaborative setting, volunteering for and completing assignments, acting as a positive team member by contributing to discussions.

Skills Required:

Essential Skills:

SRE practices (SLI/SLO/error budgets, incident and problem management); k6 performance and load engineering; Chaos Engineering; Dynatrace including Davis AI and the Davis / Dynatrace API v2. Supporting stack: AWS (EKS, EC2, native services), Kubernetes, Terraform and CI/CD pipeline automation. Refer to the scope of work for the detail.

Desired Skills:

  • Grafana, Prometheus and OpenTelemetry; experience with other load testing tools (JMeter, Gatling, Locust) in addition to k6.
  • Harness for CI/CD pipeline automation, deployment orchestration and release management; progressive delivery patterns (canary, blue-green, feature flags).
  • Certifications such as Dynatrace Associate/Professional, Certified Kubernetes Administrator (CKA), or AWS Solutions Architect / DevOps Engineer.
  • Hands-on experience and familiarity with AI-assisted “vibe coding” using tools such as GitHub Copilot, Claude Code, Cursor, or similar AI development platforms is preferred. Given the evolving nature of these technologies, practical exposure and a clear understanding of AI-assisted development concepts are acceptable.

Education Qualification:

B.E. / B.Tech / MCA / M.Sc. in Computer Science, Information Technology or an equivalent discipline

Score my resume against this job, free →

Get your ATS score for this role — free. Score my resume free →