Sr.Data Engineer

VectorMatch
Greater Hyderabad Area

Design, build, and maintain robust cloud data pipelines using Azure Data Factory (ADF), Azure Databricks, PySpark, and Spark SQL.

  • Implement and manage the Medallion Architecture — moving and transforming data through Bronze (raw/audit), Silver (cleansing/dedup/SCD), and Gold (business aggregates) layers.
  • Perform complex data transformations, cleansing, deduplication, and incremental loads using Delta MERGE, supporting both Slowly Changing Dimensions (SCD Type 1 and Type 2).
  • Optimize Spark workloads by tuning shuffle partitions, managing memory to prevent OOM errors, and leveraging Adaptive Query Execution (AQE).
  • Apply optimized join strategies, including broadcast joins for small datasets and salting techniques to handle data skew.
  • Implement robust exception handling, file dependency validation, and asynchronous batch processing to ensure pipeline reliability.
  • Ensure high data quality through schema enforcement, schema evolution handling, and validation against expected target criteria.
  • Monitor, troubleshoot, and resolve production job failures by analysing cluster scaling behaviour, Spark UI metrics, and physical query plans.
  • Build and maintain CI/CD pipelines for Databricks using Git, Azure DevOps, and Databricks Asset Bundles (DABs), with environment-specific parameterization for Dev, Test, and Prod.

Mandatory Skills

Azure Data Factory,Data Bricks,Data Engineer,SQL,Pyspark,Azure

Role

Required Qualifications & Skills

  • Extensive hands-on experience as a Data Engineer within the Azure ecosystem.
  • Strong programming proficiency in PySpark and Spark SQL.
  • Deep expertise in Delta Lake operations (ACID transactions, Time Travel, VACUUM, Deep/Shallow Clone) and Managed vs. External table management.
  • Advanced SQL skills, particularly complex window functions (ROW_NUMBER, RANK, DENSE_RANK, LEAD, LAG) for analytics and deduplication.
  • Proven experience with orchestration and monitoring using ADF and Azure Monitor.
  • Solid grounding in software engineering best practices, including version control (Git) and environment-based deployment.
  • Strong analytical and problem-solving skills, with a track record of diagnosing and resolving performance issues such as small-file proliferation, data skew, and excessive driver-side collection.

Preferred & Advanced Skills

  • Experience with Structured Streaming and integrating micro-batches (foreachBatch) into Delta targets.
  • Experience generating deterministic business/composite hashes for records lacking natural primary keys.
  • Familiarity with Infrastructure as Code for deploying ADF and Databricks resources.
  • Exposure to Unity Catalog for data governance, access control, and lineage across workspaces.
  • Familiarity with cost optimization practices — cluster right-sizing, auto-termination policies, and job cluster vs. all-purpose cluster tradeoffs.
  • Experience with data quality frameworks for automated validation.
  • Working knowledge of Python packaging/testing (pytest, unit testing for PySpark transformations).

Location

Chennai / Hyderabad / Bengaluru / Pune / New Delhi

Experience

3 to 12 years

Skills

Azure Data FactoryAzure Data BricksPysparkSpark SQLAzureAzure DevOpsDatabricks

Good to have

Databricks Asset BundlesPython

Score my resume against this job, free →

Get your ATS score for this role — free. Score my resume free →