Data Engineer

Scaletrix.AI
Gurugram, Haryana, India

## About the Role

We are looking for a highly skilled Data Engineer with 4–6 years of experience in designing, building, and maintaining scalable data platforms and data pipelines. The ideal candidate should possess strong expertise in cloud-based data engineering, ETL/ELT development, data warehousing, and big data technologies. The candidate will work closely with Data Scientists, Analysts, BI teams, and business stakeholders to enable reliable and efficient data solutions.

## Key Responsibilities

* Design, develop, and maintain scalable ETL/ELT pipelines using Python, PySpark, SQL, Pandas, and DBT.

* Build and optimize batch and real-time data processing workflows using Apache Airflow and Databricks.

* Develop and manage cloud-based data lakes and data warehouses using AWS services such as S3, EMR, Redshift, Lambda, and RDS.

* Implement and maintain Snowflake data warehouse solutions, including Snowpipe, external stages, and data-sharing capabilities.

* Develop robust data ingestion frameworks for structured and semi-structured data from multiple sources.

* Design and implement event-driven architectures using AWS Lambda, Kafka, S3 Event Notifications, and other cloud-native services.

* Collaborate with cross-functional teams including Data Science, Analytics, BI, and Product teams to support data requirements.

* Build and expose APIs for data integration and data-sharing use cases.

* Ensure data quality, governance, privacy, and compliance through validation, masking, and monitoring frameworks.

* Implement CI/CD pipelines for data engineering projects using GitHub Actions, GitLab, JFrog, and automated testing frameworks.

* Optimize data pipelines for performance, scalability, and cost efficiency.

* Create operational dashboards and reporting solutions using Tableau or similar BI tools.

## Required Skills & Qualifications

### Technical Skills

* Strong programming skills in Python and SQL.

* Hands-on experience with PySpark, Pandas, and large-scale data processing.

* Expertise in Databricks, DBT, and Apache Airflow.

* Experience with Snowflake Data Warehouse and Snowpipe.

* Strong understanding of AWS ecosystem including:

* S3

* EMR

* Redshift

* Lambda

* EC2

* RDS

* Experience working with Apache Kafka and event-driven data architectures.

* Hands-on experience with PostgreSQL, MongoDB, and relational databases.

* Knowledge of Docker and containerized deployments.

* Experience with REST APIs and data integration frameworks.

* Familiarity with GitHub, GitLab, CI/CD pipelines, and automated testing practices.

* Understanding of Data Lake and Data Warehouse architectures.

### Preferred Qualifications

* Experience in Healthcare, Life Sciences, Retail, or eCommerce domains.

* Knowledge of data governance, security, and compliance best practices.

* Exposure to cloud certifications (AWS, OCI, Azure, or GCP).

* Experience working in Agile/Scrum environments.

## Education

* Bachelor's Degree in Computer Science, Information Technology, Engineering, or a related field.

## Nice to Have

* Experience with DuckDB.

* Exposure to Generative AI, Machine Learning, or MLOps platforms.

* Experience designing enterprise-scale data platforms and modern lakehouse architectures.

## What Success Looks Like

* Deliver highly reliable and scalable data pipelines.

* Ensure high-quality and trusted data for analytics and business decision-making.

* Improve data processing efficiency and reduce operational overhead.

* Contribute to the evolution of the organization's modern data platform strategy.

Score my resume against this job, free →

Get your ATS score for this role — free. Score my resume free →