Senior Data Engineer

Naviq
Kochi, Kerala, India

Senior Data Engineer

Experience : 8 - 12 years

Location : Kochi & Trivandrum

We are looking for Data Engineers to design, develop, implement, and maintain scalable data pipelines, data products, and data processing applications.

Core / Mandatory Skills

  • Python & Pyspark – Should be able to write clean, well-tested Python code.
  • SQL – advanced SQL and query optimization (complex SQL queries for data transformation, analysis, and reporting.)
  • ETL / ELT – hands-on experience designing and implementing data workflows
  • Data Pipeline Development – building scalable, production-grade pipelines
  • Apache Spark – strong hands-on experience with large-scale data processing
  • Cloud – hands-on experience with AWS / Azure / GCP
  • Data Warehousing – experience with modern data warehouse/lakehouse platforms
  • Data Modeling – dimensional and scalable data modeling
  • Distributed Data Processing – experience handling large-volume datasets
  • Git & CI/CD – version control and deployment automation

Role Overview

The role involves working with modern data engineering technologies and tools across cloud and distributed computing environments to support analytics, reporting, and machine learning platforms.

The Data Engineer will collaborate with cross-functional and geographically distributed teams to translate business requirements into robust, reliable, secure, and efficient data solutions. As data engineering continues to evolve, continuous learning, adaptation to new technologies, and adherence to industry best practices are essential.

What the Engineer Would Do

  • Design, build, maintain, and optimize large-scale batch data pipelines, ETL/ELT workflows, and related applications using appropriate data engineering technologies and patterns.
  • Develop scalable data processing solutions in public, hybrid, multi-cloud, and multi-region environments, applying sound software engineering principles and modern data engineering practices.
  • Build, manage, and optimize cloud-based data workflows, leveraging AWS services such as Amazon S3, Amazon Redshift, AWS Glue, AWS Lambda, and Amazon EMR, or equivalent services on other cloud platforms.
  • Develop clean, modular, well-tested, maintainable, and efficient code using Python, SQL, Scala, Java, and Apache Spark, as applicable to the solution.
  • Write complex SQL queries for data transformation, processing, analysis, reporting, troubleshooting, and performance optimization.
  • Collaborate with product managers, analysts, data scientists, and other technical and business stakeholders to gather requirements and translate business needs into scalable, reliable data engineering solutions.
  • Design and implement appropriate data architectures and processing patterns, understanding concepts such as Data Mesh and Data Fabric and determining the right technologies for different use cases.
  • Work with relational database management systems (RDBMS), NoSQL databases, and distributed data storage systems to support data processing and downstream consumption.
  • Design efficient data models for both Online Analytical Processing (OLAP) and Online Transaction Processing (OLTP) systems.
  • Develop and maintain data solutions using big data technologies and platforms, including data lakes, Amazon EMR, AWS Glue, and, where applicable, Cloudera.
  • Implement and maintain data ingestion and processing solutions for streaming data and large-scale distributed datasets.
  • Orchestrate and schedule data pipeline execution using workflow management and scheduling tools such as Apache Airflow or Cron.
  • Establish data quality checks, testing frameworks, monitoring, logging, observability, and alerting mechanisms to ensure data accuracy, reliability, availability, and timely issue resolution.
  • Investigate and resolve technical, procedural, and operational issues, including data quality problems, pipeline failures, processing bottlenecks, database performance issues, and other constraints in complex data systems.
  • Optimize database queries, data processing workflows, and pipeline performance to improve scalability, efficiency, and resource utilization.
  • Contribute to team-owned data products that support downstream consumers across analytics, reporting, and machine learning platforms.
  • Develop reports and dashboards that communicate key insights, trends, and actionable recommendations to technical and non-technical stakeholders.
  • Apply data security practices, including data encryption, access control, and compliance with applicable data privacy regulations.
  • Use version control systems such as Git and understand containerization technologies such as Docker and orchestration platforms such as Kubernetes.
  • Contribute to agile development practices and a continuous integration and delivery (CI/CD) culture by identifying opportunities for technical and process improvements.
  • Participate in peer code reviews, knowledge-sharing sessions, and team retrospectives, collaborating effectively with geographically distributed Data Engineering teams.
  • Leverage AI-powered coding assistants such as GitHub Copilot and Claude, where appropriate, to enhance development productivity and code quality.
  • Stay updated on emerging data engineering technologies, tools, patterns, and best practices, promoting continuous learning and engineering excellence.

What We Expect from the Candidate

  • Experience building and maintaining at least one data pipeline or data product in a production environment, preferably on AWS, GCP, Azure, or another cloud platform.
  • Strong programming skills in Python, Scala, or Java, along with hands-on experience with Apache Spark and familiarity with other relevant programming languages.
  • Proficiency in SQL, including complex queries for data transformation, analysis, reporting, troubleshooting, and performance optimization.
  • Practical understanding of ETL/ELT processes, data pipeline architecture, data transformation, distributed processing, and scalable data engineering patterns.
  • Knowledge of data architecture concepts and frameworks such as Data Mesh and Data Fabric.
  • Hands-on experience with or exposure to cloud data engineering services such as Amazon S3, Amazon Redshift, AWS Glue, AWS Lambda, and Amazon EMR, or equivalent services on other cloud platforms.
  • Knowledge of big data technologies, data lakes, and large-scale data processing. Familiarity with Cloudera is an added advantage.
  • Experience with or understanding of relational databases, NoSQL databases, distributed data storage systems, and data warehousing solutions, preferably Amazon Redshift.
  • Understanding of data modelling principles for OLAP and OLTP systems.
  • Knowledge of streaming data technologies and data ingestion patterns.
  • Familiarity with workflow orchestration and job scheduling tools such as Apache Airflow or Cron.
  • Exposure to relevant data engineering and analytics technologies such as Hive, Iceberg, Qubole, Collibra, and Power BI.
  • Understanding of database performance optimization, data quality management, pipeline troubleshooting, and resource optimization.
  • Strong focus on software engineering best practices, including clean code, modular design, unit testing, maintainability, and performance optimization.
  • Experience with or understanding of monitoring, logging, observability, alerting, and reliability practices for data pipelines and data systems.
  • Understanding of version control systems such as Git, containerization technologies such as Docker, and container orchestration platforms such as Kubernetes.
  • Knowledge of data security practices, including encryption, access control, and data privacy compliance.
  • Ability to collaborate effectively in agile, cross-functional, and geographically distributed team environments.
  • Strong communication skills, with the ability to explain technical concepts, design decisions, and trade-offs clearly to technical and non-technical stakeholders.
  • A proactive approach to problem-solving, continuous improvement, knowledge-sharing, and adopting tools and practices that enhance engineering productivity and code quality.
  • Willingness to continuously learn and adapt to evolving data engineering technologies, tools, and industry best practices.

Score my resume against this job, free →

Get your ATS score for this role — free. Score my resume free →