Data Scientist

Deltek
Bengaluru, Karnataka, India

Key Responsibilities :

  • Architect and evolve the full medallion lakehouse — Bronze, Silver, and Gold layers — on AWS S3 with Apache Iceberg; own schema design, partitioning, compaction, and retention policies.
  • Design and implement scalable Glue ETL (PySpark) pipelines for bronze_to_silver and silver_to_gold transformations, incorporating dbt for SQL-layer transformations where appropriate.
  • Own and extend CDC ingestion via Fivetran; manage schema evolution, connector health, and sync reliability across 18+ source products.
  • Build and maintain AWS Step Functions state machines and EventBridge schedules for end-to-end pipeline orchestration; implement Lambda-based quality and drift monitors.
  • Govern the Glue Catalog and Lake Formation policies; enforce column-level security, row-level access controls, and audit logging to meet SOC2 and regulatory requirements.
  • Architect the query layer — optimize Athena workgroups and partition pruning; plan and execute Trino-on-EKS deployment for sub-second analytics workloads.
  • Partner with data science teams on SageMaker data supply: feature engineering pipelines, training dataset preparation, and model registry integration.
  • Implement real-time and near-real-time streaming solutions using Kafka or Kinesis where sub-13-minute latency is required.
  • Lead platform modernization initiatives: evaluate emerging formats (Iceberg vs. Delta Lake vs. Hudi), tooling, and cost optimization strategies.
  • Establish and enforce data engineering best practices: code reviews, CI/CD for pipeline code, IaC (Terraform / CloudFormation), and incident response runbooks.
  • Mentor and level up junior and mid-level data engineers; define team standards for pipeline design, testing, and documentation.

Required Qualifications

  • 10+ years of software or data engineering experience, with at least 4 years in an architect or technical lead capacity designing large-scale cloud data platforms.
  • Deep, hands-on expertise with AWS data services: S3, Glue (PySpark ETL), Athena, Step Functions, Lambda, EventBridge, Lake Formation, SageMaker, and CloudWatch.
  • Production experience with Apache Iceberg (or Delta Lake / Hudi) table formats — compaction, snapshot management, schema evolution, and time travel.
  • Strong PySpark and Python skills; ability to write, review, and optimize distributed data processing jobs at scale.
  • Hands-on experience with CDC-based ingestion platforms (Fivetran, Debezium, or equivalent) across heterogeneous source systems.
  • Proven experience designing and implementing medallion (Bronze/Silver/Gold) or equivalent multi-hop lakehouse architectures.
  • Experience with data pipeline orchestration: AWS Step Functions, Apache Airflow, or equivalent; event-driven pipeline design patterns.
  • Strong SQL skills; experience with Athena, Trino, Presto, or equivalent query engines for large-scale analytical workloads.
  • Familiarity with data governance tooling: catalog management (Glue Catalog, Apache Polaris/Iceberg REST), data lineage, access controls, and audit frameworks.
  • Experience with Infrastructure as Code (Terraform or CloudFormation) for data platform provisioning and drift management.
  • Solid understanding of dimensional modeling, schema design (star/snowflake), and data normalization for BI and analytics workloads.
  • Bachelor's degree in Computer Science, Engineering, or a related field; or equivalent professional experience.

Preferred Qualifications

  • Experience operating Trino or PrestoDB on Kubernetes (EKS); tuning for sub-second query latency and multi-tenant workloads.
  • Familiarity with streaming platforms (Kafka, Kinesis, or Pub/Sub) and real-time lakehouse patterns.
  • Experience with Apache Polaris or other Iceberg REST catalog implementations.
  • Exposure to SageMaker MLOps pipelines, Model Registry, and feature store patterns for ML data supply.
  • Experience with dbt (data build tool) for SQL-layer transformation and documentation in lakehouse environments.
  • Government contracting or ERP domain knowledge (Costpoint, Deltek, Oracle, or similar enterprise platforms) is a strong plus.
  • AWS certifications: Data Engineer Associate, Solutions Architect Professional, or equivalent.

Score my resume against this job, free →

Get your ATS score for this role — free. Score my resume free →