Data Scientist
Deltek
Bengaluru, Karnataka, India
Key Responsibilities :
- Architect and evolve the full medallion lakehouse — Bronze, Silver, and Gold layers — on AWS S3 with Apache Iceberg; own schema design, partitioning, compaction, and retention policies.
- Design and implement scalable Glue ETL (PySpark) pipelines for bronze_to_silver and silver_to_gold transformations, incorporating dbt for SQL-layer transformations where appropriate.
- Own and extend CDC ingestion via Fivetran; manage schema evolution, connector health, and sync reliability across 18+ source products.
- Build and maintain AWS Step Functions state machines and EventBridge schedules for end-to-end pipeline orchestration; implement Lambda-based quality and drift monitors.
- Govern the Glue Catalog and Lake Formation policies; enforce column-level security, row-level access controls, and audit logging to meet SOC2 and regulatory requirements.
- Architect the query layer — optimize Athena workgroups and partition pruning; plan and execute Trino-on-EKS deployment for sub-second analytics workloads.
- Partner with data science teams on SageMaker data supply: feature engineering pipelines, training dataset preparation, and model registry integration.
- Implement real-time and near-real-time streaming solutions using Kafka or Kinesis where sub-13-minute latency is required.
- Lead platform modernization initiatives: evaluate emerging formats (Iceberg vs. Delta Lake vs. Hudi), tooling, and cost optimization strategies.
- Establish and enforce data engineering best practices: code reviews, CI/CD for pipeline code, IaC (Terraform / CloudFormation), and incident response runbooks.
- Mentor and level up junior and mid-level data engineers; define team standards for pipeline design, testing, and documentation.
Required Qualifications
- 10+ years of software or data engineering experience, with at least 4 years in an architect or technical lead capacity designing large-scale cloud data platforms.
- Deep, hands-on expertise with AWS data services: S3, Glue (PySpark ETL), Athena, Step Functions, Lambda, EventBridge, Lake Formation, SageMaker, and CloudWatch.
- Production experience with Apache Iceberg (or Delta Lake / Hudi) table formats — compaction, snapshot management, schema evolution, and time travel.
- Strong PySpark and Python skills; ability to write, review, and optimize distributed data processing jobs at scale.
- Hands-on experience with CDC-based ingestion platforms (Fivetran, Debezium, or equivalent) across heterogeneous source systems.
- Proven experience designing and implementing medallion (Bronze/Silver/Gold) or equivalent multi-hop lakehouse architectures.
- Experience with data pipeline orchestration: AWS Step Functions, Apache Airflow, or equivalent; event-driven pipeline design patterns.
- Strong SQL skills; experience with Athena, Trino, Presto, or equivalent query engines for large-scale analytical workloads.
- Familiarity with data governance tooling: catalog management (Glue Catalog, Apache Polaris/Iceberg REST), data lineage, access controls, and audit frameworks.
- Experience with Infrastructure as Code (Terraform or CloudFormation) for data platform provisioning and drift management.
- Solid understanding of dimensional modeling, schema design (star/snowflake), and data normalization for BI and analytics workloads.
- Bachelor's degree in Computer Science, Engineering, or a related field; or equivalent professional experience.
Preferred Qualifications
- Experience operating Trino or PrestoDB on Kubernetes (EKS); tuning for sub-second query latency and multi-tenant workloads.
- Familiarity with streaming platforms (Kafka, Kinesis, or Pub/Sub) and real-time lakehouse patterns.
- Experience with Apache Polaris or other Iceberg REST catalog implementations.
- Exposure to SageMaker MLOps pipelines, Model Registry, and feature store patterns for ML data supply.
- Experience with dbt (data build tool) for SQL-layer transformation and documentation in lakehouse environments.
- Government contracting or ERP domain knowledge (Costpoint, Deltek, Oracle, or similar enterprise platforms) is a strong plus.
- AWS certifications: Data Engineer Associate, Solutions Architect Professional, or equivalent.