Azure Data Engineer (Trifacta/Adverity/Alteryx/Python)
Dentsu Global Services
Bangalore Urban, Karnataka, India
Required Details:
Total Exp : 4 to 6 years
Work Timings : 06.00 P.M. to 03.00 A.M.
Notice Period : Immediate to 15 Days
Work mode : Remote
- Responsibilities
- Onboard and normalize multi-source data through Adverity, standing up and maintaining Data Refinery pipelines that connect marketing, media, and platform sources into a single, reliable ingestion layer.
- Ingest, model, and reconcile DSP data (e.g., DV360, The Trade Desk, Amazon DSP) against CM360 ad-server delivery, applying programmatic media expertise to distinguish buy-side activation and bidding from independent ad serving/counting, and to resolve impression, spend, and attribution discrepancies between platforms.
- Manage and maintain data transformations in Trifacta, building governed, repeatable wrangling recipes that cleanse, standardize, and shape raw inputs into analytics-ready structures.
- Govern data assets in Databricks Unity Catalog, managing catalogs, schemas, permissions, lineage, and metadata to keep data secure, consistent, and discoverable across the platform.
- Develop and optimize further transformations in Databricks using SQL, PySpark, notebooks, and workflows, extending refined data through the medallion (bronze/silver/gold) architecture into curated, consumption-ready models.
- Design, run, and monitor data quality health checks in Databricks, including row-level integrity, schema and version validation, reconciliation logic, and automated error logging, to embed trust and observability at every stage of the pipeline.
- Working understanding and familiarity to apply modular, tested, and documented transformation orchestration, and GitHub for version control, CI/CD, code review, and change governance across SQL, Python, and YAML assets.
- Design modular, reusable, and well-documented data models that support analytics, reporting, and AI enablement, with clear definitions, lineage, and business context.
- Collaborate with analysts, engineers, and business leads to translate raw, multi-source data into governed, insight-ready datasets that support performance reporting and downstream decisioning.
- Contribute to data quality standards, monitoring, and the platform roadmap, defining best practices for reusable components, transformation patterns, and observability that scale across teams and clients.
- Prepare certified, well-modeled datasets for the consumption layer and reporting tools such as Power BI, ensuring metrics are consistent, governed, and traceable to source.
Required Qualifications
- 4-6+ years of experience as a Data Engineer or in a similar role building and operating scalable, production data pipelines.
- Bachelor’s Degree in Computer Science, Engineering, Information Systems, or a related field required; Graduate degree preferred.
- Hands-on experience with Adverity for data onboarding and normalization, including building and maintaining Data Refinery ingestion pipelines across multiple marketing and media sources.
- Proven experience managing data transformations in Trifacta (Alteryx Designer Cloud), including building governed, repeatable wrangling recipes for cleansing and standardization.
- Advanced expertise with Databricks, including Unity Catalog governance (catalogs, permissions, lineage, metadata), transformation development in SQL and PySpark, and Databricks Workflows, Notebooks, and Jobs.
- Demonstrated experience designing and operating data quality health checks, including integrity tests, reconciliation logic, schema validation, and automated monitoring, within a Databricks environment.
- Familiarity with dbt (dbt Labs) for transformation orchestration and testing, and with GitHub for version control, CI/CD, and code governance in a collaborative development workflow.
- Strong proficiency in SQL and Python for data engineering, transformation, and automation.
- Working knowledge of medallion (bronze/silver/gold) architecture, data modeling, lineage, and governance best practices for analytics-ready data.
- Self-starter with the ability to learn new tools quickly and deliver scalable, well-documented solutions across the data stack, driving continuous improvement and measurable impact.
- Bonus: Experience in advertising, marketing, or digital media environments, particularly performance reporting, reconciliation automation, or data quality optimization.