Senior Data Engineer / Data Architect
About AccuData Analytics
AccuData Analytics Pvt. Ltd. partners with US healthcare organizations to build
secure, scalable, AI-driven digital health platforms, with expertise in cloud
infrastructure, data engineering, AI systems, and healthcare compliance-first
Architecture.
About the role
We're hiring a Data Engineer to build and run ETL pipelines that migrate healthcare data
from legacy source systems into a backend platform — via REST APIs and direct
PostgreSQL writes into tenant-isolated schemas.
What you'll do
● Build and maintain Python ETL pipelines that read CSV/Excel source exports and
migrate them into a backend system via REST API calls and direct PostgreSQL inserts
● Design idempotent, rerun-safe migration commands (dry-run modes, mapping/result
CSVs, signature-based drift detection)
● Handle real-world messy data: inconsistent column naming, duplicate records, missing
fields, normalization of phone/date/ID formats
● Write and maintain validation layers — pre-migration contract checks and post-migration
read-only reconciliation
● Work with multi-tenant, schema-per-tenant PostgreSQL architecture
● Collaborate closely with the technical lead on migration scope, sequencing, and safety
guardrails for sensitive data
What we're looking for
● 5-10 years of experience in Python data engineering / ETL development
● Strong SQL and hands-on experience with PostgreSQL (psycopg2 or similar)
● Experience consuming/calling REST APIs (auth, pagination, error handling, retries)
● Comfortable working with large CSV/tabular datasets and writing data validation logic
● Understanding of idempotency, batch processing, and safe rerun design in data
pipelines
● Bonus: experience with healthcare or other regulated/sensitive data domains, or
multi-tenant systems
Tech stack
● Language: Python 3.10+
● Database: PostgreSQL (psycopg2, multi-tenant/schema-per-tenant, batch inserts with
execute
_
values )
● API/HTTP: REST APIs, requests , Bearer token auth
● Config/Env: python-dotenv , .env
-based configuration
● Testing: unittest / pytest
● Version control: Git● Cloud (nice to have): AWS (Secrets Manager, EKS/ECR)
● Backend frameworks (nice to have): Django REST Framework or FastAPI
Optional / good to have (pipeline evolution)●
Workflow orchestration — Airflow, Prefect, or Dagster
● Data validation frameworks — Great Expectations or Pandera
● Containerization — Docker
- ● Async/background processing — Celery + Redis