PySpark Developer

Infosys
Bengaluru East, Karnataka, India

Strong experience in PySpark. Good programming knowledge of Python. Hands-on experience with SQL. Understanding of ETL/ELT concepts and data warehousing. Experience working with large datasets and distributed processing. Knowledge of Spark SQL, DataFrames, and Spark transformations. Familiarity with Linux/Unix environment.

Design, develop, and maintain data pipelines using PySpark. Develop ETL/ELT processes for ingesting, transforming, and loading large volumes of data. Write optimized PySpark code for data processing and transformation. Work with structured and semi-structured data from multiple sources. Develop and optimize SQL queries for data extraction and validation. Troubleshoot data quality and performance issues. Collaborate with Data Engineers, Analysts, and Business teams to understand requirements. Participate in code reviews and follow data engineering best practices. Monitor and support production data pipelines.

Exposure to cloud platforms such as AWS, Azure, or GCP. Knowledge of Databricks. Experience with workflow orchestration tools such as Airflow. Understanding of CI/CD concepts.

Score my resume against this job, free →

Get your ATS score for this role — free. Score my resume free →