Role Overview
The Developer will design, build, and optimize PySpark based data solutions that support analytics and content decision making for a global media and entertainment organization. The role involves hybrid work collaborating with cross functional teams to deliver reliable data pipelines, improve audience insights, and enable data driven strategies that enhance content performance and operational efficiency.
Responsibilities
- Design robust PySpark data pipelines that efficiently ingest, transform, and aggregate large scale media and entertainment datasets.
- Implement optimized PySpark code that enhances performance of batch and near real time data processing.
- Collaborate with data engineers, analysts, and product teams to translate media business requirements into scalable technical solutions.
- Develop reusable data frameworks for streaming video on demand and advertising data.
- Apply data quality checks and monitoring mechanisms within PySpark workflows.
- Integrate data from multiple media platforms into unified data models.
- Optimize storage formats and execution configurations in PySpark to reduce costs.
- Document technical designs, PySpark jobs, and data flows.
- Troubleshoot production pipeline issues and perform root cause analysis.
Requirements
- Strong proficiency in PySpark programming and distributed data processing.
- Solid understanding of media and entertainment domain concepts (audience measurement, content metadata, streaming events).
- Experience working with big data platforms and data warehousing technologies.
- Familiarity with version control and collaborative development practices.
- Knowledge of performance tuning and resource optimization for PySpark applications.
- Experience with data quality frameworks and monitoring tools.
Skills
- PySpark
- Big Data
- Data Engineering
- Data Warehousing
- Python