This project involved designing and implementing a data pipeline using Apache Spark components to process employee data in CSV format. As a data engineer, responsibilities included creating a Spark DataFrame from CSV data, defining schemas, and leveraging Spark SQL for data analysis and transformation tasks. The project aimed to extract valuable insights such as average salaries by department, maximum salary by age, and employee demographic statistics. Key tasks included SQL queries, data filtering, self-joins, and calculating aggregate metrics. Through this project, I demonstrated proficiency in Apache Spark SQL and data engineering, contributing to enhanced data-driven decision-making capabilities within the organization.