Abstract
Azure Cloud Datapipeline EDA is an open-source Data Science project. A cloud-native data pipeline and visualization project analyzing Formula 1 racing data using Azure, Databricks, Delta Lake, Tableau, and Python for insightful EDA and interactive dashboards. On the data engineering side, we leverage Azure Cloud services to build a scalable and automated data pipeline following the medallion architecture (bronze, silver, gold layers). Raw data is ingested and stored in Azure Data Lake Storage Gen2, processed and transformed using Azure Databricks with PySpark and SparkSQL, and managed through Delta Lake to ensure ACID transactions and schema enforcement. It is built using Jupyter Notebook, Matplotlib. The complete source code is publicly available on GitHub under the MIT License, making it a useful reference for students building a Data Science mini project or final-year project.
1. Introduction
On the data engineering side, we leverage Azure Cloud services to build a scalable and automated data pipeline following the medallion architecture (bronze, silver, gold layers). Raw data is ingested and stored in Azure Data Lake Storage Gen2, processed and transformed using Azure Databricks with PySpark and SparkSQL, and managed through Delta Lake to ensure ACID transactions and schema enforcement. We incorporate Unity Catalog for data governance and access control, while Azure Data Factory orchestrates the workflow to achieve full automation. The architecture demonstrates cloud-native best practices such as decoupled storage and compute, batch-stream unification, and automated job triggering—effectively realizing a modern Lakehouse design.
This project presents an end-to-end data pipeline and analytics workflow centered around Formula 1 racing data, with a strong emphasis on exploratory data analysis (EDA) and visualization. It is structured into two core components: cloud-native data engineering and analytical data visualization.
On the data analysis and visualization front, we utilize both Tableau and Python for different analytical tasks. Tableau connects directly to the Databricks-backed gold layer, enabling real-time, interactive BI dashboards that cover historical driver and team rankings, national-level aggregations, and top driver trends over time. For deeper statistical insights, Python’s Matplotlib and Seaborn are used to explore multidimensional relationships, such as starting grid vs. final position, fastest lap vs. points, and stability of driver performance across seasons.
2. Objective
A cloud-native data pipeline and visualization project analyzing Formula 1 racing data using Azure, Databricks, Delta Lake, Tableau, and Python for insightful EDA and interactive dashboards.
This project demonstrates how Jupyter Notebook, Matplotlib can be applied to a real-world Data Science problem.
4. Technology Stack
5. System Requirements
General requirements for this technology stack — check the README for exact versions.
- Python 3.8 or later with Jupyter Notebook / JupyterLab (or Google Colab)
- pip for dependencies
- Git (to clone the repository)
6. Installation & Setup
git clone https://github.com/Smars-Bin-Hu/azure-cloud-datapipeline-EDA.git
cd azure-cloud-datapipeline-EDAFull setup instructions are in the project README.
7. Future Enhancements
Suggested extensions you can add to make this your own project.
- Turn the analysis into an interactive dashboard
- Automate data refresh with a scheduled job
- Add a predictive model on top of the analysis
8. Viva / Review Questions
Common questions examiners ask for projects in this domain.
- What is the source of the dataset and how was missing data handled?
- Which exploratory analysis steps revealed the most useful insight?
- Why were these particular charts chosen to present the data?
- Which statistical or ML technique supports the conclusions?
- How could the analysis be automated or refreshed with new data?
9. Source Code & License
This project is developed by Smars-Bin-Hu and published on GitHub under the MIT License. Please follow the license terms and credit the original author when you use or modify this code.
Want to build this as your internship project?
Work on a Data Science project like this with mentor guidance, weekly reviews and an internship certificate from Training Trains, Erode — online or offline.
Apply for Data Science Internship