Data Frame

C++ DataFrame for statistical, financial, and ML analysis in modern C++

Data ScienceC++BSD-3-Clause

Abstract

Data Frame is an open-source Data Science project. C++ DataFrame for statistical, financial, and ML analysis in modern C++. DataFrame is a high-performance C++ library for in-memory data exploration, transformation, and statistical analysis — designed for data scientists, quant traders, and C++ developers who need efficient tabular data processing without Python overhead. This library is designed to provide similar functionalities to data manipulation and analysis tools found in other languages, such as Python's Pandas or R's data.frame. It is built using C++. The complete source code is publicly available on GitHub under the BSD 3-Clause "New" or "Revised" License, making it a useful reference for students building a Data Science mini project or final-year project.

1. Introduction

DataFrame is a high-performance C++ library for in-memory data exploration, transformation, and statistical analysis — designed for data scientists, quant traders, and C++ developers who need efficient tabular data processing without Python overhead. This library is designed to provide similar functionalities to data manipulation and analysis tools found in other languages, such as Python's Pandas or R's data.frame. It aims to offer a robust and efficient way to handle tabular data in C++. The depth and breadth of functionalities offered by C++ DataFrame alone are greater than functionalities offered by packages such as Pandas, data.frame, and Polars combined. You can slice the data in many ways. You can join, merge, group-by, cross tabulate, pivot the data. You can run various statistical, summarization, financial, and ML algorithms on the data. You can add your custom algorithms easily. You can multi-column sort, custom pick and delete the data. And much more ... DataFrame also includes a large collection of analytical algorithms in form of visitors. These are from basic stats such as Mean, Stdev, Moving Averages, ... to more involved analysis such as PCA, Polynomial Fit, FFT, Eigens ... including a good collection of trading indicators. You can also easily add your own algorithms. Many of these algorithms work seamlessly with both scalar and multidimensional datasets. DataFrame also employs extensive multithreading in almost all its API’s, for large datasets. That makes DataFrame especially suitable for analyzing large datasets. For basic operations to start you off, see Hello World and/or Cheat Sheet. For a complete list of features with code samples, see documentation.

The maximum dataset I could load into Polars was 300m rows per column. Any bigger dataset blew up the memory and caused OS to kill it. I ran C++ DataFrame with 2b rows per column and I am sure it would have run with bigger datasets too. So, I was forced to run both with 300m rows to compare. I ran each test 4 times and took the best time. Polars numbers varied a lot from one run to another, especially calculation and selection times. C++ DataFrame numbers were significantly more consistent.

2. Objective

C++ DataFrame for statistical, financial, and ML analysis in modern C++

This project demonstrates how C++ can be applied to a real-world Data Science problem.

4. Technology Stack

C++

5. System Requirements

General requirements for this technology stack — check the README for exact versions.

  • Arduino IDE / PlatformIO or a C++ compiler (g++)
  • Target board (e.g. Arduino, ESP32) where applicable
  • Git (to clone the repository)

6. Installation & Setup

git clone https://github.com/hosseinmoein/DataFrame.git
cd DataFrame

Full setup instructions are in the project README.

7. Future Enhancements

Suggested extensions you can add to make this your own project.

  • Turn the analysis into an interactive dashboard
  • Automate data refresh with a scheduled job
  • Add a predictive model on top of the analysis

8. Viva / Review Questions

Common questions examiners ask for projects in this domain.

  1. What is the source of the dataset and how was missing data handled?
  2. Which exploratory analysis steps revealed the most useful insight?
  3. Why were these particular charts chosen to present the data?
  4. Which statistical or ML technique supports the conclusions?
  5. How could the analysis be automated or refreshed with new data?

9. Source Code & License

This project is developed by hosseinmoein and published on GitHub under the BSD 3-Clause "New" or "Revised" License. Please follow the license terms and credit the original author when you use or modify this code.

Want to build this as your internship project?

Work on a Data Science project like this with mentor guidance, weekly reviews and an internship certificate from Training Trains, Erode — online or offline.

Apply for Data Science Internship