Dataframe

A fast, safe, and intuitive DataFrame library.

Data ScienceHaskellMIT

Abstract

Dataframe is an open-source Data Science project. A fast, safe, and intuitive DataFrame library. Tabular data analysis in Haskell. Read CSV, Parquet, and JSON files, transform columns with a typed expression DSL, and optionally lock down your entire schema at the type level for compile-time safety. It is built using Haskell. The complete source code is publicly available on GitHub under the MIT License, making it a useful reference for students building a Data Science mini project or final-year project.

1. Introduction

Tabular data analysis in Haskell. Read CSV, Parquet, and JSON files, transform columns with a typed expression DSL, and optionally lock down your entire schema at the type level for compile-time safety.

2. Objective

A fast, safe, and intuitive DataFrame library.

This project demonstrates how Haskell can be applied to a real-world Data Science problem.

4. Technology Stack

Haskell
  • Concise, declarative, composable data pipelines using the |> pipe operator.
  • Choose your level of type safety: keep it lightweight for quick analysis, or lock it down for production pipelines.
  • High performance from Haskell's optimizing compiler and an efficient columnar memory model based onn Apache Arrow.
  • Designed for interactivity: a custom REPL, IHaskell/Sabela notebook support, terminal and web plotting, and helpful error messages.

5. System Requirements

General requirements for this technology stack — check the README for exact versions.

  • See the project README for exact requirements
  • Git (to clone the repository)

6. Installation & Setup

git clone https://github.com/DataHaskell/dataframe.git
cd dataframe
cabal update
cabal install dataframe
build-depends: base >= 4, dataframe
$ dataframe
dataframe> df = D.fromNamedColumns [("product", D.fromList [1, 1, 2, 2, 3, 3 :: Int]), ("amount",  D.fromList [100, 120, 50, 20, 40, 30 :: Int]) ]
dataframe> df |> D.groupBy ["product"] |> ["total" .= F.countAll ]
-- cabal: build-depends: dataframe, text
-- cabal: default-extensions: OverloadedStrings, TypeApplications, TemplateHaskell, DataKinds, TypeFamilies, FlexibleInstances, FlexibleContexts, ScopedTypeVariables, DeriveGeneric, UndecidableInstances
import qualified DataFrame as D
import qualified DataFrame.Functions as F
import DataFrame.Expression.Operators

sales = D.fromNamedColumns
    [ ("product", D.fromList [1, 1, 2, 2, 3, 3 :: Int])
    , ("amount",  D.fromList [100, 120, 50, 20, 40, 30 :: Int])
    ]

-- Group by product and compute totals
sales
    |> D.groupBy ["product"]
    |> D.aggregate [ F.sum (F.col @Int "amount") `as` "total"
                   , F.count (F.col @Int "amount") `as` "orders"
                   ]
    |> D.toMarkdown'

Full setup instructions are in the project README.

7. Future Enhancements

Suggested extensions you can add to make this your own project.

  • Turn the analysis into an interactive dashboard
  • Automate data refresh with a scheduled job
  • Add a predictive model on top of the analysis

8. Viva / Review Questions

Common questions examiners ask for projects in this domain.

  1. What is the source of the dataset and how was missing data handled?
  2. Which exploratory analysis steps revealed the most useful insight?
  3. Why were these particular charts chosen to present the data?
  4. Which statistical or ML technique supports the conclusions?
  5. How could the analysis be automated or refreshed with new data?

9. Source Code & License

This project is developed by DataHaskell and published on GitHub under the MIT License. Please follow the license terms and credit the original author when you use or modify this code.

Want to build this as your internship project?

Work on a Data Science project like this with mentor guidance, weekly reviews and an internship certificate from Training Trains, Erode — online or offline.

Apply for Data Science Internship