Ngboost

Natural Gradient Boosting for Probabilistic Prediction

AI & Machine LearningJupyter NotebookApache-2.0

Abstract

Ngboost is an open-source AI & Machine Learning project. Natural Gradient Boosting for Probabilistic Prediction. ngboost is a Python library that implements Natural Gradient Boosting, as described in "NGBoost: Natural Gradient Boosting for Probabilistic Prediction". It is built on top of Scikit-Learn, and is designed to be scalable and modular with respect to choice of proper scoring rule, distribution, and base learner. It is built using Jupyter Notebook, Machine Learning, Python. The complete source code is publicly available on GitHub under the Apache License 2.0, making it a useful reference for students building an AI & Machine Learning mini project or final-year project.

1. Introduction

ngboost is a Python library that implements Natural Gradient Boosting, as described in "NGBoost: Natural Gradient Boosting for Probabilistic Prediction". It is built on top of Scikit-Learn, and is designed to be scalable and modular with respect to choice of proper scoring rule, distribution, and base learner. A didactic introduction to the methodology underlying NGBoost is available in this slide deck.

2. Objective

Natural Gradient Boosting for Probabilistic Prediction

This project demonstrates how Jupyter Notebook, Machine Learning, Python can be applied to a real-world AI & Machine Learning problem.

4. Technology Stack

Jupyter NotebookMachine LearningPython

5. System Requirements

General requirements for this technology stack — check the README for exact versions.

  • Python 3.8 or later with Jupyter Notebook / JupyterLab (or Google Colab)
  • pip for dependencies
  • Python 3.8 or later
  • pip / virtualenv for dependencies
  • VS Code, PyCharm or Jupyter Notebook
  • Git (to clone the repository)

6. Installation & Setup

git clone https://github.com/stanfordmlgroup/ngboost.git
cd ngboost
via pip

pip install --upgrade ngboost

via conda-forge

conda install -c conda-forge ngboost
from ngboost import NGBRegressor

from sklearn.datasets import fetch_california_housing
from sklearn.model_selection import train_test_split
from sklearn.metrics import mean_squared_error

# Load California housing dataset
cal = fetch_california_housing()
X, Y = cal.data, cal.target

X_train, X_test, Y_train, Y_test = train_test_split(X, Y, test_size=0.2)

ngb = NGBRegressor().fit(X_train, Y_train)
Y_preds = ngb.predict(X_test)
Y_dists = ngb.pred_dist(X_test)

# test Mean Squared Error
test_MSE = mean_squared_error(Y_preds, Y_test)
print('Test MSE', test_MSE)

# test Negative Log Likelihood
test_NLL = -Y_dists.logpdf(Y_test).mean()
print('Test NLL', test_NLL)

Full setup instructions are in the project README.

7. Future Enhancements

Suggested extensions you can add to make this your own project.

  • Deploy the model as a web app with Streamlit, Flask or FastAPI
  • Compare against an additional model and report the metric difference
  • Add explainability (SHAP / Grad-CAM)

8. Viva / Review Questions

Common questions examiners ask for projects in this domain.

  1. What dataset does the project use and how was it pre-processed?
  2. Which algorithm / model architecture is used and why was it chosen over alternatives?
  3. How are training and testing data split, and how is overfitting avoided?
  4. Which evaluation metrics (accuracy, precision, recall, F1) are reported and what do they mean here?
  5. How would you deploy this model for real users?

9. Source Code & License

This project is developed by stanfordmlgroup and published on GitHub under the Apache License 2.0. Please follow the license terms and credit the original author when you use or modify this code.

Want to build this as your internship project?

Work on an AI & Machine Learning project like this with mentor guidance, weekly reviews and an internship certificate from Training Trains, Erode — online or offline.

Apply for AI & Machine Learning Internship