Paddle Vi T

:robot: PaddleViT: State-of-the-art Visual Transformer and MLP Models for PaddlePaddle 2.0+

AI & Machine LearningPythonApache-2.0

Abstract

Paddle Vi T is an open-source AI & Machine Learning project. :robot: PaddleViT: State-of-the-art Visual Transformer and MLP Models for PaddlePaddle 2.0+. PaddlePaddle Visual Transformers (PaddleViT or PPViT) is a collection of vision models beyond convolution. Most of the models are based on Visual Transformers, Visual Attentions, and MLPs, etc. It is built using Python, Computer Vision. Key capabilities include: State-of-the-art; State-of-the-art transformer models for multiple CV tasks; State-of-the-art data processings and training methods. The complete source code is publicly available on GitHub under the Apache License 2.0, making it a useful reference for students building an AI & Machine Learning mini project or final-year project.

1. Introduction

PaddlePaddle Visual Transformers (PaddleViT or PPViT) is a collection of vision models beyond convolution. Most of the models are based on Visual Transformers, Visual Attentions, and MLPs, etc. PaddleViT also integrates popular layers, utilities, optimizers, schedulers, data augmentations, training/validation scripts for PaddlePaddle 2.1+. The aim is to reproduce a wide variety of state-of-the-art ViT and MLP models with full training/validation procedures. We are passionate about making cuting-edge CV techniques easier to use for everyone.

PaddleViT provides models and tools for multiple vision tasks, such as classifications, object detection, semantic segmentation, GAN, and more. Each model architecture is defined in standalone python module and can be modified to enable quick research experiments. At the same time, pretrained weights can be downloaded and used to finetune on your own datasets. PaddleViT also integrates popular tools and modules for custimized dataset, data preprocessing, performance metrics, DDP and more.

PaddleViT is backed by popular deep learning framework PaddlePaddle, we also provide tutorials and projects on Paddle AI Studio. It's intuitive and straightforward to get started for new users.

2. Objective

:robot: PaddleViT: State-of-the-art Visual Transformer and MLP Models for PaddlePaddle 2.0+

This project demonstrates how Python, Computer Vision can be applied to a real-world AI & Machine Learning problem.

3. Key Features / Modules

  • State-of-the-art
  • State-of-the-art transformer models for multiple CV tasks
  • State-of-the-art data processings and training methods
  • We keep pushing it forward.
  • Easy-to-use tools
  • Easy configs for model vairants
  • Modular design for utiliy functions and tools
  • Low barrier for educators and practitioners
  • Unified framework for all the models
  • Easily customizable to your needs

4. Technology Stack

PythonComputer Vision

5. System Requirements

General requirements for this technology stack — check the README for exact versions.

  • Python 3.8 or later
  • pip / virtualenv for dependencies
  • VS Code, PyCharm or Jupyter Notebook
  • Git (to clone the repository)

6. Installation & Setup

git clone https://github.com/BR-IDL/PaddleViT.git
cd PaddleViT
  1. Create a conda virtual environment and activate it.
  2. Install PaddlePaddle following the official instructions, e.g.,
  3. Install dependency packages
  4. General dependencies:
  5. Packages for Segmentation:
  6. Packages for GAN:
  7. Clone project from GitHub
conda create -n paddlevit python=3.7 -y
   conda activate paddlevit
conda install paddlepaddle-gpu==2.1.2 cudatoolkit=10.2 --channel https://mirrors.tuna.tsinghua.edu.cn/anaconda/cloud/Paddle/
pip install yacs pyyaml
pip install cityscapesScripts

Full setup instructions are in the project README.

7. Future Enhancements

Suggested extensions you can add to make this your own project.

  • Deploy the model as a web app with Streamlit, Flask or FastAPI
  • Compare against an additional model and report the metric difference
  • Add explainability (SHAP / Grad-CAM)

8. Viva / Review Questions

Common questions examiners ask for projects in this domain.

  1. What dataset does the project use and how was it pre-processed?
  2. Which algorithm / model architecture is used and why was it chosen over alternatives?
  3. How are training and testing data split, and how is overfitting avoided?
  4. Which evaluation metrics (accuracy, precision, recall, F1) are reported and what do they mean here?
  5. How would you deploy this model for real users?

9. Source Code & License

This project is developed by BR-IDL and published on GitHub under the Apache License 2.0. Please follow the license terms and credit the original author when you use or modify this code.

Want to build this as your internship project?

Work on an AI & Machine Learning project like this with mentor guidance, weekly reviews and an internship certificate from Training Trains, Erode — online or offline.

Apply for AI & Machine Learning Internship