Abstract
Video Audio Face Emotion Recognition is an open-source AI & Machine Learning project. The repo contains an audio emotion detection model, facial emotion detection model, and a model that combines both these models to predict emotions from a video. This multimodal emotion detection model predicts a speaker's emotion using audio and image sequences from videos. The repository contains two primary models: an audio tone recognition model with a CNN for audio-based emotion prediction, and a facial emotion recognition model using a CNN and optional mediapipe face landmarks for facial emotion prediction. It is built using Jupyter Notebook. The complete source code is publicly available on GitHub under the MIT License, making it a useful reference for students building an AI & Machine Learning mini project or final-year project.
1. Introduction
This multimodal emotion detection model predicts a speaker's emotion using audio and image sequences from videos. The repository contains two primary models: an audio tone recognition model with a CNN for audio-based emotion prediction, and a facial emotion recognition model using a CNN and optional mediapipe face landmarks for facial emotion prediction. The third model combines a video clip's audio and image sequences, processed through an LSTM for speaker emotion prediction. Hyperparameters such as landmark usage, CNN model selection, LSTM units, and dense layers are tuned for optimal accuracy using included modules. For new datasets, follow the instructions below to retune the hyperparameters. _
2. Objective
The repo contains an audio emotion detection model, facial emotion detection model, and a model that combines both these models to predict emotions from a video
This project demonstrates how Jupyter Notebook can be applied to a real-world AI & Machine Learning problem.
4. Technology Stack
5. System Requirements
General requirements for this technology stack — check the README for exact versions.
- Python 3.8 or later with Jupyter Notebook / JupyterLab (or Google Colab)
- pip for dependencies
- Git (to clone the repository)
6. Installation & Setup
git clone https://github.com/rishiswethan/Video-Audio-Face-Emotion-Recognition.git
cd Video-Audio-Face-Emotion-Recognition- git clone https://github.com/rishiswethan/Video-Audio-Face-Emotion-Recognition.git
- cd Video-Audio-Face-Emotion-Recognition
- git clone https://github.com/rishiswethan/pytorch_utils.git source/pytorch_utils
- cd source/pytorch_utils && git checkout v1.0.3 && cd ../..
- python -m venv venv
- Activate the virtual environment
- Linux/MacOS: source venv/bin/activate
- Windows: venv\Scripts\activate
Full setup instructions are in the project README.
7. Future Enhancements
Suggested extensions you can add to make this your own project.
- Deploy the model as a web app with Streamlit, Flask or FastAPI
- Compare against an additional model and report the metric difference
- Add explainability (SHAP / Grad-CAM)
8. Viva / Review Questions
Common questions examiners ask for projects in this domain.
- What dataset does the project use and how was it pre-processed?
- Which algorithm / model architecture is used and why was it chosen over alternatives?
- How are training and testing data split, and how is overfitting avoided?
- Which evaluation metrics (accuracy, precision, recall, F1) are reported and what do they mean here?
- How would you deploy this model for real users?
9. Source Code & License
This project is developed by rishiswethan and published on GitHub under the MIT License. Please follow the license terms and credit the original author when you use or modify this code.
Want to build this as your internship project?
Work on an AI & Machine Learning project like this with mentor guidance, weekly reviews and an internship certificate from Training Trains, Erode — online or offline.
Apply for AI & Machine Learning Internship