Om Agent

[EMNLP-2024] Build multimodal language agents for fast prototype and production

AI & Machine LearningPythonApache-2.0

Abstract

Om Agent is an open-source AI & Machine Learning project. [EMNLP-2024] Build multimodal language agents for fast prototype and production. OmAgent is python library for building multimodal language agents with ease. We try to keep the library simple without too much overhead like other agent framework. It is built using Python. Key capabilities include: A flexible agent architecture that provides graph-based workflow orchestration engine and various memory type enabling contextual reasoning; Native multimodal interaction support include VLM models, real-time API, computer vision models, mobile connection and etc; A suite of state-of-the-art unimodal and multimodal agent algorithms that goes beyond simple LLM reasoning, e.g. ReAct, CoT, SC-Cot etc. The complete source code is publicly available on GitHub under the Apache License 2.0, making it a useful reference for students building an AI & Machine Learning mini project or final-year project.

1. Introduction

OmAgent is python library for building multimodal language agents with ease. We try to keep the library simple without too much overhead like other agent framework.

2. Objective

[EMNLP-2024] Build multimodal language agents for fast prototype and production

This project demonstrates how Python can be applied to a real-world AI & Machine Learning problem.

3. Key Features / Modules

  • A flexible agent architecture that provides graph-based workflow orchestration engine and various memory type enabling contextual reasoning.
  • Native multimodal interaction support include VLM models, real-time API, computer vision models, mobile connection and etc.
  • A suite of state-of-the-art unimodal and multimodal agent algorithms that goes beyond simple LLM reasoning, e.g. ReAct, CoT, SC-Cot etc.
  • Supports local deployment of models. You can deploy your own models locally by using OllamaOllama or LocalAI.
  • Fully distributed architecture, supports custom scaling. Also supports Lite mode, eliminating the need for middleware deployment.

4. Technology Stack

Python

5. System Requirements

General requirements for this technology stack — check the README for exact versions.

  • Python 3.8 or later
  • pip / virtualenv for dependencies
  • VS Code, PyCharm or Jupyter Notebook
  • Git (to clone the repository)

6. Installation & Setup

git clone https://github.com/om-ai-lab/OmAgent.git
cd OmAgent
  1. python >= 3.10
  2. Install omagent_core
  3. Run the simple VQA demo with webpage GUI:
pip install omagent-core
pip install -e omagent-core
cd examples/step1_simpleVQA
   python run_webpage.py

Full setup instructions are in the project README.

7. Future Enhancements

Suggested extensions you can add to make this your own project.

  • Deploy the model as a web app with Streamlit, Flask or FastAPI
  • Compare against an additional model and report the metric difference
  • Add explainability (SHAP / Grad-CAM)

8. Viva / Review Questions

Common questions examiners ask for projects in this domain.

  1. What dataset does the project use and how was it pre-processed?
  2. Which algorithm / model architecture is used and why was it chosen over alternatives?
  3. How are training and testing data split, and how is overfitting avoided?
  4. Which evaluation metrics (accuracy, precision, recall, F1) are reported and what do they mean here?
  5. How would you deploy this model for real users?

9. Source Code & License

This project is developed by om-ai-lab and published on GitHub under the Apache License 2.0. Please follow the license terms and credit the original author when you use or modify this code.

Want to build this as your internship project?

Work on an AI & Machine Learning project like this with mentor guidance, weekly reviews and an internship certificate from Training Trains, Erode — online or offline.

Apply for AI & Machine Learning Internship