Plant Doc Dataset

Dataset used in "PlantDoc: A Dataset for Visual Plant Disease Detection" accepted in CODS-COMAD 2020

AI & Machine LearningMulti-languageCC-BY-4.0

Abstract

Plant Doc Dataset is an open-source AI & Machine Learning project. Dataset used in "PlantDoc: A Dataset for Visual Plant Disease Detection" accepted in CODS-COMAD 2020. India loses 35% of the annual crop yield due to plant diseases. Early detection of plant diseases remains difficult due to the lack of lab infrastructure and expertise. The complete source code is publicly available on GitHub under the Creative Commons Attribution 4.0 International, making it a useful reference for students building an AI & Machine Learning mini project or final-year project.

1. Introduction

India loses 35% of the annual crop yield due to plant diseases. Early detection of plant diseases remains difficult due to the lack of lab infrastructure and expertise. In this paper, we explore the possibility of computer vision approaches for scalable and early plant disease detection. The lack of availability of sufficiently large-scale non-lab data set remains a major challenge for enabling vision based plant disease detection. Against this background, we present PlantDoc: a dataset for visual plant disease detection. Our dataset contains 2,598 data points in total across 13 plant species and up to 17 classes of diseases, involving approximately 300 human hours of effort in annotating internet scraped images. To show the efficacy of our dataset, we learn 3 models for the task of plant disease classification. Our results show that modelling using our dataset can increase the classification accuracy by up to 31%. We believe that our dataset can help reduce the entry barrier of computer vision techniques in plant disease detection.

While the Object-Detection dataset can accessed through the following link.

2. Objective

Dataset used in "PlantDoc: A Dataset for Visual Plant Disease Detection" accepted in CODS-COMAD 2020

This project demonstrates how modern tools can be applied to a real-world AI & Machine Learning problem.

4. Technology Stack

See repository.

5. System Requirements

General requirements for this technology stack — check the README for exact versions.

  • See the project README for exact requirements
  • Git (to clone the repository)

6. Installation & Setup

git clone https://github.com/pratikkayal/PlantDoc-Dataset.git
cd PlantDoc-Dataset

Full setup instructions are in the project README.

7. Future Enhancements

Suggested extensions you can add to make this your own project.

  • Deploy the model as a web app with Streamlit, Flask or FastAPI
  • Compare against an additional model and report the metric difference
  • Add explainability (SHAP / Grad-CAM)

8. Viva / Review Questions

Common questions examiners ask for projects in this domain.

  1. What dataset does the project use and how was it pre-processed?
  2. Which algorithm / model architecture is used and why was it chosen over alternatives?
  3. How are training and testing data split, and how is overfitting avoided?
  4. Which evaluation metrics (accuracy, precision, recall, F1) are reported and what do they mean here?
  5. How would you deploy this model for real users?

9. Source Code & License

This project is developed by pratikkayal and published on GitHub under the Creative Commons Attribution 4.0 International. Please follow the license terms and credit the original author when you use or modify this code.

Want to build this as your internship project?

Work on an AI & Machine Learning project like this with mentor guidance, weekly reviews and an internship certificate from Training Trains, Erode — online or offline.

Apply for AI & Machine Learning Internship