
Start with what the model will actually see.
The repository uses a Kaggle-hosted PlantVillage subset with 15 categories across tomato, potato, and bell pepper. Labels include healthy leaves and conditions such as early blight, late blight, bacterial spot, and leaf mold.
The notebook records 18,580 training images and 2,058 validation images. Images are resized to 150 × 150 pixels and rescaled by 1/255 before being passed to the network. The published generators apply rescaling; they do not configure additional image augmentation.
An inspectable convolutional baseline.
- Prepare150 × 150 RGB input
- Extract32 → 64 → 128 filters
- RegularizePooling + 25% dropout
- Classify15-way softmax
Three convolutional stages, each followed by max pooling, extract increasingly abstract features. A dropout layer precedes flattening, a 1,024-unit dense layer, and the final softmax output. Training uses Adam and categorical cross-entropy.
This straightforward architecture makes the data flow easy to understand, but the large flattened representation is expensive. A follow-up could compare a smaller global-pooling head or transfer-learning baseline under the same evaluation protocol.
The best epoch is not the last epoch.
| Checkpoint | Training accuracy | Validation accuracy |
|---|---|---|
| Epoch 13 | 98.84% | 92.66% |
| Epoch 15 · final | 98.99% | 91.69% |
Training accuracy continues to rise while validation results fluctuate. That gap is a reason to select checkpoints on validation performance and inspect per-class errors, rather than headline the training accuracy.

A leaf on a clean background is not a field test.
Lighting, clutter, camera quality, and unseen disease conditions can change the input distribution. The project’s validation split does not establish deployment accuracy on farm photographs.
PlantVillage research by Mohanty and colleagues illustrates this challenge: strong performance on controlled imagery can fall substantially on differently sourced images. That is research context, not a result measured for this model. Read the study.
- Build an independent set of field photographs and report per-class precision, recall, and a confusion matrix.
- Use checkpoint selection and early stopping to address the training–validation gap.
- Test realistic augmentations and an explicit “uncertain” outcome before designing a field-facing application.
The hardest part: finding data that matters in Nepal.
The hardest part was finding real-world imagery for the crops most relevant to Nepal. An accessible labeled dataset makes it possible to train and evaluate a classifier, but it does not automatically represent local crops, growing conditions, or the photographs a farmer would take.
This project made the dataset question central for me. The current PlantVillage-based experiment is a starting point; the next step I care most about is a locally relevant collection with realistic field conditions and reliable labels.