All selected work

Vegetation detectionLearning to read the leaves.

A convolutional classifier for healthy and diseased leaf imagery. Built around a 15-class PlantVillage subset, the project explores the full path from image preparation to validation—and why a good validation score is only the beginning.

Computer vision / Image classification
The problem
Recognize disease-related visual patterns in tomato, potato, and bell-pepper leaf images.
My contribution
Prepared image generators, implemented a three-stage CNN, trained the model, and inspected learning curves and validation behavior.
The result
The saved notebook reaches 92.66% validation accuracy at epoch 13; the final epoch records 91.69%.
The constraint
Evaluation uses a curated leaf-image split. Performance on uncontrolled field photographs is not established.
Leaf-image examples from the vegetation disease detection project
Visible variation in leaf color, texture, and lesions provides the signal for classification.Open image

Start with what the model will actually see.

The repository uses a Kaggle-hosted PlantVillage subset with 15 categories across tomato, potato, and bell pepper. Labels include healthy leaves and conditions such as early blight, late blight, bacterial spot, and leaf mold.

The notebook records 18,580 training images and 2,058 validation images. Images are resized to 150 × 150 pixels and rescaled by 1/255 before being passed to the network. The published generators apply rescaling; they do not configure additional image augmentation.

Examples of the leaf-image categories used for model development
Dataset examples from the project’s exploratory analysis.Open image

An inspectable convolutional baseline.

  1. Prepare150 × 150 RGB input
  2. Extract32 → 64 → 128 filters
  3. RegularizePooling + 25% dropout
  4. Classify15-way softmax
The sequential architecture in the published TensorFlow notebook.

Three convolutional stages, each followed by max pooling, extract increasingly abstract features. A dropout layer precedes flattening, a 1,024-unit dense layer, and the final softmax output. Training uses Adam and categorical cross-entropy.

This straightforward architecture makes the data flow easy to understand, but the large flattened representation is expensive. A follow-up could compare a smaller global-pooling head or transfer-learning baseline under the same evaluation protocol.

The best epoch is not the last epoch.

Recorded outputs in Vegetation_Disease_Detector.ipynb
CheckpointTraining accuracyValidation accuracy
Epoch 1398.84%92.66%
Epoch 15 · final98.99%91.69%

Training accuracy continues to rise while validation results fluctuate. That gap is a reason to select checkpoints on validation performance and inspect per-class errors, rather than headline the training accuracy.

Training and validation accuracy curves from the leaf-classification experiment
Learning curves from the project materials. Exact values above are taken from the saved notebook output.Open image

A leaf on a clean background is not a field test.

Lighting, clutter, camera quality, and unseen disease conditions can change the input distribution. The project’s validation split does not establish deployment accuracy on farm photographs.

PlantVillage research by Mohanty and colleagues illustrates this challenge: strong performance on controlled imagery can fall substantially on differently sourced images. That is research context, not a result measured for this model. Read the study.

  • Build an independent set of field photographs and report per-class precision, recall, and a confusion matrix.
  • Use checkpoint selection and early stopping to address the training–validation gap.
  • Test realistic augmentations and an explicit “uncertain” outcome before designing a field-facing application.

The hardest part: finding data that matters in Nepal.

The hardest part was finding real-world imagery for the crops most relevant to Nepal. An accessible labeled dataset makes it possible to train and evaluate a classifier, but it does not automatically represent local crops, growing conditions, or the photographs a farmer would take.

This project made the dataset question central for me. The current PlantVillage-based experiment is a starting point; the next step I care most about is a locally relevant collection with realistic field conditions and reliable labels.

Follow the evidence.

1 / 0

Arrows: browse · + / −: zoom · Shift + arrows: pan · Esc: close