Training, comparative evaluation, TensorFlow Lite optimization, and a deployable mobile model built with the Oxford 102 Flowers dataset.
This repository presents a professional AI engineering workflow for multiclass image classification. It covers dataset integrity checks, class-aware splitting, augmentation, imbalance handling, transfer learning, CNN feature extraction for gradient-boosted trees, comparative evaluation, and TensorFlow Lite conversion.
MobileNetV2 achieved the strongest recorded test result and was selected as the deployable model artifact.
| Project metric | Recorded value |
|---|---|
| Flower classes | 102 |
| Source images | 8,189 |
| Images retained after duplicate handling | 8,185 |
| Models evaluated | 4 |
| Best test accuracy | 95.71% |
| Optimized TensorFlow Lite size | 2.53 MB |
- Validates image availability, readability, duplication, and class distribution.
- Creates class-aware training, validation, and test splits.
- Applies augmentation and balanced class weights for neural model training.
- Compares XGBoost with CNN features, VGG16, NASNetMobile, and MobileNetV2.
- Exports and evaluates an optimized TensorFlow Lite classifier.
- Packages the final model and its 102-label mapping for mobile integration.
Oxford 102 Flowers
│
▼
Integrity checks ──► duplicate handling ──► class-aware split
│
▼
Augmentation + [0, 1] scaling + class weighting
│
├──► MobileNetV2 features + XGBoost
├──► VGG16
├──► NASNetMobile
└──► MobileNetV2 ──► fine-tuning ──► TFLite optimization
│
▼
model + 102-class labels
The pipeline uses the Oxford 102 Flowers dataset, containing 8,189 images across 102 flower categories. The official labels and split metadata are loaded, while this project creates a class-aware 60/10/30 split after duplicate handling.
Four duplicate groups were identified. One repeated image from each group was moved out of the training image pool, leaving 8,185 images:
| Split | Images | Share of retained images |
|---|---|---|
| Training | 4,872 | 59.52% |
| Validation | 774 | 9.46% |
| Test | 2,539 | 31.02% |
Every class is represented in all three generated splits. Small differences from the target proportions result from integer rounding within each class.
Class counts before splitting show the dataset's substantial imbalance.
Class-aware allocation after duplicate handling.
The ten images in benchmark_images are unmodified samples from the Oxford dataset. Dataset downloads, extracted folders, and generated split folders are excluded from Git. Included dataset samples and all other third-party dataset content are not covered by the project's MIT License; see ATTRIBUTION.md.
All neural models use 224 × 224 RGB inputs. Keras generators scale pixel values with rescale=1./255, producing the [0, 1] range for training, validation, testing, and TensorFlow Lite evaluation. The tracked TensorFlow Lite model expects the same input preprocessing.
Training augmentation includes:
- rotations up to 15 degrees;
- width and height shifts of 10%;
- shear up to 10% and zoom up to 20%;
- horizontal flips;
- brightness scaling from 0.8 to 1.2;
- nearest-neighbour fill for transformed pixels.
NumPy, TensorFlow, and generator seeds are set to 42. Balanced class weights are calculated from the training labels and supplied to neural model training.
An ImageNet-pretrained MobileNetV2 backbone removes its classification head and produces fixed feature vectors through global-average pooling. RandomizedSearchCV evaluates three XGBoost configurations with two-fold cross-validation. The recorded best configuration uses 150 estimators, maximum depth 4, and learning rate 0.1.
The ImageNet backbone begins frozen beneath global-average pooling and a 128-unit dense classifier. A second phase unfreezes the model and continues fine-tuning at a learning rate of 1e-5.
The pretrained backbone remains frozen beneath global-average pooling, a 256-unit dense layer, 50% dropout, and a 102-class softmax layer. The retained experiment is the base-training run because the attempted fine-tuning did not improve validation behaviour.
The pretrained backbone begins frozen beneath global-average pooling, 20% dropout, and a 102-class softmax layer. The selected model then undergoes full fine-tuning at 1e-5.
The neural models use categorical cross-entropy, early stopping, learning-rate reduction where configured, augmentation, and class weighting. Accuracy plus macro and weighted precision, recall, and F1 are recorded. XGBoost uses multiclass log loss during fitting and the same classification metrics for evaluation.
Combined MobileNetV2 base-training and fine-tuning history. The dashed line marks the start of fine-tuning.
The following results use the reserved test split of 2,539 images.
| Model | Test accuracy | Weighted F1 |
|---|---|---|
| XGBoost | 79.80% | 79.39% |
| VGG16 | 83.42% | 83.41% |
| NASNetMobile | 80.66% | 80.50% |
| MobileNetV2 | 95.71% | 95.72% |
| Model | Training time | Inference per image | Serialized size |
|---|---|---|---|
| XGBoost | 483.13 s | 0.000016 s | 13.52 MB |
| VGG16 | 21,641.87 s | 0.055700 s | 169.40 MB |
| NASNetMobile | 2,821.40 s | 0.008552 s | 22.30 MB |
| MobileNetV2 | 11,841.52 s | 0.005182 s | 27.74 MB |
Training and inference times come from the recorded notebook run and are hardware-dependent. Model sizes are measurements of the serialized artifacts produced by that run.
The original comparison uses logarithmic scaling because the recorded measures span different units and ranges.
The selected MobileNetV2 model is exported through SavedModel and converted with TensorFlow Lite default optimization.
| Artifact | Size |
|---|---|
| Keras MobileNetV2 | 27.74 MB |
| Optimized TensorFlow Lite | 2.53 MB |
| Reduction | 25.22 MB |
The conversion notebook evaluates the optimized model on all 2,539 test images with the same [0, 1] input scaling used during neural-model training. Its displayed macro and weighted precision, recall, and F1 values round to 0.95.
The selected MobileNetV2 model was converted to TensorFlow Lite and integrated into FlowerApp on Google Play. The Android application is maintained separately: this repository contains the machine learning pipeline, final TensorFlow Lite artifact, and labels, but it does not include Android source code.
The related Google Play developer profile lists the published application.
flower-classification/
├── assets/figures/
│ ├── dataset/ # Dataset inspection figures
│ ├── evaluation/ # Model and TFLite evaluation figures
│ └── training/ # Learning curves
├── benchmark_images/ # Ten Oxford dataset samples
├── models/
│ ├── labels.txt # Ordered 102-class mapping
│ └── mobilenetv2_optimized.tflite
├── notebooks/
│ ├── 01_flower_classification_training.ipynb
│ ├── 02_tflite_conversion_validation.ipynb
│ └── labels.txt # Conversion-notebook label output
├── .gitignore
├── ATTRIBUTION.md
├── LICENSE
├── README.md
├── requirements-conversion.txt
└── requirements-training.txt
Raw data, generated splits, local Keras models, SavedModel exports, pickled predictions, and experiment-output directories remain local and are excluded from normal Git tracking.
notebooks/01_flower_classification_training.ipynb records Python 3.12.7, TensorFlow 2.19.0, NumPy 1.26.4, pandas 2.2.2, and scikit-learn 1.5.1. The remaining package pins in requirements-training.txt match the inspected local training environment.
python3.12 -m venv .venv-training
source .venv-training/bin/activate
python -m pip install --upgrade pip
python -m pip install -r requirements-training.txtActivate it later with:
source .venv-training/bin/activatenotebooks/02_tflite_conversion_validation.ipynb was intentionally developed in a separate environment. Its saved output records Python 3.10.18, TensorFlow 2.14.0, NumPy 1.23.5, pandas 2.3.1, and scikit-learn 1.7.1. Exact recorded core versions and bounded compatible notebook dependencies are defined in requirements-conversion.txt.
python3.10 -m venv .venv-conversion
source .venv-conversion/bin/activate
python -m pip install --upgrade pip
python -m pip install -r requirements-conversion.txtActivate it later with:
source .venv-conversion/bin/activateTwo environments are retained because the training and conversion runs were recorded with different Python, TensorFlow, NumPy, pandas, and scikit-learn versions. Combining them would misrepresent the environments that produced the saved results.
- Clone the repository.
- Create the training environment above.
- Download
102flowers.tgz,imagelabels.mat, andsetid.matfrom the official Oxford dataset page. - Place those three files in a local
raw_data/directory at the repository root.
The downloaded archive and generated dataset folders are intentionally ignored by Git.
Run the notebooks in this order:
- Activate
.venv-trainingand opennotebooks/01_flower_classification_training.ipynb. - Complete training and export the selected SavedModel locally.
- Deactivate the training environment.
- Activate
.venv-conversionand opennotebooks/02_tflite_conversion_validation.ipynb. - Convert, validate, and package the TensorFlow Lite model.
source .venv-training/bin/activate
jupyter lab notebooks/01_flower_classification_training.ipynbThen, in a separate conversion session:
source .venv-conversion/bin/activate
jupyter lab notebooks/02_tflite_conversion_validation.ipynbBoth notebooks locate the repository root automatically when launched from any directory inside the repository.
- NumPy, TensorFlow, and generator seeds are set to 42.
- The two recorded software environments remain intentionally separate.
- Dataset archives, split folders, and large generated models are reproducible local artifacts and are not committed.
- Notebook outputs, execution counts, metrics, charts, and exported models are retained from the recorded runs.
- The final label mapping contains 102 entries and is packaged with the TensorFlow Lite artifact.
- The dataset is imbalanced, and the custom split differs from the official Oxford split.
- Timing measurements depend on the original hardware and software stack and should not be treated as device-independent benchmarks.
- Neural models were trained and evaluated with pixel values scaled to
[0, 1]. Architecture-specific preprocessing is a future experiment and is not claimed as part of the recorded results. - No per-class confusion matrix or systematic analysis of commonly confused categories is published.
- The repository does not include the large Keras, SavedModel, serialized prediction, or Android source artifacts.
- Evaluate architecture-specific preprocessing in new controlled training runs.
- Add per-class error analysis and confusion matrices without altering the recorded baseline.
- Benchmark the TensorFlow Lite model across documented mobile hardware.
- Add automated static validation for notebook syntax, artifact hashes, and model-interface compatibility.
Dataset, pretrained-weight, architecture, and software references are listed in ATTRIBUTION.md. Third-party dataset content and pretrained weights are not covered by this repository's MIT License.
Project-authored source and documentation are available under the MIT License.
This project was developed by Fatemeh Sabourinia and Alireza Zaeri.





