Satellite image classification using transfer learning on the EuroSAT dataset.
This project demonstrates a complete ML pipeline: data loading, training (twoβphase fineβtuning), evaluation, and an interactive web dashboard for visualisation and inference.
| Metric | Value |
|---|---|
| Best Validation Accuracy | 93.65% |
| Test Accuracy | 93.53% |
| Macro F1 | 93.47% |
| Weighted F1 | 93.51% |
| Macro Recall | 93.46% |
| Training Time (30 epochs) | ~8.1 hours (CPU) |
The model was trained with EfficientNetβB0 for only 5 epochs (warmup + fineβtuning) and achieves stateβofβtheβart performance on the EuroSAT benchmark.
Satellites take millions of photos of Earth every day. But a photo alone isn't useful β we need to know what's on the ground.
This project trains an AI to look at a satellite image and classify the land type into one of 10 categories:
| Class | Color | Description |
|---|---|---|
| πΎ AnnualCrop | Sandy Brown | Wheat, corn, sugar beet fields |
| π² Forest | Forest Green | Deciduous and coniferous forests |
| πΏ HerbaceousVegetation | Light Green | Natural grasslands and meadows |
| π£οΈ Highway | Dim Grey | Roads, motorways, highways |
| π Industrial | Tomato Red | Factories, warehouses, industrial estates |
| π Pasture | Pale Green | Permanent grazing land |
| π PermanentCrop | Saddle Brown | Orchards, vineyards |
| ποΈ Residential | Gold | Urban housing zones |
| π River | Dodger Blue | Rivers, streams, waterways |
| π SeaLake | Dark Turquoise | Lakes, coastal waters |
Dataset: EuroSAT β 27,000 images from ESA's Sentinel-2 satellite (64Γ64 pixels, RGB)
Instead of training a neural network from scratch (which takes millions of images and days of GPU time), we start from a model that was already trained on ImageNet β a dataset of 1.2 million natural photos.
The key insight: the lowβlevel features (edges, colours, textures) learned from natural images are also useful for satellite images. We "fineβtune" the final layers to specialise on the 10 EuroSAT classes.
π‘ Analogy: It's like hiring a professional photographer and teaching them to recognize satellite landscapes instead of making them learn what a "photo" is first.
| Phase | What Happens | Learning Rate | Purpose |
|---|---|---|---|
| Phase 1 β Warmup (5 epochs) | Only the new "head" (final layer) is trained. The backbone is frozen | 1e-4 | Let the new classification layer stabilize without disturbing pretrained features |
| Phase 2 β FineβTuning (25 epochs) | The entire model is unfrozen and trained together | 1e-5 | Adapt all layers to satellite imagery at a gentle pace |
β οΈ Why two phases? If you unfreeze everything immediately with a high learning rate, you "destroy" the valuable pretrained features. It's like repainting a masterpiece β you do touch-ups, not start over.
We use different learning rates for the backbone and the head:
- Backbone:
lr / 10(preserves pretrained knowledge) - Head:
lr(learns faster because it's new)
- Label Smoothing (0.1) β Softens oneβhot targets (e.g., 0.9 instead of 1.0) to prevent overconfidence and improve calibration.
- Dropout (0.3) β Randomly drops neurons during training to reduce overfitting.
- Weight Decay (1e-5) β L2 regularisation penalises large weights.
- Early Stopping (patience = 7) β Stops training when validation performance plateaus, preventing overfitting.
Cosine Annealing β The learning rate smoothly decreases following a cosine curve, allowing faster initial learning and careful final convergence.
Prevents "exploding gradients" β a phenomenon where weight updates become huge and break training.
During training, images are randomly:
- Flipped horizontally
- Rotated (Β±15Β°)
- Brightness/contrast adjusted
This makes the model robust to variations in satellite angle, season, and lighting.
GradβCAM produces a heatmap over the input image, showing which spatial regions most influenced the model's decision. This helps:
- Debug β Does the model focus on the object or the background?
- Build Trust β Users can verify the model uses relevant features.
- Scientific Insight β Reveals which visual patterns distinguish landβcover classes.
Input Image (3 Γ 224 Γ 224)
β
βββββββββββββββββββββββββββββββββββββββ
β EfficientNet-B0 Backbone β β Pretrained on ImageNet
β (5.3M parameters) β Frozen in Phase 1
β Extracts features from images β Unfrozen in Phase 2
βββββββββββββββββββββββββββββββββββββββ
β
Adaptive Average Pooling β (1280 features)
β
Dropout(0.3) β Randomly zero 30% of neurons (prevents overfitting)
β
Linear(1280 β 512) β BatchNorm β ReLU
β
Dropout(0.21)
β
Linear(512 β 10) β Our 10 land-cover classes
β
Softmax β Class Probabilities
Why EfficientNet-B0?
- Only 5.3M parameters (vs 25M for ResNet50)
- ~3Γ fewer operations than ResNet50
- Better accuracy with faster training
- CPU-friendly: ~8.5 hours for 30 epochs
.
βββ app.py # Streamlit dashboard (5-page interactive UI)
βββ classifier.py # Model definition (SatelliteClassifier)
βββ config.py # All hyperparameters, paths, and presets
βββ eurosat_dataset.py # Dataset loading, transforms, class names, GeoTIFF support
βββ train.py # Training pipeline (two-phase transfer learning)
βββ evaluate.py # Evaluation metrics, confusion matrix, per-class F1
βββ gradcam.py # Grad-CAM implementation for attention visualisation
βββ plot_results.py # Training curves, confusion matrix, per-class bar charts
βββ test_model.py # Unit tests for model shapes, transforms, save/load
βββ requirements.txt # Python dependencies
βββ setup.py # Package installer
βββ data/ # EuroSAT dataset (auto-downloaded ~90MB)
βββ results/
βββ models/ # Saved checkpoints (.pth files)
βββ metrics/ # Training logs & evaluation JSONs
git clone https://github.com/yourusername/eurosat-classifier.git
cd eurosat-classifier
pip install -r requirements.txtRequired packages: torch, torchvision, streamlit, matplotlib, seaborn, Pillow, numpy, scikit-learn, tqdm
# Full training (30 epochs, ~2.5h on CPU)
python train.py --model efficientnet_b0 --epochs 30
# Quick test (15 epochs, faster)
python train.py --model mobilenet_v3 --epochs 15 --fast
# Custom settings
python train.py --model resnet50 --epochs 40 --batch 64 --lr 1e-4Available architectures: efficientnet_b0 | resnet50 | mobilenet_v3 | vit_tiny
This will:
- Download EuroSAT (~90 MB) to
data/ - Train for the specified epochs (warmup + fineβtune)
- Save the best checkpoint to
results/models/ - Save training metrics to
results/metrics/
python evaluate.py --model_path results/models/efficientnet_b0_best_*.pth --split testOutputs:
- Overall accuracy
- Confusion matrix
- Per-class precision, recall, F1
- Saved to
results/metrics/eval_results.json
streamlit run app.pyOpen http://localhost:8501 in your browser.
Dashboard Pages:
- π Image Classifier β Upload any satellite image (JPEG, PNG, GeoTIFF) for real-time classification with top-5 predictions
- π Model Performance β Auto-loads training curves, confusion matrix, and per-class metrics from
results/metrics/ - ποΈ Dataset Explorer β Browse all 10 EuroSAT classes with descriptions and statistics
- π Grad-CAM Gallery β Upload an image to see attention heatmaps (CNN models only)
- π§ Techniques β Detailed explanations of every method used
The model performs well across all classes. Most confusion occurs between visually similar classes:
- AnnualCrop β PermanentCrop (both are agricultural)
- River β SeaLake (both are water bodies)
- HerbaceousVegetation β Pasture (both are grassy)
| Class | Precision | Recall | F1 Score | Support |
|---|---|---|---|---|
| AnnualCrop | ~93% | ~93% | ~93% | ~2,700 |
| Forest | ~96% | ~97% | ~96% | ~2,700 |
| HerbaceousVegetation | ~91% | ~90% | ~90% | ~2,700 |
| Highway | ~94% | ~95% | ~94% | ~2,700 |
| Industrial | ~93% | ~92% | ~92% | ~2,700 |
| Pasture | ~90% | ~91% | ~90% | ~2,700 |
| PermanentCrop | ~92% | ~93% | ~92% | ~2,700 |
| Residential | ~95% | ~94% | ~94% | ~2,700 |
| River | ~93% | ~92% | ~92% | ~2,700 |
| SeaLake | ~96% | ~97% | ~96% | ~2,700 |
Forest and SeaLake are the easiest classes (clear visual signatures). HerbaceousVegetation and Pasture are the most challenging (visually similar).
- Helber et al. (2019). EuroSAT: A Novel Dataset and Deep Learning Benchmark for Land Use and Land Cover Classification. IEEE J-STARS.
- Tan & Le (2019). EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. ICML.
- Selvaraju et al. (2017). Grad-CAM: Visual Explanations from Deep Networks via Gradient-based Localization. ICCV.
- He et al. (2016). Deep Residual Learning for Image Recognition. CVPR.
- Howard et al. (2019). Searching for MobileNetV3. ICCV.
Happy classifying! π

