A deep learning web app that classifies chest X-ray images as Normal or Pneumonia, using transfer learning with MobileNetV2 and an interactive Streamlit interface.
⚠️ Disclaimer: This is a student/portfolio project built for educational purposes. It is not a certified medical device and should never be used for real diagnosis — always consult a qualified doctor.
Chest X-rays are uploaded through a simple web interface, and a CNN (built on top of MobileNetV2, pretrained on ImageNet) predicts whether the image shows signs of pneumonia, along with a confidence score.
The project has two parts:
train_model.py— trains the classifier using transfer learning and saves it aspneumonia_detector.h5app.py— a Streamlit app that loads the trained model and serves real-time predictions
- 🧠 Transfer learning with MobileNetV2 (frozen base + custom classification head)
- 🔄 Data augmentation (rotation, zoom, horizontal flip) during training
- 📊 Auto-generated training curves, confusion matrix, and classification report
- 🌐 Streamlit app with image upload, live prediction, and confidence display
- 💾 Cached model loading (
@st.cache_resource) for fast repeated predictions
| Category | Tools |
|---|---|
| Language | Python |
| Deep Learning | TensorFlow / Keras |
| Base Model | MobileNetV2 (ImageNet weights, frozen) |
| Web App | Streamlit |
| Evaluation | scikit-learn (classification report, confusion matrix) |
| Visualization | Matplotlib, Seaborn |
Chest X-Ray Images (Pneumonia) — Kaggle 🔗 https://www.kaggle.com/datasets/paultimothymooney/chest-xray-pneumonia
Expected folder structure:
chest_xray/
├── train/
│ ├── NORMAL/
│ └── PNEUMONIA/
├── val/
│ ├── NORMAL/
│ └── PNEUMONIA/
└── test/
├── NORMAL/
└── PNEUMONIA/
Class indices: {'NORMAL': 0, 'PNEUMONIA': 1}
| Split | Normal | Pneumonia | Total |
|---|---|---|---|
| Train | 1,341 | 3,875 | 5,216 |
| Validation | 8 | 8 | 16 |
| Test | 234 | 390 | 624 |
⚠️ Note: the validation set is very small (16 images), which is a known quirk of this particular dataset split. This is likely why validation accuracy/loss swing around a lot between epochs — it's not a bug in the training script, just a side effect of evaluating on so few samples each epoch.
pneumonia-detection/
├── train_model.py # Trains the model, saves pneumonia_detector.h5
├── app.py # Streamlit app for inference
├── pneumonia_detector.h5 # Trained model weights
├── training_curves.png # Generated after training
├── confusion_matrix.png # Generated after training
├── classification_report.txt
├── requirements.txt
└── README.md
- Load & explore data from
chest_xray/train,val, andtestfolders. - Preprocess & augment — training images are rescaled (
1/255) and augmented with rotation (±20°), zoom (0.2), and horizontal flip. Validation/test images are only rescaled. - Build the model —
MobileNetV2(include_top=False, frozen) →GlobalAveragePooling2D→Dense(128, relu)→Dropout(0.3)→Dense(1, sigmoid). - Train for 10 epochs with the Adam optimizer and binary cross-entropy loss.
- Evaluate on the test set — generates accuracy/loss curves, a confusion matrix, and a full classification report.
- Save the trained model as
pneumonia_detector.h5.
- User uploads a chest X-ray (
jpg/jpeg/png). - Image is resized to
224x224and normalized (/255) to match the training pipeline. - The cached model predicts a probability;
> 0.5→ Pneumonia, otherwise → Normal. - Result is shown with a confidence percentage and progress bar.
# 1. Clone the repository
git clone https://github.com/<your-username>/pneumonia-detection.git
cd pneumonia-detection
# 2. Create a virtual environment (optional but recommended)
python -m venv venv
source venv/bin/activate # On Windows: venv\Scripts\activate
# 3. Install dependencies
pip install -r requirements.txt# Download the dataset via Kaggle API first, then:
python train_model.pystreamlit run app.pytensorflow
streamlit
pillow
numpy
matplotlib
seaborn
scikit-learn
Trained for 10 epochs using a frozen MobileNetV2 base (only 164K of 2.42M params trainable). Final test set performance:
| Metric | Score |
|---|---|
| Test Accuracy | 89% |
| Precision (Normal) | 0.90 |
| Recall (Normal) | 0.79 |
| F1-Score (Normal) | 0.84 |
| Precision (Pneumonia) | 0.88 |
| Recall (Pneumonia) | 0.95 |
| F1-Score (Pneumonia) | 0.91 |
Confusion Matrix (624 test images):
| Predicted: Normal | Predicted: Pneumonia | |
|---|---|---|
| Actual: Normal | 185 | 49 |
| Actual: Pneumonia | 20 | 370 |
The model leans toward catching pneumonia cases (95% recall on Pneumonia) at the cost of some false positives on Normal images (49 normal X-rays misclassified as pneumonia). For a screening tool, that's a reasonable trade-off — missing a pneumonia case is generally worse than a false alarm — but it's worth knowing before drawing conclusions from the accuracy number alone.
- Add Grad-CAM visualizations to show which regions drove the prediction
- Fine-tune deeper MobileNetV2 layers instead of keeping the base fully frozen
- Get a larger validation split (currently only 16 images) for more stable training metrics
- Deploy on Streamlit Community Cloud / Hugging Face Spaces
- Add multi-class classification (bacterial vs. viral pneumonia)
- Dataset: "Chest X-Ray Images (Pneumonia)" by Paul Mooney, Kaggle
- Base architecture: MobileNetV2 paper — Sandler et al., 2018
- Built with TensorFlow and Streamlit
Sherry