Deep Learning 2025 · Project Assignment
Few-shot adaptation of CLIP (ViT-B/16) on the Oxford Flowers 102 dataset using parameter-efficient fine-tuning methods. The goal: improve accuracy on Base classes while preserving zero-shot performance on Novel classes (base-to-novel generalization).
Given only 10 labeled samples per class (k=10 shots) for 51 Base categories, adapt a pre-trained CLIP model so that it:
- improves top-1 accuracy on Base classes,
- does not degrade accuracy on the 51 Novel categories (unseen at train time).
Metric: top-1 accuracy on Base, Novel, and their Harmonic Mean.
Three approaches were implemented and compared:
Only the LayerNorm layers of CLIP are trained — everything else stays frozen. A regularization term pulls trained weights toward the original ones to reduce forgetting. Hyperparameters (learning rate, weight decay, label smoothing, unfreeze depth) tuned with Optuna.
Image-conditioned prompt learning: a lightweight Meta-Net (MLP) generates dynamic context tokens per image, shifting prompts from class-specific to instance-specific to reduce overfitting. The CLIP backbone remains frozen.
Combines both: LayerNorm layers are fine-tuned while CoCoOp generates per-image prompts on top. Best overall generalization.
| Method | Base ↑ | Novel ↑ | Harmonic Mean ↑ |
|---|---|---|---|
| CLIP Zero-Shot (baseline) | 71.33 | 78.24 | 74.62 |
| PEFT LayerNorm | 85.56 | 73.20 | 78.90 |
| CoCoOp | 93.81 | 69.56 | 79.89 |
| LN-Tuning + CoCoOp | 93.13 | 72.91 | 81.78 |
All results on Oxford Flowers 102 with CLIP ViT-B/16.
# 1. Install dependencies
pip install torch==2.6.0 torchvision==0.21.0 torchaudio==2.6.0
pip install openai-clip==1.0.1
pip install pandas seaborn matplotlib tqdm optuna
# 2. Open the notebook in Google Colab
# 3. Run all cells sequentially (restart runtime after the first cell)
The dataset (Oxford Flowers 102) is downloaded automatically via torchvision.
├── clip_fewshot_adaptation.ipynb # Full project: code + report in one self-contained notebook
└── README.md
└── Deep_Learning_Project_Assignment_2025.pdf
- Mixed precision: vision encoder in FP16, LayerNorm and text encoder in FP32 for stability.
- Optuna for automated hyperparameter search across all three methods.
- No extra data: strictly 10 shots per base class; novel classes never seen during training.
- Backbone:
CLIP ViT-B/16(OpenAI).
- Radford et al., Learning Transferable Visual Models from Natural Language Supervision (CLIP), ICML 2021
- Zhou et al., Conditional Prompt Learning for Vision-Language Models (CoCoOp), CVPR 2022
- Zhou et al., Learning to Prompt for Vision-Language Models (CoOp), IJCV 2022
- Hu et al., LoRA: Low-Rank Adaptation of Large Language Models, ICLR 2022
- Nilsback & Zisserman, Automated Flower Classification over a Large Number of Classes (Oxford Flowers 102), 2008