Skip to content

Latest commit

 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 

Repository files navigation

CLIP Few-Shot Adaptation - Oxford Flowers 102

Deep Learning 2025 · Project Assignment

Few-shot adaptation of CLIP (ViT-B/16) on the Oxford Flowers 102 dataset using parameter-efficient fine-tuning methods. The goal: improve accuracy on Base classes while preserving zero-shot performance on Novel classes (base-to-novel generalization).


Task

Given only 10 labeled samples per class (k=10 shots) for 51 Base categories, adapt a pre-trained CLIP model so that it:

  • improves top-1 accuracy on Base classes,
  • does not degrade accuracy on the 51 Novel categories (unseen at train time).

Metric: top-1 accuracy on Base, Novel, and their Harmonic Mean.


Methods

Three approaches were implemented and compared:

1. PEFT LayerNorm Tuning

Only the LayerNorm layers of CLIP are trained — everything else stays frozen. A regularization term pulls trained weights toward the original ones to reduce forgetting. Hyperparameters (learning rate, weight decay, label smoothing, unfreeze depth) tuned with Optuna.

2. CoCoOp

Image-conditioned prompt learning: a lightweight Meta-Net (MLP) generates dynamic context tokens per image, shifting prompts from class-specific to instance-specific to reduce overfitting. The CLIP backbone remains frozen.

3. LN-Tuning + CoCoOp (Hybrid)

Combines both: LayerNorm layers are fine-tuned while CoCoOp generates per-image prompts on top. Best overall generalization.


Results

Method Base ↑ Novel ↑ Harmonic Mean ↑
CLIP Zero-Shot (baseline) 71.33 78.24 74.62
PEFT LayerNorm 85.56 73.20 78.90
CoCoOp 93.81 69.56 79.89
LN-Tuning + CoCoOp 93.13 72.91 81.78

All results on Oxford Flowers 102 with CLIP ViT-B/16.

FR_1301(2)

Setup

# 1. Install dependencies
pip install torch==2.6.0 torchvision==0.21.0 torchaudio==2.6.0
pip install openai-clip==1.0.1
pip install pandas seaborn matplotlib tqdm optuna

# 2. Open the notebook in Google Colab
# 3. Run all cells sequentially (restart runtime after the first cell)

The dataset (Oxford Flowers 102) is downloaded automatically via torchvision.


Repository Structure

├── clip_fewshot_adaptation.ipynb   # Full project: code + report in one self-contained notebook
└── README.md
└── Deep_Learning_Project_Assignment_2025.pdf

Key Design Choices

  • Mixed precision: vision encoder in FP16, LayerNorm and text encoder in FP32 for stability.
  • Optuna for automated hyperparameter search across all three methods.
  • No extra data: strictly 10 shots per base class; novel classes never seen during training.
  • Backbone: CLIP ViT-B/16 (OpenAI).

References

  • Radford et al., Learning Transferable Visual Models from Natural Language Supervision (CLIP), ICML 2021
  • Zhou et al., Conditional Prompt Learning for Vision-Language Models (CoCoOp), CVPR 2022
  • Zhou et al., Learning to Prompt for Vision-Language Models (CoOp), IJCV 2022
  • Hu et al., LoRA: Low-Rank Adaptation of Large Language Models, ICLR 2022
  • Nilsback & Zisserman, Automated Flower Classification over a Large Number of Classes (Oxford Flowers 102), 2008

About

Few-shot adaptation of CLIP on Oxford Flowers 102 via LayerNorm tuning, CoCoOp, and a hybrid approach - with Optuna HPO and base-to-novel generalization.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages