Find European road scenarios in driving data, and measure how well you found them.
A driving function trained mostly outside Europe meets things here it has rarely seen: cobblestone streets, zebra crossings without signals, hatched no-drive areas, roadworks barriers, speed bumps, electronic signs, tractors. Before you can fix how it behaves there, you have to find those frames in the fleet data. This repo does that on A2D2 (Audi, recorded in southern Germany) in two ways, and scores one against the other:
- Label scanner (C++). Reads every semantic label PNG, counts pixels per class with a 16M-entry colour lookup table and worker threads, and turns class shares into scenario tags. This is the ground truth.
- Text queries (CLIP, zero-shot). One plain-English query per scenario ("a photo of a street paved with cobblestones") ranks the camera frames with no labels at all, the way you would search unlabelled fleet data. Average precision against the label tags says how far you can trust that search.
All numbers are in docs/RESULTS.md, written by scripts/run_all.py from results/results.json.
Measured on 12,497 labelled front-camera frames from 18 A2D2 drives (full tables in docs/RESULTS.md):
| What | Result |
|---|---|
| C++ scanner vs numpy reference, 300 label PNGs (1920x1208) | 35.8x faster with 32 threads (3.4x on one thread), per-class counts identical |
| Full scan of all 12,497 labels | 14.9 s |
| Scenario tags found | cobblestone 636 frames, hatched no-drive areas 440, electronic signs 206, speed bumps / slow zones 269, zebra crossings 34, ... |
| Text-query mAP over 8 scenarios, CLIP ViT-L/14 | 0.189 vs 0.076 for random ranking (ViT-B/32: 0.169) |
| Best query: cobblestone streets (ViT-B/32) | AP 0.565, 99 of the top 100 frames correct (random: 5%) |
| Electronic signs (ViT-L/14) | AP 0.162, 15 of the top 20 correct (random: 1.6%) |
What this says: text queries are a useful first filter for scenarios that fill the frame (paving, pedestrians and cyclists, electronic signs). They are at random level for speed bumps and tractors, and weak for zebra crossings (AP 0.030 at best, 34 positives). For small road markings and rare objects you still need labels or a trained detector. The larger model helps on signs and markings but not everywhere (cobblestone AP drops from 0.565 to 0.420).
CLIP ran on CPU for this table because the shared GPU had no free memory; the device is recorded per model.
# 1. data: stream the 52 GB A2D2 3D-box subset and keep only front-camera labels, images (480 px) and box JSONs
curl -s https://aev-autonomous-driving-dataset.s3.eu-central-1.amazonaws.com/camera_lidar_semantic_bboxes.tar \
| python scripts/stream_extract_a2d2.py data/a2d2
# 2. build the scanner and the Python package
cmake -S cpp -B build -DCMAKE_BUILD_TYPE=Release && cmake --build build
pip install -e ".[dev]" # plus a torch build for your machine
ruff check esm tests scripts && pytest -q
# 3. everything else
python scripts/run_all.py --data data/a2d2/camera_lidar_semantic_bboxesOr with Docker: docker build -t esm . && docker run -v $PWD/data/a2d2:/data esm.
| Path | What |
|---|---|
cpp/scan_labels.cpp |
C++17 label scanner (stb_image, lookup table, std::thread pool) |
esm/scan_py.py |
numpy reference; the test suite requires identical counts |
esm/tags.py |
scenario tags: class groups + minimum pixel share, with the reason each matters |
esm/clip_miner.py |
CLIP image and text encoding, the query per tag |
esm/metrics.py |
average precision, precision@k |
scripts/run_all.py |
pairing, scan, tags, benchmark, CLIP evaluation, report |
results/scenario_index.csv |
one row per frame: drive, frame id, tags |
- Tags come from 2D labels and fixed thresholds; they say a scenario is visible, not how the car reacted.
- A2D2 covers southern Germany only. No class exists for yellow temporary lane markings, trams or roundabouts.
- Consecutive frames look alike, so precision@k overstates how many different scenes a query finds.
- The first roadworks tag also counted "Traffic guide obj." and matched 70% of frames, because A2D2 puts roadside
reflector posts in that class. It now uses barriers only. The change is noted in
esm/tags.py.
Code: MIT. A2D2 is licensed CC BY-ND 4.0 by Audi AG; this repo contains no A2D2 images or labels, only frame ids and derived statistics. Author: Pavan Yadav Annappa.