Training-data overlap is the benchmark's central validity finding: 447 of the 670 BBBC038 images evaluated here appear in the DSB2018 training split that StarDist distributes with its 0.1.0 release. StarDist documents 2D_versatile_fluo as trained on an unenumerated subset of DSB2018, so the exact overlap with the released weights cannot be confirmed from public materials, as is also true for Cellpose nuclei. A second methodological error initially erased non-fluorescence performance: routing dark nuclei on light backgrounds through Cellpose without inversion produced near-zero masks, whereas corrected inversion changed tracked AJI from approximately 0.000 to 0.836 on H&E-style and 0.877 on grayscale-brightfield images. After correction, Cellpose leads the H&E-style stratum, reversing the broken-run conclusion.
This benchmark compares Cellpose nuclei and StarDist pretrained 2D nuclei models on BBBC038 / DSB2018 stage1_train, stratified by visual modality cluster. BBBC038 is cited as recommended by the collection through the Data Science Bowl report [1], alongside the canonical Broad Bioimage Benchmark Collection publication [2].
The companion EMG Clinical Benchmark, maintained as a separate repository, shares the same larger lesson: the most informative result was a methodological problem in the benchmark itself. Here it is training-data overlap in the nuclei evaluation; there it is a gait-speed confound that makes nominal stroke/control accuracy an unsafe measure of pathology recognition.
Use 64-bit CPython 3.11. The benchmark has two environments because Cellpose uses PyTorch while the pinned StarDist release uses TensorFlow 2.15. The StarDist environment also contains the scoring stack. Dependencies are fully version-pinned in the three requirements files.
python3.11 -m venv .venv-cellpose
.venv-cellpose/bin/python -m pip install --upgrade pip
.venv-cellpose/bin/python -m pip install -r requirements-cellpose.txt
python3.11 -m venv .venv-stardist
.venv-stardist/bin/python -m pip install --upgrade pip
.venv-stardist/bin/python -m pip install -r requirements-stardist.txt -r requirements-score.txtDownload BBBC038 directly; no authentication is required. Clone the metric reference before scoring and pin it to the reviewed commit.
mkdir -p data reference
curl -L https://data.broadinstitute.org/bbbc/BBBC038/stage1_train.zip \
-o data/stage1_train.zip
unzip -q data/stage1_train.zip -d data
git clone https://github.com/STOmics/cs-benchmark reference/cs-benchmark
git -C reference/cs-benchmark checkout 4f7e50fbfbeeb6b9169c3040419b7da63d6a0004Prepare ground truth and freeze the modality routing before inference. Cellpose and StarDist download their named pretrained checkpoints from their official registries on first use. Both inference scripts assert that every image has a modality assignment.
.venv-stardist/bin/python prepare_gt.py
.venv-stardist/bin/python cluster_modality.py
.venv-cellpose/bin/python run_cellpose.py
.venv-stardist/bin/python run_stardist.py
.venv-stardist/bin/python score.py --figure figures/full_examples.pngThe fluorescence training-overlap analysis re-scores the cached masks; it performs no new inference. It additionally requires the DSB2018 bundle linked by the StarDist project.
curl -L https://github.com/stardist/stardist/releases/download/0.1.0/dsb2018.zip \
-o reference/stardist_dsb2018.zip
unzip -q reference/stardist_dsb2018.zip -d reference/stardist_dsb2018
.venv-stardist/bin/python split_fluorescence_overlap.pyThe primary outputs are results/per_image.csv, results/by_modality.csv, results/summary.json, and the PNG files under figures/. Raw data, model masks, reference repositories, and virtual environments are intentionally excluded from Git.
Full 670-image inference and scoring are complete after a routing fix. Data came from BBBC038 by direct HTTP download.
- Source page: https://bbbc.broadinstitute.org/BBBC038
- Download: https://data.broadinstitute.org/bbbc/BBBC038/stage1_train.zip
- Local extraction layout:
data/{image_id}/images/{image_id}.pnganddata/{image_id}/masks/*.png - Full prediction cache:
masks/cellpose/*.npy,masks/stardist/*.npy - Full metric outputs:
results/per_image.csv,results/by_modality.csv,results/summary.json - Fluorescence StarDist-training split:
results/fluorescence_overlap_split.csv - Training overlap memo:
results/training_overlap.md
Fluorescence predictions were not recomputed after the routing fix. Cached masks for he_style and grayscale_brightfield were deleted and regenerated only for those 124 images.
Before writing evaluation code, this project cloned and reviewed STOmics/cs-benchmark:
- Source: https://github.com/STOmics/cs-benchmark
- Local path:
reference/cs-benchmark - Reviewed commit:
4f7e50fbfbeeb6b9169c3040419b7da63d6a0004
The local evaluator adapts the contiguous relabeling, AJI, PQ, and binary Dice implementations from src/methods/hover_net/metrics/stats_utils.py into reference_metrics.py. StarDist's matching / matching_dataset functions are used for thresholded F1 and AP-style instance matching.
Modality assignments are computed from the images, not taken from BBBC038 metadata. cluster_modality.py converts every image to RGB in [0, 1] and extracts nine global features: mean red, green, and blue intensity; mean and 90th-percentile HSV saturation; 20th-percentile grayscale brightness; 95th-minus-5th-percentile grayscale contrast; and the fractions of pixels below 0.15 and above 0.85 brightness. Features are standardized over all 670 images. K-means uses seed 17, with 50 initializations for the k-selection sweep and 100 initializations for the final fit; the number of clusters is selected between k=3 and k=4 by the higher silhouette score.
Cluster names are assigned deterministically from cluster-average features. A red-and-blue excess over green identifies he_style; otherwise background brightness below 0.18 identifies fluorescence_dark; otherwise mean saturation below 0.08 identifies grayscale_brightfield. The resulting 106/546/18-image contact sheet in figures/modality_groups.png was manually reviewed and accepted before inference. These are appearance clusters, not verified acquisition-modality labels.
- Aggregated Jaccard Index (AJI): each annotated instance is paired with its best-overlapping prediction. Intersections and unions are summed across pairs, and unmatched true or predicted instances add their area to the denominator. AJI therefore penalizes missed objects, false objects, merges, splits, and boundary error in one image-level score.
- Panoptic Quality (PQ): instances with intersection-over-union greater than 0.5 are matched. Detection quality is
TP / (TP + 0.5 FP + 0.5 FN); segmentation quality is the mean IoU of matched pairs; PQ is their product. - mAP: for consistency with the original benchmark code, this label denotes the mean of StarDist matching
accuracy = TP / (TP + FP + FN)over IoU thresholds 0.50, 0.55, ..., 0.90. It is an AP-style threshold average, not confidence-ranked average precision. - Count error: predicted instance count minus annotated instance count. Negative values are undercounts and positive values are overcounts.
- Confidence intervals: table intervals are percentile 95% bootstrap intervals over images, using 2,000 resamples of the image-level metric mean.
- Cellpose: raw images for
fluorescence_dark; inverted grayscale images forhe_styleandgrayscale_brightfield. - StarDist:
2D_versatile_fluoforfluorescence_dark;2D_versatile_heforhe_style. - StarDist brightfield assignment:
2D_versatile_hewas selected forgrayscale_brightfieldafter testing both candidate models on the expanded smoke set. It scored higher than2D_versatile_fluoby mAP, AJI, and PQ. Evidence is saved inresults/stardist_brightfield_choice.csvandresults/stardist_brightfield_choice.json.
Values are mean scores with 95% bootstrap confidence intervals over images. mAP is the mean of StarDist matching accuracy over IoU thresholds 0.50 to 0.90.
| Modality | Model | n | AJI | PQ | mAP | Count error |
|---|---|---|---|---|---|---|
he_style |
Cellpose | 106 | 0.699 [0.670, 0.725] | 0.683 [0.659, 0.705] | 0.509 [0.482, 0.536] | -1.2 [-2.0, -0.3] |
he_style |
StarDist | 106 | 0.480 [0.458, 0.500] | 0.489 [0.468, 0.510] | 0.297 [0.279, 0.315] | -19.0 [-21.9, -16.2] |
fluorescence_dark |
Cellpose | 546 | 0.735 [0.717, 0.752] | 0.733 [0.716, 0.749] | 0.626 [0.606, 0.645] | -4.2 [-5.1, -3.4] |
fluorescence_dark |
StarDist | 546 | 0.814 [0.805, 0.823] | 0.788 [0.779, 0.798] | 0.688 [0.672, 0.703] | -2.2 [-2.8, -1.5] |
grayscale_brightfield |
Cellpose | 18 | 0.740 [0.696, 0.779] | 0.686 [0.640, 0.727] | 0.519 [0.460, 0.575] | -13.0 [-18.9, -7.7] |
grayscale_brightfield |
StarDist | 18 | 0.565 [0.507, 0.624] | 0.563 [0.519, 0.605] | 0.401 [0.353, 0.444] | -28.2 [-36.6, -20.9] |
grayscale_brightfield has n=18 and is underpowered; interpret those confidence intervals and model rankings cautiously.
After deleting and regenerating the non-fluorescence masks, the tracked smoke anchors reproduced their corrected Cellpose AJI values in the fresh full score:
| Image prefix | Modality | Cellpose AJI |
|---|---|---|
442c4eb0 |
he_style |
0.835699 |
8d05fb18 |
grayscale_brightfield |
0.877171 |
Training overlap is the largest validity caveat. StarDist 2D_versatile_fluo has confirmed partial training overlap with BBBC038: 447 of its DSB2018 release-bundle training images are present in this benchmark, all in fluorescence_dark. Cellpose nuclei also has DSB2018 involvement in the checked paper/repo sources, but the exact released-weight image-ID overlap could not be resolved from public materials. Treat this as an operational BBBC038 comparison, not a clean external generalization benchmark.
On fluorescence_dark, both tools perform well and StarDist is higher across AJI, PQ, and mAP. This is also the stratum with direct StarDist DSB2018 training overlap, so the aggregate StarDist advantage should not be read alone as evidence of better out-of-domain generalization.
Splitting fluorescence_dark by StarDist DSB2018 training membership points more to a robustness difference than a simple contamination effect. On the 447 StarDist-seen images, the models are comparatively close on AJI: StarDist 0.825 versus Cellpose 0.766, a gap of 0.058. On the 99 images unseen by StarDist's documented DSB2018 training subset, the gap triples to 0.175: StarDist 0.768 [0.732, 0.800] versus Cellpose 0.593 [0.535, 0.651]. Both models degrade on the unseen subset, indicating those images are harder, but Cellpose degrades disproportionately: Cellpose loses 22.6% of seen-subset AJI while StarDist loses 6.8%. The same pattern holds for PQ and mAP, where StarDist remains higher on unseen images. This does not remove the broader Cellpose training-overlap caveat, but it argues against StarDist's fluorescence advantage being only a train-set contamination artifact.
The corrected H&E-style full results run counter to published H&E comparisons that favor H&E-specialized StarDist weights and show that grayscale conversion plus color inversion can substantially improve initially weak models [6]: after dark-on-light inversion, Cellpose is higher on he_style by AJI, PQ, mAP, and count error. A plausible hypothesis is that failure to invert dark nuclei on light backgrounds can make Cellpose look much worse than it is on this modality. This is a hypothesis, not a conclusion, because the benchmark has training-overlap caveats and uses one modality clustering scheme.
The Cellpose inversion effect is a headline preprocessing result. On the tracked H&E-style smoke image, Cellpose AJI changed from ~0.000 to 0.836 after inversion; on the tracked grayscale-brightfield image, AJI changed from 0.000 to 0.877. Those before/after metrics are saved in results/preprocessing_before_after.csv and visualized in figures/inversion_effect.png.
The full grayscale_brightfield stratum remains underpowered at n=18. Cellpose is higher on AJI, PQ, and mAP after inversion, while both models undercount on average. The confidence intervals are wide enough that this stratum should be treated as directional evidence rather than a stable ranking.
The clearest Cellpose failures are faint, sparse fluorescence fields. On images 84eeec681..., 3d0ca349..., 07761fa3..., d7ec8003..., and 8f6597cd..., Cellpose predicts no instances and therefore has AJI 0.000, while StarDist predicts exactly the annotated count (1, 4, 5, 7, and 9) with AJI from 0.841 to 0.920. In 724b6b70..., Cellpose finds 5 of 18 annotated nuclei (AJI 0.081), whereas StarDist finds all 18 (AJI 0.886). Visual review of figures/disagreement.png suggests that Cellpose is losing dim, small objects near its detection threshold; this is an image-level hypothesis, not a demonstrated causal mechanism.
StarDist's largest H&E-style failures have the opposite signature: systematic undercounting of large, pale, irregular nuclei. It predicts 3 of 8 instances on 589f86de... (AJI 0.083), 6 of 15 on 2255d5ab... (AJI 0.231), 5 of 11 on a7f767ca... (AJI 0.253), and 3 of 7 on 01d44a26... (AJI 0.257). Cellpose reaches AJI 0.580, 0.581, 0.588, and 0.798 on those same images. The overlays in figures/worst_cases.png are consistent with missed low-contrast nuclei and occasional merging, but expert review is still needed to distinguish model error from ambiguous boundaries.
Image 7b38c917... is a shared catastrophic outlier that should not be interpreted as ordinary model failure. Its ground truth contains one connected instance, while Cellpose and StarDist predict 92 and 98 instances; both score AJI 0.000. The image visibly contains many bright nucleus-like objects, so this case appears to be an annotation-granularity mismatch: the reference treats the field as one object while both models perform nucleus-level instance segmentation. It illustrates why aggregate scores need image review and why BBBC038's heterogeneous source annotations are a limitation.
This exploratory section is pending manual review. The named cases and scores come from results/per_image.csv; the causal interpretations come from visual inspection of figures/worst_cases.png and figures/disagreement.png and should not be treated as adjudicated error labels.
- Caicedo JC et al. Nucleus segmentation across imaging experiments: the 2018 Data Science Bowl. Nature Methods. 2019;16:1247-1253. doi:10.1038/s41592-019-0612-7. Dataset: BBBC038, version 1, CC0.
- Ljosa V et al. Annotated high-throughput microscopy image sets for validation. Nature Methods. 2012;9:637. doi:10.1038/nmeth.2083.
- Stringer C et al. Cellpose: a generalist algorithm for cellular segmentation. Nature Methods. 2021;18:100-106. doi:10.1038/s41592-020-01018-x. This benchmark uses the
nucleicheckpoint distributed by Cellpose. - Schmidt U et al. Cell Detection with Star-Convex Polygons. In: MICCAI 2018, LNCS 11071. 2018:265-273. doi:10.1007/978-3-030-00934-2_30.
- Weigert M et al. Star-convex Polyhedra for 3D Object Detection and Segmentation in Microscopy. In: WACV 2020. 2020:3655-3662. doi:10.1109/WACV45572.2020.9093435. This benchmark uses
2D_versatile_fluoand2D_versatile_hefrom the StarDist pretrained-model registry. - Liu X et al. CellBinDB: a large-scale multimodal annotated dataset for cell segmentation with benchmarking of universal models. GigaScience. 2025;14:giaf069. doi:10.1093/gigascience/giaf069.







