112.9 ms
Median, single question, preprocessing included.
The entire public POPE suite, pinned to exact repository, checkpoint, and dataset revisions. Every number below comes from the reproducible local run.
All 9,000 questions across 500 real COCO images: 3,000 random, 3,000 popular, and 3,000 adversarial rows. BF16 inference, batch size 32, identity option order, and grouped 18-question calls on an RTX 3090.
| Split | n | Accuracy | Published | NLL | Brier | ECE |
|---|---|---|---|---|---|---|
| Random | 3,000 | 83.67% | 83.57% | 0.400 | 0.123 | 0.0390 |
| Popular | 3,000 | 82.20% | 81.87% | 0.424 | 0.133 | 0.0380 |
| Adversarial | 3,000 | 78.27% | 77.70% | 0.490 | 0.159 | 0.0393 |
| All | 9,000 | 81.38% | — | .438 | .138 | .0359 |
The local run is within 0.10–0.57 percentage points of published accuracy across all three splits. The repository’s other 59,427-question suite was not rerun and is not presented as local evidence.
Mean confidence was 82.16% against 81.38% accuracy. High confidence was strong overall, but 29 errors remained at or above 90% confidence.
112.9 ms median and 122.5 ms p90 for a single question. Grouping 18 questions took 154.7 ms median and delivered 112.9 questions/second, with 493.4 MB peak allocated VRAM.
Median, single question, preprocessing included.
Effective grouped throughput with shared image prefix.
Evidence hashes and published-claim checks passed.