LAYA / VISION
INDEPENDENTJEVAL.DEV ↗
API STATUS
●Evaluation / 01

Evidence, not vibes.

The entire public POPE suite, pinned to exact repository, checkpoint, and dataset revisions. Every number below comes from the reproducible local run.

●Method

Complete benchmark coverage.

All 9,000 questions across 500 real COCO images: 3,000 random, 3,000 popular, and 3,000 adversarial rows. BF16 inference, batch size 32, identity option order, and grouped 18-question calls on an RTX 3090.

SplitnAccuracyPublishedNLLBrierECE
Random3,00083.67%83.57%0.4000.1230.0390
Popular3,00082.20%81.87%0.4240.1330.0380
Adversarial3,00078.27%77.70%0.4900.1590.0393
All9,00081.38%—.438.138.0359

The local run is within 0.10–0.57 percentage points of published accuracy across all three splits. The repository’s other 59,427-question suite was not rerun and is not presented as local evidence.

●Calibration

Confidence has shape.

Mean confidence was 82.16% against 81.38% accuracy. High confidence was strong overall, but 29 errors remained at or above 90% confidence.

●Accuracy ■ vs mean confidence —
.50–.60
53.18%
.60–.70
55.56%
.70–.80
68.33%
.80–.90
86.10%
.90–1.0
96.92%
●Runtime

Small enough to be useful.

112.9 ms median and 122.5 ms p90 for a single question. Grouping 18 questions took 154.7 ms median and delivered 112.9 questions/second, with 493.4 MB peak allocated VRAM.

LATENCY

112.9 ms

Median, single question, preprocessing included.

THROUGHPUT

112.9 q/s

Effective grouped throughput with shared image prefix.

VERIFICATION

5 / 5 + 47 / 47

Evidence hashes and published-claim checks passed.