24.3% false-negative rate
It missed 1,093 of 4,500 present-object questions. Small or difficult objects remain a central weakness.
A useful small model with measurable blind spots. Treat uncertainty as operational data, not permission to automate everything.
It missed 1,093 of 4,500 present-object questions. Small or difficult objects remain a central weakness.
Absent but contextually plausible objects lowered accuracy to 78.27%, versus 83.67% on random negatives.
Errors still occurred above 90% confidence. A score is evidence, never a guarantee.
Do not use the base checkpoint as an arbitrary OCR extractor, fine-grained localizer, open-ended visual reasoner, or unattended safety-critical decision maker. Domain fine-tuning, independent calibration, threshold selection, and a review path are requirements for high-impact use.
The earlier 11-question synthetic run was only a smoke test. It is intentionally excluded from the evidence presented here.