LAYA / VISION
INDEPENDENTJEVAL.DEV ↗
API STATUS
●Limitations / 03

Know the failure surface.

A useful small model with measurable blind spots. Treat uncertainty as operational data, not permission to automate everything.

RECALL

24.3% false-negative rate

It missed 1,093 of 4,500 present-object questions. Small or difficult objects remain a central weakness.

HALLUCINATION

Adversarial plausibility

Absent but contextually plausible objects lowered accuracy to 78.27%, versus 83.67% on random negatives.

CONFIDENCE

29 confident errors

Errors still occurred above 90% confidence. A score is evidence, never a guarantee.

●Not the right tool

Boundaries are part of the contract.

Do not use the base checkpoint as an arbitrary OCR extractor, fine-grained localizer, open-ended visual reasoner, or unattended safety-critical decision maker. Domain fine-tuning, independent calibration, threshold selection, and a review path are requirements for high-impact use.

The earlier 11-question synthetic run was only a smoke test. It is intentionally excluded from the evidence presented here.