LAYA / VISION
INDEPENDENTJEVAL.DEV ↗
API STATUS
●Capabilities // bounded visual judgments

Constrain the question.
Strengthen the answer.

The model is designed for decisions with a declared output space—not free-form narration. POPE directly verifies object-presence yes/no behavior; other forms are product-fit guidance, not locally benchmarked claims.

01

Yes / no decisions

The verified use case: object-presence questions with a probability over yes and no.

02

Compact choices

The API can express a small declared label set. Validate this mode on your own held-out data before production use.

03

Ordinal scores

Bounded rubrics fit the typed contract, but this evaluation did not rerun the upstream rubric suite.

04

Question batches

Reuse an image prefix across many questions. Eighteen grouped decisions cost only 1.37× one question in this run.

05

Selective review

Use probabilities to route uncertain or high-impact cases to a person or stronger model; tune thresholds on domain data.

●Verified signal

What POPE actually proves.

It measures whether a model says an object is present when it is present—and resists saying yes when a plausible object is absent. It does not test arbitrary OCR extraction, bounding boxes, or open-ended image reasoning.

RANDOM NEGATIVES

83.67%

Highest local accuracy of the three POPE subsets.

POPULAR NEGATIVES

82.20%

Common absent objects create a modest increase in false positives.

ADVERSARIAL

78.27%

Contextually plausible absent objects expose the hallucination surface.