Yes / no decisions
The verified use case: object-presence questions with a probability over yes and no.
The model is designed for decisions with a declared output space—not free-form narration. POPE directly verifies object-presence yes/no behavior; other forms are product-fit guidance, not locally benchmarked claims.
The verified use case: object-presence questions with a probability over yes and no.
The API can express a small declared label set. Validate this mode on your own held-out data before production use.
Bounded rubrics fit the typed contract, but this evaluation did not rerun the upstream rubric suite.
Reuse an image prefix across many questions. Eighteen grouped decisions cost only 1.37× one question in this run.
Use probabilities to route uncertain or high-impact cases to a person or stronger model; tune thresholds on domain data.
It measures whether a model says an object is present when it is present—and resists saying yes when a plausible object is absent. It does not test arbitrary OCR extraction, bounding boxes, or open-ended image reasoning.
Highest local accuracy of the three POPE subsets.
Common absent objects create a modest increase in false positives.
Contextually plausible absent objects expose the hallucination surface.