Parsewave visual verifier study

Can specialist visual predicates replace VLM judges?

A controlled comparison of Gemini 2.5 Pro against objective visual predicate stacks for image and video rubric checks.

Protocol

Four splits with fixed labels and strict predicates

Aggregate Results

Accuracy and false-positive behavior

Accuracy by split

Reverse split false positives

Category Results

Each stack and its measured outcome

Real-World Stress Test

What the model stacks actually do

Full Case Audit

All images, prompts, stacks, and decisions

Interpretation

Where this supports replacement, and where it does not

The specialist stacks are compelling for objective rubric criteria: object presence, count, geometry, color, chart structure, pose structure, and frame-timed video checks. In these conditions they match or exceed Gemini while producing explainable intermediate signals.

This does not remove the need for VLM judges on subjective or holistic dimensions such as visual taste, brand fit, ambiguous intent, or open-ended semantic quality. The practical verifier design is hybrid: deterministic visual predicates for objective checks, and VLM judging only where the rubric is genuinely semantic.