Protocol
Four splits with fixed labels and strict predicates
Aggregate Results
Accuracy and false-positive behavior
Accuracy by split
Reverse split false positives
Category Results
Each stack and its measured outcome
Real-World Stress Test
What the model stacks actually do
Full Case Audit
All images, prompts, stacks, and decisions
Interpretation
Where this supports replacement, and where it does not
The specialist stacks are compelling for objective rubric criteria: object presence, count, geometry, color, chart structure, pose structure, and frame-timed video checks. In these conditions they match or exceed Gemini while producing explainable intermediate signals.
This does not remove the need for VLM judges on subjective or holistic dimensions such as visual taste, brand fit, ambiguous intent, or open-ended semantic quality. The practical verifier design is hybrid: deterministic visual predicates for objective checks, and VLM judging only where the rubric is genuinely semantic.