HarnessEval-W: Agentifying the Evaluation of Visual Worlds
A benchmark should deliver more than a scalar score: what makes an evaluation trustworthy is the reasoning that justifies the score. This is especially critical for world models, where judging a rollout requires understanding whether physics, causality, and world state evolve correctly. Humans spot such violations naturally, yet no existing benchmark automates this capability: metrics are computed brute-force, leaving no reasoning chain that can be examined or verified. Authors: Weiliang Chen, Haowen Sun, Jun Gao.
Why it matters
Read this for the paper's specific claim in Artificial Intelligence / Machine Learning: A benchmark should deliver more than a scalar score: what makes an evaluation trustworthy is the reasoning that justifies the score.
