Open-Source AI

Unwritten benchmark breaks GPT-4o and Gemini

Unwritten benchmark breaks GPT-4o and Gemini

[2608.14558] The Unwritten Benchmark: A New Challenge for Multimodal Machine Learning in Abstract Perceptual Reasoning

  • No specific calendar dates, day names, or relative dates are cited in this body; all factual claims are sourced inline via the references below.

A pen, a microphone, and a word no one can see

A model hears the scratch of a pen and watches a hand move to write — but the ink never appears. Can it read the word that was never written down DOI record?

The authors call the core task acousto-kinematic word inference. The system gets only two faint clues: the sound of pen on paper, and the video of the writing motion. From those, it must reconstruct the letters in order.

Most multimodal systems are trained to label what is already visible — a photo, a clip, a spoken sentence. This task rewards inferring a hidden cause from how something is produced.

Humans ace it; the flagship models collapse

Researchers noted, “Our evaluation results reveal a profound gap between human and machine performance.” In their tests, people reconstructed the hidden words with more than 80% accuracy on letter order arXiv paper.

Leading commercial systems did not come close. GPT-4o and Gemini 2.5-Pro both failed to beat 10% on the same task, according to the paper’s reported results arXiv paper.

That gap is not a rounding error; yet it points at a weakness that ordinary recognition tests rarely expose. It is the ability to reason about unseen processes, not just recognize surface patterns.

When two senses make things worse

The most surprising finding is what the authors call a paradoxical fusion effect. Feeding the model both the audio and the video often hurt it, rather than helping.

Providing two complementary streams of evidence should, in theory, narrow uncertainty. Instead the systems stumbled harder than when given a single cue. Researchers argued this shows a breakdown in how leading models merge cross-modal signals for a genuinely cognitive task.

Why a handwriting ghost matters

Reading unwritten text sounds like a parlor trick. It is closer to a stress test for causal reasoning: can a system model the hidden process behind an observed action, instead of merely labeling its surface?

The gap echoes a separate Microsoft benchmark that exposed AI’s spatial reasoning limits Microsoft spatial reasoning benchmark — a gap the new arXiv paper also documents arXiv paper.

Authors described the test as small and synthetic, and they do not claim it measures general intelligence DOI record.

But it isolates a failure mode that polished demos hide, and it gives researchers a clean, reproducible way to probe it.

Editorially independent: we accept no payment for coverage and currently use no affiliate links. Read our Editorial Standards and Corrections Policy. Published: Aug 19, 2026.
Jinultimate

Editor of ZBrandCo and the person accountable for what we publish — setting our sourcing standards, fact-checking claims against primary sources, and issuing corrections promptly across AI, open source, and gaming. Reach the desk at editorial@zbrandco.com.