AI

Microsoft benchmark exposes AI’s spatial reasoning gap

Microsoft benchmark exposes AI’s spatial reasoning gap

MindTopo reveals VLMs' spatial reasoning abilities - Microsoft Research

Microsoft Research has released a benchmark that shows today’s vision-language models can often recognize how objects are connected, enclosed, or knotted in a single image, but they fall apart when they have to preserve that understanding while acting. The test, called MindTopo, measures a form of spatial judgment — topology — that has largely been missing from how multimodal AI is evaluated, according to the Microsoft Research announcement.

Topology is the branch of spatial understanding concerned with relationships that survive bending and stretching: whether two rooms stay connected after a wall is added, whether an animal is inside a fence, or whether a rope is truly knotted rather than merely tangled. MindTopo organizes its tasks around five categories — continuity, separation, order, enclosure, and knots — and tests each at two levels. In reasoning tasks the model inspects rendered scenes and answers a question; in planning tasks it acts inside a simulator where illegal moves, such as passing one rope strand through another, are blocked. The benchmark design and its ground-truth controls are described in the same Microsoft Research post.

The researchers tested many closed and openly released models. In every case the systems scored higher when they only had to interpret a still scene than when they had to plan through a sequence of moves, and neither result came close to how people performed, the team reports in the MindTopo announcement. The disadvantage grew whenever a task demanded that a relationship hold steady across many steps, which indicates the real weakness is keeping state intact over time rather than reading one picture.

The error logs show where things break. When a model failed on a still image, the cause was usually perceptual — it overlooked a wall, an opening, or a crossing. Failures during planning showed up only after the model had already grasped the scene: it would take a move that looked fine locally without weighing later consequences, or suggest an action the simulator’s rules ruled out. “Seeing topology is not the same as acting on it,” the team writes, summarizing a split between perception and execution that the Microsoft Research post documents across its test runs.

The practical stakes are concrete. Robots, accessibility tools, and interactive assistants must know not just where objects are but what stays connected, enclosed, ordered, or knotted as actions unfold. The authors argue that closing the gap may require models that carry an explicit topological state, or world models whose predictions preserve topology by construction.

The work comes from a collaboration between Northwestern University and Microsoft Research, including Jianfeng Gao, a Microsoft Technical Fellow and Corporate Vice President who sets the company’s direction for multimodal reasoning and agentic AI, as listed on his Microsoft Research profile, and co-authored by researchers whose publications are indexed on Google Scholar. It joins a steady run of Microsoft AI evaluation efforts, such as MAI-Code-1.1-Flash landing in GitHub Copilot, that probe where large models help and where they still fall short.

Editorially independent: we accept no payment for coverage and currently use no affiliate links. Read our Editorial Standards and Corrections Policy. Published: Aug 12, 2026.
Jinultimate

Editor of ZBrandCo and the person accountable for what we publish — setting our sourcing standards, fact-checking claims against primary sources, and issuing corrections promptly across AI, open source, and gaming. Reach the desk at editorial@zbrandco.com.