Skip to main content

AI Models Master Word Games But Fail Spatial Reasoning Tests

26 AUGUST 2026·2 MIN READ·1 SOURCE·Trusted source

While AI models have rapidly improved at logic puzzles like the New York Times Connections, they still fail at visual and spatial reasoning tasks such as mental rotation.

AI Models Master Word Games But Fail Spatial Reasoning Tests

Key takeaways · 2

  • 01

    AI models advanced from an 18% success rate on NYT Connections puzzles in late 2024 to near-perfect scores in early 2025.

  • 02

    Language models continue to fail abysmally at spatial reasoning tasks like mental rotation.

Rapid Gains in Word Puzzles

Puzzles and games have been central to the development of artificial intelligence since the very beginning. [1] In late 2024, scientists from Columbia University demonstrated that the most capable models could only solve 18 percent of New York Times Connections puzzles. [1] By early 2025, however, certain models were able to solve these puzzles almost perfectly every time. [1]

The Spatial Reasoning Gap

Despite these advances, today's models still struggle with subtle changes in classic riddles and demonstrate a particular weakness in visual puzzles. [1] Modern language models fail abysmally at mental rotation problems, which require determining if different images show the same object from varying angles. [1] These systems still appear unable to manipulate 3D objects in the way spatial thinkers like mechanical engineers can. [1]

What it means

The rapid progression from an 18% success rate to near-perfect performance on the NYT Connections puzzle highlights how quickly language-based AI reasoning can evolve. However, the persistent failure in spatial and visual tasks like mental rotation indicates a significant divergence between human cognition and machine learning architectures. While text-based reasoning is accelerating, physical world comprehension remains a major hurdle. What the sources don't address: How exactly AI developers plan to bridge the gap between two-dimensional visual processing and genuine three-dimensional spatial manipulation.

Evaluating AI through puzzles exposes the precise cognitive gaps between machines and humans. Understanding these limitations is critical for deploying AI in spatial or physical environments.

Why it matters
Daily session

Turn this story into practical AI skill after launch.

Get the release link for daily sessions built around your role and industry.

Join the waitlist

How this developed

  1. 26 August 2026

    AI Models Master Word Games But Fail Spatial Reasoning Tests

  2. 26 August 2026

    Event created from source cluster.

Sources

AI fluency, one session a day, built for your work.