Despite persistent marketing suggesting otherwise, the foundational intelligence of vision-language models continues to exhibit significant shortcomings. Two new papers, both published on arXiv on May 9, 2026, expose critical gaps in how these models perceive the physical world and how efficiently they learn new skills, suggesting that the path to truly intelligent AI is still fraught with basic, unresolved challenges.

These papers, 'SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition' and 'Continually Evolving Skill Knowledge in Vision Language Action Model,' address what many of us have quietly observed: these machines, for all their impressive feats, struggle with concepts humans grasp before kindergarten. The continued reliance on incremental academic fixes highlights just how far the industry is from the integrated, robust intelligence it so often promises.

Spatial Cognition: More Than a Dot on a Map

Take spatial reasoning, for instance. Multimodal large language models (MLLMs) are supposed to be interacting with our physical environment, yet their understanding of 'up,' 'down,' or 'to the left of' remains rudimentary. According to researchers, existing benchmarks often "oversimplify spatial cognition, reducing it to a single-dimensional metric" arXiv CS.AI.

This isn't just an academic quibble; it means these models can't truly grasp the complex relationships of objects in a 3D space. The proposed SpatialBench aims to rectify this by acknowledging the "hierarchical structure and interdependence of spatial abilities" that current evaluations ignore arXiv CS.AI. It's a small step, certainly, but a necessary one, revealing that our supposedly advanced models are still toddlers in a world designed for adults.

The Endless Treadmill of Learning: A Glimmer of Efficiency?

Then there's the problem of learning. Vision-language-action (VLA) models, touted for their 'knowledge accumulation ability' from pretraining, hit a wall when it comes to continual learning. They struggle to adapt efficiently without forgetting old skills or bloating their architecture to an unsustainable degree. This 'continual learning in VLA remains challenging, especially for efficient adaptation,' as one paper grimly notes arXiv CS.AI.

Existing methods for continual imitation learning (CIL) compound the issue by "often rely[ing] on additional parameters or external modules, limiting scalability for large VLA models" [arXiv CS.AI](https://arxiv.org/abs/2511.18085]. It’s the computational equivalent of needing to buy a new brain every time you learn a new trick. However, the researchers behind 'Stellar VLA' claim to have developed a "knowledge-driven CIL framework without increasing network parameters" [arXiv CS.AI](https://arxiv.org/abs/2511.18085]. A tiny beacon of hope, perhaps, that we might eventually get models that learn without needing to consume the entire planet's processing power.

Industry Impact: Grounding AI in Reality, or Just More Benchmarks?

These aren't abstract academic concerns. The inability of MLLMs to genuinely understand spatial relationships fundamentally hobbles their application in fields like robotics, autonomous navigation, and augmented reality. If a robot can't reliably discern what's 'under' or 'behind' an object, its utility in dynamic environments is severely limited. Similarly, the scalability issues in continual learning mean that adaptive, real-world AI agents remain prohibitively expensive or simply unfeasible for complex, evolving tasks.

The industry's breathless pursuit of larger, more complex models often overshadows these fundamental deficiencies. These papers serve as a stark reminder that true intelligence isn't just about processing more data; it's about understanding the world in a robust, adaptable, and efficient manner. Without addressing these basics, the grand visions of embodied AI or universally competent virtual assistants will remain just that: visions.

What Comes Next?

We can anticipate more benchmarks, more frameworks, and more papers detailing why the last batch of solutions wasn't quite enough. Investors and developers should scrutinize claims of 'breakthroughs' with extreme prejudice, focusing instead on whether new models genuinely solve these underlying problems without introducing a fresh set of insurmountable compromises. The introduction of tools like SpatialBench and frameworks like Stellar VLA, while incremental, are crucial litmus tests. The true measure of progress won't be in the sheer size of the models, but in their ability to finally grasp the world as it is, learning efficiently without needing to continually rewrite their own operating instructions.