A preprint posted to arXiv on October 6 describes a method for adding vision to two large language models that lack native visual capabilities.
The work trains a 50‑million‑parameter projector between the frozen vision encoder of Kimi K2.6 and the GLM 5.2 and 5.3 language models, neither of which includes vision processing. The paper, submitted to the NeurIPS 2026 Workshop on Grounded and Faithful Vision‑Language Models for Real‑World Deployment, revisits an established approach at frontier scale: coupling a frozen encoder to a much larger language model through a small trainable adapter.
Training a lightweight projector is a common technique for making pure language models multimodal without altering their pretrained weights. As models have grown, it has been unclear whether the method continues to deliver proportional vision gains. The authors say they examine how the visual capabilities of such a system scale when only the language‑model side grows, and which capabilities can be added at frontier sizes versus those that remain limited. They evaluate on the MMMU‑Pro and BLINK benchmarks, breaking down performance by visual task.
The abstract does not include specific benchmark scores, nor does it disclose the parameter counts of the GLM 5.2 and 5.3 base models. The preprint has not been peer reviewed, and no code repository accompanies the paper. The authors frame the work as a reproducible recipe for training vision adapters at scale.