A preprint posted to arXiv on 30 September 2026 analyzes how intermediate representations in transformer language models progressively narrow toward a final prediction, comparing six pretrained models.

The work offers a fine-grained view of the inference process inside transformers, a mechanism that remains only partially understood despite the architecture's widespread use. The paper, "The Geometry of Inference in Transformer Residual Streams," appeared as a preprint on the arXiv CS.LG and cs.CL listings.

Automatica Press reported in October 2026 on a preprint probing when recurrence aids looped language models, part of a growing body of research into transformer inference dynamics. The new paper takes a geometric approach, comparing intermediate residual states with their own final state and with an empirical bank of final states from other contexts, according to the abstract.

The authors report that across the six studied models, the model's own final endpoint becomes preferable to the average alternative early in processing, though many individual endpoints remain closer. The set of competing endpoints typically shrinks with depth, but its membership changes and surviving endpoints do not necessarily become more similar to one another. Directional alignment and endpoint rank can therefore improve while Euclidean distance to the final state changes little.

The paper develops a simple high-dimensional model to separate the roles of norm, alignment, and endpoint geometry, showing how gradual directional changes can produce sharp reductions in competition. It also proves mathematically that a straight path toward the own endpoint cannot introduce new competitors under either Euclidean or cosine distance; observed trajectories therefore indicate departures from straight-line convergence. Finally, endpoints associated with lower-ranked output tokens tend to lie farther away in cosine distance across all models studied, connecting residual geometry to output organization.

The preprint has not been peer reviewed, and the authors' institutional affiliations are not specified in the submission. The abstract does not name the six language models or their sizes, nor does it provide independent test results beyond the geometric analysis described.