Two new research papers released on arXiv this week are tackling the significant computational hurdles facing large language models (LLMs), offering promising solutions for both training efficiency and inference speed.

One groundbreaking development, dubbed FOCUS, aims to drastically improve the inference performance of Diffusion Large Language Models (DLLMs) by intelligently pruning unnecessary computations. Another, TEON, presents a novel optimization technique that accelerates the pre-training phase of LLMs by refining how gradients are managed. Separately, a third paper details an advancement in speech recognition, enabling decoder-only LLMs to handle streaming audio with significantly reduced latency.

Taming the Compute Beast: FOCUS for DLLM Inference

Diffusion Large Language Models (DLLMs) represent an exciting alternative to traditional auto-regressive models, but their widespread adoption has been hampered by substantial decoding costs. This bottleneck arises because, while computation can be parallelized across token blocks, only a fraction of these tokens are actively being decoded at each diffusion step. This means a significant portion of computational resources are spent on tokens that ultimately don't contribute to the output. Researchers have observed a strong correlation between the importance of a token, as determined by its attention weights, and its probability of being decoded. This insight forms the foundation of the new system, FOCUS.

FOCUS dynamically directs computational effort towards tokens that are most likely to be decoded, while