Recent advancements in Artificial Intelligence are tackling fundamental limitations in transformer models and the training of large language models, promising greater efficiency and controllability. Researchers are introducing novel attention mechanisms and optimization strategies to overcome computational bottlenecks and enhance model robustness.

Reclaiming Global Competition in Transformers

Standard transformers, while powerful, suffer from a quadratic computational complexity that makes processing long contexts prohibitively expensive. Linear attention mechanisms have emerged as a solution, reducing complexity to linear time, but often at the cost of expressivity due to the removal of softmax normalization. This omission strips away a crucial "global competition" mechanism, hindering models' ability to focus on relevant information amidst noise. A new framework, Softmax Linear Attention (SLA), aims to restore this vital competitive selection without sacrificing efficiency. By shifting the softmax operation from the token level to the head level, SLA treats attention heads as semantic slots and employs a competitive gating mechanism. This approach dynamically selects the most relevant subspaces, reintroducing "winner-take-all" dynamics essential for precise retrieval and robust long-context understanding. Experiments show SLA consistently enhances state-of-the-art linear baselines like RetNet, GLA, and GDN, particularly in challenging retrieval tasks where it significantly boosts robustness against noise, all while maintaining linear complexity. This work, detailed in arXiv:2602.01744, demonstrates a path toward more capable and efficient long-context processing.

Stabilizing LLM Training with Dynamic Learning Rate Scheduling

Training large language models with Reinforcement Learning (RL) has long been plagued by instability, often attributed to a "training-inference mismatch." While traditional remedies like Importance Sampling can falter during extended training, new research frames this instability through an optimization lens. The study reveals that gradient noise and the training-inference mismatch escalate in tandem as training progresses. Crucially, shrinking the update size effectively suppresses this mismatch. Based on these insights, researchers propose a specialized learning rate (LR) scheduler. Unlike pre-defined decay schedules, this method dynamically triggers LR decay based on response length, a reliable early-warning signal for impending instability. By reducing the learning rate as gradient noise rises, the approach consistently stabilizes RL training and maintains the training-inference mismatch at a safe level, as documented in arXiv:2602.01826. This elegant solution offers a more robust path to training complex language models.

Enhancing Multimodal Model Control with FiLoRA

Multimodal foundation models, which integrate diverse signals across modalities like text, images, and audio, present a challenge in understanding and controlling their reliance on specific internal features. Existing methods often rely on post-hoc analysis or feature removal, offering limited ability to modulate feature reliance without altering task semantics. FiLoRA (Focus-and-Ignore LoRA) introduces a novel instruction-conditioned, parameter-efficient adaptation framework. It decomposes adaptation into feature group-aligned LoRA modules and applies instruction-conditioned gating. This allows natural language instructions to serve as computation-level control signals, guiding the model to selectively amplify or suppress core and spurious feature groups. Across text-image and audio-visual benchmarks, FiLoRA induces consistent and causal shifts in internal computation without modifying the label space or training objective. Further analyses show improved robustness under spurious feature interventions, offering a principled mechanism to regulate reliance beyond simple correlation-driven learning, as presented in arXiv:2602.02060. This work promises greater interpretability and control over complex multimodal systems.

Optimizing LLM Quantization for Efficiency

Deploying large language models (LLMs) on resource-constrained devices often necessitates quantization, a process that reduces model size and computational demands. Group-wise quantization is effective, but existing methods like GPTQ can neglect crucial factors like input statistics and inter-group correlations, leading to suboptimal accuracy. A proposed two-stage optimization framework for group scales directly addresses this by minimizing layer-wise reconstruction loss. In the first stage, group scales are initialized to minimize group-wise reconstruction loss, incorporating input statistics before GPTQ. The second stage refines these scales using a closed-form update rule derived from coordinate descent, minimizing layer-wise loss without costly numerical optimization. Notably, this refinement incorporates quantization errors from preceding layers to prevent accumulation. Experimental results show this method consistently enhances group-wise quantization, achieving higher accuracy with negligible overhead, as detailed in arXiv:2602.02126. This offers a more efficient pathway for deploying powerful LLMs on edge devices.

Accelerating DNN Deployment on RISC-V Architectures

Deep neural networks (DNNs) are essential for applications ranging from natural language processing to autonomous systems, but their deployment on resource-constrained platforms like RISC-V remains challenging due to high computational and memory demands. Low-rank factorization (LRF) offers a promising compression technique for fully connected layers, but the vast design space complicates optimization. A new methodology for LRF design space exploration and a specialized design tool aim to streamline this process for RISC-V processors. Using Tensor Train Decomposition (TTD) via TensorFlow's T3F library, the approach prunes inefficient decomposition shapes and solutions with poor inference performance. Compiler optimizations further enhance T3F layer performance. The result is TT-decomposed layers that run significantly faster—up to 3x faster than IREE and 8x faster than Pluto on the same compressed model. This work provides an efficient solution for deploying DNNs on edge and embedded devices powered by RISC-V, as reported in arXiv:2602.01996.