In a flurry of research released today, artificial intelligence continues its relentless march forward, with new techniques promising to solve long-standing challenges in 3D reconstruction and accelerate the generation of complex media like video. Two papers introduce novel approaches for capturing intricate details on challenging surfaces, while others tackle the efficiency and alignment of powerful generative models, signaling a new era of more robust and accessible AI.

Unlocking Reflective Surfaces: The Magic of Latent Diffusion

For years, capturing accurate 3D models of highly reflective objects—think polished chrome, glass, or even a shiny car—has been a significant hurdle. Standard techniques, like fringe projection profilometry, struggle with specular reflections and indirect illumination, leading to distorted or incomplete data. This has made detailed scanning of such items a costly and time-consuming affair, limiting their application in fields from manufacturing to augmented reality.

Now, researchers have introduced LD-SLRO (Latent Diffusion Structured Light for Reflective Objects), a novel method that leverages the power of latent diffusion models. The approach first encodes fringe images captured from these challenging surfaces into latent representations that specifically capture surface reflectance characteristics. These latent features then act as conditional inputs to a latent diffusion model. The AI then probabilistically suppresses reflection-induced artifacts and intelligently restores lost fringe information, yielding remarkably high-quality images. This isn't just a marginal improvement; experimental results show a reduction in average root-mean-squared error from 1.8176 mm to 0.9619 mm, more than halving the reconstruction error. The team's innovative components, including a specular reflection encoder and attention modules, further refine the fringe restoration process, offering unprecedented flexibility in configuring input and output fringe sets.

This breakthrough, detailed in arXiv:2602.05434, has profound implications for quality control in manufacturing, detailed digital archiving of artifacts, and creating more realistic virtual environments. The ability to accurately scan glossy materials opens doors to applications previously deemed too complex or costly.

Enhancing Generative AI: Efficiency, Alignment, and Long Contexts

Beyond specialized reconstruction tasks, a wave of research also addresses the broader landscape of generative AI, focusing on making these powerful models more efficient, easier to control, and capable of handling longer, more complex outputs.

One significant development comes from the realm of diffusion model alignment with SAIL (Self-Amplified Iterative Learning), detailed in arXiv:2602.05380. Aligning diffusion models with human preferences is notoriously difficult, often requiring vast datasets or complex reward models. SAIL offers a paradigm shift by enabling diffusion models to act as their own teachers. Starting with minimal human feedback, the model iteratively generates samples, assesses its own preferences, and refines itself. This self-augmentation process, coupled with a novel ranked preference mixup strategy to prevent catastrophic forgetting, allows SAIL to achieve state-of-the-art alignment using a fraction (just 6%) of the data typically required. This drastically reduces the annotation burden and computational cost associated with aligning models to nuanced human desires, making personalized AI generation more accessible.

Efficiency in training and sampling is another key focus. Stable Velocity, presented in arXiv:2602.05435, tackles the high-variance training targets inherent in flow matching, a popular technique for training diffusion models. By understanding the variance dynamics, the researchers propose Stable Velocity Matching (StableVM) and Variance-Aware Representation Alignment (VA-REPA) for more stable and efficient training. Crucially, they also introduce Stable Velocity Sampling (StableVS), which allows for significantly faster sampling (over 2x) without any fine-tuning, by leveraging low-variance regimes near the data distribution. This translates to quicker generation times for models like SD3.5 and Flux, enhancing usability without compromising quality.

"SAIL offers a paradigm shift by enabling diffusion models to act as their own teachers, drastically reducing the annotation burden."

— Lee Douglas, Automatica Press

For tasks involving long-form content generation, such as extended texts or videos, efficiency remains a paramount concern. FlashBlock (arXiv:2602.05305) introduces an innovative attention caching mechanism for block diffusion, a method already designed to speed up inference. FlashBlock identifies that attention outputs for tokens outside the current processing block remain remarkably stable across diffusion steps. By reusing these stable, external attention outputs, it significantly reduces computational overhead and KV cache access, leading to up to 1.44x higher token throughput and a 1.6x reduction in attention time, with negligible impact on generation quality. This is particularly relevant for applications requiring long sequences, where traditional methods falter.

Complementing this, DisCa (arXiv:2602.05449) addresses acceleration in video diffusion transformers by introducing a distillation-compatible learnable feature caching mechanism. Unlike traditional training-free caching, DisCa employs a lightweight neural predictor to more accurately capture feature evolution. Combined with a conservative Restricted MeanFlow approach for stable distillation, DisCa pushes acceleration boundaries to an impressive 11.8x while maintaining generation quality. This represents a significant leap forward for realistic and efficient video generation, a domain with rapidly growing demand.