A September 30 arXiv preprint introduces AlphaDLM, a training method that uses a sequence-level alpha loss as an alternative to cross-entropy to improve discrete diffusion language models. arXiv CS.LG

The approach targets a weakness of factorized diffusion models: cross-entropy training fits conditional token marginals, but parallel token generation requires consistent joint predictions. By adjusting the alpha parameter, the objective can enforce joint-mode optimality at alpha=1 while recovering cross-entropy as alpha approaches zero, according to the abstract.

Discrete diffusion models can generate tokens in parallel, but reducing denoising steps can lead to inconsistent predictions, the abstract notes. Automatica reported earlier on Simplex Diffusion Models, which address a related information-collapse problem. The new work shifts focus to the training objective itself.

When trained on the TinyGSM dataset and evaluated on the GSM8K math benchmark, AlphaDLM reaches 34.6% accuracy with only four model evaluations, the preprint says. The paper also describes scaling the method to SDAR-1.7B and testing on code and mathematics benchmarks, though specific scores for those larger models are not detailed in the abstract.

The preprint has not been peer reviewed.