The quest to accurately measure and enhance the capabilities of large language models (LLMs) has taken a curious turn: it seems the models themselves hold the key. New research reveals that how we ask questions significantly impacts their performance, and a novel approach leverages the LLM's own internal signals to choose the optimal evaluation format, leading to substantial accuracy improvements. Simultaneously, other researchers are delving into the very fabric of multilingual LLMs, uncovering how shared concept spaces emerge during training and how language-dependent alignment can complicate true cross-lingual understanding. Further explorations highlight the power of LLMs in low-resource settings for complex tasks like skill extraction, while a paradigm shift in speech AI promises to imbue models with internal reasoning, enhancing both accuracy and the quality of spoken output.

The Case for Model-Preferred Evaluation

Evaluating the true intelligence of LLMs is a surprisingly thorny problem. Take multiple-choice questions, for instance. Performance can fluctuate dramatically depending on whether the model is asked to select a symbol from a list or complete a sentence (a cloze-style format). Researchers, including those behind arXiv:2601.22699, have observed these discrepancies are not random but systematically linked to the task's inherent characteristics. Natural language continuation tasks, which mimic how LLMs typically generate text, benefit from likelihood scoring. In contrast, explicit comparison tasks, requiring direct selection, are better served by symbol-based formats.

What's truly fascinating is that these trends appear to be model-agnostic, consistent across various decoder-based LLMs. This suggests a fundamental aspect of how these models process information rather than an artifact of a specific architecture. To tackle this evaluation inconsistency, a new strategy has emerged: dynamic format-alignment. Instead of relying on human-designed heuristics, which can sometimes degrade performance, this approach trains a lightweight classifier. This classifier learns to infer the model's internal preference signals, effectively asking the LLM, "Which way of answering this question do you find more natural or effective?" This model-generated insight allows for the selection of the optimal format for each individual problem instance. The results, as detailed in arXiv:2601.22699, show substantial and consistent improvements in zero-shot accuracy across reasoning and knowledge benchmarks, offering a clearer picture of the models' latent capabilities.

Unpacking Multilingual Minds and Resource-Scarce Skills

Beyond evaluation formats, understanding how LLMs handle multiple languages is another critical frontier. As monolingual resources become scarcer, training LLMs with extensive multilingual coverage is increasingly vital. Prior work has suggested that these models process multilingual inputs within shared concept spaces, which facilitates generalization and cross-lingual transfer. However, the emergence and nature of these spaces during training have remained somewhat opaque, often lacking rigorous causal analysis or focusing solely on the final model's state.

A study published on arXiv:2601.22851 dives into this, using the causal interpretability technique of activation patching on a model called EuroLLM. The researchers isolated cross-lingual concept representations and then injected them into translation prompts to test how consistently translations could be altered, irrespective of the original language. They found that shared concept spaces indeed emerge early in the training process and continue to refine. Crucially, however, the alignment with these shared spaces proved to be language-dependent. Furthermore, a detailed manual analysis revealed that some perceived improvements in translation quality were not due to enhanced translation ability but rather behavioral shifts. These included selecting specific senses for polysemous words or translating cognates (words with shared etymology) rather than simply copying them across languages. This nuanced understanding of multilingual training dynamics is essential for building truly effective cross-lingual AI systems.

Meanwhile, the practical applications of LLMs in low-resource settings are also advancing. Skill extraction, a cornerstone of modern recruitment systems, is particularly challenging for morphologically complex languages like Turkish, which lacks a dedicated skill taxonomy and dataset. A paper on arXiv:2601.22885 addresses this by introducing the first Turkish skill extraction dataset and evaluating LLMs for this task. The dataset comprises nearly 5,000 labeled skill spans from job postings. The research found that LLMs, when used in an end-to-end pipeline, significantly outperform traditional supervised sequence labeling methods. They are more effective at aligning extracted skill spans with standardized taxonomies like ESCO. The optimal configuration involved Claude Sonnet 3.7 with dynamic few-shot prompting, embedding-based retrieval, and LLM-based reranking for skill linking, achieving a notable end-to-end performance. This work paves the way for similar advancements in skill extraction for other underrepresented languages, demonstrating LLMs' potential to democratize AI capabilities.

The Dawn of 'Silent Thought, Spoken Answer'

Finally, the very nature of speech AI is being reimagined. Current speech language models typically generate responses directly, without explicit reasoning steps. This often leads to errors that are irreversible once audio is produced. The new paradigm, termed 'Silent Thought, Spoken Answer,' aims to equip speech LLMs with internal text reasoning capabilities that inform and improve the quality of their spoken output. The core innovation, detailed in arXiv:2601.22889, is a diffusion-based speech-text language model named DiffuSpeech. This model unifies discrete text and tokenized speech under a single masked diffusion framework, allowing it to understand and generate both modalities.

Unlike autoregressive models that generate tokens sequentially, DiffuSpeech jointly produces reasoning traces and speech tokens through iterative denoising. The system also benefits from modality-specific masking schedules. To support this research, the first speech QA dataset with paired text reasoning traces, containing over 26,000 samples, has been constructed. Experiments show that DiffuSpeech achieves state-of-the-art accuracy in speech-to-speech question answering, surpassing the best baseline by up to 9 percentage points. It also delivers superior text-to-speech quality and maintains strong language understanding capabilities. Ablation studies confirm that both the diffusion architecture and the explicit generation of 'thinking traces' are instrumental in these gains. This represents a significant step towards more robust, transparent, and accurate spoken AI systems.