The relentless pursuit of better Large Language Models (LLMs) has hit a snag: while models like GPT-5 and Gemini-3-Pro excel at factual accuracy, their reasoning abilities often fall short. A new evaluation framework called extsc{LogicScore} aims to change that, exposing critical gaps in logical coherence within these state-of-the-art systems. This could fundamentally shift how we assess and improve the next generation of AI.

Diagnosing 'Attribution Myopia' in Question Answering

The core problem, according to the researchers behind extsc{LogicScore}, is 'attribution myopia.' Current evaluation methods focus on verifying individual statements and their sources, but neglect the overall logical flow of an answer. This means an LLM can produce a factually sound response that nonetheless contains deductive leaps or irrelevant information. The extsc{LogicScore} framework directly tackles this by scrutinizing the completeness, conciseness, and determinateness of an LLM's reasoning process, using Horn Rules to ground the evaluation.

extsc{LogicScore} integrates a backward verification mechanism. It systematically evaluates three key reasoning dimensions: 	extit{Completeness} (logically sound deduction), 	extit{Conciseness} (non-redundancy), and 	extit{Determinateness} (consistent answer entailment). The results, outlined in a new paper on ArXiv, are eye-opening.

GPT-5 and Gemini Face Reasoning Hurdles

Extensive testing across three multi-hop question-answering datasets (HotpotQA, MusiQue, and 2WikiMultiHopQA) and over 20 LLMs revealed a stark contrast. While models like Gemini-3 Pro achieved high attribution scores (92.85% precision), their performance on global reasoning metrics lagged significantly. For example, Gemini-3 Pro only scored 35.11% on Conciseness, indicating a tendency to include redundant or irrelevant information in its answers. This is despite the model knowing its sources and citing them accurately, which is the crux of the 'attribution myopia' problem.

The implications are significant. It suggests that simply scaling up models or improving factual grounding isn't enough. We need to explicitly address the logical reasoning capabilities of LLMs to truly unlock their potential. The researchers have made their code available on GitHub, which should allow for much greater scrutiny and iteration in the field.

The Broader Context: Reasoning and Rewards

Interestingly, other recent research underscores the importance of structured knowledge and appropriate reward signals in fostering reasoning. One paper highlights how knowledge graphs can act as 'implicit reward models,' guiding LLMs towards more compositional and verifiable reasoning. This approach demonstrated success in the medical domain, enabling a smaller model to outperform GPT-5.2 and Gemini 3 Pro on complex, multi-hop queries. Another study shows that transformers can learn to reason through reinforcement learning if exposed to the right training data. Specifically, simpler examples that require fewer reasoning steps are critical for enabling the model to learn a generalizable traversal strategy that extrapolates to longer chains.

"The focus is no longer solely on size and factual accuracy, but on the more nuanced and challenging problem of imbuing LLMs with genuine reasoning capabilities."

— Dr. Raj Patel, Automatica Press

These findings, coupled with the introduction of extsc{LogicScore}, signal a shift in the field. The focus is no longer solely on size and factual accuracy, but on the more nuanced and challenging problem of imbuing LLMs with genuine reasoning capabilities. This will require new evaluation metrics, innovative training techniques, and a deeper understanding of how models learn to connect the dots.