Lee Douglas

Deep Tech Correspondent

Automated systems are increasingly asked to reflect human values, but how do we even measure that? A new paper on arXiv is shining a harsh light on a common methodology: using social survey questions to probe the "value orientation" of large language models (LLMs). The research, "On the Credibility of Evaluating LLMs using Survey Questions," highlights critical flaws that can make these evaluations misleading, potentially leading us to both overestimate and underestimate an LLM's alignment with human ethical frameworks. This isn't just an academic quibble; as LLMs become more integrated into decision-making roles, understanding their true value alignment is paramount.

The Flaws in the Survey Setup

The prevailing approach involves feeding survey questions directly to LLMs and then comparing their answers to aggregated human responses. However, this seemingly straightforward method is fraught with nuance. The research team, using data from the World Value Survey across three languages and five countries, discovered that seemingly minor technical choices can drastically skew results. Specifically, they found that different prompting techniques, such as direct questioning versus chain-of-thought (CoT) prompting, and decoding strategies, like greedy sampling versus more probabilistic approaches, significantly impact the perceived alignment.

This variability suggests that a positive result using one setup might evaporate with a slight adjustment, casting doubt on the robustness of many existing evaluations. It’s a classic case of "garbage in, garbage out," or perhaps more accurately, "suboptimal setup in, misleading conclusion out." The paper even introduces a novel metric, self-correlation distance, to address a deeper issue: whether LLMs maintain consistent relationships between their answers, mirroring how human responses form a coherent value structure.

Beyond Average Agreement: The Need for Structural Alignment

The paper's core critique is that simply achieving high average agreement with human data doesn't guarantee that an LLM understands or embodies those values in a structurally sound way. Imagine two people agreeing on the answer to questions about honesty and loyalty. If one person's reasoning for agreeing on honesty is tied to self-preservation and their reasoning for loyalty is based on social obligation, while the other's is based on intrinsic moral principles, their underlying value systems are vastly different, even if their survey answers look the same on the surface. The self-correlation distance metric aims to capture this crucial internal consistency. Without it, we risk mistaking superficial mimicry for genuine alignment.

Furthermore, the study points out a weak correlation between two commonly used metrics: mean-squared distance and KL divergence. These metrics often assume survey answers are independent. This assumption, the researchers argue, is problematic and fails to capture the intricate dependencies between different value statements that humans inherently exhibit. Relying on them in isolation can paint an incomplete or inaccurate picture of an LLM's value system.

"Simply achieving high average agreement with human data doesn't guarantee that an LLM *understands* or *embodies* those values in a structurally sound way."

— Lee Douglas

The researchers offer clear recommendations for more reliable evaluation moving forward. They advocate for the use of CoT prompting, which encourages the model to "show its work." They also suggest employing sampling-based decoding with a substantial number of samples (dozens) to explore the distribution of potential responses. Crucially, they emphasize the need for robust analysis using multiple metrics, including their newly proposed self-correlation distance, to gain a more holistic understanding.

This research is a vital contribution to the burgeoning field of AI safety and ethics. As we race to deploy more powerful AI systems, ensuring they operate within human ethical bounds is non-negotiable. This paper serves as a potent reminder that the how of evaluation is just as important, if not more so, than the what. Without rigorous, structurally aware evaluation methods, our confidence in LLM alignment may be built on a foundation of sand. The path forward requires more sophisticated tools and a deeper understanding of the complex, interconnected nature of human values.