Evaluating LLMs Without Ground Truth: Lessons from Stated-Preference Economics
當沒有標準答案時:如何用「陳述偏好經濟學」評估大型語言模型
When LLMs face questions without objective ground truths—such as policy valuation or ethical trade-offs—traditional benchmarks fail. This paper adapts 'stated-preference' economics, a framework designed to evaluate survey responses without knowing true values (using content, construct, and criterion validity). Testing six LLMs on a water-quality valuation survey, the authors found that while the newest models passed theoretical tests (e.g., downward-sloping demand), they diverged on convergent validity, while older models failed basic consistency tests.
Key points
The Ground-Truth Dilemma
Many complex tasks given to LLMs lack definitive ground-truth answers, rendering standard benchmarking ineffective.
Stated-Preference Economics Framework
Leverages decades of economic methodology designed to judge survey responses without prior knowledge of true values.
Economic Theory Consistency
Validates model responses against economic theory predictions, such as downward-sloping demand and income sensitivity.
Coherence is Not Correctness
Passing validity tests only demonstrates that a model's answers are coherent, not that they are inherently correct.
Why it matters
This research provides a novel pathway for LLM alignment and safety evaluation. As AI increasingly assists in decision-making, policy formulation, and value trade-offs, simple binary benchmarks are insufficient. Importing economics-based validity frameworks allows systematic evaluation of model coherence in subjective spaces, helping filter out logically inconsistent models.
Who it affects
- AI Researcher
- Product Manager
- Policy Maker
How to use it
- 1Evaluating AI's value trade-off capabilities in policy-making and public choice.
- 2Designing LLM safety and preference alignment tests for subjective scenarios.
Limitations & caveats
- Coherence does not guarantee correctness; a model can be highly consistent yet logically incorrect or biased.
- The evaluation relies heavily on the assumptions of specific economic theories, which may not hold in all scenarios.
Related
Building Persistent 3D Object Memory: How Ledger Tracks Objects from Egocentric Videos
打造過目不忘的 3D 空間記憶:Ledger 如何透過第一人稱影片追蹤隱形物體
Researchers introduce Ledger, a framework that builds a persistent 3D object memory from egocentric videos, significantly improving spatial question-answering accuracy for embodied agents.
Decoupling Exploration from Optimization: How ExpDis Boosts LLM Reasoning and Solution Diversity
探索與優化解耦:全新強化學習框架 ExpDis 提升大語言模型的推理多元性
The ExpDis framework decouples exploration from optimization in RLVR. By training explorers with novelty bonuses and distilling filtered trajectories into a student model, it prevents model degradation while fostering diverse reasoning.
Clipped Decentralized SGD: Achieving Optimal Convergence and Linear Speed-Up Under Heavy-Tailed Noise
去中心化 SGD 克服重尾雜訊:梯度裁剪如何實現最佳收斂與線性加速
This study proves that clipped decentralized SGD (DSGD) achieves order-optimal convergence rates and linear speed-up under heavy-tailed noise for non-convex optimization.