Aivora
arXivAI ResearchIntermediate

Evaluating LLMs Without Ground Truth: Lessons from Stated-Preference Economics

當沒有標準答案時:如何用「陳述偏好經濟學」評估大型語言模型

2 min read
Evaluating LLMs Without Ground Truth: Lessons from Stated-Preference Economics
The 30-second version

When LLMs face questions without objective ground truths—such as policy valuation or ethical trade-offs—traditional benchmarks fail. This paper adapts 'stated-preference' economics, a framework designed to evaluate survey responses without knowing true values (using content, construct, and criterion validity). Testing six LLMs on a water-quality valuation survey, the authors found that while the newest models passed theoretical tests (e.g., downward-sloping demand), they diverged on convergent validity, while older models failed basic consistency tests.

Key points

01

The Ground-Truth Dilemma

Many complex tasks given to LLMs lack definitive ground-truth answers, rendering standard benchmarking ineffective.

02

Stated-Preference Economics Framework

Leverages decades of economic methodology designed to judge survey responses without prior knowledge of true values.

03

Economic Theory Consistency

Validates model responses against economic theory predictions, such as downward-sloping demand and income sensitivity.

04

Coherence is Not Correctness

Passing validity tests only demonstrates that a model's answers are coherent, not that they are inherently correct.

Why it matters

This research provides a novel pathway for LLM alignment and safety evaluation. As AI increasingly assists in decision-making, policy formulation, and value trade-offs, simple binary benchmarks are insufficient. Importing economics-based validity frameworks allows systematic evaluation of model coherence in subjective spaces, helping filter out logically inconsistent models.

Who it affects

  • AI Researcher
  • Product Manager
  • Policy Maker

How to use it

  1. 1Evaluating AI's value trade-off capabilities in policy-making and public choice.
  2. 2Designing LLM safety and preference alignment tests for subjective scenarios.

Limitations & caveats

  • Coherence does not guarantee correctness; a model can be highly consistent yet logically incorrect or biased.
  • The evaluation relies heavily on the assumptions of specific economic theories, which may not hold in all scenarios.

Related