arXivAI 研究
當沒有標準答案時:如何用「陳述偏好經濟學」評估大型語言模型
Evaluating LLMs Without Ground Truth: Lessons from Stated-Preference Economics
本研究引入經濟學的「陳述偏好」評估框架,在沒有客觀標準答案的情況下,透過內容、建構與效度等經濟學指標,測試並區分 LLM 決策的一致性與合理性。
2 分鐘閱讀
#ai-alignment
1 篇文章
Evaluating LLMs Without Ground Truth: Lessons from Stated-Preference Economics
本研究引入經濟學的「陳述偏好」評估框架,在沒有客觀標準答案的情況下,透過內容、建構與效度等經濟學指標,測試並區分 LLM 決策的一致性與合理性。