Aivora
arXivAI SafetyIntermediate

Beyond Single-Pair Comparison: Robust Stereotype Evaluation in LLMs via Dual Minimal Pairs

大型語言模型偏見評估新突破:引入雙重極小對立組與互資訊指標

2 min read
Beyond Single-Pair Comparison: Robust Stereotype Evaluation in LLMs via Dual Minimal Pairs
The 30-second version

Evaluating bias in LLMs by comparing log-likelihoods of single-pair contrastive sentences is often unreliable and prone to logical inconsistency. To address this, the authors introduce a 'dual minimal pair' framework offering two axes of comparison. Accompanied by a data-augmentation framework that generates paraphrases and alternate attributes in English, Russian, Spanish, and Chinese, they also present a new Mutual Information (MI)-based metric. This MI metric models the relationship between social groups and stereotyped attributes, enabling robust, aggregate-ready, and cross-lingual stereotype evaluations.

Key points

01

Unreliability of Single-Pair Evaluations

Comparing log-likelihoods of single contrastive sentence pairs often yields logically inconsistent preferences when attributes are slightly rewritten.

02

Dual Minimal Pair Framework

Introduces two axes of comparison for robust stereotype evaluation, ensuring consistency across different contexts.

03

Multilingual Data Augmentation

Features a data-augmentation framework to generate paraphrases and alternative attributes, successfully tested across English, Chinese, Spanish, and Russian.

04

Mutual Information (MI) Metrics

Introduces MI-based metrics that measure association strength between social groups and attributes, facilitating comparison of bias levels across varying architectures.

How it works

Comparison of Single-Pair vs. Dual Minimal Pair Methods
單一對立組評估 (Single-Pair)雙重極小對立組評估 (Dual Minimal Pair)
Axes of Comparison單一軸線 (比較兩句對立刻板句)雙重軸線 (結合釋義與交替屬性)
Logical Consistency較差,容易因重寫屬性而產生矛盾偏好較高,透過多維度對照抵抗局部語法擾動
Core Metrics單純對數概似度 (Log-likelihood)互資訊 (MI) 聚合指標
Cross-Model Comparability低,難以在不同基準線下進行穩定聚合高,適合進行不同語言與架構間的偏見強度對比

Why it matters

Current LLM bias benchmarks are highly sensitive to phrasing, risking false conclusions about model safety. By introducing dual minimal pairs and MI-based metrics, this work circumvents the fragility of traditional single-pair testing. This method yields a highly standardized, logically consistent, and cross-lingual standard of measurement, significantly aiding future researchers and developers in identifying and mitigating model bias more accurately.

Who it affects

  • AI Researcher
  • AI Developer
  • Product Manager

How to use it

  1. 1Evaluating and optimizing bias mitigation during LLM safety alignment
  2. 2Cross-lingual and cross-architecture benchmarking of stereotype levels

Limitations & caveats

  • The evaluation and data-augmentation frameworks have currently only been validated on English, Russian, Spanish, and Chinese.
  • While MI metrics provide more stable comparisons, the reliability still depends heavily on the generative quality of the contrastive sentence datasets.

Related

Anthropic Expands Cyber Verification Program to Give Defenders the AI Advantage
AnthropicAI Safety

Anthropic Expands Cyber Verification Program to Give Defenders the AI Advantage

Anthropic 擴大「網路安全驗證計畫」:放寬安全防護,為資安防守者提供強大 AI 武器

Anthropic has expanded its Cyber Verification Program (CVP) into a three-tier model, granting verified security professionals access to advanced Claude models with reduced safeguards for cyberdefense and red-teaming.

2 min read
NVIDIA Launches Open Agent Safety Platform for Continuous In-Silicon Agent Monitoring
NVIDIA DeveloperAI Safety

NVIDIA Launches Open Agent Safety Platform for Continuous In-Silicon Agent Monitoring

NVIDIA 推出 Open Agent Safety Platform:基於晶片與開源沙盒的 AI Agent 持續安全監控架構

NVIDIA introduced the Open Agent Safety Platform, combining open-source OpenShell sandboxing and BlueField hardware DPUs to deliver independent, zero-trust monitoring and real-time intervention for autonomous agents.

2 min read