Aivora
arXivAI ResearchIntermediate

ScholarCatalyst: A Benchmark for Testing AI's Intuition in Retrieving Inspiring Research Papers

ScholarCatalyst:評估 AI 是否擁有「科學家直覺」的學術文獻檢索基準

2 min read
ScholarCatalyst: A Benchmark for Testing AI's Intuition in Retrieving Inspiring Research Papers
The 30-second version

While AI progresses in solving open problems, it lacks the scientific intuition to identify which prior papers can inspire new ideas. The authors present ScholarCatalyst, a benchmark built with 184 lead authors of 207 computer science papers who labeled the prior works that actually advanced their projects. Given an initial research question, models must retrieve these inspiring papers from the literature available at the project's start. Surprisingly, agentic search reached only 0.51 Recall@20 (using Claude Fable 5.1), barely outperforming simple embedding retrieval (0.48), exposing a major gap in expert-level literature search.

Key points

01

Testing Scientific Intuition

Traditional retrieval focuses on surface similarity, whereas ScholarCatalyst evaluates if AI can sense which buried prior ideas can solve a new research problem.

02

Ground-Truth from 184 Authors

Using a scalable pipeline, the creators gathered direct annotations and rationales from 184 lead authors of 207 recent computer science papers.

03

Agentic Search Underperforms

Agentic search scored only 0.42 Recall@20, failing to beat a standard embedding retriever (0.48) despite having access to it as a tool.

04

Top Models Struggle

Even Claude Fable 5.1, which might have seen the completed papers during training, only achieved 0.51 Recall@20, highlighting a gap in expert-level training.

How it works

ScholarCatalyst Evaluation Task Flow
Initial ResearchQuestionPre-project LiteratureCorpusAI / Agent RetrievalSystemRetrieved CandidatesAuthor Ground-TruthMatch

Why it matters

This work highlights a critical bottleneck in AI for Science. Future scientific agents cannot rely on simple keyword or semantic similarity; they must grasp the deep conceptual links that spark scientific inspiration. ScholarCatalyst provides a challenging benchmark that will push RAG, agentic search, and specialized pre-training paradigms toward deeper logical reasoning and associative thinking.

Who it affects

  • AI Developer
  • AI Researcher
  • Enterprise Leader

How to use it

  1. 1Evaluating academic AI agents on cross-disciplinary literature retrieval.
  2. 2Developing scientific assistants that recommend potential inspiring works based on half-formed ideas.
  3. 3Testing and improving RAG systems on complex, long-context retrieval reasoning tasks.

Limitations & caveats

  • The dataset is primarily focused on computer science, which may not fully represent literature retrieval patterns in other fields like biology or physics.
  • Annotations rely on retrospective self-reporting by authors, which might introduce subjective bias or memory gaps.

Related

Ai2 Open-Sources AstaBrief: An 8B Scientific Report Generator 3.5x Faster than Claude
Hugging FaceAI Research

Ai2 Open-Sources AstaBrief: An 8B Scientific Report Generator 3.5x Faster than Claude

艾倫人工智慧研究所開源 AstaBrief:比 Claude 快 3.5 倍的 8B 科學報告生成模型

Allen Institute for AI (Ai2) has open-sourced AstaBrief 8B, a specialized model for scientific report generation that achieves a 3.5x speedup over proprietary pipelines while maintaining high citation accuracy.

2 min read