Score: 2

Evaluating Large Language Models in Scientific Discovery

Published: December 17, 2025 | arXiv ID: 2512.15567v1

By: Zhangde Song , Jieyu Lu , Yuanqi Du and more

Potential Business Impact:

Tests if AI can do real science experiments.

Business Areas:

Natural Language Processing Artificial Intelligence, Data and Analytics, Software

Large language models (LLMs) are increasingly applied to scientific research, yet prevailing science benchmarks probe decontextualized knowledge and overlook the iterative reasoning, hypothesis generation, and observation interpretation that drive scientific discovery. We introduce a scenario-grounded benchmark that evaluates LLMs across biology, chemistry, materials, and physics, where domain experts define research projects of genuine interest and decompose them into modular research scenarios from which vetted questions are sampled. The framework assesses models at two levels: (i) question-level accuracy on scenario-tied items and (ii) project-level performance, where models must propose testable hypotheses, design simulations or experiments, and interpret results. Applying this two-phase scientific discovery evaluation (SDE) framework to state-of-the-art LLMs reveals a consistent performance gap relative to general science benchmarks, diminishing return of scaling up model sizes and reasoning, and systematic weaknesses shared across top-tier models from different providers. Large performance variation in research scenarios leads to changing choices of the best performing model on scientific discovery projects evaluated, suggesting all current LLMs are distant to general scientific "superintelligence". Nevertheless, LLMs already demonstrate promise in a great variety of scientific discovery projects, including cases where constituent scenario scores are low, highlighting the role of guided exploration and serendipity in discovery. This SDE framework offers a reproducible benchmark for discovery-relevant evaluation of LLMs and charts practical paths to advance their development toward scientific discovery.

Advancing the Scientific Method with Large Language Models: From Hypothesis to Discovery

Artificial Intelligence

Helps scientists discover new things faster.

22 May 2025 0

92%

The Empowerment of Science of Science by Large Language Models: New Tools and Methods

Computation and Language

AI helps scientists discover new ideas faster.

19 Nov 2025 1

92%

A Survey of Scientific Large Language Models: From Data Foundations to Agent Frontiers

Computation and Language

AI helps scientists discover new things faster.

28 Aug 2025 1

View PDF Login to Bookmark

Country of Origin

🇺🇸 🇨🇦 United States, Canada

Repos / Data Links

github.com github.com

Page Count

54 pages

Evaluating Large Language Models in Scientific Discovery

Tests if AI can do real science experiments.

Technical Abstract

Advancing the Scientific Method with Large Language Models: From Hypothesis to Discovery

The Empowerment of Science of Science by Large Language Models: New Tools and Methods

A Survey of Scientific Large Language Models: From Data Foundations to Agent Frontiers