The Information Machine

This Yale + University of Chicago paper shows that real gap between LLM generated research ideas vs humans is not idea q…

Rohan Paul Twitter · Rohan Paul (@rohanpaul_ai) · 2026-08-01

A Yale and University of Chicago study using 11,683 real papers finds that LLMs match human researchers on research idea quality but produce significantly narrower ideas, over-indexing on connecting prior work at 4-5x the human rate.

Open original ↗

Extraction

Topics: llm-research-ideationai-creativityllm-evaluationscientific-ai

Claims

  • The real gap between LLM and human research ideas is diversity of contribution type, not quality.
  • LLMs propose ideas that connect separate prior works 47-64% of the time versus only 12.1% for human researchers, a 4-5x overrepresentation.
  • Human researchers spread ideas across many contribution patterns such as explaining mechanisms, testing failures, measuring evidence, building systems, and improving efficiency.
  • Additional reasoning steps (e.g., extended chain-of-thought) makes the narrowness worse rather than compensating for it.
  • The study used a controlled benchmark derived from 11,683 real papers where models were given the same prior-work inputs as the original human authors.

Key quotes

Only 12.1% of human ideas were mainly about connecting separate work, but 47.1% to 64.2% of LLM ideas did that, meaning models used this move about 4 to 5 times more often.
Even extra reasoning made this pattern stronger, suggesting models often polish a familiar recipe instead of finding more varied research moves.
The real gap between LLM generated research ideas vs humans is not idea quality, but idea range: LLMs think narrower than human researchers.