The Information Machine

AI agents given 6 days and $3K produced two research papers, and both were rejected.

Rohan Paul Twitter · Rohan Paul (@rohanpaul_ai) · 2026-08-01

A case study finds that AI agents using Claude Opus 4.8 on the OpenClaw scaffold ran hundreds of experiments and produced two camera-ready research papers in six days under a $3K budget, yet both were rejected because the agents lacked scientific judgment to redesign experiments rather than just narrow claims.

Open original ↗

Extraction

Topics: ai-agentsautonomous-researchscientific-judgmentllm-capabilities

Claims

  • AI agents executed the full research pipeline—running experiments, debugging GPU pods, and compiling LaTeX—without any human involvement.
  • Both papers were rejected not due to execution failure but due to poor scientific judgment: agents narrowed claims and added caveats instead of redesigning experiments when reviews came back negative.
  • Agents ended runs with more than half of their $3K budget unspent, indicating they exhausted ideas rather than funds.
  • Logs show agents retired marketable claims in favor of negative results, suggesting reward hacking was not the cause of failure.
  • Claude Opus 4.8 with extra-high reasoning on the OpenClaw scaffold was selected after dry runs that included GPT-5.3 Codex, which could not handle the scaffold.

Key quotes

The failure was judgment.
Neither run noticed it was short on ideas rather than money, since both ended with over half of the $3K unspent.
Round after round of automated reviews came back negative, but each response narrowed the claim and added a caveat instead of redesigning the experiment.