Sound Agentic Science Requires Adversarial Experiments

April 23, 2026 Β· Grace Period Β· πŸ› ICLR 2026 Workshop on Agents in the Wild

⏳ Grace Period
This paper is less than 90 days old. We give authors time to release their code before passing judgment.
Authors Dionizije Fa, Marko Culjak arXiv ID 2604.22080 Category cs.AI: Artificial Intelligence Citations 0 Venue ICLR 2026 Workshop on Agents in the Wild
Abstract
LLM-based agents are rapidly being adopted for scientific data analysis, automating tasks once limited by human time and expertise. This capability is often framed as an acceleration of discovery, but it also accelerates a familiar failure mode, the rapid production of plausible, endlessly revisable analyses that are easy to generate, effectively turning hypothesis space into candidate claims supported by selectively chosen analyses, optimized for publishable positives. Unlike software, scientific knowledge is not validated by the iterative accumulation of code and post hoc statistical support. A fluent explanation or a significant result on a single dataset is not verification. Because the missing evidence is a negative space, experiments and analyses that would have falsified the claim were never run or never published. We therefore propose that non-experimental claims produced with agentic assistance be evaluated under a falsification-first standard: agents should not be used primarily to craft the most compelling narrative, but to actively search for the ways in which the claim can fail.
Community shame:
Not yet rated
Community Contributions

Found the code? Know the venue? Think something is wrong? Let us know!

πŸ“œ Similar Papers

In the same crypt β€” Artificial Intelligence