GitGoodBench: A Novel Benchmark For Evaluating Agentic Performance On Git

May 28, 2025 Β· Declared Dead Β· πŸ› Proceedings of the 1st Workshop for Research on Agent Language Models (REALM 2025)

πŸ‘» CAUSE OF DEATH: Ghosted
No code link whatsoever

"No code URL or promise found in abstract"

Evidence collected by the PWNC Scanner

Authors Tobias Lindenbauer, Egor Bogomolov, Yaroslav Zharov arXiv ID 2505.22583 Category cs.SE: Software Engineering Cross-listed cs.AI Citations 1 Venue Proceedings of the 1st Workshop for Research on Agent Language Models (REALM 2025) Last Checked 5 months ago
Abstract
Benchmarks for Software Engineering (SE) AI agents, most notably SWE-bench, have catalyzed progress in programming capabilities of AI agents. However, they overlook critical developer workflows such as Version Control System (VCS) operations. To address this issue, we present GitGoodBench, a novel benchmark for evaluating AI agent performance on VCS tasks. GitGoodBench covers three core Git scenarios extracted from permissive open-source Python, Java, and Kotlin repositories. Our benchmark provides three datasets: a comprehensive evaluation suite (900 samples), a rapid prototyping version (120 samples), and a training corpus (17,469 samples). We establish baseline performance on the prototyping version of our benchmark using GPT-4o equipped with custom tools, achieving a 21.11% solve rate overall. We expect GitGoodBench to serve as a crucial stepping stone toward truly comprehensive SE agents that go beyond mere programming.
Community shame:
Not yet rated
Community Contributions

Found the code? Know the venue? Think something is wrong? Let us know!

πŸ“œ Similar Papers

In the same crypt β€” Software Engineering

Died the same way β€” πŸ‘» Ghosted