Measuring frontier coding agents on original, long-horizon engineering tasks
Categories & Topics
Repository Stats
Related Tools
agents-md
Drop-in AGENTS.md that makes every coding agent behave like a senior engineer instead of an eager intern. Kills sycophancy, stops drive-by refactors, forces verification loops. Synthesizes Karpathy's four principles and Boris Cherny's Claude Code workflow. Works with Claude Code, Codex, Gemini CLI, Cursor, and the open standard.
AlayaWorld
Full-stack open-source interactive long-horizon world model.
app
The GitHub Copilot app is an agent-native desktop experience for finding, running, steering, and landing software work across your GitHub repositories.
auto-harness
Bring your own agent and build a self-improving agentic system. Automatically mine failures, optimize the agent harness, and gate against regressions.
bullshit-benchmark
BullshitBench evaluates how AI models respond to nonsensical prompts, determining whether they challenge the prompts instead of providing confident answers. This helps improve the reliability and understanding of AI responses in ambiguous situations.
cmux
Open source Ghostty-based macOS terminal with vertical tabs and notifications for AI coding agents. Built for multitasking, organization, and programmability.