AI IDE Leaderboard
Same tasks. Same machine. Real numbers. Updated monthly. Last run: coming soon.
First benchmark run is scheduled for the end of this month. We are testing:
- FizzBuzz in TypeScript (baseline)
- Fastify REST API with in-memory store
- Callback → async/await refactor
Tools included in month 1: Claude Code (CLI), Aider, Cursor Agent (CLI). Cursor, Windsurf, Copilot manual results appended after human runs.
Methodology
Each tool runs the same task on a fresh workspace. We time the run and score with a deterministic verify script (usually: run the produced code, diff against expected output). No cherry-picking. All raw runs published as JSON.
Tools that only run inside a GUI IDE (Cursor, Windsurf, Copilot Chat) are timed manually by a human running the same prompt and appended to the same dataset.
Data license: CC-BY. Cite as "devtoolsniff AI IDE Leaderboard, coming soon".
Raw data
Download the full benchmark JSON: benchmarks.json
Suggest a task
Email hello@devtoolsniff.com with a task idea (must have deterministic pass/fail check).