AI IDE Leaderboard

Same tasks. Same machine. Real numbers. Updated monthly. Last run: coming soon.

First benchmark run is scheduled for the end of this month. We are testing:

Tools included in month 1: Claude Code (CLI), Aider, Cursor Agent (CLI). Cursor, Windsurf, Copilot manual results appended after human runs.

Methodology

Each tool runs the same task on a fresh workspace. We time the run and score with a deterministic verify script (usually: run the produced code, diff against expected output). No cherry-picking. All raw runs published as JSON.

Tools that only run inside a GUI IDE (Cursor, Windsurf, Copilot Chat) are timed manually by a human running the same prompt and appended to the same dataset.

Data license: CC-BY. Cite as "devtoolsniff AI IDE Leaderboard, coming soon".

Raw data

Download the full benchmark JSON: benchmarks.json

Suggest a task

Email hello@devtoolsniff.com with a task idea (must have deterministic pass/fail check).