116
Anyone can one-shot green tests now. The interesting part is how well you drive the agent.
kodwai gives you real, ticket-sized engineering challenges (a multi-tenant rate limiter, a CRDT text buffer, an idempotent ETL pipeline, a Raft-lite log) that you solve on your own machine with your own agent: Claude Code, Cursor or Codex. No browser sandbox.
How it works:
Pick a challenge at kodwai.com/challenges
Run npx @kodwai/cli@latest challenge <slug> and choose your agent
Solve it the way you actually work
Run npx @kodwai/cli@latest submit and get your score
The score is 0 to 100 across three axes. Direction: how you steer, verify and decompose. Outcome: what shipped, replayed and stress-tested. Lift: how far you beat a solo AI, not just that you passed. Every signal cites evidence from your transcript, commits and test runs.
Free for developers: 3 scored submissions on us, then bring your own Anthropic key. Public profile, leaderboard and a weekly league that resets every Monday. I read every piece of feedback myself, so tell me what feels off.
Built with