Can a coding task be graded fairly?
What I learned building fairtask, an agent that screens SWE-bench-style tasks for vague issues and tests that only accept one fix.
- agents
- evaluation
- swe-bench
Notes on coding agents, evaluation, and building products.