SWE-Race: a coding-agent benchmark of 188 real concurrency bugs, with results from three models [P]
We've been building a benchmark out of real concurrency bugs (race conditions, deadlocks, cancellation issues) taken from merged PRs in about 100 Python projects. Each task gets graded by the project's own tests, in a co
📄
This source provides headlines only. Use the button below to read the complete article on the original site.
📰 Read the original article on r/MachineLearning
Originally published by r/MachineLearning. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.