Terminal coding agents catch 61% of their wrong answers and fix half of those
TL;DR
- Across ten terminal agents on TerminalBench 2.1, the authors find that agents check their work 99.53% of the time but detect only 61.43% of their wrong candidates and repair only 49.36% of the errors they detect.
- Their method, Student-Conditioned Verification Distillation, has a stronger model continue from the student's own candidate and teaches the student that checking; Pass@1 rises 9.74 to 16.85 points over the base models.
- On SWE-bench Verified the method holds or improves while ordinary distillation drops up to 25.87 points. The paper is the only source.
Read the full story
Sign in with your email to read AI News. It’s free.