Claude Opus 5.5 leads ThinkingBox, an agent benchmark graded on database state
TL;DR
- In a joint post on 3 October, Microsoft and Hugging Face said ThinkingBox, an agent benchmark of 507 business workflows, now runs through the OpenEnv interface.
- It runs every task 20 times and grades the database state each run leaves behind; Claude Opus 5.5 and Claude Opus 5 each pass 241 tasks on all 20 tries.
- The authors label 79.9% of failures as tool usage, and every number is theirs.
Read the full story
Sign in with your email to read AI News. It’s free.