Claude Opus 5.5 leads ThinkingBox, an agent benchmark graded on database state

TL;DR

  • In a joint post on 3 October, Microsoft and Hugging Face said ThinkingBox, an agent benchmark of 507 business workflows, now runs through the OpenEnv interface.
  • It runs every task 20 times and grades the database state each run leaves behind; Claude Opus 5.5 and Claude Opus 5 each pass 241 tasks on all 20 tries.
  • The authors label 79.9% of failures as tool usage, and every number is theirs.

Read the full story

Sign in with your email to read AI News. It’s free.

Share this story

Explain like I’m 15