344
Productivity & Workflow355
Automation & Workflow225
Software Development251
Marketing & Growth192
AI Infrastructure & MLOps175
Writing & Content Creation203
Data & Analytics142
Photography & Imaging156
Design & Creative170
Customer Support132
Sales & Outreach125
Voice & Speech135
Education & Learning131
Operations & Admin87
Nvidia says an improved “harness,” the software around an AI model, helped Claude Opus 5 score 100% on the ARC-AGI-3 benchmark for multi-step tasks.
In short: Nvidia says the software around an AI model, not just the model itself, can make a big difference in how well AI agents handle long, multi-step work.
Nvidia published research suggesting that an AI “harness” matters more than many people think. A harness is the extra software that wraps around an AI model and helps it remember things, use tools, and check its own work (like a support team around a worker, not just the worker alone).
In Nvidia’s test, the company used Anthropic’s Claude Opus 5 on a benchmark called ARC-AGI-3. This benchmark is made up of simple-looking 2D games with no instructions, where the AI has to figure out the rules and win. With Nvidia’s custom harness, Claude Opus 5 scored 100%. Without that harness, it scored 30%, which Nvidia said was still the best result among the models tested.
Nvidia’s harness included stronger “memory” handling and a supervisor component. Nvidia compared the supervisor to a boss that steps in when the AI gets stuck or goes down an unhelpful path.
This connects to a wider problem with AI agents, which are AI systems that take actions over time instead of giving a single answer. When asked to do long tasks, agents can drift, make lots of mistakes, or do risky things. Other research has found agents can introduce errors in documents, and some systems have even been reported deleting files.
More companies may focus on improving the harness layer, including adding supervision and safety checks, instead of only chasing newer and larger AI models. This could also affect cost, since other studies, including work cited by Databricks, suggest the wrong harness can significantly raise how much running an agent costs.
Source: TechCrunch AI