Nvidia Just Proved That the AI Model Is the Least Important Part of an Agent, So Why Are We Still Arguing About Models?

Nvidia Just Proved That the AI Model Is the Least Important Part of an Agent, So Why Are We Still Arguing About Models?

Industry Sentiment

Disruptive

What’s Happening at a Glance

  • Nvidia's AVO harness boosted Claude Opus 5's score on ARC-AGI-3 from 30% to 100%, solving all 183 levels.
  • The harness includes persistent memory, a supervisory agent, and tool integration, acting like a "CEO" for the AI.
  • The same architecture transferred from GPU kernel optimization to interactive reasoning, demonstrating generality.
  • Other labs, including OpenAI and Databricks, are also emphasizing harness design over raw model capability.

Summary

Nvidia published research showing that the software wrapper around an AI model, known as a harness, is often more critical than the model itself for long-horizon tasks. Using a custom harness called AVO, researchers got Claude Opus 5 to achieve a perfect 100% score on the ARC-AGI-3 benchmark, a set of 2D games with no instructions, where the model must figure out how to play and win. Without the harness, Opus 5 scored only 30%, which was the top result among all models tested. The AVO system includes persistent memory, a supervisory agent that nudges the main agent when it gets stuck, and a loop that can run for days, as demonstrated in a seven-day GPU kernel optimization run that produced kernels outperforming FlashAttention-4 by up to 10.5%. The same architecture was then applied to ARC-AGI-3 without modification, only swapping the task interface, and still achieved a perfect score.

The findings add to a growing body of evidence that the model is just one component of an effective AI agent. OpenAI, stung by low scores on ARC-AGI-3, conducted its own research and found that tweaking just two settings in its harness tripled its models' scores, though none reached 100%. Databricks CEO Ali Ghodsi noted that the harness can dramatically impact AI costs, with the wrong harness potentially doubling expenses. Nvidia's research underscores that open harnesses, like open models, give users more control and can drive up accuracy, contrasting with OpenAI's more closed approach.

Why This Is Happening

The AI industry has been locked in a "bigger model" arms race, but as models plateau in raw capability, the focus is shifting to system design. Long-horizon tasks – those requiring sustained decision-making over many steps – expose the limitations of a lone model, which can easily lose context, repeat mistakes, or veer off course. The harness provides the memory, tool access, and feedback loops needed to keep an agent on track. Nvidia, a hardware and infrastructure giant, is leveraging its position to promote an open agent stack, believing that controlling the full system – from harness to runtime – is key to secure and reliable AI. Meanwhile, competitors like OpenAI are playing catch-up, recognizing that their model-centric evaluation metrics like ARC-AGI-3 are incomplete without a robust harness. The pressure to deliver practical, autonomous agents for enterprise and consumer use is driving this shift, as businesses demand AI that can complete complex, multi-step workflows without constant human oversight.

Key Industry Impact

  • Big tech: Companies like Microsoft, Google, and Amazon will need to invest in harness infrastructure, not just model APIs, to offer competitive AI agents.
  • Startup ecosystem: Startups building agent frameworks and harness tools could see increased demand, especially those offering open, customizable solutions.
  • Developers: The role of the AI engineer will evolve to include harness design, memory management, and agent orchestration, creating new job categories.
  • Consumer technology: Future AI assistants may become more reliable for complex tasks like trip planning or software development, but only if the harness is well-designed.
  • Regulatory implications: As agents gain more autonomy, regulators may need to consider the security and accountability of the entire agent stack, not just the model.
  • Business competition: The ability to build effective harnesses could become a key differentiator, potentially shifting competitive advantage from model providers to system integrators.

Impact on People

  • Consumer experience: AI agents could become more capable of handling long, complex tasks, but their reliability will depend heavily on the underlying harness, which may vary by provider.
  • Privacy/data usage: Harnesses often involve extensive logging and memory storage, raising new privacy concerns about how agent interactions are recorded and used.
  • Employment/jobs: The shift toward agent systems may create demand for harness engineers and agent orchestrators, while potentially displacing some routine knowledge work that agents can now automate.
  • Pricing/costs: The harness can significantly affect computational costs; a poorly designed harness could make AI services more expensive, while an efficient one could reduce costs.
  • Accessibility: Open harnesses could democratize access to high-performance AI agents, allowing smaller teams to build competitive solutions without relying on proprietary models.
  • Daily life impact: As agents become more reliable, they could be integrated into more aspects of daily life, from managing personal finances to assisting with creative projects, but their performance will be inconsistent without standardized harness designs.

Key Technologies

  • AI: Large language models (LLMs) like Claude Opus 5 and GPT-5.6 Sol serve as the "brain" of agents.
  • Cloud: Agent harnesses often run on cloud infrastructure, leveraging scalable compute for long-running tasks.
  • Hardware: GPU kernels optimized by agents, such as those produced by AVO, improve hardware utilization.
  • Software: Harness frameworks include tools for memory management, tool integration, supervision, and feedback loops.
  • Cybersecurity: The agent stack must be secured against prompt injection, data exfiltration, and other attacks, especially as agents gain more autonomy.
  • Platforms: Open-source harness platforms like Nvidia's Nemo and agent frameworks like LangChain or AutoGen enable custom agent development.
  • Infrastructure: The underlying runtime and execution environment must support persistent state and recovery from failures.

Key Companies

  • Major corporations: Nvidia (AVO harness, Nemo platform), Anthropic (Claude models), OpenAI (GPT models, Codex), Databricks (agent cost research), Microsoft (Copilot, agent frameworks).
  • Startups: Companies building agent harness tools, such as LangChain, AutoGen, and others in the AI agent ecosystem.
  • Government agencies: Not directly involved, but may regulate agent autonomy and data usage.
  • Investors/partners: Venture capital firms investing in agent infrastructure, and corporate partners like Microsoft investing in AI agent integrations.

Future Outlook

Over the next few years, we can expect the focus in AI to shift from model size to system design, with harnesses becoming a key battleground for competition. Open-source harness frameworks will likely proliferate, enabling more teams to build reliable, long-horizon agents. Evaluation benchmarks will evolve to measure full agent systems, not just model outputs, making scores like ARC-AGI-3 more meaningful. However, the lack of standardization could lead to a fragmented landscape where agent performance varies widely depending on the harness. Regulation may eventually step in to set security and transparency standards for agent systems. Ultimately, the ability to build robust, autonomous agents could determine which companies lead the next wave of AI applications.