The model is no longer the whole intelligence

Inherent says its AI teammate outperformed Anthropic and OpenAI; Nvidia says the harness, not the AI model, is the real hero; VentureBeat says enterprises winning with AI agents are limiting what agents can do alone; and Vero turns repo-scale verification into a benchmark. The pattern is clear: the next advantage for product builders is not just picking the smartest model, but designing the scaffold that frames, supervises and proves the work.

·3 min read

TechCrunch

Inherent, founded by DeepMind alumni, says its AI ‘teammate’ just outperformed Anthropic and OpenAI at replicating research

Inherent, founded by DeepMind alumni, says its AI ‘teammate’ just outperformed Anthropic and OpenAI at replicating research.

techcrunch.com

The model is no longer the whole intelligence

Inherent’s claim about Faraday is deliberately awkward for the model leaderboard: its research “teammate” beat Anthropic and OpenAI systems at replicating research.

That is the interesting part. TechCrunch reported that Inherent, founded by Google DeepMind alumni, says Faraday outperformed Anthropic and OpenAI at replicating scientific work. The obvious reading is “another startup claims a better agent”. I think the better reading is that the unit of competition is shifting.

The product is no longer the model call. The product is the scaffold around the model.

The scaffold is where judgement lives

Faraday is framed as a collaborative research teammate, which is a very different claim from “we have a smarter chatbot”. Research replication is not raw text generation. It involves choosing what matters, reconstructing methods, noticing missing assumptions and deciding when an answer is good enough to trust. Some of that can come from the base model. A lot of it comes from how the system is asked to work.

That is why TechCrunch’s Nvidia story lands so neatly beside Inherent. Nvidia’s argument points at the harness: tools, memory, rules and supervisory components that turn a model into an agent capable of longer-running work. The wrapper is not cosmetic. It decides what the agent remembers, what it can touch, when it retries, when it stops and how its work gets checked.

This should feel familiar to anyone who has built production software. The clever algorithm matters, but the real system lives in the queues, permissions, logs, fallbacks, tests and human approval paths. A beautiful model with a bad harness is a sports car with no brakes.

The enterprise pattern is even less romantic. VentureBeat reported that companies getting value from AI agents are limiting how much those agents can do by themselves. That cuts against the demo culture of “look how autonomous this thing is”. In production, autonomy is not an unalloyed good. It is a liability unless bounded by role design, auditability, legal review, compliance and approval flows.

The funny thing is that this makes agents sound less futuristic and more like organisations. A good company does not give every employee root access to the bank account, the legal department and production infrastructure on day one. It defines roles, escalation paths and controls. Ronald Coase’s old question, why firms exist instead of pure market transactions, was partly about the cost of coordination. Agents raise the same question inside software: when is it cheaper to let the system act, and when does supervision pay for itself?

Vero pushes this logic into one of coding’s hardest corners. AGI Hunt reported on Vero, a benchmark for repository-scale formal verification by AI agents. It targets work across whole repositories, where agentic coding claims have often been strongest in demos and weakest where real software hurts: cross-module reasoning, specifications and changes that must be proven correct rather than merely plausible.

This is the shape of the next AI product cycle. The winning teams will not simply ask, “Which model is smartest this week?” They will ask: what can the agent see, what can it change, what counts as evidence, who signs off, and how do we know it did the work?

The model still matters. But the advantage is moving to the contract around it. Builders who learn to design that contract will make their agents look smarter than the model rankings imply.


Read the original on TechCrunch

techcrunch.com

Stay up to date

Get notified when I publish something new, and unsubscribe at any time.

More news