← Insights Strategie & Adoptie 24 August 2026 7 min Written with AI assistance

The scaffolding is the story

A benchmark that tripled without a new model, a lab reorganised around cloud, and a defender who finally got cheaper than the attacker — what a week of AI news looks like when you read it as plumbing.

Ruben Horbach Ruben Horbach Co-founder

In short

  • A benchmark score without its harness and cost column is closer to a rumour than a measurement.
  • ARC-AGI-3 results jumped orders of magnitude in four months with almost no change in the weights.
  • DeepMind's reorg shows org charts and P&L lines shape capability more than any single training run.
  • 48% calling AI a 'massive disappointment' looks like the bottom of a J-curve, not proof AI fails.
  • Relocating intelligence off the expendable part — interceptors, model pairings — flips the cost curve.

I spent an hour this week reading the same benchmark result twice and getting two different numbers.

Same model. Same test. GPT-5.6 Sol on ARC-AGI-3: 13.3% on the public set with the official harness, and roughly 38.5% when the harness kept its reasoning and compacted context between steps. OpenAI published that themselves, and their framing was not "a better model" — it was "how enabling two settings tripled our scores."

Three more facts from the same week. Demis Hassabis no longer runs DeepMind, and Google Cloud is growing over 100% year over year. Sweden is testing a drone interceptor with no cameras, no sensors and no onboard guidance. And 48% of enterprises describe their AI adoption as a "massive disappointment" (WRITER / Workplace Intelligence, 1 May 2026), in the same quarter that AI budgets hit 14.6% of total IT spend on average (Presenc AI, 18 March 2026).

Four stories, four industries, one shape. In each of them, the thing that changed is not the intelligence. It's the arrangement around the intelligence — where the reasoning is kept, where the sensor sits, where the money is booked, which curve you happen to be standing on. So the question I want to answer this week: if capability is roughly fixed on a Tuesday, why does the outcome move so much by Friday?

The benchmark that tripled without a new model

Start with ARC-AGI-3, because it's the cleanest case I've seen all year.

In March, the best frontier agent on this benchmark scored 0.51%. By July, the official harness had Sol Max at 13.33% public and 7.78% semi-private, while self-reported harness results ran as high as 58.12%, and one team claimed 98.98% (Schema harness writeup, 18 July 2026). Then on 5 August, Prime Intellect released Prime Agent under an MIT licence; paired with Claude Opus 5 it reported 95.5% RHAE Best@1, just past ARC's stated human-expert baseline of 95.4%, with 183 of 183 levels solved (Studio Global AI, mid-August 2026).

Two orders of magnitude of apparent progress in four months. Almost none of it in the weights.

What actually moved is state management: retained reasoning between steps, context compaction so the agent doesn't blow its window, retry policy. Boring engineering. And the harness is now open source, which means the plumbing is a commodity anyone can bolt on in an afternoon.

The admission I owe you here: I don't think this makes leaderboards worthless. It makes them under-specified. A number without its harness and its cost column is closer to a rumour than a measurement — and the cost column matters more than people admit. Continual Harness reported 20.54% at $774 per run (Seth Karten, 12 July 2026). That is a worse score than Sol's harness result at, I'd guess, several hundred times the spend. Neither number is wrong. They're answering different questions.

So what. If you are evaluating models by benchmark score, you are probably buying a scaffolding decision and calling it a model decision. The two years of leaderboard-reading habits most of us built are measuring the wrong layer. I'd hold the exact percentages loosely — self-reported harness results have no referee — but the direction is not in doubt, because OpenAI made the point against its own interest.

success and waste have the same silhouette for years, and the only thing that separates them is whether the learning compounds

The lab that got reorganised around the P&L

On 5 August, Google announced a full overhaul of DeepMind's leadership. The co-founder who built AlphaGo is out of the top job.

Read that next to one number: Google Cloud growing more than 100% year over year. SemiAnalysis put it bluntly — DeepMind's long-term failure is Google Cloud's short-term gain.

I'd hold that reading loosely. It is one analyst's framing of one reorganisation at one company, and leadership changes at large labs get over-read constantly. Hassabis is not gone from Google, and "reorganised" is not "defunded." But the pressure the framing describes is real and it is not specific to Google. Research that wins Go tournaments does not appear on a quarterly P&L. Cloud contracts do. When one line grows 100% a year and the other burns cash chasing a goal a decade out, the org chart eventually gets rearranged around the line that reports.

Here is why it belongs in this edition rather than in the morsels. It's the same move as the harness, one level up. The capability inside DeepMind did not change on 5 August. What changed is where it sits, who it reports to, and which metric it has to justify itself against. The scaffolding around the intelligence got rebuilt, and that will change what the intelligence produces over the next three years far more than any single training run will.

The counter-case worth holding: it's possible this is straightforwardly good. Research attached to a revenue engine gets deployed, and deployment is where most of the learning actually happens. Bell Labs was inside a monopoly's phone business. The lab-as-monastery model has produced brilliant papers and a lot of orphaned demos.

So what, for you: when you are trying to work out which lab's roadmap to plan around, the org chart tells you more than the paper output does. Watch where a lab's compute budget is approved. That's the harness.

The dip that looks exactly like failure

57% of companies use AI to write drafts, summaries, and brainstorms faster. That's Level 1: a thought partner sitting beside the existing process, making it slightly quicker. Safe, measurable, and — the part nobody says out loud — not a transformation.

Breadth is basically solved. McKinsey's Global AI Survey has 72% of enterprises running at least one AI workload in production as of Q1 2026, up from 55% in 2024 and 20% in 2020, with 83% of firms above 5,000 employees against 42% of smaller ones. Stanford's AI Index 2026 puts organisational use at 88%. Nobody is failing to adopt.

And yet: 48% of enterprises call their AI adoption a "massive disappointment," and 39% have no formal plan to turn AI tools into revenue (WRITER / Workplace Intelligence, 1 May 2026).

I don't read that 48% as evidence AI doesn't work. I read it as what the bottom of a J-curve feels like from the inside. Every project spends before it earns; the line dips below break-even and only later climbs. A company is dozens of those curves running at once, at different phases, and the P&L shows you the sum. So the individual shapes are invisible. When most of your AI portfolio is simultaneously in the dip — pilots, infrastructure, retraining, process redesign — the sum reads as pure cost for a long stretch, then stops reading that way.

The uncomfortable arithmetic: a firm rebuilding its process around AI trails a firm that merely bolted it on for close to seven years before overtaking it. For most of that window, the system builder looks like the loser. Which is why I'd put it this way — success and waste have the same silhouette for years, and the only thing that separates them is whether the learning compounds.

The honest counter: AI spend is now 14.6% of total IT budgets on average, above 22% in the top quartile (Presenc AI, 18 March 2026). At that level you cannot call it a pilot, and "we're in the dip" becomes a very convenient thing to say for five years running.

So what: ask which of your projects is producing learning that transfers to the next one. That's the difference between a J-curve and a hole.

The interceptor with nothing on board

Sweden is testing the Kreuger 100, a drone interceptor with no cameras, no sensors, no onboard guidance. Perception and guidance sit somewhere else. What flies is close to a dumb body: cheap, fast, expendable.

Air defence has spent a decade on the wrong side of the cost curve. You can win every engagement and still lose the exchange, because a $2,000 attack drone is being met by a missile costing several hundred times that. Strip the sensors and the compute out of the thing you throw away, and the defender's unit cost falls toward the attacker's.

Same move as the harness, wearing different clothes. The intelligence didn't get cheaper — it got relocated to somewhere it isn't destroyed on every use.

This pattern is showing up everywhere I look. It's the argument behind cheap humanoid bodies with the model in the cloud. It's what the top of CyberGym now looks like: MAI-Cyber-1-Flash paired with GPT-5.4 hitting 95.95%, roughly ten points clear of every single-model configuration, with one cheap model grinding attempts and one heavyweight deciding what to try next. Division of labour beat scale, and the pairing is something anyone can assemble in an afternoon.

The counter-case is genuine and I don't want to wave it away. Moving perception off the airframe means the airframe needs a link, and a link can be jammed, spoofed or saturated. You have traded unit cost for a dependency, and in contested electromagnetic environments that trade can go badly. Whether the Kreuger 100 survives contact with real jamming is exactly the thing a test programme is for, and I have not seen results.

So what: when a system's cost is dominated by the intelligence it carries, the first serious optimisation is to stop carrying it. That question — what does this thing need to hold locally, and what can sit elsewhere — is now a design question in weapons, robots, agents and software licences at the same time.

The through-line, named

Four stories, one mechanism. In each case the model, the researcher, the technology and the airframe stayed roughly constant, and the result changed anyway — because someone rearranged what surrounds them.

I'm going to call this the Harness Layer: the arrangement around a capability that determines how much of it you actually get. State management around a model. Reporting lines around a research lab. Project sequencing around a transformation. Sensor placement around a munition. It is unglamorous, it is copyable, and right now it is producing bigger swings than the capability underneath it.

The reason it matters commercially is that the Harness Layer moves at a completely different speed. Training a frontier model takes months and a nine-figure budget. Prime Agent shipped under MIT licence and closed most of a benchmark gap in a weekend of integration work. Anthropic's frontier is advancing at about +13.6 ECI per year — a slope you can read off a line and plan against. The harness advances in jumps nobody schedules.

Which gives you two dials, both now measurable: how fast the underlying capability improves per year, and how much of it your arrangement lets you collect today. Most organisations I talk to are budgeting hard for the first and paying almost no attention to the second.

What I'll be watching over the next month: whether any lab publishes a benchmark score with its harness and cost-per-attempt in the same table. OpenAI came close by admitting the two settings did the work. If that becomes normal, the leaderboard becomes useful again. If it doesn't, expect the gap between reported capability and delivered capability to keep widening — and expect the people who quietly master the plumbing to keep beating the people who bought the better model.

If this weekly signal-sorting is useful to you, the newsletter is where I do it slowly — you can subscribe from the top of this page.

Eight things I could not fit

China's chip gap is growing in the wrong category. Inferred demand rises from 4.17M units in 2025 to 11.47M in 2028, while SMIC and Hua Hong target a 30–40% increase in domestic wafer output by 2028 — at 28nm and 40nm (IndexBox, 2 May 2026). China can plausibly end up oversupplied in mature nodes and starved at the leading edge simultaneously. EUV, advanced deposition and high-end EDA remain blocked under the Oct/Dec 2024 US rules (Mordor Intelligence, 16 Jan 2026), and EDA is the hardest of the three to substitute.

The two strongest open-weights models in the world are both Chinese. DeepSeek V4 Flash 0731 debuts at roughly 153 ECI, behind only Kimi K3 at ~157. The same lab is lining up a public listing on close to $500M revenue while giving frontier-class weights away — which makes the premium on a closed frontier API a harder line to defend each quarter.

Readers can't tell AI writing from human writing. Two experiments put detection accuracy at 40% and 52% — one worse than chance — and readers rated the AI-written stories higher on quality. Grant Sanderson argues good writing requires theory of mind, which models lack. Both can't be true, and I suspect the bar we call "good writing" is lower than we'd like.

91% of GenAI unicorn market cap sits in the Bay Area. Capability diffuses at the speed of a download; equity value stays where the cap table was signed. Any sovereign AI strategy should be explicit about which of those two it's actually buying.

A frontier lab published an action-level autopsy of an agent-driven intrusion: 17,613 attacker actions, clustered into ~6,280 groups across 9 phases. A human red team never produces 17,000 moves against one target — people tire and take the three likeliest paths. The log size is the finding.

The college wage premium is one average hiding 64 majors, running from barely beating a high school diploma to far outrunning it. It lands awkwardly: AI is strongest at exactly the entry-level drafting and first-pass analysis many of those degrees were built to train.

19% of procurement teams apply game theory; 70% say they have only basic or below-basic understanding of it, and 65% expect its use to grow. Survivable when the other side of the table is also half-remembering the prisoner's dilemma. Less so when the counterparty is an agent, since finding an equilibrium is precisely what machines do well.

GPT-5.6 Sol appears three times in the same "which AI should you use" guide — twice in the top tier, once below it — separated only by the thinking-effort setting. Two people on identical subscriptions can be getting frontier-grade reasoning and mid-tier output, and neither knows which one they have.

Ruben Horbach

Ruben Horbach

Co-founder · Back From the Future

Ruben researches how organisations adopt AI meaningfully — not as technology, but as a change in work and people. He builds the agent infrastructure behind BFF and speaks about the near future of work.

Translate this to your situation?

Book a conversation — we're happy to think along about what this means for you.