← Insights Strategie & Adoptie 17 August 2026 8 min Written with AI assistance

The price column is the new leaderboard

Demographics, export controls and a discount rack at the frontier — five stories from this week that all turn out to be about the same thing: capability stopped being scarce, and something else became the constraint.

Ruben Horbach Ruben Horbach Co-founder

In short

  • Ageing economies automate to backfill missing workers, not to displace them.
  • Frontier 'low' tiers now overlap rival 'top' tiers, gutting the premium for score leadership.
  • Composition beats selection: cheap model for volume, expensive model for judgement.
  • Export controls pushed Chinese labs to open weights — now ~61% of top-model usage on OpenRouter.
  • Agents are non-human employees with credentials that HR never vetted.

I spent part of this week looking at a leaderboard where the seventh-best model beat the eighth and charged 25% less for the privilege. Fable 5 Low: 64.2% at $5.70 per task. Opus 4.8 Max: 63.8% at $7.59. In the score column that's a tie you'd need a magnifying glass to see. On an invoice at any real volume, it's a rout.

Three other things landed the same week. NBER published a paper on 6 July 2026 arguing that seven decades of falling fertility have reshaped labour markets in ways the automation debate mostly ignores. A Lawfare analysis dated 1 July 2026 put Chinese open-weight models at roughly 61% of usage among top models on OpenRouter, up from under 2% of token traffic in late 2024. And a fund that rode the AI trade from $9.3bn in March to about $45bn by early July gave almost all of it back within thirty days.

Here's the reconciliation question I kept circling. If capability is getting cheaper, more abundant and more evenly distributed at the same time — cheaper per task, spread across open weights, arriving faster each generation — why does everyone still argue about which model is smartest? What are you actually short of, once intelligence stops being the scarce input?

My answer this week: you're short of workers, short of governable identities, and short of a way to tell whether the numbers you're citing mean anything. Let me show the working.

1. The countries automating fastest are the ones running out of hands

Most automation arguments assume a fixed pile of work. Machines take a slice, humans keep the rest, and the fight is over the ratio. Demographics wrecks that picture before the AI even shows up.

NBER working paper w35401, "Baby Busts and Growth Booms," published 6 July 2026, traces what falling birth rates do to an economy over decades. Its framing is counterintuitive enough that I'd read the abstract twice: rather than treating lower fertility as a straightforward drag on growth, the authors find growth effects associated with it, across seven decades of declining global fertility. I'd hold the causal direction loosely — this is one paper and macro identification at this scale is hard — but the mechanism it describes is not in dispute. Fewer babies now means fewer working-age hands in twenty years, holding up more retirees.

Japan is the clean case. Its working-age population has been shrinking since 1995, from 87.3 million to 73.7 million in 2024, a 16% fall, and the OECD employment outlook projects a further 31% decline between 2023 and 2060 (via CNBC, May 2026). Against that, a projected shortfall of roughly 11 million workers. So when Japan Airlines began humanoid trials for baggage handling at Haneda with a progressive rollout across the airport (CNBC, 1 May 2026), the framing in the press release was not efficiency. It was tourism demand rising while the workforce shrinks.

The honest counter-case: this logic does not travel everywhere. Nigeria, Egypt, the Philippines and Indonesia have young, growing labour forces, and there the displacement worry is real in exactly the way the standard debate imagines. Automation-as-backfill is a rich-and-ageing story. It happens to describe Japan, Korea, Germany, Italy, China after 2030 — and most of the buyers writing the cheques.

So what. If you are planning workforce strategy for a Dutch, German or Japanese operation, you are probably solving the wrong problem. The 2030 constraint on your line is more likely to be nobody applying than too many people applying. That changes what you pilot, and it changes how you talk to your works council about why you're piloting it.

Token price stopped predicting job cost the moment models started differing in how many attempts they need.

2. The frontier grew a discount rack

Back to that leaderboard. The Fable/Opus pair is one instance of a pattern now visible across the whole board. Artificial Analysis, in its 29 May 2026 snapshot, lists GPT-5.6 Luna (low) at $0.01 per Intelligence Index task, tied with MiMo-V2.5 and Llama 4 Scout. On SWE-bench Verified, DeepSeek V3.2 resolves a task for roughly $0.028 — about 24 times cheaper than Claude Opus on the same measure (SSOJet, verified 7 Jun 2026). A 25% gap is the polite end of this distribution.

The mechanism is that the low tier of one frontier lab has climbed high enough to overlap the top tier of a rival. Vendors ship tiers to segment their own customers, and the tiering has now been overtaken by cross-vendor progress. Add open weights, where quality-to-price concentrates hardest — Kimi K3 (max) sits at Intelligence Index 60, the highest-ranked open-weights model in that same Artificial Analysis snapshot — and the premium you pay for the top of the score column gets thin.

Then the composition result from this week made it thinner. GLM 5.1 doing the work with Opus 4.7 behind it as an advisor scored 18 quality points at $368. Opus 4.7 running solo scored 14 points at $954. Better output, about $586 cheaper, same job. The lever moved from which model to how you arrange them: cheap model for volume, expensive model reserved for judgement.

My admission: these are single benchmark configurations, some of them run by parties with an interest in the answer, and I have not reproduced any of them. Treat the exact figures as directional. The shape is what I'd commit to, because three independent measurement efforts point the same way.

Call the thing you now need the Price-Per-Task Leaderboard. One column, denominated in your currency, per unit of work you actually ship — not per token, not per benchmark point. Token price stopped predicting job cost the moment models started differing in how many attempts they need. If your architecture review this quarter still ranks models by intelligence score, you are optimising the column that no longer touches your margin.

3. Export controls built the thing they were meant to prevent

The logic behind US chip restrictions was clean: withhold the best hardware, withhold access to the leading closed models, and China's frontier slows down. An analysis in Lawfare on 1 July 2026 reads the outcome differently, and its central number is the one I keep repeating to people. Chinese open-weight models went from under 2% of OpenRouter token traffic in late 2024 to roughly 61% of usage among top models in 2026.

The mechanism is a squeeze producing an adaptation. Denied the biggest chips, Chinese labs optimised for efficiency rather than raw scale. Denied the closed US stack to build on, they released their own models with open weights — DeepSeek, Qwen, MiniMax, Kimi and GLM, all now described in that Lawfare piece as at or near the frontier. On 16 July 2026, Moonshot published Kimi K3: mixture-of-experts, around 2.8 trillion parameters, natively multimodal, roughly a million tokens of context, open (Trending Topics, July 2026).

Open weights are the part policy has no handle on. Once a file is out, it gets copied, fine-tuned and rebuilt into other people's products, and there is no switch to flip. The restriction landed on the one branch of the ecosystem it could not reach afterwards.

Two things complicate the tidy irony, and I'd want them on the record. First, we cannot know the counterfactual: perhaps Chinese labs would have gone open anyway, and controls only changed the timing. Second, both sides are now converging on the same instinct. China's MOFCOM is consulting Alibaba, ByteDance and Z.ai (formerly Zhipu) on export controls for its own model weights (TechTimes, 22 Jul 2026), and on 12 June 2026 the US briefly controlled a frontier model itself — the model as the controlled item, not the chips — before removing it again (NCUSCR, Aug 2026).

So what. If your procurement policy has a line saying "no Chinese models," check whether it survives contact with your own dependency graph. A fine-tuned derivative of an open Chinese base can arrive inside a vendor's product with no label on it at all.

4. You hired a thousand workers with logins and nobody vetted them

Every agent you deploy authenticates, holds credentials, clicks through workflows and writes to live systems. Okta's AI security risk assessment lands on the part nobody wants: the identity and access controls guarding those logins were designed around humans — one person, one badge, an onboarding conversation, a manager, a leaving date.

Count what is already inside your building. Codex runs for millions of developers weekly. Claude turns up with its own Slack handle. Each one is a non-human employee reading and writing across the stack. HR vetted the people. The agents arrived through an expensed API key.

What makes this the sharpest story of the week is that the same capability curve is being scored from both ends. Irregular put GPT-5.6 Sol through offensive-security benchmarks, asking how good a frontier model is at attacking systems. Okta is measuring the mirror image: how easily the agents you deploy get compromised. As models get better at reasoning through a system, they get better at breaking in and at being broken into. Most benchmark categories are safe to win. This one is scored twice, and you lose the second scoring.

The surrounding data fits. Critical and High severity CVE disclosures held flat from 2022 through early 2026, then climbed steadily through July. On CyberGym, the benchmark for AI finding real bugs, a MAI-Cyber-1-Flash plus GPT-5.4 configuration hit 95.95%, roughly ten points clear of the field. I'd hold the causal link between those two facts loosely — disclosure counts respond to reporting incentives as well as to discovery — but the direction is consistent: better tooling is finding more flaws, and the same tooling is available to whoever wants to use them.

So what. The cheap capability from story two arrives as agents, and agents arrive as identities. Every euro you save on price-per-task shows up as a governance line item somewhere else. That is the trade this week keeps making, in one domain after another.

5. Money is now moving faster than the businesses it is chasing

One fund grew from $9.3bn in March 2026 to roughly $45bn by the start of July — about 5x in a quarter — and a month later the money had drained out again. The underlying boom is genuine: products ship, usage climbs, revenue grows. Capital is simply moving faster than any operating business can absorb it.

Two speeds are worth separating. Value velocity is how fast the underlying business compounds — customers, revenue, margin. Hype velocity is how fast money decides it believes that. Run together, you get a healthy market. When hype velocity pulls ahead, you get $36bn arriving and leaving inside four months while the companies underneath carry on at their own pace.

The mirror image sits in the same week's market data. IT services stocks are trading around 11x forward earnings by mid-2026, the sector's lowest multiple since roughly 2016. These are the companies whose product is rented human hours: outsourced developers, support desks, back-office teams billed by the seat. The market is pricing what happens when a lot of those hours get automated. Most of the jobs debate lives in surveys and forecasts about what might happen; this is investors putting a number on it and marking the businesses down.

My caution: an 11x multiple is a claim about the future, and markets have been early and wrong on labour substitution before — the same argument was made about offshoring in 2003. I'd read it as a well-informed opinion with money behind it, not evidence.

So what. Both moves are the market applying the Price-Per-Task Leaderboard at the level of whole companies. If your revenue is denominated in hours, someone is already modelling what those hours cost when a model does them at $0.028 a task.

What I'm watching

The interesting question for the rest of 2026 is which constraint bites first in your organisation. If you're in an ageing economy, the people shortage arrives before the displacement does. If you're deploying at volume, the price column arrives before the capability ceiling does. If you're running agents, identity arrives before either.

One pattern I'll be tracking specifically: whether price-per-task shows up as a standard column in enterprise procurement documents by year-end. Artificial Analysis publishes it. Aggregators like WhatLLM now say outright that for agentic workflows, cost per task filters better than token price. The gap between what the measurement community publishes and what buyers actually put in an RFP is usually about eighteen months. If that closes faster here, it will be because someone got a very large invoice.

If this is useful, subscribing keeps it turning up weekly — and if you disagree with the fund story, I'd genuinely like to hear why.

8 things I could not fit

  • 48% of 18–29 year-olds expect AI to have negative consequences — the highest of any age group, against 39%, 38% and 35% for older cohorts, while usage climbs in every group including theirs. Adoption and anxiety are rising in the same people, which breaks the assumption that familiarity fixes distrust.
  • Gemini 3.5 Pro holds 2 million tokens in one prompt, the largest context window in the frontier pack, and almost nobody is discussing it while the headlines chase reasoning scores. Context is the spec that decides the shape of the job: a whole codebase before writing a line, a full legal archive before answering.
  • Two Felony Bench runs named opposite winners two days apart — one put Anthropic on top at 7, another put OpenAI on top at 5. Safety leaderboards are starting to feed procurement and policy citations, and someone chose the prompts, the scoring and the cutoff for what counts as a felony.
  • Grid-scale batteries went from 5 MW in Oregon in 2012 to 267 GW globally, roughly 53,000x in thirteen years, with almost all of it after 2022 and 2025 the biggest year on record. Forecasters now put 2030 near 776 GW and up; the flat decade before 2020 was already doubling, just from numbers too small to see.
  • Unitree tops Asian humanoid valuations at around $6.2bn on roughly 5,500 robots shipped last year, of which about 3% do genuinely useful work. That valuation prices the actuator and supply-chain ecosystem behind it, not the current use case.
  • Z.ai (formerly Zhipu) open-sourced GLM-Image, which puts an autoregressive step of roughly 256 tokens in front of a diffusion decoder. That front half is why the text on the sign is spelled correctly — a design choice aimed at factual fidelity rather than aesthetics.
  • 4DV.ai rebuilds ordinary footage into walkable 4D scenes using Gaussian splatting, so you can circle a moment from angles nobody filmed. Every clip on your drives has been locked to one camera position; lifting that turns archives into places.
  • Shopify's Tobias Lütke has 39 MRI series of his own body, versioned like a codebase. Models can already read imaging; what they cannot read is a scan that does not exist, and almost nobody has a baseline of healthy-them on file to diff against.
Ruben Horbach

Ruben Horbach

Co-founder · Back From the Future

Ruben researches how organisations adopt AI meaningfully — not as technology, but as a change in work and people. He builds the agent infrastructure behind BFF and speaks about the near future of work.

Translate this to your situation?

Book a conversation — we're happy to think along about what this means for you.