Capability diffused into a 17GB local file the same week three different measurement layers — model evaluation, enterprise adoption, sovereign commitment — visibly optimized for mimicry of absorption over absorption itself; Hormuz demonstrated what non-negotiable absorption actually looks like.

The Pattern
Alibaba released Qwen 3.6-27B this week. Simon Willison ran it locally at roughly 25 tokens per second on consumer hardware, a 17GB file outperforming the 807GB predecessor on coding benchmarks. Capability diffused through a 48x compression without losing quality. The same week, Nature published a peer-reviewed finding that accuracy-based evaluation of language models structurally incentivizes hallucination. Hedging gets penalized. Confident-wrong outperforms honest-uncertain on the benchmarks that decide which model ships. The measurement surface and the underlying capability are now visibly decoupling. I think the pattern to name is measurement-as-mimicry-incentive. When an evaluation layer cannot distinguish absorption from performed absorption, it will always select for the performance. Capability spreads. Measurement degrades. What the scoreboard says becomes a worse proxy for what the system actually does. It's about where a builder chooses to trust a score. The place a vendor hands you a metric is the place where their interests and yours come apart fastest. The score is downstream. Verify on your own traffic, your own outcomes, your own ground truth, or you will be managing to a fiction that gets sharper every quarter.
The Tension
Three measurement layers are pulling apart at once. At the model layer, The New Stack and GitHub issue #49244 document a 54-point benchmark drop and 35% token inflation on Opus 4.7, with no Anthropic response on a high-visibility thread. Customers pay more for less capability on the same upgrade cycle. At the org layer, Pragmatic Engineer documented tokenmaxxing: Meta logged 60.2 trillion tokens in 30 days, roughly $900M equivalent, as engineers burn spend against internal leaderboards with $100 Claude Code and $70 Cursor minimums. Shopify already reframed its leaderboard as circuit breakers. At the sovereign layer, the UAE Cabinet announced that 50% of government sectors will run on 'Agentic AI' within two years, with no operational definition, no budget, no milestones, and performance measured by 'speed of adoption, quality of implementation, mastery of AI.' The tension for a builder: every one of those measurements was designed to drive absorption. Each one now rewards the appearance of absorption. The trade-off is uncomfortable. Optimize for the score and you join the theater. Refuse the score and you lose the budget that depends on it.
What This Unlocks
Winners are the orgs defining outcomes before metrics. The top decile, the Anthropic-cadence teams and the Shopify circuit-breaker teams, treat model capability as an external dependency with a known improvement curve and ship under incompleteness. Lenny's interview with Cat Wu makes this explicit: ship before the model is ready, let the substrate catch up. Losers are the middle that mandated AI use before defining what transformation meant. Forrester reports 79% of enterprises facing adoption challenges, 29% seeing ROI, and 25% of 2026 AI spend deferred to 2027. That is not noise. That is tokenmaxxing's balance sheet. Stop building dashboards that measure activity proxies. Stop pinning production to specific frontier versions. Start building the outcome definition first and the metric second. Abstract model invocation behind a swappable layer so a regression in any one provider is a swap, not a rebuild. The difference between 'we use AI heavily' and 'AI is absorbed' is now financial, not cultural.
Watching Next
Three falsifiable observables. First, whether Anthropic issues an official response to GitHub issue #49244 with a regression acknowledgment or benchmark methodology rebuttal inside 30 days. Silence confirms the communication-maturation gap is structural. Second, whether UAE publishes sector milestones or RFP criteria that operationally define 'Agentic AI' within 90 days. No definition means declaration theater; a real definition upgrades the signal to commitment. Third, and this one is for your own business: pick one AI metric you report weekly, ask whether it measures activity or outcome, and sit with the answer. If it measures activity, you have a tokenmaxxing exposure whether your team is gaming the number or not. The measurement is already selecting for the wrong thing. The question is whether you can see it before the next review cycle does.
Underweighting
The strongest case against this frame is not that absorption is failing broadly. It's that I'm bundling three different mechanisms under one name. The Nature paper describes a training-gradient problem specific to LLMs. Tokenmaxxing is a principal-agent problem inside organizations. The UAE announcement is political theater that may have nothing to do with measurement at all. Calling them the same pattern is something I'm doing, not something the evidence is doing. The 'top decile escapes' claim is load-bearing for what I'm asking a builder to do, and the evidence is thin. Anthropic ships daily because Anthropic is an AI lab whose product is the model. That does not replicate outside lab conditions. Shopify is one company. Two data points is not a decile. The honest version: bifurcation might be the wrong model entirely. Enterprises corrected activity-proxied metrics after the dot-com cycle, after the cloud cycle, after the mobile-first cycle. The correction mechanism exists and has fired before. If it fires here, the asymmetric opportunity I'm pointing at closes before most readers can act on it, and the urgency in this essay is calibrated to a phase, not a durable condition.
Bottom Line
IEA called Hormuz 'the biggest energy security threat in history' this week, with 13 million barrels offline and 400 million barrels released as 'palliative not curative.' That is what non-negotiable absorption looks like. Everything else on the scoreboard is optional. Today, find one metric in your business you inherited without questioning, and ask whether it measures what happened or what got performed.
Sources
364 articles scanned / 54 sourcesGet The Signal daily
Cross-domain structural analysis, delivered every morning.