The human role is migrating from execution to constraint arc...
March 5, 2026Pillar Report

AI & Agents

High Confidence5 signals / 18 sources
Orchestration quality and cost efficiency are overtaking raw model capability as the binding constraints for AI builders.

Signals

C

A researcher running self-hosted experiments via vLLM reports that Qwen3.5-35B-A3B, a mixture-of-experts model with only 3B active parameters, achieves 37.8% on SWE-bench Verified Hard and 67.0% on the full 500-task benchmark. The key intervention was adding a verify-after-every-edit nudge to the agent loop, which moved the model from 22% to 38% on hard tasks. Claude Opus 4.6 scores 40% on the same hard subset.

The structural signal is not that a small model is close to a frontier model — it is that agentic scaffolding (verify-after-edit loops) is becoming the dominant variable in coding benchmarks, not raw model scale. A 3B-active-param model closing a 2-point gap with a frontier model through workflow design means the moat for coding agents is shifting from model capability to orchestration quality.

B

Kief Morris, writing on martinfowler.com, proposes a three-position taxonomy for human roles in AI-assisted development: Humans Outside the Loop (vibe coding), Humans In the Loop (humans inspect each line), and Humans On the Loop (humans design and maintain the agent harness). Morris describes an Agentic Flywheel where agents analyze test results and operational data to recommend workflow improvements.

This is the first rigorous naming convention for the three postures teams are actually adopting. The framework reframes the human role from reviewer to architect-of-the-review-system. The On the Loop position is the control surface that separates teams who scale with AI from teams who bottleneck on it.

B

Stanford economist Nick Bloom, partnering with the Bank of England and Federal Reserve, surveyed 6,000+ senior executives across four countries. 69% of CEOs, CFOs, and senior executives use AI less than one hour per week. 28% are not using AI at all. Only 5% reported AI reducing headcount, and only 10% reported productivity improvements.

This is a legitimacy/telos erosion signal. Executive leadership is making capital allocation decisions and workforce predictions based on narrative consensus rather than operational experience. The say-do gap creates two risks: over-investment with no executive understanding to course-correct, and workforce anxiety driven by proclamations from people who have not internalized the technology.

B

The Pentagon designated Anthropic as a supply chain risk. Prediction markets dropped Anthropic valuation from ~$550B to $475B then rebounded to $550B. Claude jumped from #120 to #1 on the App Store.

This is coordination-cost inflation crystallizing. The government is learning that designating an AI company as a supply chain risk has asymmetric effects: it costs the government more than the company. Tech companies at sufficient scale become functionally un-sanctionable by their own government.

B

OpenAI released GPT-5.3 Instant and Google released Gemini 3.1 Flash-Lite on the same day, March 3. Both position as speed-and-cost-optimized variants of their respective flagship families.

The simultaneous release signals the frontier competition has bifurcated: capability race now runs parallel to a cost-efficiency race. The AI cost curve is collapsing from three directions — open-source MoE, Google Flash-Lite, and OpenAI Instant tier.

Control Surfaces

LeverStatusChangeEvidence
Cost per useful tokenDropping fastThree-front collapseGPT-5.3 Instant, Gemini 3.1 Flash-Lite, Qwen3.5 3B-active
Agent scaffolding ROIRising+16 pts on SWE-bench from verification loopQwen 22% to 38% with verify-after-edit
Executive AI fluencyCritically lowStable-to-declining relative to investment pace69% <1hr/week, Stanford/Bloom survey
Government AI leverageWeakeningFirst test case absorbed by markets in 48hrsPentagon-Anthropic designation, prediction market rebound

Watchlist

  • ConfirmationSmall MoE models achieving comparable results on production codebases
  • InvalidationMajor enterprise pulling back AI investment citing executive inability to evaluate outcomes
  • ObservableOther governments attempting similar supply-chain designations

Falsifiers

  • A new frontier model release where raw capability creates a measurable gap that orchestration cannot close
  • Evidence that enterprise AI adoption is driven primarily by model intelligence rather than cost-per-token

Key Unknowns

  • Whether Qwen SWE-bench results replicate across diverse codebases
  • Actual revenue impact of the Pentagon-Anthropic designation
  • Whether executive say-do gap leads to budget pullbacks

Noise Filter

  • Offshore Wind Turbine Data Center— Pre-revenue startup, no operational data
  • Physics-simulated humanoids framework— Niche tooling, not structurally significant
  • GFlowNets for ray tracing— Domain-specific ML research
  • Vue Router 5— Framework release, not AI-related

Share this article

Get The Signal daily

Cross-domain structural analysis, delivered every morning.