Every domain is discovering that the infrastructure it assum...
March 12, 2026Pillar Report

AI & Agents

High Confidence5 signals / 49 sources
AI's production failures are now measurement failures — benchmarks, governance frameworks, and organizational metrics all lag what's actually breaking.

Signals

B

METR researchers had 4 active maintainers from scikit-learn, Sphinx, and pytest review 296 AI-generated pull requests that passed SWE-bench automated grading. Maintainers rejected approximately half. The measured gap between automated benchmark pass rate and actual maintainer merge decisions was 24 percentage points.

SWE-bench has become the primary mechanism by which AI labs claim coding competence and by which enterprises make adoption decisions. A 24pp gap means the benchmark is not measuring what decision-makers think it is.

B

The former backend lead at Manus published a post-mortem describing a shift away from function calling entirely. Core failure mode: LLMs under function calling produce tool invocations that are syntactically valid but semantically inconsistent at scale.

This represents a practitioner-level inversion of a core assumption baked into every major agent framework. If function calling degrades in precisely the multi-step contexts where agents are most valuable, the entire scaffolding layer may be misarchitected.

B

Anthropic's Claude Opus 4.6 introduces a Compaction API that automatically summarizes earlier conversation segments. 76% multi-needle retrieval accuracy at 1M tokens versus 18.5% for Sonnet 4.5. Maximum output doubled to 128K tokens.

Context compaction as a first-class API primitive changes the architectural calculus for long-running agents. Developers managing context through application-layer summarization now have a model-level solution.

B

Essay argues AI eliminates routine tasks that provided cognitive rest, replacing them with unrelenting orchestration and judgment. Firms treating AI gains as headcount reduction signals destroy institutional knowledge and create churn cycles.

If AI eliminates low-cognitive-load tasks and firms simultaneously reduce headcount, the result is an organization with less slack and higher cognitive demand per employee.

B

Dwarkesh Patel argues processing 100M U.S. CCTV cameras via AI currently costs ~$30B annually and will become trivially cheap. AI structurally favors centralized surveillance. The only barrier is political norm, not technical or economic constraint.

The asymmetry is structural and accelerating. The governing constraint on AI deployment is political will, and that is eroding in multiple jurisdictions simultaneously.

Control Surfaces

LeverStatusChangeEvidence
Benchmark validity for AI codingDegradingWorsening — 9.6pp/year gapMETR study
Context window managementImproving4x retrieval improvement at 1M tokensOpus 4.6 Compaction API
Agent framework primitivesContestedPractitioner pushback on function callingManus post-mortem
AI surveillance cost barrierDecliningCost curve acceleratingDwarkesh essay

Watchlist

  • ConfirmationSWE-bench gap replicated by larger study
  • InvalidationMultiple agent operators report function calling works fine at scale
  • ObservableAmazon mandatory AI systems meeting details surface publicly

Falsifiers

  • SWE-bench maintainer study gap closes with larger samples
  • Manus finding doesn't generalize across model sizes
  • Organizations successfully capture AI productivity gains without burnout dynamics

Noise Filter

  • IEEE Launches Global Virtual Career Fairs— Institutional marketing
  • Nvidia $26B open-weight AI investment— No architectural detail
  • AWS Strands Labs launch— Too early to signal

Share this article

Get The Signal daily

Cross-domain structural analysis, delivered every morning.