Authority without verification is the structural risk of 202...
March 9, 2026Pillar Report

AI & Agents

Medium Confidence5 signals / 21 sources
The state is asserting control over AI deployment boundaries while the empirical case for AI-assisted productivity weakens, compressing builders between political and performance risk.

Signals

A

IEEE Spectrum reports (Mar 8) that a "simmering dispute between the DOD and Anthropic has now escalated into a full-blown confrontation." Defense Secretary Pete Hegseth reportedly gave Anthropic CEO Dario Amodei a deadline to allow DOD unrestricted use of Claude, including domestic surveillance and autonomous targeting applications. Anthropic refused, citing its Acceptable Use Policy. The Pentagon designated Anthropic a supply chain risk, effectively blacklisting it from defense contracts. Anthropic filed suit on March 9.

This is the first time a frontier AI lab has been punished by the U.S. government specifically for maintaining use restrictions. Either AI companies retain the right to set deployment boundaries, or the executive branch establishes that national security procurement can override company safety policies. Every other frontier lab is now watching to see if Anthropic's position is legally defensible or commercially suicidal.

B

The METR study (randomized controlled trial with experienced open-source developers working on their own repos) found AI tools made developers 19% slower on average. Developers predicted they would be 24% faster. A separate Anthropic RCT (52 developers, January 2026) found the AI-assisted group scored 17% lower on follow-up comprehension tests, with productivity gains "not statistically significant." METR updated on Feb 24, 2026, noting developers are "likely faster now" but evidence is "very weak."

Three data points now converge: AI tools slow experienced developers on familiar codebases, reduce comprehension for learners, and have coincided with a 67-73% collapse in junior hiring since 2022. The productivity narrative justifying $100B+ in AI infrastructure spend is under empirical pressure from its most favorable test scenario.

A

OpenAI announced (Mar 9) it is acquiring Promptfoo, an open-source AI security platform widely used for LLM evaluation, red-teaming, and prompt injection testing. No deal terms disclosed.

OpenAI is vertically integrating the security testing layer. Promptfoo was one of the few independent, widely-adopted tools for adversarial testing of AI systems. This acquisition removes a neutral evaluation tool from the ecosystem and places it under the control of the company whose models it most frequently tests.

C

A security audit of LlamaIndex revealed that deep retriever classes silently fall back to OpenAI's API when llm= or embed_model= parameters are omitted. Data intended to stay 100% local gets shipped to OpenAI without warning or error.

The default-to-cloud anti-pattern exists because OpenAI was the first API and became the implicit fallback across the Python AI ecosystem. For any organization building sovereign or air-gapped AI systems, every abstraction layer must be audited for cloud fallbacks. Trust boundaries in AI toolchains are not where developers think they are.

B

Andrej Karpathy released autoresearch (Mar 9), an open-source framework where AI agents autonomously conduct LLM training experiments. Architecture: one GPU, one file (train.py), one metric (val_bpb). Runs approximately 12 experiments per hour, yielding ~100 overnight on a single H100. MIT-licensed.

This is the clearest demonstration yet of AI agents doing the actual work of AI research — running the experimental loop, not generating papers. The structural question is whether this pattern scales: if 100 experiments per night per GPU becomes standard, the bottleneck moves from who can run experiments to who can interpret results.

Control Surfaces

LeverStatusChangeEvidence
Government-AI company power balanceEscalatingFirst supply chain blacklisting of a safety-principled labAnthropic designated as supply chain risk by Pentagon
AI coding productivity evidenceWeakeningTwo RCTs showing negative or null resultsMETR (-19%), Anthropic (-17% comprehension)
AI security tooling independenceConsolidatingKey independent tool acquired by audited companyOpenAI acquires Promptfoo
Local/sovereign AI stack reliabilityFragileSilent cloud fallbacks in major frameworksLlamaIndex OpenAI default exposure

Watchlist

  • ConfirmationOther frontier labs receive or resist similar Pentagon demands
  • InvalidationAnthropic wins injunction quickly and supply chain designation reversed
  • ObservableJunior developer hiring numbers Q1 2026; METR follow-up study timeline

Falsifiers

  • Anthropic lawsuit dismissed quickly or settled with Anthropic capitulating
  • A large-N RCT (500+ devs) shows significant AI productivity gains for experienced developers on familiar codebases
  • Other frontier labs publicly side with Anthropic, forming a unified industry position

Key Unknowns

  • Whether other frontier labs have quietly accepted unrestricted military use terms
  • METR's Feb 24 update says devs are 'likely faster now' — where is the new data?
  • Promptfoo acquisition terms and whether the open-source project continues independently

Noise Filter

  • Cross-Cultural Engineering Drives Tech Advancement— General career essay, no signal
  • Experiencing lack of motivation (r/ExperiencedDevs)— Individual career post, not structural
  • Do Offshore Wind Farms Pose National Security Risks?— Energy/defense, not AI-builder
  • My first setup for local AI / RTX 6000 build— Hardware shopping, no structural signal

Share this article

Get The Signal daily

Cross-domain structural analysis, delivered every morning.