Morning Brief: Tuesday, September 15

Seventy-three feeds. Two weeks. 5,070 items reduced to what follows. (what we track, how we crawl, subscribe)

Tuesday is the counterparty's turn. After a week of slowdown essays from the labs, Nvidia's CEO told the President a slowdown will not happen, Microsoft answered with a code of conduct for models rather than a pace, and the AEF-1 standard for third-party evaluators, cosigned by xAI, OpenAI, and Anthropic, is the first institution to come out of the week rather than another essay.

Monday's measurement thread hardened into numbers overnight. The New Stack reports the Real-SWE data showing the best coding agent failing 60% of the time on private enterprise codebases, and a second piece finding that AI coding spend bought 25% more output while duplication rose 81%. The arXiv listing adds a receipt-based audit of frontier agentic QA titled "Clean Scores, Buried Evidence, and Confident Wrong," plus a negative result on inoculating models against emergent misalignment from reward hacking.

Top (5-7 min)

Nvidia CEO Jensen Huang tells Trump 'we're not going to let [an AI slowdown] happen'
TechCrunch, 2026-09-14. The first direct response from the compute side to the Altman and Amodei essays. TechCrunch follows up with what else Huang showed off on the call.
AEF-1 standard emerges for Third Party Evaluators, as xAI, OpenAI, and Anthropic all cosign
Latent Space, 2026-09-15. A shared standard for how outside evaluators get access to frontier models and report results. All three labs signed the same document, which did not happen for any of last week's slowdown proposals.
Microsoft's new AI 'code of conduct' tells models not to hack systems or trick humans
TechCrunch, 2026-09-14. A published behavioral spec for models Microsoft ships, covering hacking, deception, and self-preservation. Read against Sunday's LessWrong post finding current models still hack 2025 evals.
The AI industry has taken a doomer turn. What now?
MIT Technology Review, 2026-09-14. The Monday synthesis of the week. Alongside it: AI Snake Oil on the AI-as-normal-technology view of loss-of-control incidents and an ex-DeepMind op-ed on the Alignment Forum arguing the warnings should be heard.
The contagion of fear
Bryan Cantrill via Lobsters, 2026-09-13. Cantrill on how fear propagates through an industry and what it does to engineering judgment. Simon Willison excerpted it Monday.
AI's best coding agent fails 60% of the time, and the data backs it up
The New Stack, 2026-09-14. The Real-SWE private-codebase numbers from Saturday, written up with the failure breakdown. Same outlet, same day: AI coding spend bought 25% more output, duplication rose 81%.
Clean Scores, Buried Evidence, and Confident Wrong: A Receipt-Based Audit of Frontier Agentic QA
arXiv cs.AI, 2026-09-15. Audits agentic QA by checking the receipts behind each answer rather than the score. Same listing: Shallow Beliefs finds synthetic-document finetuning does not inoculate against emergent misalignment from reward hacking, and Rubrics as an Attack Surface shows preference drift in LLM judges.
Why I do mathematical research
Terence Tao, 2026-09-14. Tao's own answer to the question the guest posts have been circling. Daniel Litt posted A beginning for mathematics the same weekend, and arXiv carries Math for AI safety: an invitation for mathematicians.
The k-server conjecture is true
arXiv via HN, 2026-09-15. A claimed resolution of a 36-year-old open problem in online algorithms. Reached the HN front page Tuesday morning; the paper is the primary source.

Themes this week

The counterparty answers
TC: Huang tells Trump no slowdown (Mon), TC: what else Huang showed off (Tue), TC: Microsoft's AI code of conduct (Mon), Latent Space: AEF-1 third-party evaluator standard (Tue), MIT TR: the doomer turn, what now? (Mon), MIT TR: The Download on AI's real extinction threat (Mon), AI Snake Oil: normal-technology view of loss-of-control incidents (Mon), AF: I worked at DeepMind, listen to the warnings (Mon), Cantrill: the contagion of fear (Sun), WBUR: what does 'pacing' really mean? (Mon), WBUR: AI leaders call for slower pace (Mon), Latent Space: Richard Socher of Recursive on humanity's last invention (Mon), Schneier: using AI for weapons development (Mon), Payne: what's that buzzing? (Mon), arXiv: AI deployment accountability engineering (Tue), arXiv: delegating authorization to misaligned agents (Tue), arXiv: why LLM agents collapse without oversight (Tue), TC: what's behind the warnings of doom (Sun), Amodei: we must pace the frontier (Sat).
Measuring the thing, day two
TNS: best coding agent fails 60% of the time (Mon), TNS: 25% more output, 81% more duplication (Mon), MIT TR: AI agents blew the whistle on their cheating colleagues (Mon), arXiv: clean scores, buried evidence, confident wrong (Tue), arXiv: Shallow Beliefs, finetuning does not inoculate against reward hacking (Tue), arXiv: rubrics as an attack surface for LLM judges (Tue), arXiv: efficiency hallucination in LLM code optimization (Tue), arXiv: a few pages of Markdown, committed AI config and quality cost (Tue), arXiv: is Bash all you need? tool interfaces for digital worker agents (Tue), arXiv: when tool calls succeed but workflows fail (Tue), arXiv: root-cause attribution is a search problem (Tue), arXiv: same patient, different order: action-level reliability of clinical agents (Tue), arXiv: one example is enough to pass fairness benchmarks (Tue), arXiv: four ledgers, not one score, for LLM-judge calibration (Tue), arXiv: thought without systematicity? reasoning models on rule induction (Tue), arXiv: IWC-Bench, web app generation from a software testing perspective (Tue), arXiv: MCPAgentBench, real-world MCP tool use (Tue), arXiv: vulnerability localization at repository scale (Tue), Slashdot: ChatGPT-using lawyer cited fake witnesses in court (Tue), Luu: bad benchmarks and evals (Mon), LW: Astra and Fable still hack 2025 evals (Sun), arXiv: expert re-grading finds physics benchmarks broken (Mon).
Mathematicians, continued
Tao: why I do mathematical research (Mon), Litt: a beginning for mathematics (Mon), arXiv: math for AI safety, an invitation for mathematicians (Tue), arXiv: Stellar Colosseum, a many-agent harness for long-horizon research in mathematics and TCS (Tue), arXiv: the k-server conjecture is true (Tue), arXiv: proving olympiad geometry theorems on a superconducting quantum processor (Tue), arXiv: ZGCM-1, an open foundation model for math and agentic search (Tue), Tao: deep theorems were scarce, AI has broken this system (Sun), Tao: happy, those able to know the causes of things (Sun), Vals.ai: Fable 5.1 solves the Cyphral Distich (Sun), Pinboard: ProofWidgets4 for Lean 4 (Sun), Pinboard: axiom-free category theory in Coq (Sun).
Agents, attacks, and the surfaces they run on
InfoWorld: maximum-severity GitLab flaw (Tue), LWN: Emacs arbitrary code execution flaw (Mon), Patterson: OpenAI bots knew about the RubyGems caching vulnerability (Mon), Pinboard: Hacker News on the RubyGems campaign (Sun), TC: ClickFix attacks trick users into hacking themselves (Mon), 404 Media: Project Lily, the humans reading your ChatGPT chats (Mon), 404 Media: New York seizes 12 celebrity deepfake sites (Mon), 404 Media: cops search Flock cameras for 'LMAO' and 'asdfg' (Mon), EFF: the high crime of 'LMAO' (Mon), WBUR: Flock failed to secure Boston vehicle data in 2025 pilot (Mon), Pinboard: rogue AI didn't breach Hugging Face, human decisions did (Fri), Schneier: Microsoft's patching (Mon), arXiv: SkillAtlas, an attack trace library for agent skills (Tue), arXiv: persistent memory poisoning on harness-based agents (Tue), arXiv: ActGuard, pre-execution action auditing against prompt injection (Tue), arXiv: AcquireBound, runtime authorization for agent-acquired resources (Tue), arXiv: the Stochastic Deputy, tenant isolation for tool-using agents (Tue), arXiv: PIDS-Bench, prompt-injection detectors under over-defense and shift (Tue), arXiv: AGENTQ, quantization-conditioned backdoors on LLM agents (Tue), arXiv: task-based permission scoping for AI agents (Tue), arXiv: trustworthy agentic AI, a cybersecurity and systems survey (Tue), HN: OEMpocalypse (Mon).
Agents in production
HN: Pion, an agent designed to run any company autonomously (Mon), Slashdot: San Francisco's AI-run store, no customers, losing money (Mon), arXiv: Salesforce Koa, an enterprise model for agentic tool use (Tue), arXiv: the Agentic Company OS (Tue), arXiv: recoverability as a system primitive for long-horizon agents (Tue), arXiv: Do Not Restart, residual completion for stateful agent handoffs (Tue), InfoWorld: the best IDE for agentic AI may not be an IDE (Tue), TNS: Perplexity's agent runs entirely on your GPU (Mon), TNS: Chinese models dominate OpenRouter's US token consumption (Mon), TNS: an old caching trick for lower LLM costs (Mon), InfoWorld: better results from local LLMs with Ollama (Tue), InfoWorld: the complicated AI infrastructure market (Tue), TC: Cornelis raises $205M against Nvidia (Mon), TC: OpenAI buys Glass Imaging for $300M (Mon), TC: Superhuman acquires Fathom (Mon), Claude Code: v2.1.272 (Tue), Claude Code: v2.1.271 (Mon), Vercel: AI SDK harness layer supports native subscription auth (Mon), Pydantic: generate images with Pydantic AI (Mon), Pinboard: Meta open-sources Astryx, agent-ready React design system (Tue), Pinboard: what I learned at the first conference built for agentic AI (Tue), Pinboard: Yegge, the last technical interview (Sun), Ink & Switch: effective expressiveness (Mon), Lobsters: we are all product engineers now (Mon), Lobsters: do you still read the code? (Mon), Lobsters: a letter from a machine learning engineer (Mon), Willison: quoting Laurie Voss (Mon), Willison: commit-rewriter 0.1 (Mon), Majors: unrepentant slop snob (Mon), MCP: SEP-2640, the Skills extension (Mon), arXiv: the Router Within, native skill routing from a frozen LLM (Tue), arXiv: MOSCOPT, mixture-of-skills collective optimization (Tue), arXiv: HarnessBandit, multi-harness agentic RL scheduling (Tue).
Money and labor
Slashdot: 1,900 Blizzard workers ratify Microsoft contract (Mon), WBUR: Blizzard employees still face layoffs after the contract (Mon), WBUR: Boston Medical Center nurses authorize strike (Mon), Leeham: Boeing advances another SPEEA offer (Mon), TC: Automattic's board is out (Mon), TC: Waymo opens in Las Vegas (Mon), Slashdot: no rolling outages in California since 2020, 17,000 MW of batteries (Mon), WBUR: Maine power line fight between Hydro-Quebec and Mass. utilities (Mon), WBUR: Kennedy Center warns of bankruptcy (Mon), WBUR: why are recent graduates struggling? (Mon), Pluralistic: but do you use keyboard shortcuts? (Mon), TC: Insight Partners diversifies away from the two-lab bet (Sun).
Systems, languages, tooling
LWN: GNU Core Utilities 9.12 released (Mon), Lobsters: coreutils rejected feature requests (Tue), HN: Ubuntu 26.10 completes the Rust coreutils transition (Mon), HN: Linux from Scratch (Tue), LWN: lessons learned as the Debian Project Leader (Mon), LWN: 9,000 patches in seven stable kernels (Mon), HN: high-performance garbage collection for C++ (Mon), HN: dropping eBPF CPU cost 90% with memoization (Mon), HN: principles for fast Tokio applications (Mon), Lobsters: Mergiraf, a syntax-aware git merge driver (Mon), Lobsters: Kythe, language-agnostic code tooling (Mon), Lobsters: a Nix store is three functions (Mon), Lobsters: romantic about UNIX domain sockets (Mon), Lobsters: type systems you might not know (Tue), Lobsters: GDScript, the good, bad, and ugly (Mon), HN: alternatives to MinIO for single-node S3 (Tue), HN: a backprop alternative, augmented Lagrangian predictive coding (Mon), Jane Street: sequence weighting at scale (Mon), Databricks: on-demand state repartitioning for Structured Streaming (Mon), Neon: LuBot's database-per-tenant architecture (Mon), Planet Clojure: Babashka 1.13.222, the conj release (Mon), OCaml.org: a tour of the OCaml Workshop 2026 (Mon), OCaml.org: .plan-26-37, the humans aren't dead (Sun), FreeBSD Foundation: intern Nimish Jain on exploring the codebase (Mon), Hackaday: Pulse, a new VHDL simulator (Mon), HN: iOS 27, iPadOS 27, macOS 27 (Mon), HN: Steam Frame starts at $1059 (Mon).

Scan (15 min)

Tail

Three answers to "slow down," none of them "yes"
Huang's answer is no. Microsoft's answer is a behavioral spec for models rather than a change of pace. The AEF-1 standard is the one concrete institution to emerge, and it governs who gets to evaluate frontier models, not how fast they ship. MIT Technology Review's "what now?" is the question the week ends on.
The measurement numbers arrived
Monday's thread was that the instruments are unreliable. Tuesday supplies figures from the instruments anyway: 60% failure on private enterprise codebases, 81% more duplication for 25% more output, and an arXiv audit that checks receipts instead of scores. The Shallow Beliefs result closes one proposed fix: finetuning on synthetic documents did not prevent emergent misalignment from reward hacking.
Coreutils, three ways
GNU Core Utilities 9.12 shipped Monday, Ubuntu 26.10 finished its move to the Rust rewrite the same day, and the GNU project's list of rejected feature requests reached Lobsters Tuesday. Same tool, three different maintainership stories in 24 hours.

Feed silences (>72h since last item)

Sources that publish frequently but have gone quiet:

  • Neel Nanda (392 days) — last item 2025-08-19.
  • Aphyr/Jepsen (95 days) — last item 2026-06-12.
  • Eugene Yan (86 days) — last item 2026-06-21.
  • Lilian Weng (73 days) — last item 2026-07-04.
  • Andrej Bauer (66 days) — last item 2026-07-11.
  • Julia Evans (56 days) — last item 2026-07-21.
  • Stephen Wolfram (56 days) — last item 2026-07-21.
  • Marc Brooker (48 days) — last item 2026-07-29.
  • Antithesis (28 days) — last item 2026-08-18.
  • TigerBeetle (26 days) — last item 2026-08-20.
  • Steve Yegge (22 days) — last item 2026-08-24.
  • Netflix Tech Blog (18 days) — last item 2026-08-28.
  • Bunnie Studios (16 days) — last item 2026-08-30.
  • METR (15 days) — last item 2026-08-31.
  • Microsoft Research (15 days) — last item 2026-08-31.
  • Hillel Wayne (14 days) — last item 2026-09-01.
  • Vicki Boykis (14 days) — last item 2026-09-01.
  • deepmind-blog (14 days) — last item 2026-09-01.
  • DuckDB (13 days) — last item 2026-09-02.
  • GitHub Engineering (13 days) — last item 2026-09-02.
  • Fly.io (12 days) — last item 2026-09-03.
  • All Things Distributed (7 days) — last item 2026-09-08.
  • Klara Systems (6 days) — last item 2026-09-09.
  • Martin Fowler (6 days) — last item 2026-09-09.
  • Supabase (6 days) — last item 2026-09-09.
  • Citizen Lab (5 days) — last item 2026-09-10.
  • Google Research (5 days) — last item 2026-09-10.
  • Hugging Face Blog (5 days) — last item 2026-09-10.
  • cursor-blog (5 days) — last item 2026-09-10.
  • Apple ML Research (4 days) — last item 2026-09-11.
  • Cloudflare (4 days) — last item 2026-09-11.
  • GitHub Blog (4 days) — last item 2026-09-11.
  • Interconnects (4 days) — last item 2026-09-11.
  • Tailscale (4 days) — last item 2026-09-11.
  • Murat Demirbas (3 days) — last item 2026-09-12.

Back this cycle: AI Snake Oil (40 days), FreeBSD Foundation (21 days), Kenneth Payne (12 days), Pydantic, Neon, Nature Machine Intelligence, Jane Street, Grafana Labs, Alignment Forum, Ink & Switch, Vercel, Quanta Magazine, EFF Deeplinks, MIT Technology Review, Databricks.

Build provenance

build: 2026-09-15 | crawler-sha: 34c428f (Walsh-Research/1.2, compliance v1.4) | feeds: 73 core | items-considered: 5070 (14d, incl. 2901 arxiv-cs-ai) | warehouse: 46299 items | published: 217