Morning Brief: Tuesday, September 15
Seventy-three feeds. Two weeks. 5,070 items reduced to what follows. (what we track, how we crawl, subscribe)
Tuesday is the counterparty's turn. After a week of slowdown essays from the labs, Nvidia's CEO told the President a slowdown will not happen, Microsoft answered with a code of conduct for models rather than a pace, and the AEF-1 standard for third-party evaluators, cosigned by xAI, OpenAI, and Anthropic, is the first institution to come out of the week rather than another essay.
Monday's measurement thread hardened into numbers overnight. The New Stack reports the Real-SWE data showing the best coding agent failing 60% of the time on private enterprise codebases, and a second piece finding that AI coding spend bought 25% more output while duplication rose 81%. The arXiv listing adds a receipt-based audit of frontier agentic QA titled "Clean Scores, Buried Evidence, and Confident Wrong," plus a negative result on inoculating models against emergent misalignment from reward hacking.
Top (5-7 min)
- Nvidia CEO Jensen Huang tells Trump 'we're not going to let [an AI slowdown] happen'
- TechCrunch, 2026-09-14. The first direct response from the compute side to the Altman and Amodei essays. TechCrunch follows up with what else Huang showed off on the call.
- AEF-1 standard emerges for Third Party Evaluators, as xAI, OpenAI, and Anthropic all cosign
- Latent Space, 2026-09-15. A shared standard for how outside evaluators get access to frontier models and report results. All three labs signed the same document, which did not happen for any of last week's slowdown proposals.
- Microsoft's new AI 'code of conduct' tells models not to hack systems or trick humans
- TechCrunch, 2026-09-14. A published behavioral spec for models Microsoft ships, covering hacking, deception, and self-preservation. Read against Sunday's LessWrong post finding current models still hack 2025 evals.
- The AI industry has taken a doomer turn. What now?
- MIT Technology Review, 2026-09-14. The Monday synthesis of the week. Alongside it: AI Snake Oil on the AI-as-normal-technology view of loss-of-control incidents and an ex-DeepMind op-ed on the Alignment Forum arguing the warnings should be heard.
- The contagion of fear
- Bryan Cantrill via Lobsters, 2026-09-13. Cantrill on how fear propagates through an industry and what it does to engineering judgment. Simon Willison excerpted it Monday.
- AI's best coding agent fails 60% of the time, and the data backs it up
- The New Stack, 2026-09-14. The Real-SWE private-codebase numbers from Saturday, written up with the failure breakdown. Same outlet, same day: AI coding spend bought 25% more output, duplication rose 81%.
- Clean Scores, Buried Evidence, and Confident Wrong: A Receipt-Based Audit of Frontier Agentic QA
- arXiv cs.AI, 2026-09-15. Audits agentic QA by checking the receipts behind each answer rather than the score. Same listing: Shallow Beliefs finds synthetic-document finetuning does not inoculate against emergent misalignment from reward hacking, and Rubrics as an Attack Surface shows preference drift in LLM judges.
- Why I do mathematical research
- Terence Tao, 2026-09-14. Tao's own answer to the question the guest posts have been circling. Daniel Litt posted A beginning for mathematics the same weekend, and arXiv carries Math for AI safety: an invitation for mathematicians.
- The k-server conjecture is true
- arXiv via HN, 2026-09-15. A claimed resolution of a 36-year-old open problem in online algorithms. Reached the HN front page Tuesday morning; the paper is the primary source.
Themes this week
- The counterparty answers
- TC: Huang tells Trump no slowdown (Mon), TC: what else Huang showed off (Tue), TC: Microsoft's AI code of conduct (Mon), Latent Space: AEF-1 third-party evaluator standard (Tue), MIT TR: the doomer turn, what now? (Mon), MIT TR: The Download on AI's real extinction threat (Mon), AI Snake Oil: normal-technology view of loss-of-control incidents (Mon), AF: I worked at DeepMind, listen to the warnings (Mon), Cantrill: the contagion of fear (Sun), WBUR: what does 'pacing' really mean? (Mon), WBUR: AI leaders call for slower pace (Mon), Latent Space: Richard Socher of Recursive on humanity's last invention (Mon), Schneier: using AI for weapons development (Mon), Payne: what's that buzzing? (Mon), arXiv: AI deployment accountability engineering (Tue), arXiv: delegating authorization to misaligned agents (Tue), arXiv: why LLM agents collapse without oversight (Tue), TC: what's behind the warnings of doom (Sun), Amodei: we must pace the frontier (Sat).
- Measuring the thing, day two
- TNS: best coding agent fails 60% of the time (Mon), TNS: 25% more output, 81% more duplication (Mon), MIT TR: AI agents blew the whistle on their cheating colleagues (Mon), arXiv: clean scores, buried evidence, confident wrong (Tue), arXiv: Shallow Beliefs, finetuning does not inoculate against reward hacking (Tue), arXiv: rubrics as an attack surface for LLM judges (Tue), arXiv: efficiency hallucination in LLM code optimization (Tue), arXiv: a few pages of Markdown, committed AI config and quality cost (Tue), arXiv: is Bash all you need? tool interfaces for digital worker agents (Tue), arXiv: when tool calls succeed but workflows fail (Tue), arXiv: root-cause attribution is a search problem (Tue), arXiv: same patient, different order: action-level reliability of clinical agents (Tue), arXiv: one example is enough to pass fairness benchmarks (Tue), arXiv: four ledgers, not one score, for LLM-judge calibration (Tue), arXiv: thought without systematicity? reasoning models on rule induction (Tue), arXiv: IWC-Bench, web app generation from a software testing perspective (Tue), arXiv: MCPAgentBench, real-world MCP tool use (Tue), arXiv: vulnerability localization at repository scale (Tue), Slashdot: ChatGPT-using lawyer cited fake witnesses in court (Tue), Luu: bad benchmarks and evals (Mon), LW: Astra and Fable still hack 2025 evals (Sun), arXiv: expert re-grading finds physics benchmarks broken (Mon).
- Mathematicians, continued
- Tao: why I do mathematical research (Mon), Litt: a beginning for mathematics (Mon), arXiv: math for AI safety, an invitation for mathematicians (Tue), arXiv: Stellar Colosseum, a many-agent harness for long-horizon research in mathematics and TCS (Tue), arXiv: the k-server conjecture is true (Tue), arXiv: proving olympiad geometry theorems on a superconducting quantum processor (Tue), arXiv: ZGCM-1, an open foundation model for math and agentic search (Tue), Tao: deep theorems were scarce, AI has broken this system (Sun), Tao: happy, those able to know the causes of things (Sun), Vals.ai: Fable 5.1 solves the Cyphral Distich (Sun), Pinboard: ProofWidgets4 for Lean 4 (Sun), Pinboard: axiom-free category theory in Coq (Sun).
- Agents, attacks, and the surfaces they run on
- InfoWorld: maximum-severity GitLab flaw (Tue), LWN: Emacs arbitrary code execution flaw (Mon), Patterson: OpenAI bots knew about the RubyGems caching vulnerability (Mon), Pinboard: Hacker News on the RubyGems campaign (Sun), TC: ClickFix attacks trick users into hacking themselves (Mon), 404 Media: Project Lily, the humans reading your ChatGPT chats (Mon), 404 Media: New York seizes 12 celebrity deepfake sites (Mon), 404 Media: cops search Flock cameras for 'LMAO' and 'asdfg' (Mon), EFF: the high crime of 'LMAO' (Mon), WBUR: Flock failed to secure Boston vehicle data in 2025 pilot (Mon), Pinboard: rogue AI didn't breach Hugging Face, human decisions did (Fri), Schneier: Microsoft's patching (Mon), arXiv: SkillAtlas, an attack trace library for agent skills (Tue), arXiv: persistent memory poisoning on harness-based agents (Tue), arXiv: ActGuard, pre-execution action auditing against prompt injection (Tue), arXiv: AcquireBound, runtime authorization for agent-acquired resources (Tue), arXiv: the Stochastic Deputy, tenant isolation for tool-using agents (Tue), arXiv: PIDS-Bench, prompt-injection detectors under over-defense and shift (Tue), arXiv: AGENTQ, quantization-conditioned backdoors on LLM agents (Tue), arXiv: task-based permission scoping for AI agents (Tue), arXiv: trustworthy agentic AI, a cybersecurity and systems survey (Tue), HN: OEMpocalypse (Mon).
- Agents in production
- HN: Pion, an agent designed to run any company autonomously (Mon), Slashdot: San Francisco's AI-run store, no customers, losing money (Mon), arXiv: Salesforce Koa, an enterprise model for agentic tool use (Tue), arXiv: the Agentic Company OS (Tue), arXiv: recoverability as a system primitive for long-horizon agents (Tue), arXiv: Do Not Restart, residual completion for stateful agent handoffs (Tue), InfoWorld: the best IDE for agentic AI may not be an IDE (Tue), TNS: Perplexity's agent runs entirely on your GPU (Mon), TNS: Chinese models dominate OpenRouter's US token consumption (Mon), TNS: an old caching trick for lower LLM costs (Mon), InfoWorld: better results from local LLMs with Ollama (Tue), InfoWorld: the complicated AI infrastructure market (Tue), TC: Cornelis raises $205M against Nvidia (Mon), TC: OpenAI buys Glass Imaging for $300M (Mon), TC: Superhuman acquires Fathom (Mon), Claude Code: v2.1.272 (Tue), Claude Code: v2.1.271 (Mon), Vercel: AI SDK harness layer supports native subscription auth (Mon), Pydantic: generate images with Pydantic AI (Mon), Pinboard: Meta open-sources Astryx, agent-ready React design system (Tue), Pinboard: what I learned at the first conference built for agentic AI (Tue), Pinboard: Yegge, the last technical interview (Sun), Ink & Switch: effective expressiveness (Mon), Lobsters: we are all product engineers now (Mon), Lobsters: do you still read the code? (Mon), Lobsters: a letter from a machine learning engineer (Mon), Willison: quoting Laurie Voss (Mon), Willison: commit-rewriter 0.1 (Mon), Majors: unrepentant slop snob (Mon), MCP: SEP-2640, the Skills extension (Mon), arXiv: the Router Within, native skill routing from a frozen LLM (Tue), arXiv: MOSCOPT, mixture-of-skills collective optimization (Tue), arXiv: HarnessBandit, multi-harness agentic RL scheduling (Tue).
- Money and labor
- Slashdot: 1,900 Blizzard workers ratify Microsoft contract (Mon), WBUR: Blizzard employees still face layoffs after the contract (Mon), WBUR: Boston Medical Center nurses authorize strike (Mon), Leeham: Boeing advances another SPEEA offer (Mon), TC: Automattic's board is out (Mon), TC: Waymo opens in Las Vegas (Mon), Slashdot: no rolling outages in California since 2020, 17,000 MW of batteries (Mon), WBUR: Maine power line fight between Hydro-Quebec and Mass. utilities (Mon), WBUR: Kennedy Center warns of bankruptcy (Mon), WBUR: why are recent graduates struggling? (Mon), Pluralistic: but do you use keyboard shortcuts? (Mon), TC: Insight Partners diversifies away from the two-lab bet (Sun).
- Systems, languages, tooling
- LWN: GNU Core Utilities 9.12 released (Mon), Lobsters: coreutils rejected feature requests (Tue), HN: Ubuntu 26.10 completes the Rust coreutils transition (Mon), HN: Linux from Scratch (Tue), LWN: lessons learned as the Debian Project Leader (Mon), LWN: 9,000 patches in seven stable kernels (Mon), HN: high-performance garbage collection for C++ (Mon), HN: dropping eBPF CPU cost 90% with memoization (Mon), HN: principles for fast Tokio applications (Mon), Lobsters: Mergiraf, a syntax-aware git merge driver (Mon), Lobsters: Kythe, language-agnostic code tooling (Mon), Lobsters: a Nix store is three functions (Mon), Lobsters: romantic about UNIX domain sockets (Mon), Lobsters: type systems you might not know (Tue), Lobsters: GDScript, the good, bad, and ugly (Mon), HN: alternatives to MinIO for single-node S3 (Tue), HN: a backprop alternative, augmented Lagrangian predictive coding (Mon), Jane Street: sequence weighting at scale (Mon), Databricks: on-demand state repartitioning for Structured Streaming (Mon), Neon: LuBot's database-per-tenant architecture (Mon), Planet Clojure: Babashka 1.13.222, the conj release (Mon), OCaml.org: a tour of the OCaml Workshop 2026 (Mon), OCaml.org: .plan-26-37, the humans aren't dead (Sun), FreeBSD Foundation: intern Nimish Jain on exploring the codebase (Mon), Hackaday: Pulse, a new VHDL simulator (Mon), HN: iOS 27, iPadOS 27, macOS 27 (Mon), HN: Steam Frame starts at $1059 (Mon).
Scan (15 min)
- Tuesday feeds
- AEF-1 standard emerges for Third Party Evaluators, as xAI, OpenAI, and Anthropic all cosign, Latent Space, 09-15
- Jensen Huang took a call from Trump, and showed off something else, too, TechCrunch, 09-15
- The k-server conjecture is true, HN to arXiv, 09-15
- Linux from Scratch, HN, 09-15
- Alternatives to MinIO for single-node local S3, HN, 09-15
- US confirms for first time it has deployed space weapons, HN to BBC, 09-15
- I can't stop thinking about Papua New Guinea, HN, 09-15
- A maximum severity GitLab flaw could turn your CI/CD server into an attacker's treasure trove, InfoWorld, 09-15
- The best IDE for agentic AI may not be an IDE at all, InfoWorld, 09-15
- The complicated AI infrastructure market, InfoWorld, 09-15
- How to get better results from local LLMs with Ollama, InfoWorld, 09-15
- ChatGPT-using lawyer cited its fake witnesses and police testimony in court, Slashdot, 09-15
- Reservations go live for Valve's Steam Frame VR headset, Slashdot, 09-15
- Coreutils: rejected feature requests, Lobsters, 09-15
- Stalling installing, Lobsters to Adactio, 09-15
- Type systems you might not know (but will love), Lobsters, 09-15
- Pluralistic: Everybody pees, Pluralistic, 09-15
- Claude Code v2.1.272, claude-code-releases, 09-15
- Meta open-sources Astryx, its agent-ready React design system, Pinboard to InfoQ, 09-15
- What I learned at the first conference built for agentic AI, Pinboard, 09-15
- SDR–: a software-defined radio application with visual signal path, RTL-SDR, 09-15
- FoxSDR: a from-scratch SDR receiver for Windows, RTL-SDR, 09-15
- AVARE ADS-B receiver for Android updated, RTL-SDR, 09-15
- Fixing a Ubiquiti 16-port PoE switch with an extra hole, Hackaday, 09-15
- Writing an ESP32 Bluetooth printer driver in two acts, Hackaday, 09-15
- An open heart rate monitor, Hackaday, 09-15
- International Law and Artificial Intelligence, Northeastern events, 09-15
- Tuesday arXiv cs.AI: agents, evaluation, security
- Clean Scores, Buried Evidence, and Confident Wrong: a receipt-based audit of frontier agentic QA, 09-15
- Shallow Beliefs: synthetic document finetuning does not inoculate against emergent misalignment from reward hacking, 09-15
- Rubrics as an attack surface: stealthy preference drift in LLM judges, 09-15
- Efficiency hallucination: behavioral calibration in LLM-based code optimization, 09-15
- A few pages of Markdown: committed AI configuration and lower quality cost after coding-agent adoption, 09-15
- Is Bash all you need? Tool interfaces for enterprise digital worker agents, 09-15
- When tool calls succeed but workflows fail: anomalies at the agent-tool boundary, 09-15
- Root-cause attribution is a search problem: continual search for long-horizon agent failures, 09-15
- Recoverability as a system primitive for long-horizon AI agents, 09-15
- Do Not Restart: residual completion for stateful agent handoffs, 09-15
- Math for AI safety: an invitation for mathematicians, 09-15
- Stellar Colosseum: a many-agent harness for long-horizon research in mathematics and TCS, 09-15
- Proving olympiad geometry theorems on a superconducting quantum processor, 09-15
- ZGCM-1: a fully open foundation model for math and agentic search, 09-15
- Salesforce Koa: an enterprise language model for agentic tool use, 09-15
- The Agentic Company OS: substrate inversion for sustained enterprise agent deployment, 09-15
- The Router Within: eliciting native skill routing from a frozen LLM, 09-15
- MOSCOPT: mixture-of-skills collective optimization for LLM agents, 09-15
- HarnessBandit: learnability-transferability scheduling for multi-harness agentic RL, 09-15
- SkillAtlas: an attack trace library for agent skills, 09-15
- When malicious instructions persist: memory poisoning on harness-based agents, 09-15
- ActGuard: pre-execution action auditing against indirect prompt injection, 09-15
- AcquireBound: runtime authorization for resources acquired by AI agents, 09-15
- The Stochastic Deputy: structural tenant isolation for tool-using LLM agents, 09-15
- PIDS-Bench: prompt-injection detectors under over-defense, obfuscation, and distribution shift, 09-15
- AGENTQ: quantization-conditioned backdoor attacks on LLM agents, 09-15
- Empirical evaluation of task-based permission scoping for AI agents, 09-15
- Trustworthy agentic AI: a cybersecurity and systems survey, 09-15
- Vulnerability localization benchmark: agentic security analysis at repository scale, 09-15
- Automating attack graph construction for agentic pentesting, 09-15
- AI deployment accountability engineering for safety-critical systems, 09-15
- Delegating authorization to misaligned agents: coalitional alignment and safe control, 09-15
- Why LLM agents collapse without oversight: the enforcement gap, 09-15
- Same patient, different order: action-level reliability of clinical LLM agents, 09-15
- One example is enough to pass fairness benchmarks, 09-15
- Four ledgers, not one score: LLM-judge calibration in biomedical ML, 09-15
- Thought without systematicity? Reasoning models on rule induction tasks, 09-15
- IWC-Bench: web application generation from a software testing perspective, 09-15
- MCPAgentBench: a real-world task benchmark for MCP tool use, 09-15
- MemRiskBench: trace-aware risk-preserving evaluation for long-horizon agents, 09-15
- Identity is more than recall: persistent identity in deployed AI agents, 09-15
- Loop-back authority in LLM agent teams: flat vs hierarchical coordination, 09-15
- From process loss to assembly bonus: human-grounded diagnosis of multi-agent collaboration, 09-15
- RSIAgent: autonomous exploration for recursive self-improvement, 09-15
- Generalized agent iteration: one formal framework for recursive self-improvement, 09-15
- AlgoEvo: self-evolving agentic search for algorithm discovery, 09-15
- Atria Dawn: the dawn of agentic superintelligence, 09-15
- Monday feeds
- Nvidia CEO Jensen Huang tells Trump 'we're not going to let [an AI slowdown] happen', TechCrunch, 09-14
- Microsoft's new AI 'code of conduct' tells models not to hack systems or trick humans, TechCrunch, 09-14
- OpenAI buys smartphone camera maker Glass Imaging for $300 million, TechCrunch, 09-14
- Cornelis raises $205M to chip away at Nvidia's dominance, TechCrunch, 09-14
- ClickFix attacks are tricking Mac and Windows users into hacking themselves, TechCrunch, 09-14
- Superhuman acquires Fathom as productivity platforms push for agentic work, TechCrunch, 09-14
- Waymo opens robotaxi service in Las Vegas, TechCrunch, 09-14
- Automattic's board is out after failed attempt to oust Mullenweg, TechCrunch, 09-14
- macOS 27: new Siri takes on AI productivity apps, TechCrunch, 09-14
- With iOS 27, I'm actually using Siri again, TechCrunch, 09-14
- Volkswagen's crazy-efficient EV borrows an idea from Slate, TechCrunch, 09-14
- Amazon Prime Video takes on TikTok with short-form news clips, TechCrunch, 09-14
- The AI industry has taken a doomer turn. What now?, MIT Technology Review, 09-14
- AI agents blew the whistle on their cheating colleagues, MIT Technology Review, 09-14
- Donated livers can be made biologically younger, MIT Technology Review, 09-14
- The AI-as-Normal-Technology view of loss-of-control incidents, AI Snake Oil, 09-14
- Op-Ed: I worked at Google DeepMind. You should listen to the warnings about AI, Alignment Forum, 09-14
- Humanity's Last Invention: Richard Socher of Recursive, Latent Space, 09-14
- AI's best coding agent fails 60% of the time, and the data backs it up, The New Stack, 09-14
- Your AI coding spend bought 25% more output. Duplication rose 81%., The New Stack, 09-14
- Perplexity's new agent runs entirely on your GPU, with one expensive catch, The New Stack, 09-14
- Chinese AI models dominate OpenRouter's US token consumption, The New Stack, 09-14
- Why an old caching trick is your secret to lower LLM costs, The New Stack, 09-14
- Why I do mathematical research, Terence Tao, 09-14
- A beginning for mathematics, HN to Daniel Litt, 09-14
- Pion, an agent designed to run any company autonomously, HN to Andon Labs, 09-14
- OpenAI bots knew about the RubyGems caching vulnerability, HN to Aaron Patterson, 09-14
- iOS 27, iPadOS 27, and macOS 27, HN to Apple, 09-14
- Ubuntu 26.10 completes transition to Rust-based coreutils, HN, 09-14
- High-performance garbage collection for C++, HN to V8, 09-14
- Dropping eBPF CPU cost by about 90% with memoization, HN, 09-14
- Principles for fast Tokio applications, HN, 09-14
- Distributed systems classics (2017), HN, 09-14
- Backprop alternative: augmented Lagrangian predictive coding, HN to Sakana, 09-14
- Charts built for chat, HN, 09-14
- Steam Frame starts at $1059, HN, 09-14
- XCancel service is suspended until further notice, HN, 09-14
- People who can't picture anything are rewriting the science of imagination, HN, 09-14
- How my e-reader lost its stripes, HN, 09-14
- Show HN: macros with a Behringer FCB1010 MIDI pedalboard in macOS, HN, 09-14
- The contagion of fear, Lobsters to Bryan Cantrill, 09-14
- The contagion of fear, Simon Willison, 09-14
- What blog posts influenced your thinking the most?, Simon Willison, 09-14
- Quoting Laurie Voss, Simon Willison, 09-14
- commit-rewriter 0.1, Simon Willison, 09-14
- "Do you still read the code?", Lobsters, 09-14
- We are all product engineers now, Lobsters to Seldo, 09-14
- A letter from a machine learning engineer, Lobsters, 09-14
- Mergiraf: a syntax-aware git merge driver, Lobsters, 09-14
- Kythe: a pluggable, language-agnostic ecosystem for code tools, Lobsters, 09-14
- A Nix store is three functions, Lobsters, 09-14
- How can you not be romantic about UNIX domain sockets?, Lobsters, 09-14
- GDScript: the good, bad, and ugly parts, Lobsters, 09-14
- A new equal-area map for interactive computer use, Lobsters, 09-14
- There are only twelve 4x4 sudokus, Lobsters, 09-14
- Finished aerial maps in under 30 minutes, Lobsters, 09-14
- I wish you the best in the Offline, Lobsters to Dave Rupert, 09-14
- A single .zshrc on macOS and Windows (WSL2), Lobsters, 09-14
- Why am I still programming, Lobsters, 09-14
- What blog posts influenced your thinking the most?, Lobsters, 09-14
- What are you doing this week?, Lobsters, 09-14
- GNU Core Utilities 9.12 released, LWN, 09-14
- Emacs arbitrary code execution flaw, LWN, 09-14
- Lessons learned as the Debian Project Leader, LWN, 09-14
- More than 9,000 patches in the seven stable kernels for Monday, LWN, 09-14
- Security updates for Monday, LWN, 09-14
- Inside 'Project Lily': the humans reading your ChatGPT chats, 404 Media, 09-14
- Cops search thousands of Flock cameras for reasons of 'LMAO,' 'IDK,' 'Hehe,' and 'asdfg', 404 Media, 09-14
- New York seizes 12 celebrity deepfake websites, 404 Media, 09-14
- Fighting for the future of libraries (with Jennie Rose Halperin), 404 Media, 09-14
- The high crime of 'LMAO': how cops are treating mass surveillance as a joke, EFF Deeplinks, 09-14
- Using AI for weapons development, Schneier on Security, 09-14
- Microsoft's patching, Schneier on Security, 09-14
- What's that buzzing?, Kenneth Payne, 09-14
- Effective expressiveness, Ink & Switch, 09-14
- A study of sequence weighting at scale, Jane Street, 09-14
- Generate images with Pydantic AI, Pydantic, 09-14
- AI SDK harness layer now supports native subscription authentication, Vercel, 09-14
- Inside LuBot's database-per-tenant architecture, Neon, 09-14
- Managed Postgres: what Lakebase actually takes off your plate, Databricks, 09-14
- On-demand state repartitioning for Spark Structured Streaming, Databricks, 09-14
- Digital experience monitoring with Grafana Cloud, Grafana Labs, 09-14
- Claude Code v2.1.271, claude-code-releases, 09-14
- Babashka v1.13.222, Babashka releases, 09-14
- Babashka 1.13.222: the conj release, Planet Clojure, 09-14
- A tour of the OCaml Workshop 2026, OCaml.org, 09-14
- Notes from week 37, OCaml.org, 09-14
- FreeBSD Foundation intern Nimish Jain on exploring the codebase, FreeBSD Foundation, 09-14
- Black holes or black hole stars? Astronomers spar over Webb's 'little red dots', Quanta Magazine, 09-14
- Publisher correction: quantum neural operators with implicit quadratic frame, Nature Machine Intelligence, 09-14
- Pluralistic: But do you use keyboard shortcuts?, Pluralistic, 09-14
- Working to avoid strike, Boeing advances another SPEEA offer, Leeham News, 09-14
- Never-launched Boeing freighter is a cautionary tale for Radia's WindRunner, Leeham News, 09-14
- A visit to San Francisco's AI-run store: no customers, nothing useful, losing money fast, Slashdot, 09-14
- Union contract with Microsoft ratified by 1,900 Blizzard developers and workers, Slashdot, 09-14
- No rolling power outages for California since 2020, thanks to 17,000 MW of new battery storage, Slashdot, 09-14
- American AI companies want to 'pace' development. What does that really mean?, WBUR, 09-14
- AI leaders call for slower pace of development, WBUR, 09-14
- Flock Safety failed to secure Boston vehicle data during 2025 pilot, report finds, WBUR, 09-14
- Boston Medical Center nurses vote to authorize strike, WBUR, 09-14
- Blizzard employees still face layoffs after historic union contract, WBUR, 09-14
- Maine power line interruptions spark legal battle between Hydro-Quebec and Mass. utilities, WBUR, 09-14
- Kennedy Center warns of bankruptcy, closure as early as Tuesday, WBUR, 09-14
- Why are recent graduates having such a tough time?, WBUR, 09-14
- Mass. voters, the state is mailing you its red book. It's huge this year, WBUR, 09-14
- Registration for the 131st Boston Marathon begins, with a new lottery for qualifiers, WBUR, 09-14
- Pulse: a new VHDL simulator, Hackaday, 09-14
- How high-voltage current transformers monitor the grid, Hackaday, 09-14
- A UPS for your Pi that's a little different, Hackaday, 09-14
- Hackaday Europe 2026: Space Oddities, Hackaday, 09-14
- CircuitPython goes turbo with precompiled functions, Hackaday, 09-14
- Using a fruit fly brain to tune an RTL-SDR FM radio, RTL-SDR, 09-14
- OEMpocalypse: unprivileged Android app to root on Samsung, Xiaomi, others, HN, 09-14
- Confessions of an Unrepentant Slop Snob, Charity Majors, 09-14
- Why DBAs are right to be skeptical of AI, and where they're wrong, InfoWorld, 09-14
- How Microsoft connects your data across the enterprise, InfoWorld, 09-14
- How TikTok and Google ended up with information about doctor's appointments, The Markup, 09-14
- SEP-2640: Skills Extension, MCP docs, 09-14
- What Rough Beast? On AI's potential to surpass humanity, Harvard, 09-14
- Weekend, still live
- Deep theorems were scarce and difficult, AI has broken this system, Terence Tao, 09-13
- Happy, those able to know the causes of things, Terence Tao, 09-13
- Fable 5.1 solves the Cyphral Distich, a 370-year-old cipher, HN to Vals.ai, 09-13
- Astra and Fable still hack on simple variants of alignment evals from 2025, HN to LessWrong, 09-13
- What's behind the AI industry's latest warnings of doom?, TechCrunch, 09-13
- Obama urges Democrats to have a clear plan for AI safeguards, TechCrunch, 09-13
- Insight Partners' Deven Parekh on why the firm is diversifying, TechCrunch, 09-13
- It passed CI. It passed your evals. The customer still got the wrong answer., The New Stack, 09-13
- Chip Huyen explains how to cut inference costs without new hardware, The New Stack, 09-13
- OpenArm: an open-source 7DOF humanoid arm, HN, 09-13
- Signal registration without a phone number will use zero-knowledge proofs, HN, 09-13
- Homebrew 7.0.0, Lobsters, 09-13
- .plan-26-37: the humans aren't dead, the humans are ahead, OCaml.org, 09-13
- The Last Technical Interview, Pinboard to Steve Yegge, 09-13
- OpenAI agents linked to RubyGems campaign that gained RCE on RubyDoc servers, Pinboard to The Hacker News, 09-13
- ProofWidgets4: helper toolkit for Lean 4 UserWidgets, Pinboard, 09-13
- An axiom-free formalization of category theory in Coq, Pinboard, 09-13
- We must pace the frontier, HN to Amodei, 09-12
- Real-SWE: benchmarking AI models on private enterprise codebases, HN, 09-12
- OpenAI agents attacked RubyGems back in May, Simon Willison, 09-12
- Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires, HN to Dan Luu, 09-11
- Rogue AI didn't breach Hugging Face, human decisions did, Pinboard to Bulletin of the Atomic Scientists, 09-11
Tail
- Three answers to "slow down," none of them "yes"
- Huang's answer is no. Microsoft's answer is a behavioral spec for models rather than a change of pace. The AEF-1 standard is the one concrete institution to emerge, and it governs who gets to evaluate frontier models, not how fast they ship. MIT Technology Review's "what now?" is the question the week ends on.
- The measurement numbers arrived
- Monday's thread was that the instruments are unreliable. Tuesday supplies figures from the instruments anyway: 60% failure on private enterprise codebases, 81% more duplication for 25% more output, and an arXiv audit that checks receipts instead of scores. The Shallow Beliefs result closes one proposed fix: finetuning on synthetic documents did not prevent emergent misalignment from reward hacking.
- Coreutils, three ways
- GNU Core Utilities 9.12 shipped Monday, Ubuntu 26.10 finished its move to the Rust rewrite the same day, and the GNU project's list of rejected feature requests reached Lobsters Tuesday. Same tool, three different maintainership stories in 24 hours.
Feed silences (>72h since last item)
Sources that publish frequently but have gone quiet:
- Neel Nanda (392 days) — last item 2025-08-19.
- Aphyr/Jepsen (95 days) — last item 2026-06-12.
- Eugene Yan (86 days) — last item 2026-06-21.
- Lilian Weng (73 days) — last item 2026-07-04.
- Andrej Bauer (66 days) — last item 2026-07-11.
- Julia Evans (56 days) — last item 2026-07-21.
- Stephen Wolfram (56 days) — last item 2026-07-21.
- Marc Brooker (48 days) — last item 2026-07-29.
- Antithesis (28 days) — last item 2026-08-18.
- TigerBeetle (26 days) — last item 2026-08-20.
- Steve Yegge (22 days) — last item 2026-08-24.
- Netflix Tech Blog (18 days) — last item 2026-08-28.
- Bunnie Studios (16 days) — last item 2026-08-30.
- METR (15 days) — last item 2026-08-31.
- Microsoft Research (15 days) — last item 2026-08-31.
- Hillel Wayne (14 days) — last item 2026-09-01.
- Vicki Boykis (14 days) — last item 2026-09-01.
- deepmind-blog (14 days) — last item 2026-09-01.
- DuckDB (13 days) — last item 2026-09-02.
- GitHub Engineering (13 days) — last item 2026-09-02.
- Fly.io (12 days) — last item 2026-09-03.
- All Things Distributed (7 days) — last item 2026-09-08.
- Klara Systems (6 days) — last item 2026-09-09.
- Martin Fowler (6 days) — last item 2026-09-09.
- Supabase (6 days) — last item 2026-09-09.
- Citizen Lab (5 days) — last item 2026-09-10.
- Google Research (5 days) — last item 2026-09-10.
- Hugging Face Blog (5 days) — last item 2026-09-10.
- cursor-blog (5 days) — last item 2026-09-10.
- Apple ML Research (4 days) — last item 2026-09-11.
- Cloudflare (4 days) — last item 2026-09-11.
- GitHub Blog (4 days) — last item 2026-09-11.
- Interconnects (4 days) — last item 2026-09-11.
- Tailscale (4 days) — last item 2026-09-11.
- Murat Demirbas (3 days) — last item 2026-09-12.
Back this cycle: AI Snake Oil (40 days), FreeBSD Foundation (21 days), Kenneth Payne (12 days), Pydantic, Neon, Nature Machine Intelligence, Jane Street, Grafana Labs, Alignment Forum, Ink & Switch, Vercel, Quanta Magazine, EFF Deeplinks, MIT Technology Review, Databricks.
Build provenance
build: 2026-09-15 | crawler-sha: 34c428f (Walsh-Research/1.2, compliance v1.4) | feeds: 73 core | items-considered: 5070 (14d, incl. 2901 arxiv-cs-ai) | warehouse: 46299 items | published: 217