Morning Brief: Tuesday, July 21

Sixty-five feeds. Two weeks. 4,596 items reduced to what follows. (what we track, how we crawl, subscribe)

Tuesday morning. Kimi K3 is the story of the week: The New Stack benchmark has it matching Claude Fable 5's coding results at one-third the cost, four times slower. Simon Willison spent the Monday afternoon posting a companion pair — Who's Afraid of Chinese Models? and Reverse-engineering is cheap now — and Interconnects framed it as the open-weights escalation. TechCrunch adds that Anthropic's $1.5B copyright settlement has been approved, and (in a separate piece) that OpenAI is scared of open-weight models. Security-wise, the WordPress-RCE thread has closed into a live incident: Slashdot and TechCrunch report hackers are exploiting recently patched WordPress bugs — the exact continuation of Friday's Cloudflare WAF disclosure and Sunday's SLCyber "$25 GPT-5.6 → $500k RCE" writeup, now on the wire as mass exploitation. Hugging Face confirms a breach affecting internal datasets and credentials. The Tuesday arXiv drop lands 454 items, heavy on agent infrastructure: Deterministic Replay for AI Agent Systems, PlanFlip: Planning-Phase Prompt Injection on Multi-Agent LLMs, Self-State Attacks on Self-Hosted AI Agents, Autonomous Agency Scale, A Diagnostic Framework for AI Agent Behavior, and — pointed at chip design — Can AI Agents Really Complete RTL-to-GDS? Aviation: Farnborough Day 1 lands with Boeing/SPEEA sounding optimistic.

Top (5-7 min)

Claude Fable 5 vs. Kimi K3: Same results, one-third the cost, 4x slower
The New Stack, 2026-07-20. New Stack benchmark: Kimi K3 matches Fable 5 coding results at ~1/3 the cost but ~4x latency. The concrete numbers behind the open-weights argument the rest of Monday debated in prose.
Kimi K3: The open-weights escalation
Interconnects, 2026-07-20. Nathan Lambert on Kimi K3 as an escalation, not a one-off. Pairs with the New Stack benchmark and Willison's two Monday posts.
Who's Afraid of Chinese Models?
Simon Willison, 2026-07-20. Willison's Monday-afternoon piece, published alongside Reverse- engineering is cheap now. Read as a pair.
Reverse-engineering is cheap now
Simon Willison, 2026-07-20. Companion post. The economic argument sitting under Sunday's SLCyber "$25 GPT-5.6 → WordPress RCE" story and today's arXiv Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents.
Anthropic's landmark $1.5B copyright settlement is approved
TechCrunch, 2026-07-20. The settlement approved. Sits with (separately) TechCrunch's OpenAI is scared of open-weight models. Should the US be? as Monday's two big first-party-affecting news items.
Hackers Are Exploiting Recently Patched WordPress Bugs, Putting Millions of Websites at Risk
Slashdot, 2026-07-21. Mass exploitation now live. Closes the Fri Cloudflare WAF → Sun SLCyber $25-GPT5.6-RCE → Mon Beyond Success Rate arXiv chain into an actual incident.
Hugging Face confirms breach affected internal datasets and credentials
TechCrunch, 2026-07-20. HF confirms internal datasets + credentials compromised, tells users to rotate. Load-bearing for anyone using HF-hosted secrets.
Deterministic Replay for AI Agent Systems
arXiv, 2026-07-21. Tuesday arXiv, headline agent-infrastructure paper. Replay as an agent-debugging primitive — the missing complement to Willison's Claude Code uses Bun in Rust observability thread from Sunday.
PlanFlip: Attacking Multi-Agent LLM Systems via Planning-Phase Prompt Injection
arXiv, 2026-07-21. Planning-phase attack surface for multi-agent LLM systems. Companion to Self-State Attacks on Self-Hosted AI Agents and Adaptive Adversaries, both in today's drop.
Can AI Agents Really Complete RTL-to-GDS? Lessons from Benchmarking Tool-Interactive EDA Workflows
arXiv, 2026-07-21. Coding-agent question aimed squarely at EDA / chip-design workflows. Companion to yesterday's Alipay-PIBench and ARC-AGI-3 world models papers.
Farnborough Air Show, Day 1
Leeham News, 2026-07-21. Farnborough opens. The aviation week's calendar tent-pole.

Themes this week

Kimi K3 / open-weights escalation
New Stack: Kimi K3 vs Fable 5 — same results, 1/3 cost, 4x slower (Mon), Interconnects: Kimi K3 open-weights escalation (Mon), Willison: Who's Afraid of Chinese Models? (Mon), Willison: Reverse-engineering is cheap now (Mon), TechCrunch: OpenAI is scared of open-weight models. Should the US be? (Mon), HN / Alibaba_Qwen: Qwen3.8 launching and going open-weight soon (Sun, carryover), HN / Qwen-Image-3.0 (Tue).
WordPress-RCE thread → live mass-exploitation incident
Slashdot: Hackers Are Exploiting Recently Patched WordPress Bugs (Tue), TechCrunch: Hackers exploiting recently patched WordPress bugs, millions at risk (Mon), SLCyber: WordPress RCE found with GPT-5.6 for $25 (Mon, carryover), Cloudflare: WAF protects WordPress from two high-severity vulnerabilities (Fri, carryover), arXiv: Beyond Success Rate — Cost-Aware Evaluation of Offensive and Defensive Security Agents (Mon, carryover).
Anthropic corporate & Fable arc
TechCrunch: Anthropic $1.5B copyright settlement approved (Mon), New Stack: Anthropic employees worked "literally around the clock" to keep Fable 5 from disappearing (Mon), Anthropic: Apply for AI for Science rare disease research grants (Mon), claude-code v2.1.216 (Mon).
Agent infrastructure & debugging (Tue arXiv)
arXiv: Deterministic Replay for AI Agent Systems (Tue), arXiv: A Diagnostic Framework for AI Agent Behavior (Tue), arXiv: Lomekwi — Resource-Bounded Tool Discovery in LLM Agents (Tue), arXiv: Verify, Repair, Repeat, or Stop? Robust Stopping for Noisy Verify-Repair Loops in LLM Agents (Tue), arXiv: RAIL Guard — Closing the Evaluation-to-Remediation Gap in Responsible AI for LLM Agents (Tue), arXiv: Composable Verification Pipelines for Multi-Agent Systems (Tue).
Agent-security attacks (Tue arXiv)
arXiv: PlanFlip — Attacking Multi-Agent LLM Systems via Planning-Phase Prompt Injection (Tue), arXiv: Self-State Attacks on Self-Hosted AI Agents — How Far Can OS Defenses Go? (Tue), arXiv: Adaptive Adversaries — A Multi-Turn, Multi-LLM Benchmark for LLM Agent Security (Tue), arXiv: Stress Testing Concept Erasure with Large Language Model Agents (Tue), arXiv: Dynamic Defense Profiling Enables Cognitive Jailbreak of Text-to-Image Models (Tue).
Coding agents & harness engineering (Tue arXiv)
arXiv: Can AI Agents Really Complete RTL-to-GDS? Lessons from Benchmarking Tool-Interactive EDA Workflows (Tue), arXiv: TRIM — Reducing AI-Generated CodeSlop via Agent Trajectory Minimization (Tue), arXiv: Harness Engineering for LLM-Driven GPU Kernel Generation (Tue), arXiv: DataFlow-Harness — A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines (Tue), arXiv: Automated Discovery Has No Universally Superior Harness (Tue), arXiv: A Systematic Evaluation of Trajectory Data Curation for LoRA Fine-Tuning of Code Agents (Tue).
Agent behavior measurement (Tue arXiv)
arXiv: The Autonomous Agency Scale — A Behavioral Framework for Measuring Self-Directed Behavior in AI Systems (Tue), arXiv: ProEvent — An Event-centric Benchmark for Proactive Agents (Tue), arXiv: Otap — Structure-Aware Optimal Transport for Evaluating Planning and Execution in Agent Trajectories (Tue), arXiv: RECON — Benchmarking Agent Memory for Compositional Reasoning over Long Contexts (Tue), arXiv: Retain or Consolidate? Budget-Dependent Operator Selection for Language Agent Memory (Tue).
First-party AI infra (Mon)
OpenAI: Safety and alignment in an era of long-horizon models (Mon), Cloudflare: Cloudflare Internal DNS is now generally available (Mon), Hugging Face: Introducing Cosmos 3 Edge (Mon), GitHub Blog: $100 million for open source (Mon), Databricks: AI Transparency — Governance, Explainability, and Data Practices (Mon).
Aviation (Farnborough week)
Leeham: Farnborough Air Show, Day 1 (Tue), Leeham: Boeing, SPEEA optimistic over contract talks, so far (Tue), Leeham: Pratt & Whitney adding CMC composite blades to next GTF engine (Mon, carryover), The Air Current: Boeing quietly scrapped a 777X over rework (Sun, carryover).

Scan (15 min)

Tail

  • The Kimi K3 story is not a one-off. Willison publishing two companion posts (Chinese-models + reverse-engineering-is-cheap) the same afternoon Interconnects publishes the open-weights escalation and New Stack publishes the coding-benchmark numbers is the shape of a genuinely coordinated week — not editorial coordination, but the same story landing on four independent wires within hours. Track OpenAI/Anthropic Q&A tomorrow.
  • WordPress-RCE reads three ways this morning depending on where you started: (a) Cloudflare Friday WAF post — a routine vulnerability disclosure; (b) Sunday SLCyber post — a proof of concept with striking cost asymmetry ($25 tokens vs $500k broker payout); (c) Slashdot/TechCrunch today — a live mass exploitation event. The economic argument arXiv's Beyond Success Rate formalises is now measured against reality, not hypothetically.
  • Hugging Face breach is the third HF-touched security story in a month. If you're using HF-hosted secrets, tokens, or private datasets in production, rotate today.
  • arxiv-cs-ai = 454 items today, a heavy Tuesday drop. Continuation of Monday's governance/audit cluster now joined by agent-infrastructure (Deterministic Replay, Diagnostic Framework, RAIL Guard, Lomekwi tool discovery) and agent-security-attacks (PlanFlip, Self-State Attacks, Adaptive Adversaries, Concept Erasure stress test, Cognitive Jailbreak of T2I). Coding-agent papers extend into chip design (RTL-to-GDS) and codeslop reduction (TRIM). Harness Engineering shows up as a phrase in three separate papers — worth watching.
  • Farnborough opens today; expect the full-week Leeham /Air Current cadence on OEMs, engines, and orders.
  • Nikita Tonsky's Monday Looking for work post continues to fanout through Planet Clojure. No first-party responses on the wire yet.

Feed silences (diagnostic)

  • Alignment Forum, Anthropic first-party (beyond the AI-for-Science grants), Apple ML Research, Charity Majors, Citizen Lab, EFF Deeplinks, FreeBSD Foundation, Grafana Labs, Hillel Wayne, Interconnects (beyond the Kimi K3 Monday post), Kenneth Payne, Martin Fowler, METR, Microsoft Research, Nature Machine Intelligence, Netflix Tech Blog, Neon, Northeastern events (beyond calendar), Pydantic, Quanta Magazine, Schneier on Security, Supabase, Tailscale, Terence Tao, Vercel, 404 Media: no fresh Tuesday-morning posts. Latent Space AINews for the day published as not much happened today — a real editorial signal, not silence.
  • James Bornholt: DNS/TLS errors continue (unchanged from prior days).

Build provenance

build: 2026-07-21 | crawler-sha: 5fe7ab8 (Walsh-Research/1.2, compliance v1.3) | feeds: 65 active (78 configured, incl. 12 corp-eng, 6 bitsavers, 5 generated) | items-considered: 4596 (14d, incl. 2736 arXiv) | warehouse: 27134 items | published: 11 | note: Tuesday morning. Top story: Kimi K3 vs Fable 5 open-weights escalation — New Stack coding benchmark (same results, 1/3 cost, 4x slower), Interconnects /open-weights escalation/ post, Willison companion pair (/Who's Afraid of Chinese Models?/ + /Reverse-engineering is cheap now/), TechCrunch /OpenAI is scared of open-weight models/. Anthropic's $1.5B copyright settlement approved. WordPress-RCE thread closes into a live mass-exploitation incident (Slashdot Tue + TechCrunch Mon), the exact continuation of Friday's Cloudflare WAF disclosure and Sunday's SLCyber $25-GPT5.6-RCE writeup. Hugging Face confirms breach affecting internal datasets and credentials. Tuesday arXiv drop: 454 items, heavy on agent infrastructure (Deterministic Replay, Diagnostic Framework, RAIL Guard, Lomekwi tool discovery, Composable Verification Pipelines, Verify-Repair-Repeat robust stopping), agent-security-attacks (PlanFlip planning-phase injection on multi-agent, Self-State Attacks on Self-Hosted AI Agents, Adaptive Adversaries multi-turn multi-LLM benchmark, Concept Erasure stress test, Cognitive Jailbreak of T2I), coding agents (RTL-to-GDS EDA workflows, TRIM codeslop reduction, Harness Engineering for GPU kernels, DataFlow-Harness, Automated Discovery Has No Universally Superior Harness), and agent-behavior measurement (Autonomous Agency Scale, ProEvent, Otap trajectory OT, RECON memory, Retain-or-Consolidate). Aviation: Farnborough Day 1 opens; Boeing/SPEEA optimistic on contract talks. First-party carryover (Mon): OpenAI /Safety and alignment in an era of long-horizon models/, Cloudflare Internal DNS GA, Anthropic AI-for-Science rare disease grants, GitHub $100M for open source, Databricks AI transparency piece, HF Cosmos 3 Edge. HN Tue: Linux kernel $ORIGIN support, Doom on custom CPU (viral), Qwen-Image-3.0, Jane Street Incremental library, VTubing worldwide, Ex Situ displaced cultural artifacts map.