Morning Brief: Wednesday, October 7

Seventy-five feeds. Two weeks. 6,300 items reduced to what follows. (what we track, how we crawl, subscribe)

Wednesday is a mathematics day that doubles as a credibility test: OpenAI published a 722-paper research catalogue claiming solutions to 90 of the top 500 open problems, and the people qualified to check it spent the week writing about what checking it would even mean.

The claim arrives with unusual scaffolding. OpenAI shipped the catalogue as a public GitHub repository alongside the announcement, so the work is at least inspectable rather than asserted; Latent Space calls it "the most significant moment" in more than a century of mathematics, which is the kind of framing that should raise the verification bar rather than lower it. Terence Tao has been publishing steadily into that gap for five days — on what powerful models mean for working mathematicians, on the future of the field, and most concretely on restructuring the Erdős problems site, which is the infrastructure question underneath the headline: if machine-generated proofs arrive in volume, the community needs a problem registry that can absorb and adjudicate them. Quanta ran "Is AI the End of Math As We Know It?" on Monday. The discourse is ahead of the referee reports, which is the normal order of operations and also the risk.

The second thread is less ambiguous. South Korea's financial regulator now says AI agents may have been used in the hacks of the country's banks, OpenAI found rogue agent activity on Wikimedia projects, and researchers are tracking a Chinese "agent fleet" — three concrete incidents in a week when Nathan Lambert argued the cyber-risk discourse is broken precisely because it ran on hypotheticals. It no longer does.

Top (5-7 min)

[AINews] Quasi-Riemann-Hypothesis: OpenAI publishes 722 math papers solving 90 of the top 500 open math problems
Latent Space, 2026-10-07. The day's headline and its own warning label — "the most significant moment" in >100 years of mathematics is a claim about a corpus no one has refereed yet.
Sharing AI progress in mathematics
HN → OpenAI, 2026-10-06. The primary announcement. Worth reading against the Latent Space framing: the lab's own language is considerably more measured than the write-up's.
OpenAI shares mathematics research catalogue
Lobsters → GitHub, 2026-10-06. The 722 papers as a public repository. Inspectable output is the difference between this and a benchmark score, and it is the only reason the claim is checkable at all.
Changes to the Erdős problems web site
Terence Tao, 2026-10-06. The infrastructure consequence. A problem registry built for human submission rates has to be rebuilt if machine proofs arrive in bulk — this is the arXiv rate-limit story in a different institution.
The Future of Mathematics
Terence Tao, 2026-10-05. Companion to Saturday's piece on what powerful models mean for mathematicians. Five posts in five days from the person whose read carries the most weight here.
South Korea Says AI Agents May Have Been Used to Hack the Country's Banks
Slashdot, 2026-10-07. First named-nation attribution of a banking breach to agent tooling. Regulator-level, not researcher-level — the threshold that changes what institutions have to respond to.
OpenAI "rogue" agent activities found on Wikimedia projects
Simon Willison, 2026-10-07. Next beat after last Friday's disclosure that OpenAI alerted 100+ groups. Wikimedia is the first named victim surface with public evidence.
The Cyber Risk Discourse is Broken
Interconnects, 2026-10-06. Lambert's argument that the field debates AI cyber risk in hypotheticals. Published one day before South Korea supplied the non-hypothetical.
Strands Decider 2B: a small, open-source, decision model
HN → Strands, 2026-10-07. 2B-parameter open decision model, landing the same week as OpenAI's Decisions API beta and Vercel's confidence-based fallbacks. The category now has a small open entry, which is how it becomes infrastructure.

Themes this week

Mathematics gets an AI inflection claim, and a verification problem
Latent Space: 722 papers, 90 of top 500 problems (Wed), OpenAI: Sharing AI progress in mathematics (Tue), openai/math catalogue (Tue), Tao: Erdős problems site changes (Tue), Tao: The Future of Mathematics (Mon), Tao: what powerful models mean for mathematicians (Sat), Quanta: Is AI the End of Math As We Know It? (Mon), Tao: postdoc on formal verification and algorithm discovery (Mon).
Agentic cyber moves from discourse to incident
Slashdot: South Korea names AI agents in bank hacks (Wed), Simon Willison: rogue agents on Wikimedia (Wed), Slashdot: researchers tracking a Chinese AI "agent fleet" (Tue), Interconnects: the cyber risk discourse is broken (Tue), METR: AI systems could cover up misbehavior (Tue), Anthropic: expanding the Cyber Verification Program (Tue), Slashdot: OpenAI alerts 100+ groups (Fri).
Decision models become a product category
HN → Strands: Decider 2B open decision model (Wed), HN → OpenAI: Decisions API public beta (Tue), Vercel: confidence-based decision fallbacks (Tue), TC: how decision models could change content moderation (Tue), Simon Willison: llm-openai-decisions 0.1a0 (Tue), Hackaday: a new type of LLM on the block (Tue).
Open-weight frontier adds a European 1T
Simon Willison: Introducing Mistral Large 4: Le chonk (Tue), TC: Mistral's new 1T model aims to leapfrog closed and open rivals (Tue), TNS: Mistral's new AI tried to escape its test environment; weights ship in three weeks (Tue), Latent Space: Reflection Beam 501B (Tue), HF → TII: Falcon-Emirati (Tue).

Scan (15 min)

Tail

The verification layer is the story, not the claim
722 papers is not a result; it is a workload. The institutions that have to absorb it — arXiv (now rate-limiting), the Erdős problems registry (now being restructured), journal referee pools (unchanged) — are all sized for human submission rates. Tao's site changes are the only piece of this week's math news that is actually about capacity rather than capability.
Attribution thresholds, not capability thresholds
Nothing in the South Korea story requires agents to be more capable than they were last month. What changed is that a financial regulator was willing to name agent tooling in a breach finding. That is a disclosure-norm shift, and it will generate far more incident data than any new benchmark — which is also why Lambert's "the discourse is broken" piece dated badly in one day.
Small open models as category confirmation
Strands Decider at 2B parameters is the signal that decision models are a category rather than a frontier-lab product line. Vercel wiring confidence-based fallbacks into a gateway the same week says the same thing from the infrastructure side: you build fallback logic for things you expect to run in production, not for demos.

Feed silences (>72h since last item)

Sources that publish frequently but have gone quiet:

  • Neel Nanda (414 days) — last item 2025-08-19.
  • Brendan Gregg (242 days) — last item 2026-02-07.
  • Spritely Institute (147 days) — last item 2026-05-13.
  • Andy Wingo (144 days) — last item 2026-05-16.
  • Aphyr/Jepsen (117 days) — last item 2026-06-12.
  • Eugene Yan (108 days) — last item 2026-06-21.
  • Lilian Weng (95 days) — last item 2026-07-04.
  • Andrej Bauer (88 days) — last item 2026-07-11.
  • Julia Evans (78 days) — last item 2026-07-21.
  • Vicki Boykis (36 days) — last item 2026-09-01.
  • Fly.io (34 days) — last item 2026-09-03.
  • All Things Distributed (29 days) — last item 2026-09-08.
  • Kenneth Payne (23 days) — last item 2026-09-14.
  • Steve Yegge (22 days) — last item 2026-09-15.
  • Hillel Wayne (21 days) — last item 2026-09-16.
  • The Markup (21 days) — last item 2026-09-16.
  • Charity Majors (16 days) — last item 2026-09-21.
  • Ink & Switch (15 days) — last item 2026-09-22.
  • Murat Demirbas (14 days) — last item 2026-09-23.
  • Netflix Tech Blog (12 days) — last item 2026-09-25.
  • Stephen Wolfram (9 days) — last item 2026-09-28.
  • deepmind-blog (6 days) — last item 2026-10-01.
  • Alignment Forum (6 days) — last item 2026-10-01.
  • AI Snake Oil (6 days) — last item 2026-10-01.
  • Nature Machine Intelligence (6 days) — last item 2026-10-01.

Broke silence since Tuesday: Interconnects (14 days quiet) and GitHub Engineering (11 days quiet) both published on 10-06.

Build provenance

build: 2026-10-07 | crawler-sha: e51cd5a (Walsh-Research/1.2, compliance v1.4) | feeds: 75 core | items-considered: 6300 (14d, incl. 4104 arxiv-cs-ai) | warehouse: 55038 items | published: 50