P99 CONF 2026: Performance Engineering Opens With Two Keynotes About Agents

Table of Contents

Event

Field Value
Event P99 CONF 2026, the sixth edition
Organiser ScyllaDB — "the organization that curates P99 CONF"
Dates Wednesday 21 — Thursday 22 October 2026
Hours 08:00–13:00 Pacific both days, plus a 01:00–04:00 Europe/Asia/India track on day two
Format fully virtual, three concurrent stages; talks recorded in advance, Q&A live
Cost free; registration carries 30 days of the O'Reilly platform
Scale 62 talks, 14 of them still unscheduled; "thousands of engineers attend each year"
Registration lp.scylladb.com — no deadline stated
Attending registered, confirmed 2026-10-08
Sponsors brought by ScyllaDB, supported by Manning

Its own framing, from the homepage:

There's no other event like this -- a conference for engineers by engineers,
where we'll share novel approaches for solving complex problems efficiently
and at speed. Vendor and tool agnostic, this conference will be for a highly
technical audience only. Your boss's boss is not invited.

The CFP adds the constraint that makes the vendor-neutrality claim testable: "It's vendor-neutral, and we don't allow shallow overviews or product pitches." Talks are 15–20 minutes.

Why this one matters here

A conference about tail latency has put coding with agents in both day-one keynote slots and the day-two opener. That is the finding, and it is about where the subject has landed rather than about any single talk.

  • Effective Coding and Working with Agents — Chip Huyen, author of AI Engineering. On making "the most out of working with multiple agents".
  • Hardwood: Building a Parquet Parser From Scratch (With a Little Help From AI) — Gunnar Morling, Confluent. The abstract names the tool: "practical learnings from using AI (specifically, Claude Code) as a coding companion".
  • How to Improve Your Cache Algorithm Using AI — Dor Laor, ScyllaDB's CEO, opening day two. "Everybody knows that AI can code. It can also convert your code to Rust. But can AI write a database? … my journey of 'vibe coding' the ScyllaDB cache to implement an alternative LRU algorithm and improve our p99 latency."

Roughly eighteen of sixty-two talks touch AI. Two are close enough to the work here to be worth the timeslot.

The agent-in-production talks

The Autonomous Performance Agent: A Netflix Production Story — Rajat Shah, AI Platform at Netflix. An agent that "continuously hunts performance inefficiencies across live production services, traces them to source code, proposes fixes, and validates results through canary deployment — grounding every decision in measured production outcomes, not model confidence". The abstract's own framing of the hard part: "what it actually took to make an autonomous agent trustworthy enough to act in production: where it earns autonomy, where it doesn't".

That is a calibration question wearing production clothes. An agent that proposes a fix and validates it against a canary has an oracle. The interesting claim is where the boundary sits.

Give the Agent a Cluster: Effective AI for Performance Engineering at Scale — Filipe Oliveira, Redis. Currently unscheduled. "What happens when you give an AI agent real power: access to performance profiles via MCP servers, a full multi-core cluster for controlled experiments, and a dedicated budget for iterative code generation and optimization. Instead of asking, 'Can AI write code?', this talk explores a more interesting question: Can AI run your performance lab?"

Profiles over MCP is the same move as exposing a REPL over MCP, applied to a different instrument. Worth reading against the local notes on expanding what an agent can see of a running system.

And one number to check

Vectorless RAG: How Removing the Vector Database Cut P99 by 40% — Jubin Soni, Yahoo. Unscheduled. "For a meaningful class of RAG applications, the vector database is the slowest part of the pipeline… In some production systems, we've seen, 40% P99 reductions after removing vector search, with recall improving rather than degrading."

A claim that removing the component improves both latency and recall is the kind that deserves its methodology read before it is repeated.

The full talk list

61 talks from the homepage featured-speaker list, read 2026-10-08, co-presenters merged. Day and slot per talk are not recorded here: the list carries none, and the full agenda grid has not been re-read since 2026-09-26. The agenda counted 62; the one-talk gap is unreconciled.

The Related here column links a note on this site. A plain link means the note covers the talk's named subject. "guess:" means the overlap is inferred from the title alone. A blank means nothing here is close.

Agents and AI for performance work

Talk Speaker(s) Org Related here
Effective Coding and Working with Agents Chip Huyen author, AI Engineering guess: Agentic Software Engineering
Hardwood: Building a Parquet Parser From Scratch (With a Little Help From AI) Gunnar Morling Confluent guess: Rebuilding from the spec
The Prompt is the Platform Dominik Tornow Resonate HQ guess: Context Surfaces
How to Improve Your Cache Algorithm Using AI Dor Laor ScyllaDB guess: The Oracle Is the Deliverable
AI-Driven Optimization: The Good, the Bad and the Ugly Yichen Wei Disney+/Hulu guess: The Oracle Is the Deliverable
Closed-Loop Performance Engineering Tomás Senart Perfloop guess: The Oracle Is the Deliverable
AI-Assisted Profiling for Rust Services in Production Hayden Stainsby HERE Performance Engineering Tools: profiling
The Autonomous Performance Agent: A Netflix Production Story Rajat Shah Netflix The Oracle Is the Deliverable
Give the Agent a Cluster: Effective AI for Performance Engineering at Scale Filipe Oliveira Redis Context Surfaces

I/O, io_uring and storage

Talk Speaker(s) Org Related here
The Difficult Life of an I/O Request Avi Kivity ScyllaDB  
Building a Real-World 14 GB/sec io_uring + tokio + Arrow + Rust Pipeline Evan Chan Conviva  
An Asymmetric io_uring Backend for Seastar Jakub Czyszczoń, Witold Formański, Marcin Szopa ScyllaDB  
io_uring: past, present, future Jens Axboe, Glauber Costa Turso  
The Fastest Object Storage Client in the World* Georg Kreuzmayr TigerBeetle  
Memory is Slow, Disk is Fast Jared Hulbert BitFlux  
Managing 500 Billion+ Files for AI Workloads Joe Zhou Juicedata  

Axboe is listed as the creator of io_uring; Turso in the org column is Costa.

Databases, data structures and data serving

Talk Speaker(s) Org Related here
The Vertical Scaling Wall, and How Valkey Gets Over It Madelyn Olson AWS Performance Engineering Tools: caching
Billion-Query-Per-Second Lookup Tables in C++: Compile-Time Perfect Hashing Francisco Geiman Thiesen, Daniel Lemire Microsoft; U. Quebec  
The SimdQuickHeap: The Fastest Priority Queue by 2x Ragnar Groot Koerkamp KIT Python Priority Queues with heapq
Breaking SQLite's Single-Writer Bottleneck Pere Diaz Turso  
How ScyllaDB Eliminates Read Amplification at Scale Michael Litvak, Felipe Cardeneti Mendes ScyllaDB  
Keeping Latency Under Pressure: Surviving Connection and Request Storms Marcin Maliszkiewicz ScyllaDB  
Well Designed Databases are CPU Bound Tyson Brown Thoughtworks  
Low Latency at Global Scale: Isolation, Optimistic Concurrency, and Precise Time Raluca Constantin AWS guess: TLA+ for System Design
Eliminating Query Latency at Cloud Scale: Heuristic Search Space Partitioning for Multi-Tenant Data Rama Teja Repaka, Prashant Pathak Palo Alto Networks  
Writing a TSDB from Scratch: Performance Optimization Roman Khavronenko VictoriaMetrics  
Can Everything be an LSM? Almog Gavra Responsive  
PostgreSQL Goes Columnar: How Far Can It Really Push Analytics? Daniel Seybold benchANT  
Building Ultra-Performant Message Streaming Piotr Gankiewicz LaserData  
Tiling the Hot-Path: Low-Latency Feature Serving at High Fanout Piyush Narang Zipline.ai  
Escaping the Gossip: How DoorDash Built an Infinite-Scale Feature Store Luigi Tagliamonte DoorDash  
Mo Requests, Mo Problems: Managing Correctness in Asynchronous Systems Benjamin Cane American Express guess: TLA+ for System Design

Tracing, observability and eBPF

Talk Speaker(s) Org Related here
Tracegrams: Tracking Tail Latency Propagation Without Storing Traces Ivan Goncharov Azul crowsnest OTLP ingest; Performance Engineering Tools; guess: Agent Telemetry Systems
From Spans to Answers: Trace-Level Aggregation at 12 Billion Spans per Hour Sudeep Kumar, Thomas Varley Salesforce crowsnest OTLP ingest; Performance Engineering Tools
Unlocking the Go Execution Tracer: Programmatic Analysis with the Trace API Cristian Velazquez Uber  
Holistic Profiling with Stax Amos Wenger bearcove Performance Engineering Tools: profiling
Optimizing eBPF Performance: Avoiding Production Latency, Throughput & Reliability Pitfalls Tanel Poder not stated Performance Engineering Tools: eBPF
eBPF for Metrics Collection on HPC Systems Ershaad Ahamed Basheer LBNL Performance Engineering Tools: eBPF
Networking Observability Essentials - How to Know if the Network is at Fault with eBPF Peter Zaitsev Percona Performance Engineering Tools: eBPF
Your P(od)99 Isn't Slow… It's Queued Enzo Venturi CNCF Ambassador  

Networking

Talk Speaker(s) Org Related here
HTTP/3 in Production: What the Numbers Actually Say Elvis Chidera Delivery Hero  
How Debugging Nginx Throughput Led to a Linux Kernel Contribution Daniel Sedlak CDN77  
The localhost Tax: Transparent TCP Splicing for Co-located Services Cong Wang Multikernel Technologies  
Integrating QUIC into Seastar Kamil Dalidowicz, Piotr Korcz ScyllaDB  
Chasing 8 ms: Snap's Tail Latency Hunt Through Envoy's Redis Proxy Kishor Yadav Kommanaboina Snap  
How We Hit Bare-Metal Tail Latency Inside Cloud VMs: Machnet Vahab Jabrayilov Columbia  

Web and frontend

Talk Speaker(s) Org Related here
Keep the Browser's Main Thread Free Den Odell Manning Performance Engineering Tools Reference
The Age of Slow Apps Is Over: Topcoat Brings Rust to Web Apps Carl Lerche AWS  

Vector search, RAG and quantization

Talk Speaker(s) Org Related here
Searching 100B Vectors on Object Storage @ 200ms P99 Nathan VanBenschoten turbopuffer  
Vector Search: Benchmarking HNSW at 1B Scale: Performance, Recall, and Quantization Tradeoffs Szymon Wasik ScyllaDB  
Benchmarking ScyllaDB Native Full-Text Search vs. OpenSearch Karol Nowacki ScyllaDB guess: pocket-es
Vectorless RAG: How Removing the Vector Database Cut P99 by 40% Jubin Soni Yahoo guess: pocket-es
Fearless Quantization Sam Rose ngrok guess: Qwen3.6 and the KV Cache Constraint

The HNSW title carries "1B"; the abstract says 100M. See Discrepancies.

Runtimes, concurrency and scheduling

Talk Speaker(s) Org Related here
From Snakes to Crabs: Rewriting Feature Flag Evaluation Infrastructure at Web Scale Dylan Martin PostHog  
Multi-Core Without the Trilemma: Escaping Async/Await, Mutexes, and GC Peter Mbanugo not stated  
My Fair Shares: Teaching a Flat Scheduler About Hierarchy Pavel Emelyanov ScyllaDB  
Capturing Lighting in a Bottle Francesco Nigro IBM  

Talks to watch

The talks with the most overlap with work here, beyond the two agent-in-production talks above. Where an abstract has not been read, the entry says so and stops at the title.

Effective Coding and Working with Agents

Chip Huyen. A day-one keynote; the abstract's stated subject is getting "the most out of working with multiple agents". Multiple agents on one codebase is the problem Agentic Software Engineering maps, though there it is framed as merge contention on a single trunk. Whether the keynote reaches the trunk at all is unknown; the link is a guess.

Sandboxmaxxing at Lovable: Every Prompt Gets a Sandbox in < 1s

Jonathan Grahl and Adrien Delorme, Lovable. Abstract not read. The title makes the sandbox per prompt rather than per session or per agent, with a sub-second budget. Agent Sandbox Architectures treats lifecycle (ephemeral, snapshot, volume) as a fifth concern and records sub-second microVM boot as a selling point of Deno Sandbox; a per-prompt cell is the far end of that lifecycle axis. The question to bring: what persists between prompts, and who holds the secrets while it does.

Lessons Learned from Building Crazy Fast, Open Source Infrastructure for AI Agents

Catherine Jue, KERNEL. The homepage abstract cites sub-30ms browser boot times. Hosted browsers are the agent's tool here, not its container, so the nearest note is a guess: Agent Sandbox Architectures treats the browser as a sandbox for agentic file work, with egress as the clause it cannot meet. A browser booted per task inherits that egress problem. Absent from the agenda feed as of the 2026-09-26 read.

Cloud Infra Done Right: Sub-millisecond Cold Starts at Million-VM Density

Felipe Huici, Unikraft. Abstract not read. Sub-millisecond is three orders of magnitude under the sub-second microVM boot the sandbox note records for Deno Sandbox. If the number holds, the compute boundary costs nothing per call and the remaining design questions in Agent Sandbox Architectures are custody and egress.

Closed-Loop Performance Engineering

Tomás Senart, Perfloop. Abstract not read. "Closed loop" implies a measured signal fed back into the change, which is the oracle-against-generator setup in The Oracle Is the Deliverable. Whether an agent is in the loop is not stated in the title; the link is a guess.

AI-Assisted Profiling for Rust Services in Production

Hayden Stainsby, HERE. Abstract not read. Production profiling is the subject of the profiling section of Performance Engineering Tools, which lists perf, async-profiler, py-spy and eBPF but no profiler aimed at Rust services. A gap that talk may fill.

Tracegrams: Tracking Tail Latency Propagation Without Storing Traces

Ivan Goncharov, Azul. Abstract not read. The crowsnest receiver takes the opposite position at small scale: OTLP ingest maps every span to a stored sighting, parent span id as the causal edge. A method that recovers tail propagation without keeping the spans is the scaling argument against that design. Guess: metrics.jsonl over OTel makes a related keep-less choice for agent runs.

From Spans to Answers: Trace-Level Aggregation at 12 Billion Spans per Hour

Sudeep Kumar and Thomas Varley, Salesforce. Abstract not read. 12 billion spans an hour is about 3.3 million a second. The same OTLP span shape that crowsnest folds into sightings one at a time; this talk is what that fold looks like when it has to be a query rather than a list.

The Prompt is the Platform

Dominik Tornow, Resonate HQ. Abstract not read. Resonate builds durable execution, so the title plausibly argues the prompt becomes the program a durable runtime executes. Guess: Context Surfaces, which separates a surface from an agent harness; this talk may collapse the two.

Mo Requests, Mo Problems: Managing Correctness in Asynchronous Systems

Benjamin Cane, American Express. Abstract not read. Correctness in asynchronous systems is what model checking is for. Guess: TLA+ for System Design. The AWS talk on isolation, optimistic concurrency and precise time sits next to it on the same guess.

Keep the Browser's Main Thread Free

Den Odell, Manning. The author of Performance Engineering in Practice, the book Performance Engineering Tools Reference maps tool by tool. The talk is the book's author on its browser half.

Optimizing eBPF Performance: Avoiding Production Latency, Throughput & Reliability Pitfalls

Tanel Poder. Abstract not read. The eBPF entry in Performance Engineering Tools says several BCC tools have "low enough overhead for continuous 24/7 use". A talk on the production cost of eBPF itself is the check on that claim. Two more eBPF talks, from LBNL and Percona, are in the tracing table.

The archive is the durable part

p99conf.io/on-demand carries 264 session pages spanning 2021 through 2025, filterable by year, type and topic, each with abstract, speaker bio and slides. Ungated — an attendee quote the site chooses to display reads "Videos available on-demand afterwards with no gating or games."

For a conference whose talks are recorded in advance anyway, the archive rather than the live event is the artifact. ScyllaDB also runs a separate gated on-demand landing page for 2025, which is worth knowing before handing over an email.

Discrepancies on the published material

Recorded because a reader planning around this page would hit them.

  • The UTC conversion is wrong. The site says "8:00am – 1:00pm Pacific Time / 16:00 – 20:00 UTC". October Pacific is UTC-7, so 08:00 PT is 15:00 UTC, not 16:00. Sessionize's own timezone field says UTC-07. Unchanged on 2026-10-08: the page still pairs 08:00–13:00 Pacific with 16:00–20:00 UTC. Pacific Daylight Time makes it 15:00–20:00 UTC; the end is right and the start is an hour late. Recorded as found.
  • The stated hours exclude a whole track. "8:00am – 1:00pm" on both days, while the grid schedules nine day-two sessions from 01:00 to 03:55 Pacific for Europe, Asia and India. Day two actually runs about 01:00 to 12:35.
  • "Full agenda will be announced soon" still sits on the homepage while /agenda/ renders a complete two-day, three-stage grid.
  • Three talks carry different titles on the homepage and in the agenda, including one where the vector count differs — 1B on the homepage, 100M in the abstract.
  • The lower speaker wall is a legacy reel. Andy Pavlo, Michael Stonebraker, Gil Tene, Liz Rice, Charity Majors and others appear with bios and no 2026 talk. They are past speakers used as social proof and read as this year's lineup.
  • One talk exists only on the homepage — Catherine Jue of KERNEL on "building crazy fast, open source infrastructure for AI agents", with sub-30ms browser boot times — and appears nowhere in the agenda feed. It is still on the homepage list on 2026-10-08; the feed was not re-checked.
  • 61 featured, 62 scheduled. The homepage list read on 2026-10-08 has 61 talks against the 62 the agenda grid held on 2026-09-26. Which one is missing from which list is not established.

Sources

  • p99conf.io — server-rendered; the homepage carries the speaker list and abstracts
  • CFP — closed 29 May 2026; carries the track list verbatim
  • On-demand archive, 2021–2025
  • 2025 recap — "the fifth in this incredible series", titled Latency to LLMs