P99 CONF 2026: Performance Engineering Opens With Two Keynotes About Agents
Table of Contents
- Event
- Why this one matters here
- The full talk list
- Talks to watch
- Effective Coding and Working with Agents
- Sandboxmaxxing at Lovable: Every Prompt Gets a Sandbox in < 1s
- Lessons Learned from Building Crazy Fast, Open Source Infrastructure for AI Agents
- Cloud Infra Done Right: Sub-millisecond Cold Starts at Million-VM Density
- Closed-Loop Performance Engineering
- AI-Assisted Profiling for Rust Services in Production
- Tracegrams: Tracking Tail Latency Propagation Without Storing Traces
- From Spans to Answers: Trace-Level Aggregation at 12 Billion Spans per Hour
- The Prompt is the Platform
- Mo Requests, Mo Problems: Managing Correctness in Asynchronous Systems
- Keep the Browser's Main Thread Free
- Optimizing eBPF Performance: Avoiding Production Latency, Throughput & Reliability Pitfalls
- The archive is the durable part
- Discrepancies on the published material
- Related research on this site
- Sources
Event
| Field | Value |
|---|---|
| Event | P99 CONF 2026, the sixth edition |
| Organiser | ScyllaDB — "the organization that curates P99 CONF" |
| Dates | Wednesday 21 — Thursday 22 October 2026 |
| Hours | 08:00–13:00 Pacific both days, plus a 01:00–04:00 Europe/Asia/India track on day two |
| Format | fully virtual, three concurrent stages; talks recorded in advance, Q&A live |
| Cost | free; registration carries 30 days of the O'Reilly platform |
| Scale | 62 talks, 14 of them still unscheduled; "thousands of engineers attend each year" |
| Registration | lp.scylladb.com — no deadline stated |
| Attending | registered, confirmed 2026-10-08 |
| Sponsors | brought by ScyllaDB, supported by Manning |
Its own framing, from the homepage:
There's no other event like this -- a conference for engineers by engineers, where we'll share novel approaches for solving complex problems efficiently and at speed. Vendor and tool agnostic, this conference will be for a highly technical audience only. Your boss's boss is not invited.
The CFP adds the constraint that makes the vendor-neutrality claim testable: "It's vendor-neutral, and we don't allow shallow overviews or product pitches." Talks are 15–20 minutes.
Why this one matters here
A conference about tail latency has put coding with agents in both day-one keynote slots and the day-two opener. That is the finding, and it is about where the subject has landed rather than about any single talk.
- Effective Coding and Working with Agents — Chip Huyen, author of AI Engineering. On making "the most out of working with multiple agents".
- Hardwood: Building a Parquet Parser From Scratch (With a Little Help From AI) — Gunnar Morling, Confluent. The abstract names the tool: "practical learnings from using AI (specifically, Claude Code) as a coding companion".
- How to Improve Your Cache Algorithm Using AI — Dor Laor, ScyllaDB's CEO, opening day two. "Everybody knows that AI can code. It can also convert your code to Rust. But can AI write a database? … my journey of 'vibe coding' the ScyllaDB cache to implement an alternative LRU algorithm and improve our p99 latency."
Roughly eighteen of sixty-two talks touch AI. Two are close enough to the work here to be worth the timeslot.
The agent-in-production talks
The Autonomous Performance Agent: A Netflix Production Story — Rajat Shah, AI Platform at Netflix. An agent that "continuously hunts performance inefficiencies across live production services, traces them to source code, proposes fixes, and validates results through canary deployment — grounding every decision in measured production outcomes, not model confidence". The abstract's own framing of the hard part: "what it actually took to make an autonomous agent trustworthy enough to act in production: where it earns autonomy, where it doesn't".
That is a calibration question wearing production clothes. An agent that proposes a fix and validates it against a canary has an oracle. The interesting claim is where the boundary sits.
Give the Agent a Cluster: Effective AI for Performance Engineering at Scale — Filipe Oliveira, Redis. Currently unscheduled. "What happens when you give an AI agent real power: access to performance profiles via MCP servers, a full multi-core cluster for controlled experiments, and a dedicated budget for iterative code generation and optimization. Instead of asking, 'Can AI write code?', this talk explores a more interesting question: Can AI run your performance lab?"
Profiles over MCP is the same move as exposing a REPL over MCP, applied to a different instrument. Worth reading against the local notes on expanding what an agent can see of a running system.
And one number to check
Vectorless RAG: How Removing the Vector Database Cut P99 by 40% — Jubin Soni, Yahoo. Unscheduled. "For a meaningful class of RAG applications, the vector database is the slowest part of the pipeline… In some production systems, we've seen, 40% P99 reductions after removing vector search, with recall improving rather than degrading."
A claim that removing the component improves both latency and recall is the kind that deserves its methodology read before it is repeated.
The full talk list
61 talks from the homepage featured-speaker list, read 2026-10-08, co-presenters merged. Day and slot per talk are not recorded here: the list carries none, and the full agenda grid has not been re-read since 2026-09-26. The agenda counted 62; the one-talk gap is unreconciled.
The Related here column links a note on this site. A plain link means the note covers the talk's named subject. "guess:" means the overlap is inferred from the title alone. A blank means nothing here is close.
Agents and AI for performance work
Sandboxes, VMs and cold starts
| Talk | Speaker(s) | Org | Related here |
|---|---|---|---|
| Sandboxmaxxing at Lovable: Every Prompt Gets a Sandbox in < 1s | Jonathan Grahl, Adrien Delorme | Lovable | Agent Sandbox Architectures; guess: FreeBSD jails |
| Lessons Learned from Building Crazy Fast, Open Source Infrastructure for AI Agents | Catherine Jue | KERNEL | guess: Agent Sandbox Architectures |
| Cloud Infra Done Right: Sub-millisecond Cold Starts at Million-VM Density | Felipe Huici | Unikraft | Agent Sandbox Architectures; guess: Serverless Architecture |
| Performance has Layers | Steve Karam | Oxide Computer |
I/O, io_uring and storage
| Talk | Speaker(s) | Org | Related here |
|---|---|---|---|
| The Difficult Life of an I/O Request | Avi Kivity | ScyllaDB | |
| Building a Real-World 14 GB/sec io_uring + tokio + Arrow + Rust Pipeline | Evan Chan | Conviva | |
| An Asymmetric io_uring Backend for Seastar | Jakub Czyszczoń, Witold Formański, Marcin Szopa | ScyllaDB | |
| io_uring: past, present, future | Jens Axboe, Glauber Costa | Turso | |
| The Fastest Object Storage Client in the World* | Georg Kreuzmayr | TigerBeetle | |
| Memory is Slow, Disk is Fast | Jared Hulbert | BitFlux | |
| Managing 500 Billion+ Files for AI Workloads | Joe Zhou | Juicedata |
Axboe is listed as the creator of io_uring; Turso in the org column is Costa.
Databases, data structures and data serving
| Talk | Speaker(s) | Org | Related here |
|---|---|---|---|
| The Vertical Scaling Wall, and How Valkey Gets Over It | Madelyn Olson | AWS | Performance Engineering Tools: caching |
| Billion-Query-Per-Second Lookup Tables in C++: Compile-Time Perfect Hashing | Francisco Geiman Thiesen, Daniel Lemire | Microsoft; U. Quebec | |
| The SimdQuickHeap: The Fastest Priority Queue by 2x | Ragnar Groot Koerkamp | KIT | Python Priority Queues with heapq |
| Breaking SQLite's Single-Writer Bottleneck | Pere Diaz | Turso | |
| How ScyllaDB Eliminates Read Amplification at Scale | Michael Litvak, Felipe Cardeneti Mendes | ScyllaDB | |
| Keeping Latency Under Pressure: Surviving Connection and Request Storms | Marcin Maliszkiewicz | ScyllaDB | |
| Well Designed Databases are CPU Bound | Tyson Brown | Thoughtworks | |
| Low Latency at Global Scale: Isolation, Optimistic Concurrency, and Precise Time | Raluca Constantin | AWS | guess: TLA+ for System Design |
| Eliminating Query Latency at Cloud Scale: Heuristic Search Space Partitioning for Multi-Tenant Data | Rama Teja Repaka, Prashant Pathak | Palo Alto Networks | |
| Writing a TSDB from Scratch: Performance Optimization | Roman Khavronenko | VictoriaMetrics | |
| Can Everything be an LSM? | Almog Gavra | Responsive | |
| PostgreSQL Goes Columnar: How Far Can It Really Push Analytics? | Daniel Seybold | benchANT | |
| Building Ultra-Performant Message Streaming | Piotr Gankiewicz | LaserData | |
| Tiling the Hot-Path: Low-Latency Feature Serving at High Fanout | Piyush Narang | Zipline.ai | |
| Escaping the Gossip: How DoorDash Built an Infinite-Scale Feature Store | Luigi Tagliamonte | DoorDash | |
| Mo Requests, Mo Problems: Managing Correctness in Asynchronous Systems | Benjamin Cane | American Express | guess: TLA+ for System Design |
Tracing, observability and eBPF
| Talk | Speaker(s) | Org | Related here |
|---|---|---|---|
| Tracegrams: Tracking Tail Latency Propagation Without Storing Traces | Ivan Goncharov | Azul | crowsnest OTLP ingest; Performance Engineering Tools; guess: Agent Telemetry Systems |
| From Spans to Answers: Trace-Level Aggregation at 12 Billion Spans per Hour | Sudeep Kumar, Thomas Varley | Salesforce | crowsnest OTLP ingest; Performance Engineering Tools |
| Unlocking the Go Execution Tracer: Programmatic Analysis with the Trace API | Cristian Velazquez | Uber | |
| Holistic Profiling with Stax | Amos Wenger | bearcove | Performance Engineering Tools: profiling |
| Optimizing eBPF Performance: Avoiding Production Latency, Throughput & Reliability Pitfalls | Tanel Poder | not stated | Performance Engineering Tools: eBPF |
| eBPF for Metrics Collection on HPC Systems | Ershaad Ahamed Basheer | LBNL | Performance Engineering Tools: eBPF |
| Networking Observability Essentials - How to Know if the Network is at Fault with eBPF | Peter Zaitsev | Percona | Performance Engineering Tools: eBPF |
| Your P(od)99 Isn't Slow… It's Queued | Enzo Venturi | CNCF Ambassador |
Networking
| Talk | Speaker(s) | Org | Related here |
|---|---|---|---|
| HTTP/3 in Production: What the Numbers Actually Say | Elvis Chidera | Delivery Hero | |
| How Debugging Nginx Throughput Led to a Linux Kernel Contribution | Daniel Sedlak | CDN77 | |
| The localhost Tax: Transparent TCP Splicing for Co-located Services | Cong Wang | Multikernel Technologies | |
| Integrating QUIC into Seastar | Kamil Dalidowicz, Piotr Korcz | ScyllaDB | |
| Chasing 8 ms: Snap's Tail Latency Hunt Through Envoy's Redis Proxy | Kishor Yadav Kommanaboina | Snap | |
| How We Hit Bare-Metal Tail Latency Inside Cloud VMs: Machnet | Vahab Jabrayilov | Columbia |
Web and frontend
| Talk | Speaker(s) | Org | Related here |
|---|---|---|---|
| Keep the Browser's Main Thread Free | Den Odell | Manning | Performance Engineering Tools Reference |
| The Age of Slow Apps Is Over: Topcoat Brings Rust to Web Apps | Carl Lerche | AWS |
Vector search, RAG and quantization
| Talk | Speaker(s) | Org | Related here |
|---|---|---|---|
| Searching 100B Vectors on Object Storage @ 200ms P99 | Nathan VanBenschoten | turbopuffer | |
| Vector Search: Benchmarking HNSW at 1B Scale: Performance, Recall, and Quantization Tradeoffs | Szymon Wasik | ScyllaDB | |
| Benchmarking ScyllaDB Native Full-Text Search vs. OpenSearch | Karol Nowacki | ScyllaDB | guess: pocket-es |
| Vectorless RAG: How Removing the Vector Database Cut P99 by 40% | Jubin Soni | Yahoo | guess: pocket-es |
| Fearless Quantization | Sam Rose | ngrok | guess: Qwen3.6 and the KV Cache Constraint |
The HNSW title carries "1B"; the abstract says 100M. See Discrepancies.
Runtimes, concurrency and scheduling
| Talk | Speaker(s) | Org | Related here |
|---|---|---|---|
| From Snakes to Crabs: Rewriting Feature Flag Evaluation Infrastructure at Web Scale | Dylan Martin | PostHog | |
| Multi-Core Without the Trilemma: Escaping Async/Await, Mutexes, and GC | Peter Mbanugo | not stated | |
| My Fair Shares: Teaching a Flat Scheduler About Hierarchy | Pavel Emelyanov | ScyllaDB | |
| Capturing Lighting in a Bottle | Francesco Nigro | IBM |
Talks to watch
The talks with the most overlap with work here, beyond the two agent-in-production talks above. Where an abstract has not been read, the entry says so and stops at the title.
Effective Coding and Working with Agents
Chip Huyen. A day-one keynote; the abstract's stated subject is getting "the most out of working with multiple agents". Multiple agents on one codebase is the problem Agentic Software Engineering maps, though there it is framed as merge contention on a single trunk. Whether the keynote reaches the trunk at all is unknown; the link is a guess.
Sandboxmaxxing at Lovable: Every Prompt Gets a Sandbox in < 1s
Jonathan Grahl and Adrien Delorme, Lovable. Abstract not read. The title makes the sandbox per prompt rather than per session or per agent, with a sub-second budget. Agent Sandbox Architectures treats lifecycle (ephemeral, snapshot, volume) as a fifth concern and records sub-second microVM boot as a selling point of Deno Sandbox; a per-prompt cell is the far end of that lifecycle axis. The question to bring: what persists between prompts, and who holds the secrets while it does.
Lessons Learned from Building Crazy Fast, Open Source Infrastructure for AI Agents
Catherine Jue, KERNEL. The homepage abstract cites sub-30ms browser boot times. Hosted browsers are the agent's tool here, not its container, so the nearest note is a guess: Agent Sandbox Architectures treats the browser as a sandbox for agentic file work, with egress as the clause it cannot meet. A browser booted per task inherits that egress problem. Absent from the agenda feed as of the 2026-09-26 read.
Cloud Infra Done Right: Sub-millisecond Cold Starts at Million-VM Density
Felipe Huici, Unikraft. Abstract not read. Sub-millisecond is three orders of magnitude under the sub-second microVM boot the sandbox note records for Deno Sandbox. If the number holds, the compute boundary costs nothing per call and the remaining design questions in Agent Sandbox Architectures are custody and egress.
Closed-Loop Performance Engineering
Tomás Senart, Perfloop. Abstract not read. "Closed loop" implies a measured signal fed back into the change, which is the oracle-against-generator setup in The Oracle Is the Deliverable. Whether an agent is in the loop is not stated in the title; the link is a guess.
AI-Assisted Profiling for Rust Services in Production
Hayden Stainsby, HERE. Abstract not read. Production profiling is the subject
of the profiling section of
Performance Engineering Tools,
which lists perf, async-profiler, py-spy and eBPF but no profiler aimed at
Rust services. A gap that talk may fill.
Tracegrams: Tracking Tail Latency Propagation Without Storing Traces
Ivan Goncharov, Azul. Abstract not read. The crowsnest receiver takes the opposite position at small scale: OTLP ingest maps every span to a stored sighting, parent span id as the causal edge. A method that recovers tail propagation without keeping the spans is the scaling argument against that design. Guess: metrics.jsonl over OTel makes a related keep-less choice for agent runs.
From Spans to Answers: Trace-Level Aggregation at 12 Billion Spans per Hour
Sudeep Kumar and Thomas Varley, Salesforce. Abstract not read. 12 billion spans an hour is about 3.3 million a second. The same OTLP span shape that crowsnest folds into sightings one at a time; this talk is what that fold looks like when it has to be a query rather than a list.
The Prompt is the Platform
Dominik Tornow, Resonate HQ. Abstract not read. Resonate builds durable execution, so the title plausibly argues the prompt becomes the program a durable runtime executes. Guess: Context Surfaces, which separates a surface from an agent harness; this talk may collapse the two.
Mo Requests, Mo Problems: Managing Correctness in Asynchronous Systems
Benjamin Cane, American Express. Abstract not read. Correctness in asynchronous systems is what model checking is for. Guess: TLA+ for System Design. The AWS talk on isolation, optimistic concurrency and precise time sits next to it on the same guess.
Keep the Browser's Main Thread Free
Den Odell, Manning. The author of Performance Engineering in Practice, the book Performance Engineering Tools Reference maps tool by tool. The talk is the book's author on its browser half.
Optimizing eBPF Performance: Avoiding Production Latency, Throughput & Reliability Pitfalls
Tanel Poder. Abstract not read. The eBPF entry in Performance Engineering Tools says several BCC tools have "low enough overhead for continuous 24/7 use". A talk on the production cost of eBPF itself is the check on that claim. Two more eBPF talks, from LBNL and Percona, are in the tracing table.
The archive is the durable part
p99conf.io/on-demand carries 264 session pages spanning 2021 through 2025, filterable by year, type and topic, each with abstract, speaker bio and slides. Ungated — an attendee quote the site chooses to display reads "Videos available on-demand afterwards with no gating or games."
For a conference whose talks are recorded in advance anyway, the archive rather than the live event is the artifact. ScyllaDB also runs a separate gated on-demand landing page for 2025, which is worth knowing before handing over an email.
Discrepancies on the published material
Recorded because a reader planning around this page would hit them.
- The UTC conversion is wrong. The site says "8:00am – 1:00pm Pacific Time
/ 16:00 – 20:00 UTC". October Pacific is UTC-7, so 08:00 PT is 15:00 UTC,
not 16:00. Sessionize's own timezone field says
UTC-07. Unchanged on 2026-10-08: the page still pairs 08:00–13:00 Pacific with 16:00–20:00 UTC. Pacific Daylight Time makes it 15:00–20:00 UTC; the end is right and the start is an hour late. Recorded as found. - The stated hours exclude a whole track. "8:00am – 1:00pm" on both days, while the grid schedules nine day-two sessions from 01:00 to 03:55 Pacific for Europe, Asia and India. Day two actually runs about 01:00 to 12:35.
- "Full agenda will be announced soon" still sits on the homepage while
/agenda/renders a complete two-day, three-stage grid. - Three talks carry different titles on the homepage and in the agenda, including one where the vector count differs — 1B on the homepage, 100M in the abstract.
- The lower speaker wall is a legacy reel. Andy Pavlo, Michael Stonebraker, Gil Tene, Liz Rice, Charity Majors and others appear with bios and no 2026 talk. They are past speakers used as social proof and read as this year's lineup.
- One talk exists only on the homepage — Catherine Jue of KERNEL on "building crazy fast, open source infrastructure for AI agents", with sub-30ms browser boot times — and appears nowhere in the agenda feed. It is still on the homepage list on 2026-10-08; the feed was not re-checked.
- 61 featured, 62 scheduled. The homepage list read on 2026-10-08 has 61 talks against the 62 the agenda grid held on 2026-09-26. Which one is missing from which list is not established.
Sources
- p99conf.io — server-rendered; the homepage carries the speaker list and abstracts
- CFP — closed 29 May 2026; carries the track list verbatim
- On-demand archive, 2021–2025
- 2025 recap — "the fifth in this incredible series", titled Latency to LLMs