Dashboarding Agent Sessions from Hook Events

Table of Contents

1. Overview

crowsnest is a loopback receiver for telemetry sightings. Since v2.5 it folds Claude Code hook events into one record per session and serves them at GET /sessions; the dashboard shows them as a rail beside the hook feed (sessions spec). On 2026-10-06 the rail was loaded with fourteen live sessions: ten started in tmux for the purpose, across ten repos, plus four that were already running. Three versions shipped in one night (2.5.0 to 2.5.2). This note records what that load showed: which hooks matter, what the event stream looks like as a Markov chain, where the state machine was wrong, and what a session list needs in order to be readable.

The evidence base is the raw hook log ~/.claude/crowsnest-hooks.jsonl: 1,897 records, 19 sessions, 11.6 hours, ending 2026-10-07T04:01:34Z. A separate review agent read it independently. It read only event names, session ids, timestamps, cwd basenames, notification types and the presence of agent_id, and replayed every record through the production fold.

2. The hooks that carry state

Nineteen of the thirty registered hook events fired. Four of them decide every main-thread state transition:

event n what it says maps to
UserPromptSubmit 79 a turn begins; the only reliable idle to busy edge busy
Stop 45 the turn is over idle
StopFailure 1 the turn is over, with an error idle, +1 error
SessionStart 14 a process started or resumed idle
SessionEnd 5 terminal ended

That is 144 events, 7.6% of the stream. The other 92% carry only liveness and detail:

event n value for a dashboard
PreToolUse / PostToolUse 526/525 the current tool, open-tool count, liveness inside long turns
PostToolBatch 309 redundant with PostToolUse
MessageDisplay 155 none
SubagentStop 123 none; mostly fires after the parent's Stop
Notification (idle_prompt) 32 none; always 60 to 61 s after a Stop
InstructionsLoaded, CwdChanged 29/26 none
ConfigChange 19 heartbeat: one settings edit reached 13 sessions in a second

The one state the dashboard most wants is the one never observed. waiting needs PermissionRequest, Elicitation or a permission_prompt notification, and all three counted zero. Every session ran in auto mode, which does not stop to ask.

3. The stream as a Markov chain

Counting first-order transitions per session, in log order, gives 1,878 transitions over 89 distinct pairs. The full chain is blurred: 37% of events come from subagents, which report under the parent's session_id and interleave with it. With subagent events removed (1,170 transitions), the main thread reads as a grammar:

SessionStart InstructionsLoaded+
( UserPromptSubmit ( PreToolUse PostToolUse+ PostToolBatch | MessageDisplay )*
  MessageDisplay Stop [ Notification @ +60 s ] )*
from to P(to given from)
PreToolUse PostToolUse 0.77
PostToolUse PostToolBatch 0.73
PostToolBatch PreToolUse 0.68
UserPromptSubmit PreToolUse 0.80
SessionStart InstructionsLoaded 0.86
Stop Notification 0.42
Notification UserPromptSubmit 0.77

The same method, applied to browser ad events, is in Event flow: empirical Markov chain from real logs. Here the chain is useful mainly as a check: transitions that the grammar does not predict are where the fold goes wrong.

4. What the replay found

The reference for "what the session was doing" is built from main-thread events only: UserPromptSubmit or a main-thread tool event means busy, Stop means idle. The final state of every session matched it. The errors are transient, and they are what a person watching the rail sees.

anomaly evidence status
SubagentStop after Stop reads busy follows 37 of 45 stops, median 11 s later; 25 of 123 occur in sessions with no subagent at all fixed in 2.5.1: no longer sets state
subagent tool calls after Stop read busy 186 sightings, 198 s of wrong state, one session open: the hook must forward whether agent_id is present
Stop then UserPromptSubmit in the same second reads idle once, 4 s; the tie rule ranks Stop first open: millisecond timestamps
SessionStart inside a compaction flips busy, idle, busy once, 1 s open: ignore a SessionStart that follows PreCompact
open_tools never returns to zero 5 orphaned main-thread tools; one is a PermissionDenied the fold does not close on open
waiting unreachable 0 permission events in 1,897 inherent to auto mode

Two findings correct the spec. PermissionDenied was written up as "the human answered", but the only one observed had no request before it: an automatic denial. And the notification counts the spec cites (130 idle_prompt, 8 permission_prompt) come from an earlier window that the log no longer holds.

The first row was found by watching, before the replay. Ten sessions were started and given one task each. Nine finished and still read busy, all with a last event of SubagentStop. The replay then explained it and found the second row, which watching could not have separated from the first.

5. What a session list needs

5.1. Sort on a key that does not move

2.5.0 sorted by attention (waiting, busy, idle, ended), then by last_seen descending. That follows iTerm2's session list. With ten live sessions it thrashed. last_seen changes on every hook event, so busy sessions swapped places on every 2 s poll, and every turn moved a session from the busy block to the idle block. The rail grouped sessions by project in order of first appearance, so whole projects jumped.

2.5.2 uses three buckets: waiting, live (busy and idle together), and ended. Within a bucket, the newest session to start comes first, and that start time never changes. Projects are alphabetical, except that a project with a waiting session goes first. The test was a poll while all ten sessions ran a turn:

for i in $(seq 1 30); do
  curl -s localhost:8127/sessions | jq -r '[.[].origin] | join(",")'
  sleep 2
done | awk -F, 'NF==11' | uniq -c
  29 trace-spine,hydra-setup,standard-change,wharfinger-002,sicp-minus-sicp,portclaim,order-optics,wharfinger-008,github-skills-search-guide,skills,www.wal.sh

That is 29 polls with one order. Recency is still shown, as the age on each row; it no longer decides position. Removing long-idle sessions from the list (decay) is a separate change and is still open.

5.2. Duplicates inflate counts, not sets

Two settings files register the crowsnest hook here, so every event from this repository reaches the receiver twice with different sighting ids. The fold keeps tools, errors and turns as sets keyed by meaning (tool_use id, start), so they are unaffected. Only the raw events count doubles. The hook log itself has no duplicates: 526 PreToolUse rows carry 526 distinct tool_use ids. The copies appear only at the receiver.

5.3. The table lives in memory

The session table is folded on ingest and is not persisted. Restarting the receiver, which three releases in one night required, empties the rail until each session emits again. An idle session stays invisible until its next prompt. A snapshot reporter that polls claude agents --json would refill the table on restart. The contract already reserves the payload for it (session.snapshot).

5.4. Calibrate each fix against the old code

Both fixes shipped with a test that was run against main first. The SubagentStop assertion read busy on the old fold. The stability test failed two of its four assertions under the last_seen sort. A test that cannot fail on the bug it names does not show the fix works.

6. Operating a fleet of sessions

Driving ten interactive sessions from one operator session through tmux turned up these practical points:

  • A new repo stops at Claude Code's folder-trust prompt. The cursor starts on "No, exit", so an unattended Enter quits the session.
  • tmux send-keys -t '=name' fails with "can't find pane"; the target needs a trailing colon, '=name:'.
  • gh auth switch changes state shared by every session. With repos owned by two accounts, each gh call takes GH_TOKEN=$(gh auth token --user <owner>) instead.
  • A brief that ends in a fixed one-line result (SWEEP <repo>: branch=… prs=…) turns ten panes into one table you can grep.
  • The permission classifier refused a brief that told ten sessions to merge their own pull requests without review. Read, pull and report went through.
  • Every nudge to a session produces a turn on the rail. The demo screenshot and the measurement were both taken during turns that the operator caused.

7. Open

  • Tag subagent events in the hook and keep them out of the session's state.
  • Millisecond timestamps in the hook (macOS date has no %N).
  • Close open tools on PermissionDenied and on Stop.
  • Forward notification_type so a permission_prompt can map to waiting.
  • Decay: remove a session from the list after a long idle.
  • A snapshot reporter, so the rail survives a receiver restart.

Related: crowsnest contract · sessions spec · Property-testing the crowsnest client · Gastown: multi-agent orchestration