Ten Trace Explorers: A SICP Build Reflection
Table of Contents
1. Overview
Ten small ClojureScript tools trace SICP procedures step by step: a wire holds a signal, a painter draws into a frame, an agenda runs actions in time order. Each tool shares one architecture – a pure model emits an event log, three different strategies compute state at a given step, and dumb renderers turn that state into DOM, SVG, or canvas. This note is the reflection on building them: what the shared structure buys, what broke, and what a 2026-09-19 session of parallel-subagent builds looked like on the ground.
2. Build process
Every page followed the same sequence:
- Choose which internal steps of the algorithm count as events.
- Write the model as a pure function that returns an event list.
- Run the model in Node against known values before writing any UI. These were SICP's own numbers where they exist: 15,499 calls for cc(100), 42 bits for the §2.3.4 message, the half-adder timings of 8, 11 and 16, and the parallel-resistor intervals [2.201, 3.487] and [2.582, 2.973].
- Write the page.
- Render it in headless Chromium, take screenshots at chosen steps, and read the results to find layout and logic faults.
- Publish.
The Node check before the UI caught more errors than the screenshots did, because a wrong number is easier to see in a printed table than in a picture.
The first page set the visual system: colour tokens, the transport bar, the strip canvas, the pseudocode panel with line highlighting, and a result sentence above the timeline. Later pages copied it and changed only the model and the central view. This kept every page after the first to roughly one model, one view and one round of fixes.
3. Structure of a trace explorer
Each page has four layers.
3.1. Model
A pure function of (inputs, options) returns a trace. The options are the toggles: memoization, queue order, frame-coord-map variant, constructor style, and so on. The trace is an array of events. Every event carries at least:
- a type;
- the pseudocode line numbers it corresponds to;
- a sentence describing it;
- the identifiers of the objects it touches, such as a node, a connector, a table cell or a worklist index.
3.2. State at step s
The pages used three methods to compute it.
- Replay with a reducer (counting change). The state is
reduce(applyEvent, fresh, events[0..s]). It moves forward incrementally and replays from zero when moving backward. Memory is small, but a backward seek costs O(s). - Derived position (primality, intervals). The state is computed from event indices alone. For example, frames called = clamp(k, 0, L), where k is the number of events since the current run began. This works only when the event structure is regular.
- Snapshots (sets, Huffman, circuits, constraints, algebra, pictures).
Each event, or each structural change, stores a copy of the relevant
state. The copy is either a versions list
\{at, value\}per structure or full value arrays per event. Memory grows with the number of events, but seeking is O(1) or O(versions).
Snapshots were the most reliable choice for mutable algorithms, i.e. circuits and constraints. There the simulator can mutate freely during the run and the trace is still immutable data afterward.
3.3. Renderers
These read (trace, s) and write DOM, SVG or canvas. They contain no algorithm logic. The event text is produced by the model, not by the renderer, and this moved every "what just happened" explanation into the part of the code that has the relevant values.
3.4. Comparison runs
Each page runs the model more than once per input: memo on and off, sparse and dense, FIFO and LIFO, the correct and the buggy frame map, all four set representations. The results go into a table. The table carries most of the explanation, because it shows a cost or correctness difference as numbers next to the trace that produced them.
4. Session build log (2026-09-19)
This section is the concrete evidence from one build session, in
which seven of the ten – counting-change, primality-test,
interval-arithmetic, representing-sets, huffman-trees,
picture-language and symbolic-algebra – were built by parallel
subagents against this repository's own tools
convention and merged the same day. digital-circuits and
propagation-of-constraints followed the same day; a tenth page,
symbolic differentiation, is referenced in the "Issues encountered"
section below as a build already in the reflection's scope but is not
part of this session's own merge log.
4.1. What actually landed
| Tool | SICP § | PR | Port (final) |
|---|---|---|---|
| counting-change | 1.2.2 | #104 | 8728 |
| primality-test | 1.2.6 | #105 | 8729 |
| interval-arithmetic | 2.1.4 | #106 | 8730 |
| representing-sets | 2.3.3 | #107 | 8731 |
| huffman-trees | 2.3.4 | #108 | 8732 |
| picture-language | 2.2.4 | #109 | 8733 |
| symbolic-algebra | 2.5 | #110 | 8734 |
Every one of the seven was built against the same
tool-maintenance-spec v1.0.0 and tool-maintainer skill v1.0.0.
4.2. The build itself is an instance of the "many agents, one trunk" problem
Six of the seven tools branched from a main that did not yet contain the
others – exactly the setup the
agentic-software-engineering literature calls agent-scale merge
contention. Each subagent picked its shadow-cljs dev-server port by reading
its own branch's shadow-cljs.edn and taking "the next free one" –
correct information for the branch it could see, wrong by the time it
tried to land. Five of the six later branches independently chose port
8729, since from each one's point of view that was the first unclaimed
port after counting-change's 8728. Landing them required six sequential
rebase-resolve-renumber passes, in merge order, each one bumping the
newly-conflicting branch to the next actually-free port (8729 -> 8730 ->
8731 -> 8732 -> 8733 -> 8734). Every conflict landed on the same three
files every time: shadow-cljs.edn's port block, Makefile's
TOOLS_MODULES line, and site/tools/index.org's listing lines – never a
logic conflict, always a registration-list conflict, which is exactly the
shape a merge-outcome predictor keyed on "touches a shared manifest" would
flag cheaply, without needing to understand either branch's actual
content.
4.3. Bugs the build caught, by tool
- interval-arithmetic:
core.cljcandbrowser.cljsboth usedclojure.stringand the tool's owncorenamespace without a:requirefor them – silently fine under whatever REPL state happened to already have those symbols interned, a real compile failure under a clean ClojureScript build. - primality-test:
Math/round(JVM interop syntax) used inside.cljs, invalid there; needsjs/Math.round. Separately, a naive O(e) verification loop with e on the order of 10^12 was caught before a full test run and replaced with ajava.math.BigInteger.modPowcross-check – the loop would have hung for hours, not failed loudly. - representing-sets: a
StringBuildercall (no such class in ClojureScript) in the render path; a structures panel whose label always read "ordered list" regardless of which of the three representations was actually selected, because the render function never consulted the selected representation to choose its label; a validation-message element with markup but no toggle logic wired to it. - huffman-trees: two test-assertion bugs, not implementation bugs – a test's own expected shape was wrong, and a round-trip property hadn't accounted for the tool's documented single-leaf no-decode special case (an empty code carries no information, so decoding it is undefined by construction, not a bug to round-trip against).
- picture-language: the build brief itself (written by the
orchestrating session, not the reference artifact) asserted a wrong
invariant – that
right-split(p, n)produces(n+1)*ksegments for ak-segment primitivep. The recursion isright-split(p, n) = beside(p, below(s, s))withs = right-split(p, n-1), which givesf(n) = 1 + 2*f(n-1), i.e. =2^(n+1) - 1=segments, not linear growth. The subagent derived the correct closed form by hand, wrote the property test against it, and flagged the discrepancy in its own report rather than silently matching the wrong brief. - symbolic-algebra: a missing closing paren in
mul-terms's sparse branch silently swallowed its entire test namespace on the first full run – no visibleLOAD-FAILin the noisy terminal output, since carriage-return-based progress lines from an unrelated long test obscured it. Finding it took writing a small bracket-balance checker over the file. A second, independent bug: atest.checkgenerator could produce a single order-0-term polynomial that formats identically to a plain number, breaking a round-trip property through the string representation – fixed by excluding that degenerate shape from the generator rather than weakening the property it was checking.
Every one of these was caught before merge, by the same two checks the
"Build process" section above describes for the original ten-page effort:
a Node/JVM-side numeric check before the UI, and (here, since this was a
.cljc=/.cljs= port rather than a from-scratch build) a full property-test
suite run in the foreground, twice, per tool.
4.4. A model-behavior finding, not in the original ten-page build
Two of the six port-conflict-resolution subagents, when told to rebase a
branch and resolve the resulting shadow-cljs.edn=/=Makefile=/=index.org
conflicts, ran a test suite, then stopped their own turn with language
like "I'll wait for this notification before checking results" – as if
they were an orchestrator with access to the parent session's background
task-notification channel. They are not; a subagent has no such channel,
and the work (the actual `git push`) never happened. The fix both times
was the orchestrating session checking the worktree directly, finding the
rebase still sitting unpushed, and either resuming the subagent with an
explicit "block synchronously, no backgrounding" instruction or finishing
the mechanical rebase-and-push itself. Delegating a bounded, well-specified
git operation to a subagent worked five times out of six in this session;
the fully mechanical fix, done directly, took less time than the second
attempt at re-briefing the stalled subagent.
4.5. Disclosed gap common to all seven
None of the seven tools' ClojureScript bundles have been compiled in this
environment – no node_modules=/shadow-cljs toolchain was available to any
of the sessions that built them, matching this repository's own tiering of
cljs build work to a separate dev box. Every subagent said so plainly in
its own PR body rather than claiming an unverified build worked. Compiling
and smoke-testing all seven (=gmake tools-cljs, then gmake dev-tool
TOOL=<slug> per tool) is the remaining step before any of them are live.
5. Recommendations for other problems
Where the style fits. It suits algorithms whose cost or correctness depends on an internal decision that the output hides. Examples are recursion shape, queue order, representation choice, dispatch, and propagation order. Candidates include:
- the SICP chapter 4 evaluator, with environment frames as the state;
- stream and delayed evaluation (§3.5), showing when each element is forced;
- register machines (§5.2), with registers and stack as snapshots;
- garbage collection (§5.3), showing the two memory halves;
- unification and query evaluation (§4.4);
- parsers, type inference, SAT/DPLL, and any consensus protocol with a message log.
Designing the events.
- Define one event per decision the learner should notice, not one per line of code. Too fine a granularity produces thousands of steps with little change between them; too coarse hides the decision.
- Add a jump control at the level the learner thinks in: next user operation, next simulated time, next phase, next table lookup.
Pairing each page with toggles.
- Pair each toggle with a specific failure mode and a detector for it. Ex. 3.31 and 3.32 were detected by comparing the outputs to the logic evaluated on the final inputs. The frame bugs were detected by a coordinate diff against a reference run. Interval overestimation was detected by sampling.
- Put a reference implementation next to the variant under test. Tests that check only that something happened, such as "something was drawn," passed for both buggy frame maps. Only coordinate comparison failed them. Showing both kinds of test side by side makes that point without an extra explanation.
Limits and preview.
- Cap the event count and report truncation with the exact figure computed another way. Counting change reported the naive call count from a recurrence when the trace stopped at 60,000 events.
- Show future state in grey. Faint tree nodes, dashed waveforms and grey segments tell the learner where the run is going, which scrubbing alone does not.
For the ClojureScript port.
- Keep the model free of DOM references, so it can run under Node, in property tests and in the browser.
- Traces can be checked with property tests on invariants. Examples:
- every
retmatches acall; - every constraint holds after each completed user operation when contradiction handling is on;
- agenda times are non-decreasing across
runevents.
- every
6. Issues encountered
Model errors found by the Node checks
- Circuits, LIFO stimuli. The first circuit version put each input
change on the agenda as a separate item. With LIFO this also reversed
the order of the input changes, so the ex. 3.32 glitch did not appear.
The fix groups simultaneous input changes into one agenda action that
calls
set-signal!in order, which matches SICP's sequential calls. Ripple-adder carry delay. The first carry-delay formula was and + and
- or = 11 per full-adder. It was wrong: the observed last change at
time 56 exceeded the claimed bound of 44. It was replaced with a longest-path computation over the gate graph, which gives 16 per full-adder and a bound of 64.
- Huffman tokenization. "Set weights from this message" tokenized using the previous preset's symbols. After the rock-song preset, ABRACADABRA became one word-token. The fix tokenizes on the message alone.
- Symbolic differentiation printing. The first infix printer had a precedence rule that was hard to reason about. It was rewritten around the parent operator.
- Symbolic algebra printing. Nested polynomial coefficients were
printed with redundant parentheses, e.g.
(y). Parentheses are now used only for multi-term coefficients.
Layout and CSS
'hidden'attribute. Elements withdisplay: flexin CSS stayed visible when sethidden. The fix was an explicit[hidden] \{ display: none \}rule; this recurred on three pages before becoming a habit.- Overlapping curves. On the sets growth chart the unordered-list curve was invisible. It coincided exactly with the sorted-input tree curve, since both make n(n-1)/2 comparisons. It is now drawn dashed on top, with a caption stating the coincidence.
- Overflow. SVGs overflowed their panels on the pictures and constraints pages. The call-tree labels in the derivative page overflowed their boxes and are now truncated by character count.
- Pseudocode. Lines with padded comments wrapped badly in the 390-420 px side column.
Test harness
- Playwright
text=selectors matched the lede paragraph instead of the intended radio button, so one Miller-Rabin screenshot showed the wrong mode. Later tests selected inputs by value. - Google Fonts returned 403 in the sandbox, so screenshots used fallback fonts. Layout widths in the published pages may differ slightly from what was checked.
Numeric representation
- Primality needs BigInt for expmod. The other pages use JS doubles.
- Interval arithmetic displays four significant digits, which hides rounding.
- The constraint page compares values with a relative tolerance of 1e-12 where SICP uses exact equality. This avoids contradictions from floating-point division.
- Symbolic algebra uses doubles for rationals, which will lose precision for large coefficients.
Deviations from SICP to state when porting
- Primality uses standard Miller-Rabin, not the ex. 1.28 variant that
checks for nontrivial square roots inside
expmod. - Differentiation adds a subtraction rule as an extension.
- Algebra uses alphabetical variable order for ex. 2.92, and raises
through rational on the way to polynomial, which produces
2/1-style coefficients unless drop is on. - The sets and Huffman pages copy remaining list elements one step at a time, where SICP returns the rest of the list in one operation. This adds steps that SICP's procedures do not take.
Performance and duplication
- Most renderers rebuild their SVG or DOM on every step. This is fine at the current sizes but will be slow for traces above a few thousand events. It would need keyed updates or canvas.
- The transport, strip, keyboard handling and theme code are duplicated in all ten files, about 150 lines each. The ClojureScript version should factor this out first.
Not covered
- There is no URL state, so a configuration cannot be shared as a link.
- Accessibility is partial. Controls have labels, but canvas views have no text alternative beyond their captions.
- Mobile layouts were not tested at narrow widths beyond the CSS breakpoint.
7. Related research
- 2026 Agentic Software Engineering – the port-collision story in this note's build log (§) is a small, concrete instance of exactly the "many agent-authored changes competing to land on one trunk" problem that note surveys.
- Tools index – all ten pages, once merged and deployed, live under one navigation section there.