Ten Trace Explorers: A SICP Build Reflection

Table of Contents

1. Overview

Ten small ClojureScript tools trace SICP procedures step by step: a wire holds a signal, a painter draws into a frame, an agenda runs actions in time order. Each tool shares one architecture – a pure model emits an event log, three different strategies compute state at a given step, and dumb renderers turn that state into DOM, SVG, or canvas. This note is the reflection on building them: what the shared structure buys, what broke, and what a 2026-09-19 session of parallel-subagent builds looked like on the ground.

2. Build process

Every page followed the same sequence:

  1. Choose which internal steps of the algorithm count as events.
  2. Write the model as a pure function that returns an event list.
  3. Run the model in Node against known values before writing any UI. These were SICP's own numbers where they exist: 15,499 calls for cc(100), 42 bits for the §2.3.4 message, the half-adder timings of 8, 11 and 16, and the parallel-resistor intervals [2.201, 3.487] and [2.582, 2.973].
  4. Write the page.
  5. Render it in headless Chromium, take screenshots at chosen steps, and read the results to find layout and logic faults.
  6. Publish.

The Node check before the UI caught more errors than the screenshots did, because a wrong number is easier to see in a printed table than in a picture.

The first page set the visual system: colour tokens, the transport bar, the strip canvas, the pseudocode panel with line highlighting, and a result sentence above the timeline. Later pages copied it and changed only the model and the central view. This kept every page after the first to roughly one model, one view and one round of fixes.

3. Structure of a trace explorer

Each page has four layers.

3.1. Model

A pure function of (inputs, options) returns a trace. The options are the toggles: memoization, queue order, frame-coord-map variant, constructor style, and so on. The trace is an array of events. Every event carries at least:

  • a type;
  • the pseudocode line numbers it corresponds to;
  • a sentence describing it;
  • the identifiers of the objects it touches, such as a node, a connector, a table cell or a worklist index.

3.2. State at step s

The pages used three methods to compute it.

  • Replay with a reducer (counting change). The state is reduce(applyEvent, fresh, events[0..s]). It moves forward incrementally and replays from zero when moving backward. Memory is small, but a backward seek costs O(s).
  • Derived position (primality, intervals). The state is computed from event indices alone. For example, frames called = clamp(k, 0, L), where k is the number of events since the current run began. This works only when the event structure is regular.
  • Snapshots (sets, Huffman, circuits, constraints, algebra, pictures). Each event, or each structural change, stores a copy of the relevant state. The copy is either a versions list \{at, value\} per structure or full value arrays per event. Memory grows with the number of events, but seeking is O(1) or O(versions).

Snapshots were the most reliable choice for mutable algorithms, i.e. circuits and constraints. There the simulator can mutate freely during the run and the trace is still immutable data afterward.

3.3. Renderers

These read (trace, s) and write DOM, SVG or canvas. They contain no algorithm logic. The event text is produced by the model, not by the renderer, and this moved every "what just happened" explanation into the part of the code that has the relevant values.

3.4. Comparison runs

Each page runs the model more than once per input: memo on and off, sparse and dense, FIFO and LIFO, the correct and the buggy frame map, all four set representations. The results go into a table. The table carries most of the explanation, because it shows a cost or correctness difference as numbers next to the trace that produced them.

4. Session build log (2026-09-19)

This section is the concrete evidence from one build session, in which seven of the ten – counting-change, primality-test, interval-arithmetic, representing-sets, huffman-trees, picture-language and symbolic-algebra – were built by parallel subagents against this repository's own tools convention and merged the same day. digital-circuits and propagation-of-constraints followed the same day; a tenth page, symbolic differentiation, is referenced in the "Issues encountered" section below as a build already in the reflection's scope but is not part of this session's own merge log.

4.1. What actually landed

Tool SICP § PR Port (final)
counting-change 1.2.2 #104 8728
primality-test 1.2.6 #105 8729
interval-arithmetic 2.1.4 #106 8730
representing-sets 2.3.3 #107 8731
huffman-trees 2.3.4 #108 8732
picture-language 2.2.4 #109 8733
symbolic-algebra 2.5 #110 8734

Every one of the seven was built against the same tool-maintenance-spec v1.0.0 and tool-maintainer skill v1.0.0.

4.2. The build itself is an instance of the "many agents, one trunk" problem

Six of the seven tools branched from a main that did not yet contain the others – exactly the setup the agentic-software-engineering literature calls agent-scale merge contention. Each subagent picked its shadow-cljs dev-server port by reading its own branch's shadow-cljs.edn and taking "the next free one" – correct information for the branch it could see, wrong by the time it tried to land. Five of the six later branches independently chose port 8729, since from each one's point of view that was the first unclaimed port after counting-change's 8728. Landing them required six sequential rebase-resolve-renumber passes, in merge order, each one bumping the newly-conflicting branch to the next actually-free port (8729 -> 8730 -> 8731 -> 8732 -> 8733 -> 8734). Every conflict landed on the same three files every time: shadow-cljs.edn's port block, Makefile's TOOLS_MODULES line, and site/tools/index.org's listing lines – never a logic conflict, always a registration-list conflict, which is exactly the shape a merge-outcome predictor keyed on "touches a shared manifest" would flag cheaply, without needing to understand either branch's actual content.

4.3. Bugs the build caught, by tool

  • interval-arithmetic: core.cljc and browser.cljs both used clojure.string and the tool's own core namespace without a :require for them – silently fine under whatever REPL state happened to already have those symbols interned, a real compile failure under a clean ClojureScript build.
  • primality-test: Math/round (JVM interop syntax) used inside .cljs, invalid there; needs js/Math.round. Separately, a naive O(e) verification loop with e on the order of 10^12 was caught before a full test run and replaced with a java.math.BigInteger.modPow cross-check – the loop would have hung for hours, not failed loudly.
  • representing-sets: a StringBuilder call (no such class in ClojureScript) in the render path; a structures panel whose label always read "ordered list" regardless of which of the three representations was actually selected, because the render function never consulted the selected representation to choose its label; a validation-message element with markup but no toggle logic wired to it.
  • huffman-trees: two test-assertion bugs, not implementation bugs – a test's own expected shape was wrong, and a round-trip property hadn't accounted for the tool's documented single-leaf no-decode special case (an empty code carries no information, so decoding it is undefined by construction, not a bug to round-trip against).
  • picture-language: the build brief itself (written by the orchestrating session, not the reference artifact) asserted a wrong invariant – that right-split(p, n) produces (n+1)*k segments for a k-segment primitive p. The recursion is right-split(p, n) = beside(p, below(s, s)) with s = right-split(p, n-1), which gives f(n) = 1 + 2*f(n-1), i.e. =2^(n+1) - 1=segments, not linear growth. The subagent derived the correct closed form by hand, wrote the property test against it, and flagged the discrepancy in its own report rather than silently matching the wrong brief.
  • symbolic-algebra: a missing closing paren in mul-terms's sparse branch silently swallowed its entire test namespace on the first full run – no visible LOAD-FAIL in the noisy terminal output, since carriage-return-based progress lines from an unrelated long test obscured it. Finding it took writing a small bracket-balance checker over the file. A second, independent bug: a test.check generator could produce a single order-0-term polynomial that formats identically to a plain number, breaking a round-trip property through the string representation – fixed by excluding that degenerate shape from the generator rather than weakening the property it was checking.

Every one of these was caught before merge, by the same two checks the "Build process" section above describes for the original ten-page effort: a Node/JVM-side numeric check before the UI, and (here, since this was a .cljc=/.cljs= port rather than a from-scratch build) a full property-test suite run in the foreground, twice, per tool.

4.4. A model-behavior finding, not in the original ten-page build

Two of the six port-conflict-resolution subagents, when told to rebase a branch and resolve the resulting shadow-cljs.edn=/=Makefile=/=index.org conflicts, ran a test suite, then stopped their own turn with language like "I'll wait for this notification before checking results" – as if they were an orchestrator with access to the parent session's background task-notification channel. They are not; a subagent has no such channel, and the work (the actual `git push`) never happened. The fix both times was the orchestrating session checking the worktree directly, finding the rebase still sitting unpushed, and either resuming the subagent with an explicit "block synchronously, no backgrounding" instruction or finishing the mechanical rebase-and-push itself. Delegating a bounded, well-specified git operation to a subagent worked five times out of six in this session; the fully mechanical fix, done directly, took less time than the second attempt at re-briefing the stalled subagent.

4.5. Disclosed gap common to all seven

None of the seven tools' ClojureScript bundles have been compiled in this environment – no node_modules=/shadow-cljs toolchain was available to any of the sessions that built them, matching this repository's own tiering of cljs build work to a separate dev box. Every subagent said so plainly in its own PR body rather than claiming an unverified build worked. Compiling and smoke-testing all seven (=gmake tools-cljs, then gmake dev-tool TOOL=<slug> per tool) is the remaining step before any of them are live.

5. Recommendations for other problems

Where the style fits. It suits algorithms whose cost or correctness depends on an internal decision that the output hides. Examples are recursion shape, queue order, representation choice, dispatch, and propagation order. Candidates include:

  • the SICP chapter 4 evaluator, with environment frames as the state;
  • stream and delayed evaluation (§3.5), showing when each element is forced;
  • register machines (§5.2), with registers and stack as snapshots;
  • garbage collection (§5.3), showing the two memory halves;
  • unification and query evaluation (§4.4);
  • parsers, type inference, SAT/DPLL, and any consensus protocol with a message log.

Designing the events.

  • Define one event per decision the learner should notice, not one per line of code. Too fine a granularity produces thousands of steps with little change between them; too coarse hides the decision.
  • Add a jump control at the level the learner thinks in: next user operation, next simulated time, next phase, next table lookup.

Pairing each page with toggles.

  • Pair each toggle with a specific failure mode and a detector for it. Ex. 3.31 and 3.32 were detected by comparing the outputs to the logic evaluated on the final inputs. The frame bugs were detected by a coordinate diff against a reference run. Interval overestimation was detected by sampling.
  • Put a reference implementation next to the variant under test. Tests that check only that something happened, such as "something was drawn," passed for both buggy frame maps. Only coordinate comparison failed them. Showing both kinds of test side by side makes that point without an extra explanation.

Limits and preview.

  • Cap the event count and report truncation with the exact figure computed another way. Counting change reported the naive call count from a recurrence when the trace stopped at 60,000 events.
  • Show future state in grey. Faint tree nodes, dashed waveforms and grey segments tell the learner where the run is going, which scrubbing alone does not.

For the ClojureScript port.

  • Keep the model free of DOM references, so it can run under Node, in property tests and in the browser.
  • Traces can be checked with property tests on invariants. Examples:
    • every ret matches a call;
    • every constraint holds after each completed user operation when contradiction handling is on;
    • agenda times are non-decreasing across run events.

6. Issues encountered

Model errors found by the Node checks

  • Circuits, LIFO stimuli. The first circuit version put each input change on the agenda as a separate item. With LIFO this also reversed the order of the input changes, so the ex. 3.32 glitch did not appear. The fix groups simultaneous input changes into one agenda action that calls set-signal! in order, which matches SICP's sequential calls.
  • Ripple-adder carry delay. The first carry-delay formula was and + and

    • or = 11 per full-adder. It was wrong: the observed last change at

    time 56 exceeded the claimed bound of 44. It was replaced with a longest-path computation over the gate graph, which gives 16 per full-adder and a bound of 64.

  • Huffman tokenization. "Set weights from this message" tokenized using the previous preset's symbols. After the rock-song preset, ABRACADABRA became one word-token. The fix tokenizes on the message alone.
  • Symbolic differentiation printing. The first infix printer had a precedence rule that was hard to reason about. It was rewritten around the parent operator.
  • Symbolic algebra printing. Nested polynomial coefficients were printed with redundant parentheses, e.g. (y). Parentheses are now used only for multi-term coefficients.

Layout and CSS

  • 'hidden' attribute. Elements with display: flex in CSS stayed visible when set hidden. The fix was an explicit [hidden] \{ display: none \} rule; this recurred on three pages before becoming a habit.
  • Overlapping curves. On the sets growth chart the unordered-list curve was invisible. It coincided exactly with the sorted-input tree curve, since both make n(n-1)/2 comparisons. It is now drawn dashed on top, with a caption stating the coincidence.
  • Overflow. SVGs overflowed their panels on the pictures and constraints pages. The call-tree labels in the derivative page overflowed their boxes and are now truncated by character count.
  • Pseudocode. Lines with padded comments wrapped badly in the 390-420 px side column.

Test harness

  • Playwright text= selectors matched the lede paragraph instead of the intended radio button, so one Miller-Rabin screenshot showed the wrong mode. Later tests selected inputs by value.
  • Google Fonts returned 403 in the sandbox, so screenshots used fallback fonts. Layout widths in the published pages may differ slightly from what was checked.

Numeric representation

  • Primality needs BigInt for expmod. The other pages use JS doubles.
  • Interval arithmetic displays four significant digits, which hides rounding.
  • The constraint page compares values with a relative tolerance of 1e-12 where SICP uses exact equality. This avoids contradictions from floating-point division.
  • Symbolic algebra uses doubles for rationals, which will lose precision for large coefficients.

Deviations from SICP to state when porting

  • Primality uses standard Miller-Rabin, not the ex. 1.28 variant that checks for nontrivial square roots inside expmod.
  • Differentiation adds a subtraction rule as an extension.
  • Algebra uses alphabetical variable order for ex. 2.92, and raises through rational on the way to polynomial, which produces 2/1-style coefficients unless drop is on.
  • The sets and Huffman pages copy remaining list elements one step at a time, where SICP returns the rest of the list in one operation. This adds steps that SICP's procedures do not take.

Performance and duplication

  • Most renderers rebuild their SVG or DOM on every step. This is fine at the current sizes but will be slow for traces above a few thousand events. It would need keyed updates or canvas.
  • The transport, strip, keyboard handling and theme code are duplicated in all ten files, about 150 lines each. The ClojureScript version should factor this out first.

Not covered

  • There is no URL state, so a configuration cannot be shared as a link.
  • Accessibility is partial. Controls have labels, but canvas views have no text alternative beyond their captions.
  • Mobile layouts were not tested at narrow widths beyond the CSS breakpoint.

7. Related research

  • 2026 Agentic Software Engineering – the port-collision story in this note's build log (§) is a small, concrete instance of exactly the "many agent-authored changes competing to land on one trunk" problem that note surveys.
  • Tools index – all ten pages, once merged and deployed, live under one navigation section there.