Tomer Ullman: The Bare Stage Before the Mind's Eye
Harvard SEAS Computer Science Lecture Series, 8 October 2026
Table of Contents
Event
| Field | Value |
|---|---|
| Talk | The Bare Stage Before the Mind's Eye |
| Speaker | Tomer Ullman, Morris Kahn Associate Professor of Psychology, Harvard |
| Series | Computer Science Lecture Series, Harvard SEAS |
| Date | EDT |
| Schedule | Lunch 11:45, lecture 12:00, ends 13:30 |
| Venue | Science and Engineering Complex (SEC) 3.301, 150 Western Avenue, Allston MA 02134 |
| Format | Colloquium, open to the general public |
| Registration | Google Form |
| URL | events.seas.harvard.edu |
The argument
Imagining a scene feels like producing the whole scene: full, detailed, vivid. Perception gives the same impression — "a full, rich world before us in perception", in the abstract's words — and that impression is one cognitive science has spent decades taking apart. The talk applies the same correction to imagination. Recent work suggests the mind's eye builds partial scenes, leaves some parts uncommitted, and does much of its work with rough, good-enough representations rather than rendered detail.
The title carries the claim. A bare stage holds the actors and the blocking; the set dressing is never built.
Background: Ullman's work on the topic
Non-commitment in mental imagery
Bigelow, McCoy and Ullman (2023) (Bigelow, McCoy, and Ullman 2023) asked people to imagine a specified scene, then asked about basic properties of it. Across five studies (N > 1,800), most people were non-committal about features any real image would have to show. They reported non-commitment rather than uncertainty or forgetting, people with vivid imaginations did the same, and when "not committed" was not offered as an answer, people confabulated one.
The obvious objection is that the detail was there and simply went unnoticed, as it does in real scenes. Li, Hammond and Ullman (Li, Hammond, and Ullman 2026) test that directly in a preprint whose second version was posted the day before this talk. Across three experiments (N = 3,895), the properties left uncommitted in imagination did not correlate with the properties missed in briefly shown real images, so non-commitment is not perceptual noise. Imagery also unfolds in a fixed order: entities and spatial relations, then physical properties, then surface detail like colour and texture. Trying to interrupt or constrain construction did not change the order. The authors read this as hierarchical partial-scene construction. This paper is the most likely spine of the talk.
Capacity limits of imagined motion
Balaban and Ullman (2025) (Balaban and Ullman 2025) measured how many moving objects people can simulate in imagination. Nine preregistered experiments (N = 313) used an Imagined Objects Tracking task: watch objects move, then continue their motion mentally. One object is tracked in line with ground truth. With two, responses fit a serial model that simulates one object at a time better than a parallel one. Strong grouping cues reduce the bottleneck without removing it. Perception tracks several objects at once; imagination, on this evidence, moves one.
The game engine in the head
The earlier frame is Ullman, Spelke, Battaglia and Tenenbaum (2017) (Ullman et al. 2017): intuitive physics as something like a game engine, approximate, object-based simulation tuned for speed over accuracy. The imagery work reads as the next question. If the mind runs a game engine, which assets does it actually load? Game engines also cull what is off screen and skip textures nobody is looking at; non-commitment is the cognitive version of that.
The child as hacker
Rule, Tenenbaum and Piantadosi (2020) (Rule, Tenenbaum, and Piantadosi 2020) (Ullman is not an author) cast learning as program construction: children write, debug and refactor mental programs. It sits in the same research program as Lake, Ullman, Tenenbaum and Gershman (2017) (Lake et al. 2017) on building machines that learn like people, and it supplies the representational bet behind this talk: a scene in the head looks more like a program with unbound variables than like a bitmap.
LLM theory of mind
Ullman (2023) (Ullman 2023) took a reported success of large language models on false-belief tasks and showed that small changes preserving the logic of the task (a transparent container, for example) turned the results over. His conclusion was methodological: the default stance toward a model's intuitive psychology should be skeptical, and outlying failures should outweigh average success. Hu, Sosa and Ullman (2025) (Hu, Sosa, and Ullman 2025) widen this into a review: the field disagrees partly because it has not decided whether models should match human behaviour or the computation behind it.
That evaluation stance is the same one the imagery work applies to people. Self-report says the scene is complete; the probe says otherwise.
The perception literature the abstract alludes to
- Change blindness: large changes to a scene go unnoticed across a disruption (Simons and Levin 1997).
- Inattentional blindness: a person in a gorilla suit walks through a basketball game unseen by many observers counting passes (Simons and Chabris 1999).
- The "grand illusion" reading: the richness of the visual world is a feeling, not a stored representation; the world serves as its own outside memory, sampled on demand (O’Regan and Noë 2001). Noë, Pessoa and Thompson (Noë, Pessoa, and Thompson 2000) and Noë (Noë 2002) argue the illusion framing itself overstates the case. The bibliographic details of the 2002 paper are unconfirmed.
The 2026 preprint is the bridge: it is built to tell imagination's gaps apart from perception's.
Speaker
Tomer Ullman is the Morris Kahn Associate Professor of Psychology at Harvard. Per the SEAS listing: B.Sc. in Cognitive Science and Physics, Hebrew University, 2008; Ph.D. in Brain and Cognitive Sciences, MIT, 2015; postdoctoral work at the Center for Brains, Minds, and Machines, 2015–2018. His research covers commonsense reasoning, intuitive theories of agents and objects, and how children and adults learn them; the aim is a functional and algorithmic account, in service of more human-like AI.
The biography is taken from the listing and not checked independently.
Questions to bring
- Does the hierarchical order (entities, relations, physics, surface) hold when the task requires a surface property early, or does the task reorder construction?
- Is the serial bottleneck in imagined motion a property of the simulator, or of attention directed at the simulation?
- Do image and video models show anything like non-commitment, or do they commit to every pixel by construction? Ullman has a 2024 paper on vision language models seeing illusions where there are none, so he may have a view.
- What would a probabilistic program with deliberately unbound variables predict about confabulation when "not committed" is not an allowed answer?
Related repositories and books
Repositories (local clones under ~/ghq/github.com/):
| Path | GitHub | Visibility | Why |
|---|---|---|---|
jwalsh/webppl-example |
jwalsh/webppl-example | private | minimal WebPPL setup; WebPPL is the language of ProbMods |
jwalsh/racketcon-2025 |
jwalsh/racketcon-2025 | public | experiments 086–095: probabilistic programming with Roulette |
jwalsh/saylent-notes |
jwalsh/saylent-notes | private | 2017–2018 notes on Gamble, WebPPL, and Mansinghka's probcomp talk |
defrecord/journal |
defrecord/journal | private | POPL 2025 LAFI notes: memo (a PPL for reasoning about reasoning; Chandra, Chen, Tenenbaum, Ragan-Kelley) and lazy knowledge compilation for discrete PPLs |
No local clone of Gen, Church, an intuitive-physics engine, or a theory-of-mind evaluation suite was found.
Books on this site's reading list:
- Eric Kandel, The Age of Insight — the "beholder's share": the viewer completes the picture, the perceptual analogue of an uncommitted scene.
- Daniel Kahneman, Thinking, Fast and Slow — confident impressions built from little evidence.
- Kevin P. Murphy, Machine Learning: A Probabilistic Perspective — the inference machinery under the Bayesian cognition program.
Canonical reading, not yet on the shelf:
- Goodman, Tenenbaum and the ProbMods Contributors, Probabilistic Models of Cognition, 2nd ed. (Goodman, Tenenbaum, and The ProbMods Contributors 2016) — intuitive physics and social cognition as WebPPL programs.
- Lake, Ullman, Tenenbaum and Gershman, "Building machines that learn and think like people", Behavioral and Brain Sciences (Lake et al. 2017).
- Elizabeth Spelke, What Babies Know (Oxford, 2022) (Spelke 2022) — core knowledge of objects and agents.
Notes
Notes pending.