Giving the workspace something to do
Partial. 2026-06-26 · a win and its second honest negative · experiment AL
The parked serial workspace finally reaches a two-hop target where System 1 and one-step deferral both score zero, trapped on the intermediate. Reachability is real and new. But on a single deterministic chain the elaborate machinery only ties a blind "apply-twice." Its focus and inhibition are built to select among competing chains, and this probe has none. The next axis is named.
The question
Two experiments ago we built a System 2 out of counts. It had two parts. A gate that decides when to stop and think. And a workspace: a small, capacity-four serial scratchpad that holds a few candidates, suppresses the loud one, and deliberates. The gate worked beautifully. The workspace did not: on a character next-token task it lost to the most trivial possible rule, just defer to the wider context, because predicting one character is a one-step decision. There was nothing to hold, nothing to manipulate. A scratchpad with nothing to scratch.
We parked it rather than kill it, and wrote down the one condition under which it might earn its keep: a problem that takes more than one step. This experiment builds that problem and asks the workspace to show up.
What we tried
We made a two-hop question out of pure counts. For every concept we counted its single strongest content associate, a relation we can apply like a function: R of X gives Y. Then we chained it: found triples where R of X is Y, R of Y is Z, and Z is neither X nor Y. The question: starting at X, what is R of R of X? The answer is Z. Real concepts, not function-word filler: supernatural → beings → human, carnegie → andrew → jackson.
The trap is the point. The loud, salient associate of X is Y, the intermediate. Apply the relation once and you land on Y, which is wrong. To reach Z you must hold Y in mind and apply the relation a second time, the thing the workspace was built to do and never got to. Three contestants: System 1 (the fast associative blurt), the one-step deferral (last time's winner), and the multi-step workspace (hold X, apply, hold the result, apply again, read the focus).
What happened
| contestant | reached the two-hop target | landed on the trap |
|---|---|---|
| System 1 (the loud associate) | 0.00 | 100% |
| one-step deferral (last time's winner) | 0.00 | 100% |
| multi-step workspace | 1.00 | 0% |
The workspace reached the target every time; the two one-step contestants reached it never. The rule that captured the entire win on the character task now scores zero, because one step cannot get you two hops away. Set the workspace's step budget to zero and it collapses cleanly back to the one-step answer: no capacity to think, ship the fast default, never an empty answer.
Then we checked ourselves, because the result was suspiciously clean. The target is defined as apply the relation twice, so anything that applies it twice will get it. We raced the full workspace against a naive blind double-application: no focus, no inhibition, no gate. On the clean relation they tie. We added noise to the relation's edges. They still tie, at every noise level. The gate correctly declines to chase a corrupted chain, but its fallback is the one-step answer, which is also wrong for a two-hop target. Declining buys nothing here.
So the win is real and specific. The workspace's value is reachability: it is the only operator that can reach a two-hop target at all. The elaborate machinery wrapped around it (the capacity-four focus, the inhibition-of-return, the suppress-not-erase floor) adds nothing over applying the relation twice in a row. That is the workspace's second honest negative, with a sharper reason: focus and inhibition are built to choose between competing chains, and a single deterministic chain gives them nothing to choose. A wrong first hop poisons every answer equally. There is no selection to make.
The lesson
A count-native System 2 needs two parts that win on two different axes. The gate decides whether to think, and wins on knowing when the fast answer is wrong. The workspace decides what to compose, and wins on reaching answers more than one step away, now demonstrated, online, no gradient. But the workspace's elaborate inner machinery is still no better than the minimal version of its operator. The metacognition is load-bearing. The serial bookkeeping is not, yet. Its untested home is the next probe: not one chain to follow but several candidate chains that compete, where suppressing the loud wrong answer should finally pay.
Lineage
Grew from thinking slow, by counting, which parked this workspace and named the multi-step task as its home, and uses the redescribed concepts of when a habit becomes a thought as the operands it chains, the same role-filler bundles reading structure back out of a sum decodes.
Thread: System 2, the reasoning route. The cliffhanger is resolved; the next axis is competing-chains selection.