From mimicry to production

The plan. 2026-06-26 · why a reader is not a speaker, and the loop that forces the crossing · the production library

A reader predicts the next token to match the stream. A speaker emits a token to move something: a listener, a goal, the world. That is a different thing, and reading does not become it for free. We already crossed the single hardest threshold the gap predicts: a reply that answers you teaches more than the same words overheard cold. That is the engine. What is missing is the steering: which utterance to repeat, and for whom. This is the library of count-native organs that add the steering, every one tied to a standing rule and an experiment we can kill.

The question

The cortex can read. Three years of experiments built that: counts beat gradients, surprise carves a boundary, a chunk lexicon hands the cortex whole units to say, and a coverage-competition producer says them more well-formed than gibberish. The acquisition round ended with the cortex taking its first rough step from reader toward speaker.

But there is a quiet assumption hiding in "production is comprehension read the hard way," and it is worth saying out loud. A model that reads predicts the next token to match the stream. That is mimicry, fidelity to the input. It is not yet communication. Communication is when an emitted token is selected for its effect: because it will move a listener toward a referent, because it will get you the thing you want, because it fills a gap you noticed in what the other person knows. A parrot mimics. A child who says "milk" and gets milk is doing something else. The question this library answers: what does the environment have to force for the second thing to appear, and what count-native organs ride that force?

What the environment must force: you produce to get something

The answer is not a richer model. It is a loop. A child does not learn to talk by reading more; it learns by talking and getting a reply. So the first thing the harness had to grow was a mouth that speaks into a world that reacts, and the first thing we had to prove was that the reaction teaches.

It does. Exp BE is the load-bearing result under this whole library: with the same tokens and only their timing scrambled in a yoked control, a genuinely contingent reply wins by +0.45 bpc, and the gap vanishes exactly when timing stops predicting content, the signature of a real mechanism, not an artifact. A token that closed a turn is worth more than the same token read passively. Contingency is the engine that turns reading into producing-for-a-response.

And the engine now runs against a live model. No key was needed: the machine's logged-in claude CLI is an authenticated partner, so the loop talks to claude-haiku-4-5 for one short turn per reply. Across forty real replies, contingency-ON beats the scrambled-timing yoked control at 12 of 12 sweep settings (Exp BM). One line we hold firmly: the agent's utterances are still gibberish, so the model often answered them as nonsense. What is live-confirmed is the contingency, that a reply timed to answer you teaches more than the same reply re-timed to ignore you, not yet a real conversation.

What is missing is the steering

Contingency rewards any warm token equally. That is not enough, and the gap is exactly the interesting part. The same word "ball" can be three different acts: an echo (you just heard it), a name (the ball is in front of you), or a request that got you the ball. Right now the cortex learns all three identically, by surface. The functional layer above raw timing (Skinner's old insight that an operant is defined by what precedes and follows it, not by its shape) is unbuilt. The library's production organs build it as three count-views of the same emitted chunk, routed by which antecedent matches now. Reward stops attaching to a whole warm turn and starts attaching to a producible unit that got something.

That is the steering wheel. Three more loops turn it.

Metacognition: when to emit, when to think, when to ask

The cortex already produces a confidence scalar. The deliberate pass reads the conflict between the top two competitors (high only when two answers are both live) and the calibrated f·c truth value tells it how sure it is, with no tuning. Today that scalar is spent defensively: override System 1 when it's about to be wrong.

A speaker needs to read the same scalar three ways, not one. Emit when sure. Deliberate when the uncertainty is in the form: you know what to say, you're competing over how. And ask when the uncertainty is in the goal: the context itself is missing, and the right move is not to guess but to make the next input answer you. The new dial is that one routing key: is the gap in the goal or the form? Everything else is machinery the cortex already has, rewired into the act of speaking. The "ask" is the genuinely new action: a first-class competitor to emission, the count-native version of an infant who reaches for help precisely when it knows it doesn't know.

A world model, but only the part that survived

Comprehension builds a running picture of who is where and what just changed. The obvious move is to feed that picture back as a richer predictor. We tried it. The situation model lost: a persistent who/where/topic state plus a narrative-event chain bought nothing beyond the rarest 1% of contexts, and a static prior built from the same counts matched it. So the situation model does not return as a predictor.

It returns as a salience signal for production. The dimension that just changed the most is what is worth saying. That is the only honest use left for the running picture: not "predict better," but "of all the things you could say, say the one that just moved." It picks the message before the producer formulates it: the difference between continuing the text and saying a thing.

A model of the listener: recipient design, not mind-reading

The last loop is the one that recurred in nearly every research angle, and the one most easily overclaimed: a second count table that approximates the listener. The idea is simple and count-native. Keep a second table over what this partner has actually said and heard. Score your candidate utterances not by your own fluency but by how well the listener's table would recover what you mean, and prefer the utterance that closes a gap the listener has, falling silent on what they already know. Audience design, as the divergence between two count tables.

Here is where we are careful, because there is a real category error waiting. A frequency table of what the listener has heard is recipient design: tailoring to your audience's exposure. It is not theory of mind. The load-bearing part of theory of mind is tracking that the listener believes something you know to be false, and a co-occurrence table cannot represent that. So we scope the organ exactly to what it is: it is judged only on whether it suppresses redundant mentions and helps a listener resolve the right referent, measured against a static listener prior, never on held-out prediction, where this kind of state already died. We say recipient design. We do not say mind-reading. The one proposal that tried to reify belief tables routed by a verb frame, the genuine category error, we cut.

And this is the one piece of top-down state with a license to change the common case rather than the rare tail. Every persistent-state mechanism before it landed on the same dead 1% slice. The listener table escapes that grave for one reason only: it is not scored as a predictor. It is scored by whether a modeled listener understood you.

Where the motivation comes from before any listener exists

One more honest gap. Over 90% of an infant's early vocalizations are produced to no one. Production cannot wait for a reply to bootstrap it. So the library's first motivational organ needs no listener at all: an intrinsic learning-progress drive that makes the cortex babble exactly where its own counts are still sharpening: practice the nearly-mastered, open the under-explored, abandon both the solved and the impossible. It is two leaky accumulators per region and a vote, no global objective. It breaks the cold-start without inventing a reward, and it is the one organ here that needs no contingent partner to run.

The honest open questions

This is a plan, not a scorecard. The organs are designed and queued ([BN onward], continuing the experiment lineage); they are not yet run, and several are expected to be partials or clean negatives. The questions we cannot yet answer:

Where this leaves us

The spine did not move; it grew a mouth, and now we are teaching the mouth to mean it. Reading is mimicry. Production becomes communication only when the loop forces a token to be chosen for its effect, and the loop is real, with first evidence that contingency teaches. The steering is the work ahead: tag emission by its function, read confidence three ways, let the situation model pick what to say, and model the listener honestly as recipient design rather than reaching for a theory of mind the counts cannot hold. Each organ ships with the baseline that could kill it. That is the only way we know to cross from echo to intent without fooling ourselves about which one we have built.

Lineage

Grew from the harness substrate (Exp AT) and the acquisition library it asked for, and from the contingency win, now confirmed against a live model. The situation-model guardrail is the meaning-map that did not predict; the confidence scalar is the deliberate pass and calibrated f·c; the producer it steers is the generation turn.

Thread: acquisition and generation as one mechanism read in two directions, and now read a third way, for its consequence. This is where the cortex stops echoing and starts speaking to someone.