The acquisition queue, built

The round. 2026-06-26 · sixteen experiments, run at once · the build queue from AU to BJ

We turned a library of language-acquisition science into a build queue and built the whole thing in one shot: sixteen experiments, run as parallel agents. The honest scorecard is seven wins, six partials, three clean negatives. The spine of it: a chunk lexicon gives the cortex whole units to say; a generation turn says them, and the words come out more well-formed than gibberish and now keep their frame nine times in ten; a reply that answers you teaches more than the same words overheard cold; and each negative sharpens a standing rule. The cortex took its first step from reader toward speaker. The step is real and the step is rough.

The question

Every experiment before this one only ever read a corpus and scored it. The cortex was a reader. A child is not: a child learns to talk by talking, and by hearing a reply. So we read the acquisition literature, distilled it into count-native organs, and lined them up as a queue: chunk the stream into whole units, ground a word to a thing in the world, let production read the same counts the hard way, and let a contingent reply teach harder than passive reading. Then we built the queue. All of it, at once.

What we tried

Sixteen experiments, AU through BJ, each one a named result from cognitive science asked to survive translation into counts: online, single pass, no gradient, bounded memory. Each carries a kill-condition fixed before it ran. We report every one, wins and partials and negatives alike, because a clean negative that sharpens a rule is a real result. Three follow-ups (BK, BL, BM) chased the loose threads the first sixteen left.

The spine: a lexicon to say, a turn to say it

The headline lever is the chunk lexicon (Exp AU). The cortex used to keep a backoff table of fixed-order letter sequences, exactly the pure transitional-probability table the acquisition literature says is wrong. AU gives it whole committed units instead. Cover the stream with the longest confident chunk, mint the pair that recurs, and let the transitions inside a committed whole decay. The decay is the Isbilen splice signature, and it lands clean: the within-word B-to-C transition falls to 0.0003 where pure transitional probability holds it at 1.000 forever: the chunk re-routes the mass into the whole, and the decay dial sharpens it another three-hundred-fold. Boundaries fall out for free: greedy cover with no entropy model at all reaches F1 0.758, a whisker from the branching-entropy first boundaries at 0.775.

AU is a partial, and it is the good kind. It loses raw next-letter cost by +0.20: completing by whole chunk throws away the calibrated backoff's per-letter sharpness. But the kill needed both the splice to fail and the cost to lose, so it did not fire. The win that matters is in hand: a variable-length emission vocabulary, whole units the cortex can commit to and say.

Saying them is the generation turn (Exp BD). The producer retrieves a frame, lets candidate constructions compete on coverage, takes the best by validity, and fills the slot, and every emitted word is spelled out of AU's committed chunks. Against a flat word-sampler floor, on a held-out oracle that never saw the producer's grammar, it is +18.5 points more well-formed and over-generates 28.5% less. The chunk lexicon and the construction producer compose end-to-end. This too is a partial: BD's own frame-survival sub-claim came in at 61%, short of its 80% bar. The producer picked a defensible category, then said the category's most common word, often a flat function word the oracle refuses.

A follow-up fixed it. Exp BL diagnosed the slip (say the filler this frame actually hosts, not the category's global favorite) and pushed frame-survival to 87.5%, clearing the 80% bar, while lifting well-formedness from 53.5% to 80.3% and never touching the flat floor far below. It was a selection slip, not a grammar error, exactly as BD predicted. Generation is rough, but it is measurably more than noise, and it keeps its frame nine times in ten.

A reply teaches more than a word overheard

The deepest bet under the whole turn is that a reactive loop teaches harder than passive reading. Exp BE installs the dial: weight each count by how recently the agent spoke, so a reply that lands right after it speaks counts loud. The control is yoked: the same tokens with their timing scrambled, so contingency is destroyed but the words are identical. Contingency-on beats yoked by +0.45 bpc and beats reading-everything-equally by +0.24. And the gap shrinks to zero exactly as timing stops predicting content, the signature of a real mechanism, not an artifact. A token that closed a turn teaches more than the same token read cold. This is the first empirical support for the reactive-loop bet the harness substrate (Exp AT) was built to test.

And the live test ran. Exp BM wired the contingency dial into a genuinely reactive loop, and it turned out no API key was needed at all: the machine's logged-in claude CLI is itself an authenticated partner, so the loop shells out to it for one short turn per reply. Against claude-haiku-4-5, forty real replies captured live, the win holds: contingency-ON beats the scrambled-timing yoked control at 12 of 12 sweep settings on the agent's surprise at the replies, and at 12 of 12 on turn-overlap. The one caveat we keep plainly: the agent's utterances are still per-char gibberish, so the model often answered them as nonsense. What is confirmed is that a contingent reply teaches more than the same replies re-timed cold, not that the cortex held a conversation. The mechanism is live; the fluency is not.

The other wins

Five more cleared their kill. Exp AV grounds a word to a thing in the world from co-occurrence alone, and its decisive variant collapses to chance the moment a guess is disconfirmed: the human propose-but-verify signature, reproduced in counts. Exp BC recovers an over-regularized form by predict-then-decrement, faster than passive reading and with no explicit correction: the first acquisition use of the reactive contract. Exp BH builds the honesty bar the field uses, minimal-pair grammaticality, and the count band beats the bigram 60.2% to 53.4% macro (the margins are thin on a small hand-built set, so it is a lean win, but it clears its axis). Exp BG spends a bounded sleep budget by inverse count and protects the rare tail better than uniform replay. Exp AY shows the comprehension-before-production lag widens with competitor density, and Exp BB's variation-set miner helps compositional generalization exactly where the literature says it should (in syntax, not world knowledge) an honest split.

The negatives that sharpen

Three came back clean negatives, and each one tightened a rule rather than dented it.

Exp BK is the sharpest. AU lost raw cost by +0.20, so the obvious fix is to vote the chunk distribution into the pool as one more expert. It does not work. Every positive weight makes the cost monotonically worse; the only break-even is the weight set to zero. A calibrated backoff already owns the chunk's confident completions, so the chunk expert is either noise that hurts or silence that does nothing. The +0.20 was never a property of the chunk organ; it was a weak floor in AU's read-out. So the verdict is nailed down: the chunk lexicon's value is segmentation and an emission vocabulary, not prediction.

Exp AX found the function-word anchor mathematically redundant with the construction frame voter (+0.000, and the agent refused the unfair full-eval win). Exp BI tried throttling writes on a Goldilocks curve and lost to writing everything down, while hurting the rare tail it skipped: the write-side echo of starting small: a count learner can't get stuck, so it shouldn't throttle its writes. Exp BJ ordered the stream by embedding depth and tied full-input-from-start on center-embedded agreement: the same rule, now extended to structural curricula.

The lesson

The cortex moved from reader toward speaker, and the move is honest about its roughness. The chunk lexicon is the lever: it gives whole units to say, and segments for free, but it does not predict (BK nails that down). The generation turn says those units more well-formed than gibberish and now keeps its frame 87.5% of the time, but generation is still coarse. Contingency teaches more than passive reading: the central bet has its first support, now confirmed against a live Haiku partner through the logged-in CLI (no key), the win holding at 12 of 12 settings. And the three negatives each tightened a standing rule: a count learner needs no write-throttle and no structural curriculum, and the anchor cue is already subsumed. Seven wins, six partials, three clean negatives, and not one of them hidden.

Lineage

Grew from the harness substrate (Exp AT), the reactive shell this whole queue plugs into, and from the acquisition library distilled to feed it. The chunk lexicon answers the first boundaries one altitude up; the producer fills the constructions; the comprehension-then-production split reads the calibrated truth value the hard way.

Thread: acquisition and generation as one mechanism read in two directions. The cue that predicts the next word while reading is the guardrail that constrains it while speaking, and this round is the first time the cortex read it backwards and spoke.