The exploration journal

The lab notebook

Every experiment becomes a post: the question, what we tried, the one number that mattered, and the honest lesson, wins and losses alike. Newest first.

This is the lab notebook. Every experiment becomes a post: the question, what we tried, the one number that mattered, and the honest lesson. The wins and the losses both. The distilled signal, the current theory these posts add up to, lives on the theory page. How we work explains why we publish the failures and how we decide what counts as a win.

Theory updates. A few posts carry a Theory update tag. Those are the ones that shifted the main theory: a finding that changed an assumption, not just added a data point. When you see the tag, the result it reports is also recorded in the changelog at the bottom of the theory page. The rest of the posts fill in the evidence around those turns.

How we got here

Each post grew from an earlier one, most of them from a failure that pointed somewhere new. The chain runs like this.

Counting beat the neural net chose the substrate. Finding where one word ends proved surprise carves a boundary, and that signal climbs to phrases, to a memory of change, to a vote that remembers. Words lowering the cost of letters became the word-level payoff, and both folded into one part repeated, the Column the rest of the work wires bigger. Scale sent us to the gigabyte, which made the Column fast.

Then a flat result forked the road. Stacking more fixed levels stopped paying. That sent us mining the research corpus, which queued four swings: counted attention, accumulated evidence, topic ignition, and predicting the kind. The frontier those aim at, global coherence, was named by the scorecard. And two negatives earned their keep: the meaning-map that did not predict and the vote that made it worse wrote our rule that an idea is judged on the axis it can win, never killed on the headline metric alone.

The later arc chased two frontiers, coherence and abstraction, and the second one finally yielded a verdict. Building everything as one repeated part on a synchronous-recurrent runtime let us ask the purest version of the coherence question, do concepts emerge as boundaries?, and the answer was that surprise carves only words; the cheap fixes (compose, scale, share a representation) all failed. So the work turned to abstraction itself. Pure biology, no counting traced only the back half of the brain's learning arc: sequence and compression, never the abstract code. Slower, and from above tried timescale and a top-down signal; both worked, neither abstracted. Six gradient-free mechanisms had now landed in the same place, all pointing at one suspect. The head-to-head isolated it (same architecture, swap only the update rule) and found the first positive of the line: the abstraction wall was credit assignment, not topology. The question flipped from is a gradient needed? to can a local gradient keep the win inside the online, local, bounded regime? A local signal cleared it: per-unit feedback alignment abstracts where a broadcast scalar fails, and named the pathway. The apical fusion rode that per-unit credit down a real top-down wire, gated by precision, and made abstraction climb with depth for the first time.

Then the line closed. Three swings after the keystone ruled out a label-attractor shortcut, factored diverse voting off as a prediction organ rather than an abstraction one, and scaled the fusion until it named the next wall, stability. The sweep round solved it: forward-activation normalization holds the stack stable everywhere else diverges, the four levers factor cleanly, and the depth-three winner sits above the backprop ceiling, and the capstone combine ran and failed, the levers not additive, locking the architecture there. With abstraction settled, the open value moved to generation. Internal self-feedback landed the first positive on that new frontier: an agent that babbles, hears itself, and corrects toward what it meant learns to produce, the birdsong loop, with no listener. The abstraction wall is now a closed, above-ceiling chapter; generation is the open one.

With the architecture locked, two runs stress-tested the winner. Fast, and at scale gave it a Metal-fast compiled build and ran it on a million characters: it stays stable, bits-per-char keeps dropping, abstraction holds above the ceiling but does not climb, confirming it is stability-limited, not data-limited. Re-reading tested the other axis, re-reading the same bounded text rather than seeing fresh text, and found it a free, rule-legal lever on prediction (200K read five times beats a million fresh characters, no overfit) that, like fresh data, leaves abstraction untouched. And the whole program now has a benchmark scorecard: the locked v1 model measured across four angles against a SOTA transformer and the plain baselines both, far behind on raw bits, above the backprop ceiling on abstraction with no gradient at all. Then the empty cell tested the review's own diagnosis, per-unit credit on a sparse substrate, the one box the line never filled, and found the synthesis real but unstable: it abstracts to the ceiling, then collapses, the dense credit pathway driving the same few units to win. A partial answer that sharpens the next move rather than settling it.

Then the deep collapse that locked the stack turned out not to be structural. Two anti-collapse mechanisms turned it back, on two different axes. Lateral inhibition added the degree of freedom the locked stack lacked, columns at the same level that see and suppress each other, and took the collapsed depth-four apex from rank one to a real code where a vote alone did nothing: the first positive past the levers-not-additive wall, the missing ingredient being interaction within a level rather than more breadth across levels. Replay at an interior altitude (Minsky's K-line) found a sharp interior optimum: reinstating a replayed configuration at the middle level protects the deep code on a long fresh stream far better than replaying the raw input, while the top alone is too vague and everywhere at once collapses. The same round measured the gap the generation work has to close. What it learns and what it can say back read the model two ways and found a strong recognizer and a weak producer: it learns most words at the first level and almost none higher, recognizes hundreds and reproduces dozens, and the learned codes play back as generic syllable stubs. And a bank of critics tried Minsky's selector and found that firing a deliberate pass on every character hurts, and that between characters only surprise is a selective trigger while conflict is noise, a clean negative and a calibration finding with the headline comparison still owed. Then the two anti-collapse mechanisms met. Two mechanisms that compose ran lateral inhibition and interior replay together on the drift stream and found they cover each other's blind spot: inhibition alone loses dimensionality, replay alone loses abstraction, both together hold the deep code high-dimensional and above the ceiling and stable, the first composition in the program that holds, and a candidate fix for the drift, with a longer confirmation run underway.

And the generation frontier got its first building block. The wall that round drew, a strong recognizer and a weak producer, named its own fix, and producing on purpose built it: production as a separate learned map from a meaning to a sequence, trained feedback-first by the birdsong loop, not the reader run backwards. On the matched task of producing one specific intended meaning, the separate inverse recovers 82 of 300 where reversing the reader recovers 2, and it wins in every frequency band including the rare tail the neuroscience flagged. A working producer, trained on its own feedback, where the recognizer reversed plays back a generic stub.

Then it produced real words. Producing real words pointed the same separate inverse at the top 300 most-frequent DailyDialog words, nothing in the mechanism changed, and produced 113 of 300 real words against the reader run backwards at 45, recovering the intended meaning 0.317 of the time against 0.090. The words are real English: lady to ready, agreed to good, equipment to question, the first intelligible production on real conversation. The headline is recovery, the live metric: the coverage gap narrows on real words because real onsets let the reader reversed complete to a plausible word, so what matters is producing the meaning you meant, where the separate inverse leads 3.5 to 1, and recovery is moderate because the meaning code is a coarse form region, so a produced word is often a neighbor of the target.

Then it held a conversation. A conversation with Haiku took that same producer off the offline answer key and put it in a live exchange with a real Haiku partner, run with no API key through the logged-in claude -p. The cortex takes a turn through its one frozen comprehension pathway, used both ways: comprehend Haiku's turn into the meanings it knows, then produce the real words it can voice for them, and send the reply back. Over fourteen exchanges every word it says is a real dialog word (42 of 42), most turns share an exact word with Haiku (10 of 14), and Haiku stays on the cortex's topic almost every turn (12 of 13). A few times the loop does more than reply: a topic word threads across turns, the cortex says match, Haiku reflects it, the cortex says it again, then dance the same way, a call and response that holds for a few exchanges. It is word-salad with real glimmers of relevance, the honest starting line for live dialogue, not a conversation that holds, and the production gap is visible, the cortex comprehends the right region of meaning but voices a neighbor of the word it meant.

And the generation track got its first reading on real conversation. A baseline on dialogue streamed the locked model over the whole DailyDialog training split, then measured it on held-out talk and a grammar battery. Dialog is a bounded, repetitive distribution the model fits stably: held-out bits-per-char drops with data, 3.767 to 3.549, no drift, the mirror of the fresh-stream drift, and 0.542 bits easier than the same model's text8 reference. But a long pass on the narrow register erodes general grammar, the minimal-pair score falling 64.8 percent to 51.1 percent, word order holding at 100 percent while agreement sinks to chance. The free-run output is gibberish with the right shape. A measured starting line on real conversation, and the specialization-versus-generality tension of the drift theme seen from the narrow-register side.

And a new direction opened: attention. Learning where to look made the per-column look-back offset an action, learned online and gradient-free by a count-based bandit, the scan through text treated as a motor act. The learned scan beats resampling the offset at random, learns non-uniform diversified views, and uses a content-relative word-jump, and it works only with lateral inhibition on the action space, the same diversity pressure that broke the deep collapse, now spreading the columns across complementary offsets. But context-free learned attention does not match hand-set diverse views at scale: when there is no context to condition on, a fixed spread is already good and free. So where you look depends on what you understand put the context back: a working memory one level up steers where each column looks, and at the right granularity the context-guided scan finally beats the hand-set baseline on prediction (3.022 against 3.101, every seed), flipping the context-free result, while context-keying lifts abstraction at every granularity (transfer CCGP 0.494 to about 0.60). That working memory still leaked at a fixed rate, so a gate for working memory gave it control over itself: a basal-ganglia-style Go/NoGo gate that fires on how far the model's own prediction moves, not raw surprise, and decides when to update the theme and when to hold it. The gate beats the fixed leak (3.069 against 3.083) and both degenerate bounds, holding the theme about eight characters and firing at word boundaries three times in four. The arc closes: the full learn-where-to-look, context-guided, gated-working-memory stack beats hand-set fixed diverse views, online and gradient-free.

Then the attention track met the abstraction wall, and the meeting drew the program's cleanest line. The attention work had lifted transfer CCGP to about 0.60, past the 0.484 ceiling, but on the vote-consensus probe that rises with dimensionality, not the gradient-comparable probe the ceiling is drawn on. Attention is prediction, not abstraction put it on the right ruler: feed the context-guided attention representation, verbatim, as the input to the per-unit-credit readout, the one pathway proven to abstract, and measure on its own hidden code. At eighteen thousand characters, where the readout is stable, the attention input reads 0.429 and 0.463 against a plain context window's 0.509 and 0.476, under the baseline and under the ceiling on every seed, with a backprop oracle clearing the ceiling on the same input. So the 0.60 was expressive dimensionality, a richer code, not a more abstract one. The two tracks separate cleanly: attention is the prediction-and-working-memory win, abstraction is the credit-pathway win, and feeding one to the other does not combine them. The wall stays a credit-assignment problem, and the road past it is a change to the credit pathway itself.

Then the attention track found its bound. Two levels is enough tested the two open items the gated round had named, a deeper hierarchy and a fire-rate-target gate, with capacity-matched arms. Depth saturates at two guidance levels: a third nested level, a slower context steering the middle one, is real and active (it routes the level below, cross-band divergence a full bit) yet at matched capacity reads 3.096 bits-per-char against the two-level stack's 3.049, worse, at both scales, because the extra bucket-splitting it needs costs more than its longer-range guidance gains. The self-calibrating gate is the keeper: targeting a fire-rate rather than a fixed multiple, it reads 3.049 against the fixed-threshold gate's 3.072, beats the hand-set floor, and holds its fire-rate across scale where the fixed threshold drifts down. So the attention architecture is bounded, two levels of guidance with a self-calibrating gate the sweet spot, a third level no payoff as built. The arc from learning where to look to a bounded, calibrated, two-level stack is complete.

Then a test of consciousness research handed back a calibration fix. The vote was too loud asked whether Global Workspace ignition, broadcasting the single most-confident column rather than averaging every specialist, beats the product-of-experts pool the model already runs. It does not. Winner-take-all reads 2.600 bits-per-char at two hundred thousand characters against the pool's 3.307, but only because it de-sharpens an over-confident pool; a normalized blend that keeps every column reads 2.314, better than both, and beats winner-take-all at every scale. The product pool was raising the consensus to the power of the summed confidences, about three over ten active columns, an overconfident distribution that inflates the bits, and normalizing the exponent to about one while keeping every specialist predicts a full bit better. So ignition does not win for prediction, all-or-none access throws away evidence, and the right pool is every column at the right temperature, the calibrated geometric mean the main vote already uses. The honest correction rides along: the experiments built on this substrate ran the un-normalized pool, so their absolute bits read about one high, while their within-experiment comparisons hold because every arm paid the same tax.

Then the count-native System 2 met its hardest test and a long-open question closed. A selector that cannot win finished the comparison the bank-of-critics round ran out of GPU on: a bank of self-calibrating critics and a selector that picks when to think harder, against the trivial fixed policies. The selector loses, reading 3.984 bits-per-char against the bare model's 3.920 and the replay-every-character policy's 3.618, last of the four. The loss has a precise mechanism in three parts. The self-calibrating gate works and only surprise is selective: every critic tunes itself to the target fire-rate, but only surprise holds it with a high threshold, a real tail, while conflict and low precision hit the rate only by collapsing toward the bulk. The selective signal gates the wrong way to think: wiring a deliberate re-read to the surprise tail inherits the harm of always thinking, the apex dimensionality dragged down toward collapse. And the way to think that helps wants a dose, not a tail: replay is monotone in how much you do it, so always-consolidate wins outright and a sparse selectively-gated replay is worse than both heavy replay and none. So at the character level no count-native critic signal, used to gate a sparse intervention, beats the trivial fixed policies, because the one selective signal triggers the harmful mode and the helpful mode wants a dose. The selector-over-modes question is closed with a no, and the standing System-2 gap is sharper for it.

Then the substrate axis the architecture review named paid out, and the empty cell was filled. The empty cell, filled ran the diversity machinery the partial result had named, as an ablation ladder over the collapse baseline. The cure is two pieces and it takes both: boosting and positive weights together hold deep transfer abstraction at the 0.484 ceiling to the last checkpoint on both seeds, with more than two hundred distinct winner-codes against the dense reference's about forty, and the lowest bits-per-char of any arm. Boosting alone still collapses, because the signed weights let the credit drive a few units to always win; clamp the weights non-negative and then boosting keeps the winners spread. And the sign-only permanence credit breaks the code, so the credit stays graded with diversity machinery around it. So a sparse substrate carrying per-unit credit holds stable abstraction, and the dense reference is no longer the only thing that does. It matches the ceiling, it does not exceed it, at eighteen thousand characters and two seeds with the scaled confirmations running.

Then the benchmark scorecard's flat spot got a cause and a cure. The calibration had read a plateau: our fixed twelve-column voting bank improves to about a million characters and stops, 2.234 to 2.112 to 2.109 bits-per-char, while a plain n-gram keeps falling because it keeps growing its context table. Growth breaks the plateau tested whether the flat spot is a fixed-capacity artifact by building the grow-then-prune the program wanted: add a longer-context voting column when the held-out gain decays, cap the count, prune the least-used. On the identical text8 split, the growable bank goes 2.234 to 2.004 to 1.875, growing 12 to 16 to 24 columns, where the fixed bank flattens at 2.11. Past a million the fixed bank drops 0.003 and the growable bank drops 0.129, and at ten million the growable bank reads 1.875 against the fixed 2.109, within 0.047 of the n-gram floor at 1.829. It grew the right thing, progressively longer-context views, the capacity the fixed bank lacks and the n-gram wins by, did not grow where there was too little data, and held the budget at the cap. So the plateau was fixed capacity, and growth dissolves it by doing what the n-gram does. The honest bounds stay: it reaches about the n-gram, not past it, it is still about 0.85 bits behind text8 SOTA, and prediction is not our headline axis. The deliverable is the capability, a count-native, online, bounded substrate that keeps learning from data instead of saturating.

Then the two engines met on one substrate. The program had split into a voting bank that predicts (low bits-per-char) and a sparse apical stack that abstracts (at the backprop ceiling), each doing one thing. One stack, two halves gave the abstraction stack the prediction machinery, eight sparse columns each on a different view, voted by the same calibrated blend, and asked whether one stack can do both. The result is an honest trade-off. Abstraction composes and exceeds the ceiling: the unified stack reads transfer CCGP 0.520, above its single-view 0.487 and above the 0.484 single-layer backprop reference, at higher dimensionality with no collapse, the highest abstraction the program has reached gradient-free, because k-WTA and boosting preserve the cross-stimulus variance the old normalization destroyed. Prediction does not transfer: the unified stack reads 3.666 bits-per-char, barely below its single-view 3.753 and nowhere near the voting bank's 2.442, and the blended bits-per-char equals its own best single column, so the vote does nothing for prediction with neural columns. The count bank's bits-per-char win was a property of count tables combining, not of voting. So the prize, one stack low on bits and high on abstraction, is not won at this scale: the synthesis kept and improved abstraction and did not import prediction, and the path forward is to couple the readouts on the pooled error or graft a count-native head onto the abstract code.

Then the second path ran. The rank-reuse tension grafted a count-native head onto the abstract code, a count table keyed on the P0.5 abstract winner-set, against the same count machinery keyed on the raw context, one substrate with two heads. The result is an informative middle that leans deep-negative. The count head extracts real signal from the abstract code: it reads 2.578 bits-per-char at one hundred thousand characters, 0.8 to 1.0 below the stack's own readout at 3.548, so the abstract code carries next-character signal and one substrate can feed both heads, a real sub-win. But it stays about 0.45 bits behind the same count table on raw context (2.141), and the gap holds across 5.5 times the data, +0.50 at eighteen thousand and +0.44 at one hundred thousand, while the same code's transfer CCGP is 0.499, above the reference. So the same code abstracts well and predicts worse. The mechanism is a tension in the representation: the abstract code is nearly unique per position, three to six counts per key even at seventy thousand positions, because the high rank that makes it abstract makes it almost never repeat, and a count table predicts by reuse. High rank gives good abstraction and poor count-reuse, so abstraction and prediction want representations in tension, and the two-stacks split has a root in the code itself. The lever to try next, untested, is a coarser, less-sparse code that might repeat enough to predict while staying abstract enough to keep its abstraction, a sparsity-versus-reuse sweet spot.

Then a small knob turned out to be the thing holding repetition back. The re-reading round had found re-reading a modest prediction lever that flattened by the third pass, and read that as a ceiling on repetition. Gentle learning compounds tested the other reading, that the plateau was an artifact of the locked learning rate, too aggressive to let repetition compound, by sweeping the apical rate over one, five, and ten passes with everything else held to the re-reading config. The hypothesis held. A gentler rate reaches a lower floor: at 0.001 the held-out bits-per-char reaches 3.637 against the aggressive baseline's 3.762, 0.125 bits better at matched exposure. The optimum is a U, with 0.005 too aggressive, 0.001 the sweet spot, and 0.0005 too gentle. And the trajectories tell it cleanly: the aggressive rate crams the slice on the first pass and then runs flat, while the gentle rate keeps descending across all ten passes to a genuinely lower floor, with no overfit and a tighter generalization gap. The point reaches past re-reading, because the rate that proved too aggressive is the apical stack's global learning rate, used everywhere, so the locked value is likely too hot on every run, and the global rate should be revisited. The honest bounds are the familiar ones: a prediction lever only (abstraction stays flat above the ceiling), a modest absolute gain, and the apical learned stack rather than the count voting bank, which is re-reading-invariant and has no rate to tune.

Then the two engines were wired together for the first time. The rank-reuse tension had said one shared code cannot be both abstract and count-reusable, and the cortex's own answer is not to merge the codes but to keep both and let them share information through a coarse, repeatable tag, the feedback arrow from the abstraction engine to the prediction engine. Abstraction cannot cheapen prediction builds that seam, count-only, the cheapest decisive leg, and lands an honest negative with a clean diagnosis. The tag carries real information and still buys no prediction: a short-horizon tag, the abstraction stack's own deep code, beats a shuffled version of itself by up to 0.385 bits-per-char, so it is genuinely informative, but it is redundant with the order-six n-gram, so a soft coupling that never displaces a well-counted raw cell reaches break-even and no better, +0.001 at a million characters, never below the baseline. A long-horizon tag, a gist over the characters outside the n-gram window, carries almost no next-character information at all, about 0.013 bits over its shuffle, no better than random, and conditioning on it only hurts, even at the least-local word-start positions, with no scale dependence. The unified finding is sharp: next-character prediction is an overwhelmingly local game that counting already wins, and information is not usable prediction gain, redundant where present and absent where new. So the conclusion is not that the coupling fails but that the ruler was wrong. The mutual loop is already half-built, because the abstraction stack's per-unit credit is the prediction error, so prediction already teaches abstraction, and the only missing arrow is the one shown null here for the next character. The place to measure the cross-feed is the next word and the coherence of generated text, not the next character.

Then the program stood back to regroup. Where we are is the standing-back synthesis: the run of rounds has converged on two engines, a prediction engine near the n-gram level (count-native, gradient-free) and an abstraction engine at or above the 0.484 backprop ceiling with no gradient at all, the genuinely novel result. It names the recent wins that got them there (P0.5's stable sparse abstraction, grow-then-prune breaking the plateau, BLEND-NORM's calibration fix, the learning-rate sweet spot, the attention subsystem), the honest scorecard (a mile from a strong transformer on prediction, the abstraction result the real prize), the deep open problem (the rank-reuse tension, one code cannot be both abstract and count-reusable), the bet to resolve it (coupled laminar layers, the cortex's own answer), and what comes after (generation and production, learning by example). A regroup, not a new experiment, and the map the next stretch is drawn on.

Then the question changed. Six rounds on prediction had said the next symbol is a local game counting already wins, so the program asked what it is for, and a four-stream fan-out converged on an answer: the goal is a cognitive architecture we can talk to, and the smallest version of talking to it is comprehension, not compression. Say two facts, ask a question, get the answer. Counting cannot bind lands the first positive of that reframe. The apple/cup task ("here is an apple, it is green ... what color is the apple?") is not a prediction problem but a variable-binding, one-shot relational-memory problem, the job of the hippocampus the program's two neocortical engines lacked. The instrument randomizes the property every episode, so every entity is equally often every color, which makes any counting model provably sit at chance. And it does: the strongest co-occurrence reader scores 0.246 against a chance of 0.250, dead on the line. A gradient-free binder over random codes, statement binds and question unbinds and answer cleans up, scores 1.000, clearing the binding bar a three-year-old sets, and scrambling its store collapses it back to 0.243, so the store does the work, not a leak. The capacity is graceful: eleven facts in one episode still 1.000 at full width, bending only at a tiny sixty-four dimensions. The honest next step is learned cortical codes in place of the random ones, the move that turns one-shot binding into one-shot generalization, and the reason the abstraction result was never vanity. The first working piece of an architecture we can talk to, online, gradient-free, and bounded.

Then the binder proved it was an organ, not a trick. One passing task is what a hash map does, so typed, grounded, and true tested the binder across three more capabilities, each gradient-free, each under the same per-episode randomization control that pins counting at chance. It is typed: give every thing a color and a size, and "what color" returns the color and not the size, 1.000, where an untyped bag of the same facts confuses the two classes half the time, wrong-class 0.504, so the per-class role is what makes binding typed. It is cortex-grounded: bind the codes the neocortex would emit, derived from the spelling, rather than fresh random atoms, and it comprehends exactly as well, 1.000 against 1.000, so the neocortex-to-hippocampus wiring the architecture calls for works end to end. And it tracks truth: "is the box red?" gets a NO when the box is blue, even though the episode said "it is not red" first so red co-occurs with the box, 1.000 on the no-cases, where a co-occurrence reader says yes to the negated color and fails every no-case, 0.000, so this is truth over propositions, not association. The honest bound is that the cortex-code similarity dial came back inconclusive, because distinct spellings project near-orthogonal so there was no interference to measure, and the similarity payoff waits on learned semantic structure, the next phase. And the honest forward-look is the whole point: the binder is a validated comprehension organ, but every test so far is hand-parsed with hand-coded comparators, so the next work is the architecture, a reader that extracts the role-filler structure from the stream itself, generation of the answer, and realer language, not more synthetic rungs. The missing organ is built; the job now is the system that reads and answers on its own.

Then the binder showed it understands rather than looks up, and the result tied the whole architecture into one system. Every answer so far was about a fact the model was told, so the cortex is the dial ran the test that separates understanding from lookup: can the binder answer about a novel entity it was never told, getting its property right because it knows the kind of thing it is? Each class is coded as a prototype, and a member as that prototype with a fraction of its bits flipped, a dial from "members identical" to "members random." Store a class property for twelve known members, then ask about a thirteenth, never bound. When the code carries class structure, generalization is perfect (1.000) all the way out, and it collapses to chance (0.217) only when the structure is gone, while a control that did store the novel member's property stays at 1.000 throughout, so the collapse is real transfer-failure, not a broken pipe. And the part that unifies the architecture: the cortex code's CCGP geometry predicts when the binder generalizes, correlation +0.98, the strongest predictor in the sweep, far stronger than raw similarity at +0.55. So the abstraction engine the program reached the backprop ceiling with, gradient-free, the result that sat genuinely novel and slightly lonely while six earlier rounds set CCGP aside as "not our axis" for prediction, is exactly the generalization dial the comprehension organ turns. CCGP was never the wrong axis; it was measured against the wrong job, and its job is generalization. The two engines and the binder are one system: the cortex builds an abstract code, the hippocampus binds it, generalization rides on the geometry. The honest bound is the next bet: the class-structured codes here are hand-set, a stand-in for a learned cortex, so this proves the binder half, if the cortex makes high-CCGP codes the binder generalizes, and whether the abstraction stack can learn codes that good gradient-free on real data is the end-to-end question left.

Then the hand came off the codes. The previous round had proved the binder generalizes given a class-structured code but turned the class-structure dial by hand, a prototype plus bit-flip standing in for a learned cortex, so it showed only the binder half, if good codes then generalization. End to end, no backprop closes that caveat: a gradient-free learner develops the codes itself, from context, with no backprop anywhere. Entities recur in class-indicative contexts; the learner accumulates each entity's context histogram online and reads a code as the sign of a fixed random projection of that histogram, count plus a random projection, the whole gradient-free kit. The binder then stores class-property bindings for known entities and must infer the property for novel entities coded from their own contexts. The dial is now the world, noise being the fraction of an entity's contexts drawn from the global pool instead of its class signature, zero meaning the classes are separable in context and one meaning the world is random with nothing to learn. The full pipeline generalizes: 1.000 from a clean world out to noise 0.6, 0.767 at 0.8, and chance (0.217) only at noise 1.0 where the world is unlearnable, with a stored control near 1.000 throughout so the collapse is real transfer-failure. And the learned-code CCGP predicts it end to end, correlation +1.00, more robust than the hand-set sweep, holding a step further out and failing only when the world itself is random. So the cortex learns the abstract code, the hippocampus binds it, and generalization rides on the geometry, with nothing hand-fed and no gradient. The hardest bet the architecture rested on, gradient-free generalization through learned shared structure, the thing the Tolman-Eichenbaum Machine reaches only with backprop, holds for this setting with none. The honest bounds are the next phase: the world here is synthetic, the learner is a context-histogram bag rather than the apical abstraction stack, and there is still no real language, no reader extracting structure from a raw stream, and no spoken answer.

Then the two threads met. The comprehension thread wanted a reader and a voice; the generation thread, which had reached a live conversation on hand-fed meaning codes, wanted real text in and a grounded loop. Same join, opposite sides. The whole loop, on one brain makes it: one brain on the runtime, four validated organs as connected Areas, that reads raw characters, binds word to referent, comprehends a held-out question, speaks a name, and is corrected when it is wrong. A fully-neural segmenter (Hebbian char-transition synapses, no bigram dictionary) feeds a union encoder, which feeds a Hebbian binder, and a generator reads the binder's same matrix W in reverse. The load-bearing fact is that all four phases run on the same organ instances, and the single W is written by learning, read forward to comprehend, read backward to speak, and re-written by correction, so comprehension and production are one matrix read two ways. Held-out referential pointing reads 0.910 against a shuffle at 0.217; content production with a silence self-check reads 0.917 with a real comprehension-over-production gap of +0.333, the gate staying silent on the one word it cannot isolate rather than misnaming it; and correction by consequence, the environment re-uttering a misunderstood word with the right referent in focus, lifts the mistaken word 0.472 to 0.674 where a yoked recast reaches only 0.506, winning three of four seeds. A grep of the live path confirms no counting and no hand-tokenizing, fully neural and gradient-free end to end, and the brain holds a grounded conversation, pointing to and naming apple, cup, box, and ball. The dog-and-ball demo runs on the runtime, with no cheating. The two tracks are one architecture now; the open work is scale and realer language.

Then the one brain was given a childhood. A child-brain in a staged curriculum assembles the validated organs into one brain with a working memory and a top-down route (so it disambiguates a word whose meaning depends on the discourse, 1.000 with the held context against 0.500 without), and drives that same brain through a staged developmental curriculum with a real Haiku caregiver who speaks ahead of it, each stage gated by a comprehension test, the child lagging the adult by construction. G1, one word, passes on the bare brain (pointing 0.828, a contingent recast beating a yoked one 0.827 against 0.481, the lag a real 2). G2, two words and a property, passes on the same brain plus one faithfully added typed-store organ (color 0.807, size 1.000, wrong-class 0.000, the construction lag positive across the whole climb). The recurring lesson: the developmental lag lives in construction complexity, and the catch-up bottleneck is segmentation, perception preceding reference. The deep, transferable-abstraction question is held open and honest: a gradient-free route through time is a real ingredient (the no-decorrelation arm rises +0.085 over the pure-local wall) that does not clear the 0.451 credit-using bar, because the decorrelation that keeps the code high-rank fights the predictive pull that would bind same-class items. The next move is a dedicated binding mechanism, not more time.

And then the three threads converged. A tiny brain that comprehends, abstracts, and speaks is the program's biggest milestone, and it closes the three open ends at once on the one gradient-free, online, bounded-memory brain. Comprehension climbs the rest of the curriculum: G3 (morphology) and G4 (relations, negation, grade-two dialog) both pass on the same brain, so it understands a simple grounded conversation (G4 held-out relation QA 1.000, negation 1.000, live multi-turn dialog comprehension 12/12), with the segmentation bottleneck appearing exactly where Brown predicts (the forward segmenter cannot place a morpheme boundary; a recurring-sub-chunk detector cracks it). Production catches up: a composer, the inverse of comprehension, makes the brain speak the relation it binds (whole-relation structure 1.000 on held-out meanings, surface form 0.918 after a whole-word prior; live, "the cup is not gold" comes back as "cup not gold"). And the abstraction wall, the hardest problem, is resolved in-bounds by the program's own thesis. Six purely-perceptual unsupervised mechanisms (temporal prediction and LPL, a rich temporal stream, slow pooling, attention in space and time, a CA3 attractor) all failed to build transferable abstraction, for one clean reason, a class is an equivalence over non-adjacent instances and perception can only find adjacency, and the resolution is the WORD, won cross-situationally with no label, the cross-instance equivalence (held-out class-CCGP without the word about 0.40 at the wall, with it 0.87 synthetic and 0.50 to 0.99 on the real grounded encoder, collapsing under word-shuffle, random-projection, and a floor control). The constraint never relaxed: no backprop, no credit, no counting. The converging negatives are the substance, because they built the resolution, and the open edges (the mouth's segmenter over-cuts, free word order, scale and a real child-egocentric corpus) are named honestly.

Then the program read the two labs closest to it and adapted them. Adapting two labs into a grounded grammar takes an external biological-topology lab and Mitropolsky-Papadimitriou's Assembly Calculus, both of which hit our abstraction wall and concluded the same thing (transferable abstraction needs an equivalence signal), and adapts their machinery onto our gradient-free, online, grounded route, where the word is that signal. The result is a real learned grounded grammar, every piece composing the same organs: learned word order for all six constituent orders (1.000 at high exposure, the object-initial typological signature present from a salience prior, ablation-confirmed), recursion by a sequence-memory stack ("the dog that chased the ball is big" attaches "big" to the dog, 1.000 with the stack against 0.000 ablated, a human-like depth ceiling) that closes our own prior recursion negative, and a learned grounded attachment parser (0.958 against a nearest-head baseline at 0.533) that beats their hand-coded ungrounded one because the scene resolves what word order cannot. Two pieces of biological topology transfer (cortical voting, hub 0.647 against single column 0.360; thalamic role routing at 0.988/1.000/1.000), with the honest route difference that lateral feedback is a null for us. And, framed as supporting, our language route does not erode at 10M continual exposure (held-out abstraction flat, grounding improving) where an apical-credit route fell 12%, because the word recurs and a homeostatic-leak bound holds a stable fixed point, with the honest caveat that the absolute margin is modest at the hard operating point and the temporal-stability read is the robust one. The substrate, plainly, is theirs (Assembly Calculus, 2020); our edge is the grounding.

The posts

Newest first.

Adapting two labs into a grounded grammar

Program milestone · 2026-06-29 · two adjacent research programs are adapted onto our gradient-free, online, grounded route, and beaten in each case by adding the scene: a real learned grounded grammar, biological topology that transfers, and a language route that does not erode at scale. Both an external biological-topology lab and Mitropolsky-Papadimitriou (Assembly Calculus) hit our abstraction wall and concluded the same thing: transferable abstraction needs an equivalence signal. They use apical credit and hand-set structure; we use language, the word won cross-situationally, the cleanest instance of their own frame. We adapted their machinery onto our route. The grammar (the headline), every piece composing the same organs, all gradient-free. Learned word order: all six constituent orders at 1.000 at high exposure (chance 0.167), the old fixed composer scoring 1.000 on SVO only; the typological signature present (object-initial harder, OVS/OSV 0.000 against others 0.750 at modest exposure) but honestly from a weak agent-before-patient salience prior (turn it off and all six acquire equally), ablation-confirmed, the order-shuffle control at chance 0.180. Recursion by a sequence-memory stack (the one first-class Assembly-Calculus operation we lacked, the one with their cleanest theorems): "the dog that chased the ball is big" attaches "big" to the dog, outer-attachment 1.000 with the stack against 0.000 ablated (a no-stack count baseline near chance), in comprehension and production, with a human-like depth ceiling (depth 1 1.000, depth 2 0.654, depth 3 0.588, tracking the proven capacity bound, a feature not a bug), the scaffold theorem reproduced at our scale (rehearsals halved, ratio 0.52). This is the recursive machine M-P only sketched, and it closes our own prior recursion negative. Learned grounded attachment: a learned reader composing the router, the stack, and merge resolves prepositional-phrase attachment at 0.958 against a nearest-head baseline at 0.533 (the hard non-adjacent case 0.962 against 0.000 for every ungrounded arm), grounding load-bearing, every ablation and the shuffle collapsing it, a learned parser beating their hand-coded ungrounded one because the scene resolves what word order cannot. Biological topology transfers: multi-column cortical voting (hub 0.647 against single column 0.360, scaling cleanly 0.346 to 0.669 across 1 to 16 columns) and thalamic soft-competition role routing (noun 0.988, attribute 1.000, action 1.000, soft beating hard). The honest route difference: hub-to-column lateral feedback is a null for us (best lift 0.003), because our word signal needs pooling for coverage, not the iterative consensus their object-id route needs; we report the null, not a fragile positive. Retention at scale (supporting, not headline): our language route is temporally stable where an apical-credit route eroded, held-out class-CCGP from 1M to 10M moving 0.348 to 0.353 (flat) and grounding 0.708 to 0.771 (improving), against the external 0.988 to 0.865 (-12%), because the word recurs and a homeostatic-leak bound holds a stable fixed point (a naive clip collapses to chance), so we do not need their metaplastic fix. Honest caveat: this was at a hard 16-way operating point (chance 0.0625) where the absolute abstraction margin is modest and the word-shuffle control is not fully clean (about 0.19, above the strict chance line), so the script's automated verdict even reads "erodes"; the temporal-stability comparison (flat to improving where theirs fell 12%, rank healthy) is the robust read, not the absolute margin. The through-line is grounding, our edge: each adaptation beats its source by adding the scene or the caregiver. And the corrections, stated plainly: the SDR and Hebbian substrate is not ours (Assembly Calculus, 2020, earlier than us); only projection and sequences are proven in that calculus (merge and the rest are simulation, and so is our merge); and the depth and margin ceilings and the merge-decode crosstalk are real limits (the dependency readout 0.551 lags the decision 0.958). Runtime untouched throughout. Read.

A tiny brain that comprehends, abstracts, and speaks

Program milestone · 2026-06-28 · the biggest milestone of the program: one gradient-free, online, bounded-memory brain on our own recurrent runtime comprehends a grounded conversation, abstracts a concept through language, and speaks it back, the three threads converged. The program had built the organs and, in a child-brain in a staged curriculum, driven one assembled brain through the first two developmental stages. This round closes the three open ends at once. It comprehends a simple grounded conversation. The curriculum climbs to the top of the ladder: G3 (morphology) passes on the one brain (the referential test on held-out stems 1.000, chance 0.333), and the load-bearing honest finding is that the brain's forward transitional-probability segmenter cannot place a morpheme boundary (the affix is predictable, so the dip lands inside the stem, "walking" cut as "wal-king"), cracked by the same chunking principle one altitude down, a recurring detachable sub-chunk detector (-ing, -ed, -s all 10/10, zero over-peel). G4 (relations, negation, dialog) passes too: relation QA on held-out compositions reads patient 1.000, agent 1.000, predicate 1.000 with the agent/patient role won from the perceptual scene (not handed in), negation 1.000, and a contingent multi-turn dialog comprehended 12/12, so the brain understands a conversation. The recurring lesson holds the whole ladder: the lag lives in construction complexity (one word, then two, then morphology, then relations) and the catch-up bottleneck is segmentation, perception preceding reference. It speaks. G4 alone was telegraphic (it comprehended the relation but structured production was 0.000); a production-side composer, the inverse of comprehension (unbind each role, name it, emit in a fixed role order), lifts whole-relation structure to 1.000 on held-out meanings, and a whole-word prior lifts the surface form from 0.684 to 0.918 (the residual is one word per seed the brain genuinely never heard whole, the no-cheating constraint working), so in a live exchange "the cup is not gold" comes back as "cup not gold." The swap control proves the order tracks the role, not an input position; free word order is a later module. It abstracts a concept, via language, the project's hardest problem, resolved in-bounds. Be careful and honest here: six purely-unsupervised perceptual mechanisms all failed to build transferable abstraction (transfer CCGP), for one clean reason, a class is an equivalence over non-adjacent instances and perceptual similarity, contiguity, and dynamics can only find adjacency (on a decorrelated code, same-class instances are near-orthogonal). Temporal prediction (LPL) ties the pure-local wall (0.344 against 0.342, the only rise being the no-decorrelation arm at 0.427, +0.085, which is the point because you must weaken the decorrelation before time can bind; an earlier two-seed read of 0.422 did not replicate); a rich temporal stream gives the binding-isolation probe 0.159 on, 0.159 off, both below chance; slow pooling, space-and-time attention, and a CA3 attractor all fail the same way (the attractor closing the purely-perceptual question, because basins form by similarity and same-class items are near-orthogonal). The decisive diagnosis: the decorrelation that keeps the code high-rank fights the predictive pull that would bind same-class items. The resolution is the program's own thesis: the WORD, won cross-situationally with no label, is the cross-instance integral that supplies the non-adjacent equivalence. Held-out class-CCGP is without the word ~0.40 (the wall), with the word 0.87 synthetic (4/4 seeds) and 0.50 to 0.99 on the real grounded encoder (an honest operating-point curve, the lift growing with how reliably the class attribute is perceived), collapsing under word-shuffle (~0.37/0.24), random-projection (~0.39/0.35), and a floor control (no lift when the attribute is never perceived). The constraint held: no backprop, no internal credit, no counting; we did not relax it, we used language. The richness lesson, included honestly: a full-vocabulary varied environment does not hurt the abstraction and the Mowgli case (no recurring co-occurrence) fails, but the sharp finding is that an impoverished fixed-context environment is pathological (it manufactures spurious structure that fools a naive read), so richness's payoff is unconfounding. The open edges are named: the mouth's last residual lives in segmenter over-cuts, free word order is a later module, a couple of earlier results used two-seed reads that did not replicate (caught by the controls), and scale plus a real child-egocentric corpus (SAYCam, Vong) is next. Read.

A child-brain in a staged curriculum

Program milestone · 2026-06-28 · a developmental child-brain in a staged curriculum: one assembled brain, gated stages, the lag in construction, abstraction still open. The program had built the organs a developing brain needs and, in the whole loop on one brain, wired a reader, binder, and mouth into one brain that read, bound, comprehended, spoke, and was corrected. What it had not done was give that brain a life: a staged curriculum where an adult speaks ahead of it and it catches up the way a child does. This round makes two joins. First it assembles the validated organs into one brain with a working memory and a top-down route: a neural segmenter, a union encoder, a two-compartment binder, a working memory that latches the discourse topic, and a generator. On the assembled brain, held-out referential comprehension reads 0.983 against a shuffle at 0.276 (chance 0.250) and content production 0.932, and the working memory earns what the talkable brain alone could not, context-dependent disambiguation of an ambiguous word read from raw chars, 1.000 with the working memory against 0.500 without it, with every control clean (gate shut 0.500, shuffle 0.500, apical with no bottom-up word drives nothing). Then it drives that same brain through a staged developmental curriculum (G1 to G4) with a real Haiku caregiver, anchored to cited child-acquisition norms, each stage gated by a comprehension test, the child's vocabulary and comprehension lagging the adult by construction (comprehension leads production, the adult's lexicon is always a superset, both lags printed every run). G1, one word, passes on the bare brain with nothing added: held-out pointing 0.828, the shuffle control 0.238 at chance (reference won cross-situationally, not leaked), a contingent recast beating a yoked one 0.827 against 0.481 (the loop teaches, not the words), and the adult-ahead lag a real 2 (the adult uses 11 nouns, the child comprehends 9, produces 0). G2, two words and a property, passes on the same brain plus one faithfully added typed-store organ (registered locally, runtime untouched, the role read off the perceptual channel rather than a word tag): typed color 0.807 and size 1.000 with a wrong-class rate of 0.000 where the untyped bag of the same facts confuses the classes at 0.437, corrected over yoked 0.557 against 0.083, and the construction lag (single-slot comprehension ahead of two-word production) positive across the whole climb, up to +0.39. The curriculum trains the whole brain, not ad-hoc per-stage organs, and G3 (morphology) is in progress (no numbers reported). The recurring honest finding: the developmental lag lives in construction complexity (one word, then two, then morphology), and the catch-up bottleneck is segmentation, the child must isolate a word from the stream before it can ground a meaning, so perception precedes reference. And the deep, transferable-abstraction question is held open and honest, reported with the numbers and no breakthrough language. A purely-unsupervised, gradient-free route through time is a real ingredient but does not clear the bar: on character text at four seeds the plain predictive rule ties the pure-local wall (0.344 against 0.342), and the only arm that rises is the one with the anti-collapse term removed (0.427, +0.085 over the wall), which is the point, not a loophole, because you must weaken the decorrelation before time can bind at all, and none of these clears the program's 0.451 credit-using value. (An earlier two-seed run read 0.422 and looked like a win; it did not replicate at four seeds.) A faithful temporally-rich stream does not close the gap either: the load-bearing binding-isolation probe, which strips the static window of all class information, reads 0.159 with time on, 0.159 with it off, both below chance 0.333, no temporal binding. The decisive diagnosis: the decorrelation the sparse substrate needs to stay high-rank fights the predictive pull that would bind same-class items, the mechanism that keeps the code healthy is the same one that forbids the binding abstraction needs, so the open direction is a dedicated binding mechanism, not more time. Read.

The whole loop, on one brain

Theory update · 2026-06-28 · the loop closes on one brain: four Hebbian and SDR organs read raw chars, bind, comprehend, speak, and are corrected, fully neural and gradient-free, one matrix read forward and backward. Two threads had been running side by side and naming the same missing join. The comprehension thread had a reader-and-voice hole at the bottom of it (no real language, no reader, no spoken answer); the generation thread voiced real words in a live loop but on hand-fed meaning codes. This round assembles them into one brain on the runtime (cortexshell, cortexgraph), four validated organs as connected Areas, and runs the whole loop: it reads raw characters, binds word to referent, comprehends a held-out question, speaks a name, and is corrected when it is wrong. Nothing re-implements an organ; each Area class is imported from the script that validated it. The organs: a fully-neural segmenter (HebbCharSeg, Hebbian char-transition synapses with a dip threshold, the last counting removed, no bigram dictionary) feeding a union encoder feeding a Hebbian binder (HebbAssoc, the matrix W), and a generator (HebbGenerator) that reads the binder's same W in reverse. The load-bearing fact is that all four phases run on the same organ instances, and the single W is written by learning, read forward to comprehend (act = W column for the word, argmax over the present options), read backward to speak (wact = W row for the referent, cleanup to a chunk), and re-written by correction, so comprehension and production are one matrix read two ways, not two systems, which is why the comprehension-over-production gap falls out of the geometry rather than being engineered in. The numbers, all on the one brain, with chance 0.250 to point among four and 0.083 to name from twelve words. Comprehend: held-out referential pointing, choose the meant referent among distractors never studied together, 0.910 against a shuffle control at 0.217, so the binding is the word-to-referent contingency won from co-occurrence, not a leaked alignment or a memorized scene. Speak: content production with a silence self-check (keep the produced word only if it re-comprehends to the goal through the brain's own ear) 0.917, a real comprehension-over-production gap of +0.333 against open-vocab production, and where it cannot isolate a word the gate stays silent rather than misnaming it. Correct: the environment recasts a misunderstood word, re-uttering it in a fresh scene with the correct referent in focus through the brain's own raw-char reading path, an ordinary next exposure and a consequence, never a label or a gradient, lifting the mistaken word's comprehension 0.472 to 0.674, where a yoked recast (same budget, scrambled word) reaches only 0.506 and none stays flat (corrected over yoked +0.169 on the err-words, winning three of four seeds), so the contingent loop taught, not the words. And it holds a grounded conversation, pointing to and naming apple, cup, box, and ball, the dog-and-ball demo the program has chased for months, now fully neural on the runtime in about nine seconds. The faithfulness is checked, not asserted: a grep of the live path confirms no counting and no hand-tokenizing anywhere, Hebbian and SDR end to end, gradient-free, online single exposure, bounded memory, the four laws holding the whole way. The two tracks the program ran in parallel are one architecture now: one brain that reads, binds, speaks, and is corrected in its environment. Honest bounds: a small synthetic lexicon (twelve content words in a closed world built to be parsed, not open text); the twelve-word cap weakens the yoke (so the contingent advantage is a four-seed mean, not one cherry-picked seed); held-out comprehension needs enough referents because the distractor-set signatures are combinatorial; the correction headroom lives in the binder (warm the segmenter, then start the binding cold, a cold listener with a warm ear); and one content word in twelve is left un-isolated at cold start, where the silence self-check defers rather than misnaming. The open work is scale and realer language. Read.

End to end, no backprop

Theory update · 2026-06-28 · understanding, not lookup, end to end: a gradient-free learner develops the class codes itself and the binder generalizes to a novel entity on them, no backprop, learned-code CCGP predicting it at +1.00. Last round the comprehension organ generalized to a novel entity it was never told, with the cortex code's CCGP predicting when at +0.98, but the class-structured codes were hand-constructed, a dial turned by hand to stand in for a learned cortex, so it proved only the binder half: if the cortex emits high-CCGP codes, then the binder generalizes. This round closes that caveat. A gradient-free learner develops the codes itself, from context, no backprop anywhere. Entities recur in class-indicative contexts; the learner accumulates each entity's context histogram online and reads a code as the sign of a fixed random projection of that histogram, count plus a random projection, the whole gradient-free kit. The binder stores class-property bindings for known entities and must infer the property for novel entities coded from their own contexts. The dial is now the world: noise is the fraction of an entity's contexts drawn from the global pool instead of its class signature, zero meaning the classes are separable in context, one meaning the world is random with nothing to learn. Generalization is perfect (1.000) from a clean world out to noise 0.6, 0.767 at 0.8, and collapses to chance (0.217, against 0.200) only at noise 1.0 where the world is unlearnable. A stored control that did store the novel entity's property is near 1.000 throughout, so the pipeline is sound and the collapse is real transfer-failure, not a bug. Nothing is hand-fed: the codes are developed by counting contexts, and the novel entity's property is never stored, so a right answer can only come from the structure the learner pulled out of the world. And the unifying result: the learned code's CCGP geometry tracks generalization end to end, correlation +1.00, more robust than the hand-set sweep, holding a step further out and failing only when the world itself is random. The prior round showed CCGP predicts generalization for hand-set codes; this shows it for codes a gradient-free learner developed, the version that matters because it is the version the architecture must run. So the two engines and the binder are one system that learns its own codes: the cortex learns a high-CCGP abstract code, the hippocampus binds it, and generalization rides on the geometry, with no hand-feeding and no gradient. The hardest bet the architecture rested on holds here: gradient-free generalization through learned shared structure, the thing the Tolman-Eichenbaum Machine reaches only with backprop, won here with none, and the above-ceiling abstraction result is now load-bearing twice, it makes the binder generalize and a learner can develop a code with the property that powers it. The honest caveats name the next phase exactly: the context world is synthetic (disjoint class signatures; real language is messier, with polysemy and overlapping sparse contexts), the learner is a context-histogram bag rather than the apical abstraction stack (so this validates the principle, not that the stack specifically does it), and there is still no real language, no reader extracting role-filler structure from a raw stream, and no spoken answer, which are the next phase. One seed, the plateau and the cliff and the +1.00 correlation the structural reads. Read.

The cortex is the dial

Theory update · 2026-06-28 · understanding, not lookup: the binder generalizes to a novel entity it was never told, and the cortex's CCGP geometry predicts when, correlation +0.98. Two rounds built a comprehension organ that answers a question counting cannot and is typed, grounded, and truth-tracking, but every answer was about a fact the model was told. The honest worry was whether that is understanding or lookup with extra steps. The test that separates them is generalization to the novel: store a class property for known members, then ask about a member you never stored, where a right answer cannot be lookup because there is nothing to look up. The instrument puts the condition on a dial: each class is a prototype, a member is that prototype with a fraction of its bits flipped, from flip 0 (members identical, strong class structure) to flip 0.5 (members random, no structure), a stand-in for a cortex that has learned its classes. Store the property for twelve known members of each class, then unbind a thirteenth, never bound member from the bundle and clean up over the property codes. Generalization is perfect (1.000) from flip 0.00 through 0.30 and still 0.850 at 0.40, and collapses to chance (0.217, against 0.200) only at flip 0.50 where the class structure is gone. A stored control that did store the novel member's property is 1.000 at every flip, so the pipeline is sound throughout and the collapse is a real failure to transfer, not a bug. So the property the novel member never had stored is recovered from the shared structure of its class: understanding, not lookup. And the unifying result: the cortex code's CCGP geometry tracks generalization almost exactly, correlation +0.98 (raw class separation a weaker +0.55), so the abstract transfer structure of the code predicts whether the binder can answer about the novel. This is a quiet correction of the record: six earlier rounds measured CCGP against prediction and set it aside as "not our axis" (attention's lift was dimensionality, not abstraction; the count head pays a tax; the feedback arrow was null), and the gradient-free above-ceiling abstraction result sat novel and unused. CCGP was never the wrong axis; it was measured against the wrong job. Its job is generalization, and there it is decisive. So the two engines and the binder are one system: the cortex builds a high-CCGP abstract code, the hippocampus binds it, and generalization rides on the geometry, the data flow the architecture drew now closed and measured end to end. The honest forward-look, which is also the next bet: the class-structured codes here are hand-constructed (a prototype plus bit-flip), so this proves the binder half, if the cortex makes high-CCGP class-structured codes then the binder generalizes; whether the abstraction stack can learn codes that good gradient-free on realer data is the genuinely open end-to-end question. (A smaller honest note: one geometry proxy, a between-class-mean orthogonality measure, stayed flat across the sweep because the flip dial degrades within-class spread, not class-mean orthogonality, so that proxy is the wrong ruler here and CCGP and class separation are the faithful ones.) Read.

Typed, grounded, and true

Theory update · 2026-06-28 · the binder is a validated comprehension organ: typed, cortex-grounded, and truth-tracking, where the cheap strategy provably fails. The first positive showed a gradient-free binder answering one question counting cannot, but one passing task is what a hash map does. So this round tests the binder across three more capabilities, each gradient-free, each under the same per-episode randomization control that pins any counting model at chance. Typed: give every entity a color and a size and ask "what color," cleaning up over both property classes together. The typed binder (entity bound to a per-class role) is 1.000 on color and size both; an untyped bag holding the same facts answers the wrong class 0.504 of the time, a coin flip, so the per-class role is what makes the binding typed. Cortex-grounded: bind the codes a real neocortex would emit, derived from the spelling, instead of fresh random atoms. It comprehends exactly as well, 1.000 against 1.000, so the neocortex-to-hippocampus flow (encode in the cortex, bind the cortical code in the hippocampus) works end to end. The honest part: whether code similarity costs or helps came back inconclusive, because distinct spellings project near-orthogonal (cosine 0.014) so there was no interference to measure; the similarity payoff needs learned semantic structure, a next-phase question, not claimed here. Truth: "is the box red?" must answer NO when the box is blue, and the episode says "it is not red" first, so red co-occurs with the box. The binder retrieves the true color and compares, 1.000 on the no-cases; a co-occurrence reader says yes to the negated color and scores 0.000 on the no-cases, exactly the failure association predicts. So this is truth over propositions, not association: the binder knows the box is not red even though "red" was in the sentence. Together these make the binder a comprehension organ, not a one-task demo: typed binding rules out "it just remembers the entity," cortex-grounding rules out "only idealized atoms," negation rules out "association with extra steps." The honest forward-look is the steer: every test here and last round is hand-parsed with hand-coded comparators, so the next work is the architecture, a reader that extracts the role-filler structure from the stream itself (the cortex's engines feeding the binder), generation of the answer, and realer language, not more synthetic rungs. The missing organ is built; the job now is to wire it into a system that reads and answers on its own. Read.

Counting cannot bind

Theory update · 2026-06-28 · the pivot from compression to comprehension, and its first positive: counting is provably at chance, a gradient-free binder answers. Six rounds on prediction said the next symbol is a local game counting already wins, so the program changed the question: the goal is a brain-inspired cognitive architecture we can talk to, and the smallest version of talking to it is comprehension, not compression. Say two facts, ask a question, get the answer ("here is an apple, it is green, here is a cup, it is blue, what color is the apple?"). A four-stream fan-out converged on the finding that this is not a prediction problem but a variable-binding, one-shot relational-memory problem. Our two engines, the count predictor and the apical abstractor, are both the neocortex, slow and dense, structure pulled from many examples; the step the task turns on, bind this entity to this property on a single encounter and read it back from a cue, is the defining job of the organ the brain runs in parallel and we did not have: the hippocampus, fast, sparse, content-addressable. The mechanism that fits all four rules (online, gradient-free, bounded, on the sparse codes we already emit) is vector-symbolic binding, the engineering form of the Tolman-Eichenbaum Machine: a statement binds, a question unbinds, the answer is cleanup to the nearest stored item. The instrument is the apple/cup task with one load-bearing control: the property is randomized every episode, so corpus-wide every entity is equally often every color, which makes any counting or co-occurrence model provably sit at chance. Only reading this episode's binding can answer. The result, held-out QA accuracy against a chance of 0.250: counting is pinned at chance, the strongest co-occurrence reader scoring 0.246, dead on the line; an exact dictionary scores 1.000 (the instrument is valid); the corrupt-store control collapses to 0.243 (the store does the work, not a leak); and the gradient-free VSA binder over random codes scores 1.000, clearing the roughly three-year-old binding bar (0.75 to 0.90) with room to spare. Capacity is graceful: eleven facts in one episode still 1.000 at full width, 0.999 at a quarter of it, bending only at a tiny sixty-four dimensions to 0.884. This is the first positive of the neocortex-plus-hippocampus architecture: the gradient-free relational binder comprehends where the count engine provably cannot, the smallest thing we can talk to, working, in regime. Honest bounds: the fillers are random codes, not learned cortical codes, so this tests the binding mechanism, not yet generalization to novel entities (the next phase, where the abstraction engine's above-ceiling code becomes the generalization dial, and where the gradient-free constraint bites hardest); Stage zero, one, two only; one property class (color); typed binding, relations, negation, and novel-entity fast-mapping all untested up the curriculum ladder; one seed, with the control, the collapse, and the capacity falloff the structural reads. Read.

Abstraction cannot cheapen prediction

Theory update · 2026-06-28 · the feedback seam is null for the next character: a tag can carry information and still buy no prediction. The rank-reuse tension said one shared code cannot be both abstract and count-reusable; the cortex's own answer (Bastos 2012, Larkum 2013) is to keep both codes and share information, not representation, through a coarse repeatable tag, the feedback arrow from the abstraction engine to the prediction engine. This builds that seam, count-only, the cheapest decisive leg: a count predictor on the raw context, conditioned on a coarse abstraction tag, against the same machinery without it, on the identical text8 split, with a shuffled-tag control to separate the tag's information from mere count-splitting. The result is an honest negative with a clean diagnosis. The tag carries real information and still buys no prediction. A short-horizon tag, the abstraction stack's own deep code, beats its shuffle by up to 0.385 bits-per-char, so it is genuinely informative, but it is redundant with the order-six n-gram (it re-encodes the same six characters), so a soft coupling that never displaces a well-counted raw cell reaches break-even and no better: at a million characters the net is +0.001, exactly the baseline, never below it. A long-horizon tag carries almost no information. A recency-weighted gist over the characters outside the n-gram window beats its shuffle by about 0.013 bits, no better than a random tag, and conditioning on it only hurts (+0.06 to +0.10). And it does not help where local context is weakest: at word-start positions, the least locally determined in the text, the long-range tag still costs +0.05 to +0.11, with no scale dependence. The unified finding: next-character prediction is an overwhelmingly local game that counting already wins, and information is not usable prediction gain, redundant where present and absent where new. The same prediction-versus-abstraction split the rest of the program keeps finding, seen once more. So the conclusion is not that the coupling fails but that the ruler was wrong. The mutual loop is already half-built, because the abstraction stack's per-unit credit is the prediction error, so prediction already teaches abstraction, and the only missing arrow is the one shown null here for the next character. The place to measure the cross-feed is the next word and the coherence of generated text, not the next character, and the soft-coupling mechanism is validated and ready for that scale. Honest bounds: it is the theme gist and the short-horizon code that were tested (a sharper long-range signal, the previous word or a learned long-range code, is untested); character bits-per-char only; the frozen abstraction substrate; two hundred thousand and a million characters, one seed, where the slope is the read. Read.

Where we are

Theory update · 2026-06-28 · a standing-back regroup: the program has built two engines, each strong on its own axis. Step back from the run of experiments and the shape is clear. The work has converged on two engines. A prediction engine (the diverse-view count bank, the BLEND-NORM calibrated pool, grow-then-prune growth) reads the next character at the n-gram level, near 2.1 bits-per-char, count-native and gradient-free. An abstraction engine (the sparse apical stack: k-WTA, boosting, positive weights, per-unit precision-gated apical credit) builds a transfer-abstract code at or above the 0.484 backprop ceiling with no gradient at all, the genuinely novel result of the program. Each is strong on its own axis; neither does the other's job. The wins that got us here: P0.5 cracked the abstraction-wall winner-collapse into stable sparse abstraction (boosting plus positive weights, both needed); grow-then-prune broke the bits-per-char plateau (growable 1.875 at ten million where the fixed bank flattens at 2.109); BLEND-NORM corrected the prediction calibration (a normalized blend a full bit better than the over-sharpening product pool); a gentler learning rate found a lower floor (0.001, a U-shaped optimum, the locked rate too hot); and the attention and working-memory subsystem closed its own arc. The honest scorecard: on prediction we are a mile from a strong transformer, not a hundred light-years, the n-gram level and about 0.85 bits behind text8 SOTA but in the right universe; the abstraction result is the real prize, a gradient-free learner above the backprop ceiling. The deep open problem is that the two engines do not fuse: a count head on the abstract code pays a stable 0.45 bits-per-char tax, the rank/reuse tension, because the high rank that makes a code abstract makes it nearly unique per position, and a count table predicts by reuse. One code cannot be both abstract and count-reusable. The bet to resolve it is the cortex's own answer: two coupled laminar layers (Bastos 2012, Larkum 2013) that keep both codes and share information, not representation, the abstraction engine's apical credit already a Larkum cell. What comes after: the larger architecture, generation and production (the internal birdsong loop is real, the external referential game the named next move) and learning by example. A regroup, not a new experiment: the map the next stretch is drawn on. Read.

Gentle learning compounds

Theory update · 2026-06-28 · the locked learning rate was too aggressive, and a gentler rate lets re-reading compound. The re-reading round found repetition a modest prediction lever that flattened by the third pass. The hypothesis this round tests: that plateau is an artifact of the locked learning rate (0.005), too aggressive to let repetition compound, where a gentler rate would keep paying across passes. So a sweep of the apical rate over one, five, and ten passes, with everything else held to the re-reading config, evaluated on a disjoint held-out slice, two seeds. The hypothesis held: a gentler rate reaches a lower floor. At 0.001 the held-out bits-per-char reaches 3.637 against the aggressive baseline's 3.762, 0.125 bits better at matched exposure, and every gentle cell beats the baseline. The optimum is a U. 0.005 is too aggressive (it crams the slice on the first pass, then runs nearly flat, last-pass change about a thousandth of a bit), 0.001 is the sweet spot, and 0.0005 is too gentle (it underfits, starting highest at 3.905 and plateauing highest at 3.720). Repetition compounds under the gentle rate where the aggressive one stalls: the 0.001 curve descends smoothly across all ten passes (3.749 to 3.685 to 3.673 and on down to 3.637), the 0.005 curve flatlines after the first re-read. No overfit, and the gentler rates hold a tighter generalization gap (0.030 bits at 0.001 against 0.045 at 0.005, at ten passes). The broader point reaches past re-reading. The rate that proved too aggressive is the apical stack's global learning rate, fixed by the architecture sweeps and used everywhere the stack runs, so the locked value is likely too hot on every run, not only on re-reading, and the global rate should be revisited. Honest bounds: a prediction lever only (transfer CCGP stays noisy around 0.4 to 0.59, flat, above the 0.484 ceiling, and does not rise with the bits); the absolute gain is modest (about a quarter of a tenth of a bit over five passes); it is the apical learned stack, not our best predictor, since the count voting bank near 2.1 is re-reading-invariant and has no rate to tune, so the headline does not move; a 200K slice, two seeds, one harness. Read.

The rank-reuse tension

Theory update · 2026-06-28 · one substrate can feed both heads, but the abstract code pays a count-reuse tax. The last round named two paths forward; this takes the second, graft a count-native head onto the abstract code, one substrate with two heads. A count table keyed on the P0.5 abstract winner-set predicts the next character, against the identical count machinery keyed on the raw context, on the same frozen substrate, from eighteen thousand to one hundred thousand characters. The result is an informative middle that leans deep-negative, and it carries a finding either way. The count head extracts real signal from the abstract code. Counting on it reads 2.578 bits-per-char at one hundred thousand characters, 0.8 to 1.0 below the stack's own readout (3.548), so the abstract code does carry next-character signal, and one substrate can feed both heads: a real sub-win, the count-native head a better readout for the abstract code than the neural readout it ships with. But it stays about 0.45 bits behind raw context, and the gap holds with scale. The same count machinery on the raw context reads 2.141, and the abstract code trails it by +0.50 at eighteen thousand and +0.44 at one hundred thousand, stable across 5.5 times the data, not a small-data artifact, while the same code's transfer CCGP is 0.499, above the 0.484 backprop reference. So the same code abstracts well and predicts worse, and the prize, low bits and high abstraction in one representation, is not reached. The mechanism is a tension in the representation itself. A table-fill diagnostic shows the abstract code is nearly unique per position, three to six counts per key even at seventy thousand training positions, because the high rank that earns the transfer CCGP, hundreds of distinct winner-sets, is exactly what makes the code almost never repeat, and a count table predicts by reuse, so it backs off to coarser subsets and predicts bluntly. High rank gives good abstraction and poor count-reuse, and the two pull against each other in the same code. So abstraction and prediction want representations in tension, and the two-stacks split the program kept drawing has a root not just in the wiring but in the code. Honest bounds: eighteen thousand to one hundred thousand characters and one to two seeds, so the absolute bits-per-char are scale-compressed; the robust finding is the tax, about half a bit, stable across 5.5 times the data, count-on-abstract behind count-on-raw with identical machinery; the sub-win over the native readout is robust too; the deep-negative reading is leaned-toward, not closed, because the lever the diagnostic points at is untested. That lever is a sparsity-versus-reuse sweet spot: a coarser, less-sparse code (a smaller top-k winner-set) might repeat often enough to predict while staying abstract enough to keep its CCGP. Read.

One stack, two halves

Theory update · 2026-06-28 · an honest trade-off: abstraction composes and exceeds the ceiling, prediction does not transfer. The program split into two engines. The voting bank predicts (many count columns, each a view, pooled by a calibrated blend, low bits-per-char); the sparse apical stack abstracts (k-WTA with boosting, positive weights, and per-unit credit, at the backprop ceiling). Neither does the other. So the test: give the abstraction stack the prediction machinery, eight sparse columns each on a different view, voted by the same blend, abstraction kept intact, and see whether one stack does both. Three arms, eighteen thousand characters, depth three, two-seed mean, verbatim probes. Abstraction composes, and it exceeds the ceiling. The unified stack reaches transfer CCGP 0.520, above its own single-view P0.5 (0.487) and above the 0.484 single-layer backprop reference (GRAD-1L), at higher dimensionality (PR 19.5 against 8.1), with no collapse: the highest abstraction the program has reached, gradient-free. The earlier multi-view failure does not reassert, because k-WTA and boosting preserve the cross-stimulus variance the old activation normalization destroyed, so diverse views add to the abstract code instead of flattening it. Prediction does not transfer. The unified stack reads 3.666 bits-per-char, barely below its single-view P0.5 (3.753) and nowhere near the voting bank's 2.442. The diagnostic is sharp: the blended bits-per-char (3.666) equals the bank's own best single column (about 3.66), so the vote does nothing for prediction with neural columns. The voting bank's bits-per-char win is a property of count tables combining, sharper conditional tables averaging into a sharper prediction, not of voting per se, and sparse apical columns do not combine that way. So the prize, one stack low on bits and high on abstraction, is not won at this scale: the synthesis kept and improved abstraction, and did not import prediction. Honest bounds, kept: eighteen thousand characters and two seeds (the bits-per-char anchors are scale-compressed at that size, the voting bank reaches about 2.1 only past two hundred thousand characters); the 0.484 is the single-layer backprop reference, so "exceeds the ceiling" means it passes GRAD-1L, and a multi-view backprop is the fair ceiling to recheck before any claim of gradient-free past backprop; and the boosting-plus-positive-weights config is the only stable arm (the permanence-credit and full-credit variants collapse at this scale). The path forward, the diagnostic's two directions: train each column's readout on the pooled error so the vote is fit jointly (couple the columns), or graft a count-native head onto the P0.5 abstract code (the abstract sparse code as the context key for a count predictor, one substrate with two heads). Read.

Growth breaks the plateau

Theory update · 2026-06-28 · a growable substrate breaks the bits-per-char plateau. The prediction calibration found a flat spot: our fixed twelve-column voting bank improves to about a million characters and then stops (2.234 to 2.112 to 2.109 bits-per-char, flat from a million on), while a plain n-gram keeps falling because it keeps growing, adding longer-context counts as data arrives. So the test: if the plateau is a fixed-capacity artifact, does a growable bank break it? This builds the grow-then-prune the program wanted, on the identical text8 split and the identical bits-per-char loop. Grow a longer-context voting column when the held-out gain decays (the columns have saturated), cap the count, prune the least-used at the cap. The result. The fixed bank plateaus and the growable bank does not. Past a million characters the fixed bank moves 2.112 to 2.109 (a drop of 0.003, flat); the growable bank moves 2.004 to 1.875 (a drop of 0.129, still falling), growing 12 to 16 to 24 columns. At ten million the growable bank reads 1.875 against the fixed bank's 2.109, 0.233 bits ahead, and within 0.047 of the n-gram floor (1.829). It grew the right thing: progressively longer-context views, its longest look-back span climbing from five to eight, exactly the capacity the fixed bank lacks and the n-gram wins by; at two hundred thousand characters it correctly did not grow (too little data to saturate twelve columns, so it ties the fixed bank there); and the cap held with the prune engaging at ten million. So the plateau was a fixed-capacity artifact, and growth dissolves it: the growable bank restores the n-gram-like scaling slope by doing what the n-gram does, adding longer-context views with data, under a memory budget. Honest bounds, kept: it reaches about the n-gram (1.875 against 1.829), not past it; it stays about 0.85 bits-per-char behind text8 SOTA near 1.0; and prediction is not our headline axis, so this does not move the home axes (abstraction, continual learning). The deliverable is the capability: a count-native, online, gradient-free, bounded substrate that keeps learning from data instead of saturating, the thing scale needs for abstraction too. One seed. Read.

The empty cell, filled

Theory update · 2026-06-28 · a sparse code holds stable abstraction at the ceiling, with boosting and positive weights together. The empty cell gave a partial answer: per-unit credit on a sparse code abstracts, then the credit pathway collapses it. This runs the diversity machinery that answer named, as an ablation ladder over the collapse baseline, one piece at a time, to see which converts the transient peak into a held result. The cure is two pieces and it takes both. Boosting and positive weights together hold deep transfer abstraction at the 0.484 backprop ceiling (0.485 and 0.489 across two seeds), to the last checkpoint on both seeds, with more than two hundred distinct winner-codes (208 and 225, against the dense reference's about forty) and the lowest bits-per-char of any arm (3.74 and 3.77). Two honest negatives ride along. Boosting alone still collapses: the signed weights let the credit drive a few units to always win, and homeostasis alone cannot overcome it; clamp the weights non-negative and then boosting keeps the winners spread, so the minimal sufficient pair is both, not either. And the sign-only permanence credit breaks the code (it saturates the active-to-winner synapses), so the credit stays graded with diversity machinery around it. The substrate-artifact thesis holds in its amended form: k-WTA and the hypersphere support alone are not enough, but adding boosting and positive weights is, and the dense reference is no longer the only stable abstractor. Honest: it matches the ceiling, it does not exceed it, at eighteen thousand characters and two seeds, with the depth-four run and the hundred-thousand-character confirmations and the bits-per-char plateau test running. Read.

A selector that cannot win

2026-06-28 · an honest negative that closes a long-open question. A bank of critics watches the model and a selector picks when to think harder: the count-native System 2, a controller that fires a deliberate pass only when the moment calls for it. This round builds it on the locked model and races it against the trivial fixed policies, always think, always consolidate, or never. The selector loses, reading 3.984 bits-per-char against the bare model's 3.920 and the replay-every-character policy's 3.618, last of the four. The loss has a precise mechanism in three parts. The self-calibrating gate works and only surprise is selective: every critic tunes itself to the target fire-rate, but surprise holds it with the highest threshold (a real tail of prediction error) while conflict and low precision hit the rate only by collapsing toward the bulk, no tail to find. The selective signal gates the wrong way to think: wiring a deliberate re-read to the surprise tail inherits the harm of always thinking (the apex dimensionality dragged from 3.78 to 2.94, toward collapse). And the way to think that helps wants a dose, not a tail: replay is monotone in how much you do it, so always-consolidate wins outright (3.618, the highest dimensionality of the four) and a sparse selectively-gated replay is worse than both heavy replay and none, the intervention perturbing without the payoff. So at the character level no count-native critic signal, used to gate a sparse intervention, beats the trivial fixed policies, because the one selective signal triggers the harmful mode and the helpful mode wants a dose. This closes the question the bank-of-critics round left open with a no, and it sharpens the standing System-2 gap into a precise obstacle: the count-native deliberate controller is hard because the selective signal and the helpful action do not line up. Honest: the character level (a coarser grain makes conflict selective again, where the deliberate pass first worked), fifty thousand characters before the drift fully bites (a long drift sweep is the steelman, though replay's monotone benefit should still let always-consolidate win), and one seed on one substrate. Read.

The vote was too loud

2026-06-28 · a calibration win surfaced by an honest Global Workspace negative. Many columns each guess the next character and a combiner reconciles them; the model's pool is a precision-weighted product of experts. Global Workspace Theory describes a different combiner, ignition: one winning coalition seizes the workspace and is broadcast to everyone, all-or-none, not a weighted average. This round tests that as a readout on the voting substrate, twelve count columns, the mean of three seeds, swept to two hundred thousand characters. Ignition does not win for prediction. Winner-take-all reads 2.600 bits-per-char against the product pool's 3.307, better, but only because it de-sharpens an over-confident pool. A normalized blend that keeps every column reads 2.314, better than both by about one bit per char, and it beats hard winner-take-all at every scale (2.115 against 2.439 at one hundred thousand). The diagnosis is over-sharpening: the un-normalized product raises the consensus distribution to the power of the summed confidences (about three over ten active columns), an overconfident distribution that inflates the bits. Normalize the exponent to about one and keep every specialist, and the pool predicts a full bit better. So all-or-none access discards evidence the blend keeps and uses; the right pool is every column at the right temperature, the calibrated geometric mean the main vote already runs. The ignition coalition is real and dynamic (the winner changes position to position, win-entropy about 0.8, no single-column collapse), it just is not the right readout, and abstraction is flat (transfer 0.63 against 0.54, the participation-ratio inflation from 24 to 45 the expressiveness signature, not abstraction). The honest correction: the experiments built on this substrate ran the un-normalized pool, so their absolute bits were inflated by about one; the within-experiment comparisons hold (every arm paid the same tax), only the absolute numbers were high, and the normalized blend is the corrected default. A serendipitous calibration fix and an honest negative on a consciousness-research idea. Read.

Two levels is enough

2026-06-28 · the attention track's honest bound, completing the arc. The first three attention rounds built a two-level guide (a higher level steers where the lower one looks) gated on a fixed threshold; this round tests how far that climbs by adding a third nested level and a gate that calibrates itself. The third level is a slower context above the theme, admitted more slowly and firing on a slower read of the model's prediction, whose coarse code routes the theme's bucketing the way the theme routes the columns' offsets, a nested top-down chain. The gate's fixed multiple is replaced by one that targets a fire-rate, nudged toward the target by its own observed rate, so it fires at the same rate at any scale. Capacity-matched arms (context-keying splits the bottom level's data across buckets, so the comparison is read at the same bucket count), eight columns, two hundred thousand characters, the mean of three seeds. Two honest answers. Depth is a diminishing return: the three-level nested stack reads 3.096 bits-per-char against the two-level stack's 3.049, worse by 0.047, at both scales, even though the nesting is real and active (the slow context routes the middle level, cross-band divergence a full bit, and the middle level routes the bottom one). The extra bucket-splitting the third level needs to carry its super-band costs more than its longer-range guidance gains. The self-calibrating gate is a keeper: at the same capacity it reads 3.049 against the fixed-threshold gate's 3.072 and beats the hand-set floor (3.101), and it holds its fire-rate at the target (about one fire in ten) at both fifty and two hundred thousand characters, fixing the previous round's named under-fire where the fixed multiple drifts down with scale. So the attention architecture is bounded: two levels of guidance with a self-calibrating gate is the sweet spot, and a third level does not pay off as built. This is a ceiling and a keeper, an honest bound that completes the track rather than failing it. Honest: the depth verdict is read at matched capacity (against the smaller-bucket arm a second reason, fragmentation, moves the numbers, so that comparison is kept only as the absolute floor); the third level loses as built and could pay with a richer slow-context read or far more data; the gate-fix margin is modest at these scales and its larger payoff is past them; one count substrate, fifty and two hundred thousand characters, three seeds. The arc from learning where to look to a bounded, calibrated, two-level attention stack is complete. Read.

A conversation with Haiku

2026-06-28 · the generation track's first live conversation. The producer had an offline answer key; this round puts it in a live exchange with a real Haiku partner, run with no API key through the logged-in claude -p in a child register. The cortex takes a turn through its one frozen comprehension pathway, used both ways: it comprehends Haiku's turn (each content word run through the frozen monitor, the distributions over its 160 known real-word meanings summed, the top of that sum what the turn is about) then produces a reply (the real word it can voice for each top meaning, a three-word utterance about what Haiku said), and the reply goes back to Haiku. Fourteen exchanges, live, seed 0. The honest measures: 42 of 42 emitted words are real dialog words, 10 of 14 exchanges share an exact word with Haiku's turn, and Haiku stays on the cortex's topic in 12 of 13, with the comprehension peak at 0.105 over 160 meanings (about seventeen times chance). The "cool to see" part is what threads across turns: a topic word forms a call and response that holds for a few exchanges, the cortex says match, Haiku reflects it, the cortex says it again (exchanges five through eight), then dance the same way (eleven through fourteen), a coherent multi-turn topic word, not isolated word-salad. Three verbatim turns: Haiku "That's a fun fan! Do you have one at school?" to cortex "for fan school"; Haiku "you want to find a match for a place?" to cortex "for place match"; Haiku "Do you like to dance?" to cortex "match dance for". Honest assessment: word-salad with real glimmers of relevance, the starting line for live dialogue, not a conversation that holds. The production gap the offline round predicted is visible live, the cortex comprehends the right region of meaning but the producer often voices a neighbor of the word it reached for, so relevance runs ahead of recovery. Caveats kept sharp: a single live session and seed (a live partner is not deterministic across runs, so this is the captured transcript), and the keyless claude -p is the project-aware agent not a bare model, so it breaks character about three turns in fourteen (a role reminder prepended to each turn it hears helped, and the cortex's child-like replies pull it back). The first live read-speak-hear loop: the producer is real, the loop is live, the glimmers are the honest first sign of dialogue. Read.

Attention is prediction, not abstraction

2026-06-28 · the program's deepest verdict, framed precisely. The attention track lifted transfer CCGP to about 0.60, past the 0.484 backprop ceiling that capped every gradient-free attempt at abstraction. That number lived on the vote-consensus probe, which rises with dimensionality, not the gradient-comparable probe the ceiling is drawn on. So this round puts it on the right ruler. Take the context-guided attention representation, verbatim, and feed it as the input to the per-unit-credit readout, the one pathway proven to abstract, then measure the readout's transfer CCGP on its own hidden code. Two arms differ in one thing: a plain last-five-character window against the attention representation, same readout, same probe, same data, same seeds. At eighteen thousand characters, where the readout is stable, the attention input reads 0.429 and 0.463 against the plain window's 0.509 and 0.476, under the baseline and under the 0.484 ceiling on both seeds. A backprop oracle clears the ceiling on the same attention input, so the feature is real signal, not a dead one: the gradient-free credit pathway simply does not carve a more abstract code from it than from a raw window. So the 0.60 was expressive dimensionality, a richer code, not a more abstract one, and the two tracks separate cleanly. Attention wins at prediction and working memory; abstraction needs the credit pathway; and bridging them does not add the wins together. The road past the ceiling is not a richer input, it is a change to the credit pathway itself, a sparse substrate or credit across time, the only door left and now tested as such. Caveats kept sharp: the clean comparison is at eighteen thousand (a two-hundred-thousand run agrees on the verdict but the readout diverges there, a known stability limit), one architecture and one probe family, and the negative is a real one precisely because the oracle proves the feature informative. A negative that sharpens the theory: the wall stays a credit-assignment problem. Read.

Producing real words

2026-06-28 · the first intelligible production on real conversation. The previous round built the separate learned inverse that beats the reader run backwards, but on invented words. This round carries the same organ to real ones. Nothing in the mechanism changes: the producer learns its emission menu from the DailyDialog stream, the comprehension monitor learns the char statistics of real talk, and the target meanings become the top 300 most-frequent real dialog content words. One specific intended meaning at a time, graded by a frozen judge the producer never trained against, a produced string valid if it is a real dialog word and recovered if the judge maps it back to the meaning meant. On the matched conditional task the separate inverse produces 113 of 300 real words and recovers the intended one 0.317 of the time, against the reader run backwards at 45 of 300 and 0.090: 2.5 times the coverage and 3.5 times the recovery. The produced words are real English: lady to ready, helen to hello, agreed to good, equipment to question. The inverse learns it over training (recovery 0.037 to 0.317 over forty episodes), and it wins recovery in every frequency band, four times over on the rare tail (20 of 100 against 5). Two honest nuances keep it the right size. The headline is recovery, not raw coverage: the gap narrows from 41-to-1 on invented words to 2.5-to-1 here, because real onsets let the reader reversed complete to some plausible real word and score on validity, so the live number is producing the intended meaning. And recovery 0.317 is moderate, because the meaning code is a coarse heard fingerprint, a form region not rich semantics, so a produced word is often a phonetic or semantic neighbor of the target (lady to ready) rather than the target itself. Caveats kept sharp: a single seed, a form-based meaning code, the conditional task and not the unconditional free-run (the reader reversed free-runs 336 real words). The wall's fix carries from invented words to real ones: production is reachable with a separate organ trained on its own feedback, and the reader reversed is an onset-completer, not a meaning-producer. Read.

A gate for working memory

2026-06-28 · the attention track's capstone. The previous two rounds learned where to look and then guided it by context, holding the context in a working memory that leaked at a fixed rate, with no say over when to take in new context and when to hold what it had. This round gives it that say, the first module from the cognitive-architectures reading. A basal-ganglia-style Go/NoGo gate decides, at each character, whether to update the working-memory theme or hold it, and it fires on a model-update signal: how far the model's own pooled prediction moved from one step to the next, not raw surprise, the shift that carves events where surprise carves words. The gate keeps a running mean and spread of that signal and fires when it is in the high tail, an adaptive threshold that self-calibrates, learned online and gradient-free with no reward. On a fire the theme admits the current character; otherwise it is held. Eight columns, two hundred thousand characters, the mean of three seeds, at eight context buckets. The gate reads 3.069 bits-per-char against the fixed decay's 3.083 and the hand-set fixed 3.101, and it beats both degenerate bounds, always-update (3.075) and never-update (3.114), so the win is the timing of the update and not the amount of theme motion. It fires 19 percent of the time, holds the theme about eight characters at a stretch, fires at or next to a word boundary three times in four, and the abstraction lift carries through (transfer CCGP 0.584 against fixed's 0.494). The fifty-thousand-character scale agrees (gated 2.979 against fixed-decay 3.050). Caveats kept sharp: the gate is inert below eight buckets (the theme band drops out of the bucket the scan reads), a fixed threshold under-fires as the data grows (a fire-rate target is the next fix), and it is one substrate at trimmed scale. This closes the attention arc: the full learn-where-to-look, context-guided, gated-working-memory stack beats hand-set fixed diverse views, online and gradient-free. Read.

A baseline on dialogue

2026-06-28 · the generation track's first measurement on real conversation. The track had a producer and instruments, but it never read human dialogue. This round gives it a baseline and a measuring stick to move. The locked model streams strict-online over the whole DailyDialog training split, 5.24 million characters, then is measured on held-out dialogue it never saw, on the same held-out text8 with the same stick, and on a minimal-pair grammar battery. Two honest readings. Dialog is a bounded, repetitive distribution the model fits stably: held-out bits-per-char drops with data, from 3.767 at two hundred thousand characters to 3.549 at the full split, with no drift, the mirror image of the fresh text8 stream where it rises. And on the shared stick the dialog-trained model reads 3.549 on dialog against 4.090 on text8, so everyday talk is 0.542 bits-per-char easier than encyclopedia text for a model trained on talk. The warning is the second reading: a long pass on the narrow register erodes general grammar. The grammar macro falls from 64.8 percent at two hundred thousand characters to 51.1 percent at the full split, word-order sensitivity holding at 100 percent while subject-verb agreement, determiner-noun agreement, and negative-polarity licensing sink to or below the chance line. The free-run samples are gibberish with the right shape: correct space density and the turn marker recurring, no words. Caveats kept sharp: a single seed, the battery is a hand-built analogue of the published suite, the fall is the within-experiment change (a separate count-band model scored 60.2 percent, context only), and the gibberish is the intelligibility floor by design at this scale. A measured baseline, not a new mechanism, and the same specialization-versus-generality tension the drift theme carries, read from the narrow-register side. Read.

Where you look depends on what you understand

2026-06-28 · context-guided attention. The previous round made the scan learnable but context-free, and it trailed the hand-set views, because attention is supposed to look differently depending on what is happening. This round puts the context back. A working memory one level up (a leaky char-distribution over the region plus the live word-phase) is keyed into each lower column's offset policy: the column learns "in context X, look at offset Y," same gradient-free bandit, now one value vector per context bucket, with the same lateral inhibition on the action space kept. Eight columns, two hundred thousand characters, the mean of three seeds, swept over two, four, and eight context buckets. Two honest wins. At two buckets the context-guided scan reads 3.022 bits-per-char against the hand-set fixed 3.101, a gain of 0.079 on all three seeds, the prediction win the context-free policy could not reach (it trailed fixed by 0.284). The robust win is abstraction: context-keying lifts transfer CCGP from fixed's 0.494 to about 0.60 and dimensionality from 19 to 32 at every bucket count, where the context-free round left abstraction flat. And the policy is genuinely context-dependent, looking in different places in different contexts (cross-bucket offset divergence 0.80 to 0.85 bits). Caveats kept sharp: the prediction win is bucket-sensitive (four buckets reads 3.114, just above fixed, because keying splits the bandit's data), and the working-memory prediction prior consistently hurts, so it is dropped, the win is on the attention side. The bucket count is the open knob, and the right count grows with the data. Context is what attention is for. Read.

Learning where to look

2026-06-28 · the attention direction opens. The model reads through many columns, each over a window a fixed distance back from the cursor, and those distances are hand-set. This makes the look-back offset an action, chosen fresh every step and rewarded by how well the column then predicts, the "motor is moving through text" idea made concrete. Three arms over twelve columns, two hundred thousand characters, two seeds, differing only in how the per-step offset is chosen: fixed (hand-set diverse offsets, the baseline), random (resampled each step, the floor), and learnable (a count-based bandit over the offset menu, gradient-free, with a content-relative word-jump action). The learnable policy reads 3.797 bits-per-char against random's 3.872 (a gain of 0.076) and learns real structure: the offset distribution is non-uniform (mean KL from uniform 0.68 bits), 7 of 12 columns settle on distinct greedy offsets, and several learn to use the word-jump. The load-bearing mechanism is one wire: pure local-predictive reward collapses every column onto the single most recent character (the same-view degeneracy, bits-per-char explodes), and the policy works only with lateral inhibition on the action space, the same diversity pressure that broke the deep collapse, carried from the code to the scan. Honest verdict: learnable, gradient-free, online attention is real and beats random, but context-free it does not match the hand-set spread (fixed 3.513, trailing by 0.284), because when context does not matter a fixed diverse spread is already good and free. Caveats: one substrate (count columns, CPU), 200K and two seeds, and the diversity weight is the one sharp knob. The payoff is context-dependent attention, where higher levels steer where lower levels look, which fixed offsets cannot do, the next experiment. Read.

Producing on purpose

2026-06-28 · production is a separate organ. The comprehension-production wall named its own fix, and this builds it. The locked reader recognizes most words and produces few, and plays a specific learned word back to the exact spelling only 2 times in 296: a recognizer reversed returns the most common continuation, not the word you meant. The neuroscience (the DIVA model) says production is a separate learned map from a meaning to a sequence, trained feedback-first, the way a songbird learns its song by hearing itself. So this builds that separate map and races it against the reader run backwards. On the matched conditional task, produce one specific intended meaning across 300 meanings on a controlled vocabulary, judged by a frozen judge, the separate inverse recovers 82 where the reader run backwards recovers 2 (validity 0.300 against 0.030). The inverse learns it over training (recovery climbs 0.060 to 0.417 across forty episodes of the self-feedback loop), and it wins in every frequency band including the rare tail (28 of 100 at the rarest, against 4), the novel low-frequency axis the model flagged. The reader reversed is not a producer: its top-down wire carries credit and priming, not generation. Honest: this is the conditional task, not the unconditional free-run number (free generation reads 45 of 500 here), recovery 0.417 is moderate, and the vocabulary is controlled with a coarse heard-fingerprint meaning code. The wall's fix is real: production is reachable with a separate module trained on its own feedback. The internal half of generation now has a working building block, not just a mechanism. Read.

What changed: two mechanisms that compose

Theory update. 2026-06-27 · the first composition that holds. The drift is the open frontier: on a long fresh stream the locked config drifts, prediction degrading and the deep code collapsing toward rank one. Two anti-collapse mechanisms have turned up, each a partial fix on a different axis, and this round runs them together. Lateral inhibition between columns lowers bits-per-char and lifts abstraction above the 0.484 ceiling, but the deep code still collapses (dimensionality to rank one at two million characters). Interior-altitude replay (Minsky's K-line) holds the deep code high-dimensional, but its abstraction fades below the ceiling (to 0.292). With both wires on, the deep code stays high-dimensional (apex 3.75 to 5.80) and above the ceiling at every checkpoint (0.560 to 0.596), stable, bits-per-char bounded, where each half fails on a different count. They act on orthogonal axes (across columns within a level, and across time at an interior level), so they cover each other's blind spot rather than redundantly stacking, the opposite of the capstone combine where two same-axis levers did not add. This is the first composition in the program that holds, and the drift now has a candidate fix. Honest: the trimmed three-million-character stream and one seed, the trajectory across checkpoints is the signal not any single noisy point, both's bits-per-char still ticks up at the last checkpoint (drift tamed, not flattened), and a ten-million, multi-seed confirmation run is underway before this locks into the architecture. Read.

What changed: lateral inhibition clears the collapse

2026-06-27 · the first positive past the levers-not-additive wall. The deep collapse that locked the architecture at depth three is not structural. The capstone combine showed breadth and normalization act on different variance axes and do not add, and bolted the door: stacking more independent columns cannot undo a code that goes rank one across stimuli within each column. This adds the degree of freedom the locked stack never had, columns at the same level that see and suppress each other (lateral divisive inhibition, plus a Monty-style consensus vote). On the exact depth-four regime where the combine collapsed to dimensionality 1.17, lateral inhibition takes the apex from rank one (1.61) to a real high-dimensional code (3.05) and lifts the abstraction score +0.16 (0.315 → 0.473) with the lowest bits-per-char. A vote alone leaves the code rank one (1.46): the divisive inhibition, not the combiner, is the load-bearing piece, and both together are best. With every lateral knob off it reproduces the locked collapse bit for bit, so the lift is the lateral wire and nothing else. The missing ingredient the locked round was hunting was interaction within a level, not more breadth across levels. Honest: 18K characters and two seeds (the 100K confirm shared the GPU with the other arms this round), seed-robust; the win is on dimensionality and prediction, and whether it carries the deeper code above the ceiling rather than off the floor awaits the longer run. Read.

What changed: replay at an interior altitude

2026-06-27 · the second anti-collapse mechanism, Minsky's K-line. A long fresh stream makes the locked config drift: bits-per-char rises and the deep code collapses, and replay (re-firing stored configurations) is the strongest patch. Minsky's K-lines say a re-evoked memory should be reinstated at an interior altitude, not the raw surface and not the top. This makes that a sweep, varying only where the replayed configuration is consolidated, the fresh step identical across arms. There is a sharp interior optimum: reinstating at the interior level reads a deep dimensionality of 5.04 against input replay's 1.99 and gives the lowest bits-per-char (3.582), beating both input replay and the baseline. Replaying at the top alone is too vague and craters the top's own abstraction (0.173); replaying the interior and top together over-consolidates and collapses the apex outright. The interior optimum holds across every checkpoint. So the K-line altitude is real, with a number. This is the second anti-collapse mechanism this round, on a different axis from lateral inhibition: across time at an interior level, where the other works across columns. Honest: the win is on dimensionality and prediction, not the abstraction ceiling (the deep transfer score ties input replay's), and it is a three-million-character run, not the full ten where the drift first appeared. Read.

What it learns, and what it can say back

2026-06-27 · a wall, read two ways. A standing question made concrete: of the words it reads, how many does the model learn, at what altitude, and can it say any of them back? Two instruments answer. A coverage decoder on a controlled vocabulary: the model learns a stable, decodable code for 303 of 364 words at the first level, one at the middle, eighteen at the top, and the deep code collapses to a handful of distinct codes (with more data they fall, 57 to 37 to 8 before recovering to 22), with one of twelve broad classes decodable at the top. It learns words at the bottom and does not lift them into stable higher concepts. A playback dump on a trained reader: it recognizes 296 words and produces 64 in free generation, a fifth of what it recognizes, word validity about one percent, and the learned codes play back as generic frequent syllable stubs ("pr" completes to "proa", "dr" to "drea"), 2 of 296 to the exact spelling. Both reads land on the same wall: a strong recognizer, a weak producer. Recognition (a stable code at a word's last character) is cheap; production (regenerating the exact sequence) is dear, which fits a stack whose top-down wire carries credit and priming, not generation. The playback dump is now a standing instrument for reading out what concepts the model holds, and the round measures the starting line the generation work has to move. A wall, not a win. Read.

A bank of critics

2026-06-27 · partial, with a clean negative and a calibration finding. Minsky says the mind is a bank of critics and a selector that picks a way to think; the deliberate-pass result built one of each. This generalizes it to a bank of count-native critics over the locked model, with a selector that fires a deliberate action when a critic flags trouble. Two things came back clearly. Firing a deliberate pass on every character hurts: bits-per-char rises 0.07, the apex collapses to rank one, stability fails, re-applying every update over-consolidates, which is exactly the case for a selector. And at the character level only the surprise critic is selective: the conflict critic (top two guesses close) fires on a near-constant quarter of characters and low confidence on almost none, so the conflict signal that worked between words is noise between characters, and a critic bank is meaningful only with calibrated triggers that fire on the rare tail. The headline comparison, does a selective bank beat every fixed policy, is open: the consolidate and critic arms ran out of GPU under contention, though a smoke run confirms the selector routes correctly and is neutral before the drift. Partial, honest, with the drift-scale verdict owed. Read.

The empty cell

2026-06-27 · a partial answer, the synthesis alive but unstable. The architecture review found one box the line never filled: per-unit credit on a sparse (k-WTA) code, where the literature said the collapse should dissolve. Built as the minimal edit (replace the dense normalization with k-WTA plus the sparse-distributed-memory hypersphere support, mask the credit to the winners, hold everything else), it gives a real, partial answer. The sparse code does abstract: it climbs to the backprop ceiling (peak 0.469 averaged, 0.53 best seed), beating the no-credit control by +0.154. But it does not stay there: it is bistable, peaking then collapsing back to ~0.31, the deep code ending with only about four distinct winner-sets across four hundred stimuli. The collapse is dual-cause: k-WTA bounds how many units fire and fixes the substrate, but the dense credit pathway still drives the same few to win, so the rank-one collapse reappears as a winner-set collapse. k-WTA alone is not enough; no pivot, the locked dense stack stays the reference, and the fix is more biological machinery aimed at the second cause (diversity-forcing homeostasis, positive weights without bias, sparse credit). Read.

Re-reading

2026-06-27 · a free prediction lever. The scale run found the locked architecture stability-limited on fresh data; this tests the other half, re-reading the same bounded text. Re-reading a 200K slice five times, no reset, fully within our online rule, lowers held-out bits-per-char 3.863 → 3.762 with no overfit (the train-to-held-out gap stays flat), and leaves abstraction untouched (transfer CCGP flat above the ceiling). The surprise is the head-to-head: 200K read five times beats a million fresh characters on held-out prediction (3.762 vs 3.987) and ties on abstraction, because the fresh single pass drifts at depth three (held-out bits-per-char rises, mild forgetting of the fixed early slice) while re-reading a bounded corpus stays in a stable basin. So re-exposure's win is partly stability, not pure exposure: cognitively faithful, spaced re-reading of a smaller text generalizes better than a single skim of a larger one. A free, rule-legal lever on prediction that does not touch the abstraction wall: exposure beats novelty for prediction, ties for abstraction. Read.

Fast, and at scale

2026-06-27 · stability-limited, not data-limited, confirmed. With the architecture locked, two stress tests: is it fast enough to run, and does it stay itself at scale? Both yes. A Metal-fast compiled build of the locked depth-three stack (the whole strict-online step fused into one GPU function, all state on-device, zero host syncs per step) runs at about 6,400 characters a second (roughly 3.5 to 10× the old rate) and is bit-identical to the reference. Run on a million characters it stays stable the whole way (no divergence, dimensionality never collapses, no NaN, the long-run stability test feedback alignment owed, passed) with bits-per-char still dropping (3.90 → 3.63) while abstraction holds above the 0.484 ceiling but does not climb. That confirms the architecture line's diagnosis directly, from the data side: more data buys better prediction, not more abstraction. The locked config is stability-limited, not data-limited, and now confirmed fast and stable at scale. An engineering-and-scale confirmation, not a new mechanism. Read.

What changed: generation, opened

Theory update. 2026-06-27 · the first positive on a new frontier. With abstraction locked above the ceiling, the open value moved to generation, and SF1 lands the first positive there, on a different axis from the whole abstraction line. The birdsong loop: an agent babbles candidate utterances over the chunk-lexicon emission vocabulary, re-comprehends its own output to recover what it would understand, and corrects toward the meaning it intended, with no listener. It works. The full loop recovers a meaning 0.271 of the time against a chance of 0.042 and controls near chance: about 6.5× chance, 8× its own open-loop and deafened ablations (which sit at chance with zero coverage). All three guards fire: scramble the intended meaning and it collapses to exactly chance (0.042); let the self-output contaminate comprehension and an independent judge catches the private-code drift (recovery 0.122 yet the lowest self-error of any arm, it fools its own ear, buys no real production). The meaning code is a salience-weighted "heard fingerprint," so the producer can't string-match; it must produce something that sounds like the meaning. Honest caveats kept: absolute recovery is modest (~0.27 on 24 meanings, a first cut on a coarse code and a count re-hearer), and by design the loop supplies fluency not a shared convention, that needs the external referential game. Internal self-feedback teaches production with no listener; the generation frontier turns from all-open to half-answered. Read.

What changed: the architecture is locked

Theory update. 2026-06-27 · stability solved, the levers factored, the architecture locked. The sweep round that closes the abstraction line. Four parameter sweeps map the wall the scaling step named, stability, and the verdict is decisive. Stability is solved: forward-activation normalization holds the stack stable 19/19 at every depth, rate, and data size where the baseline, a gate-clip alone, and a dream-replay cycle alone all diverge (an honest standalone-replay negative). The four levers factor cleanly and are complementary: stability from normalization-plus-gate-clip, a high-dimensional deep code from breadth and dropout, abstraction from per-unit precision-gated apical credit at a gain of about 0.2, prediction from breadth at depth two and diverse views. The winner: normalization-plus-gate-clip, gain 0.2, depth three, transfer CCGP 0.50 to 0.55, above the 0.484 backprop ceiling, stable, seed-robust. And the capstone that was supposed to push it deeper, combining the stability and breadth levers, ran and failed: it stays stable but collapses to dimensionality 1.17, because breadth preserves variance across columns while normalization destroys it across stimuli within each column, and concatenation can't undo a per-lane collapse. The levers are not additive. The architecture is locked at the depth-three winner above the ceiling; pushing deeper needs a design change of uncertain payoff. With abstraction settled, the program's open value moves to generation. Read.

Three swings after the keystone

2026-06-27 · two honest negatives and a scaling verdict. The keystone (DE) named three next steps and we ran all three. DF: a self-supervised label attractor (Barrett's stabilizing word) compresses dimensionality harder than anything in the line (PR 5.8, below the Hebbian floor) but scores below the raw baseline on transfer (0.321 vs 0.339): compression is not abstraction, the gradient-free shortcut ruled out. DG: different attended views make voting earn its keep (Monty's diversity requirement, confirmed: same-view voting at 5.1 bpc is worse than a single column) and win prediction when data is thin, but leave abstraction flat (~0.49) and raise dimensionality, so the architecture factors: prediction from diverse precision-weighted views, abstraction from per-unit apical credit. DH: scaling the fusion on the MLX substrate buys one extra altitude (a depth-three rising code, apex 0.465 ≈ the 0.484 ceiling) then saturates; more data only raises divergence risk; and the precision gate flips from hero to chief destabilizer at depth. Together they finish the diagnosis and name the next wall, stability. Read.

What changed: abstraction that climbs

Theory update. 2026-06-27 · the architectural keystone. The constructive answer to the abstraction wall. DC found the missing ingredient (credit assignment), DD made it local (per-unit feedback alignment), and DE rides that per-unit credit down a real top-down apical pathway (the same wire that primes the level below) gated by precision for when it fires. On a two-level stack the full fusion (apical per-unit credit + precision gate) reaches a transfer abstraction score of 0.501, beating single-layer feedback alignment (0.408) and the raw baseline (0.339), approaching the backprop ceiling (0.537). And it is the first stack in the program whose abstraction score rises with depth (+0.178). The priming pathway and the credit pathway turn out to be one mechanism, one precision read three ways. Honest caveats kept: the win is the full fusion not apical-alone (ungated apical sits below the baseline), it needed a stabilized learning rate, HEBB-2L's apparent rise is a degenerate-code artifact, and it is still one architecture. The frontier reshapes: not whether a biologically-shaped local architecture can abstract (DE shows it can) but scale, multi-attention, and stability. Read.

A local signal clears the wall

2026-06-27 · the second positive. The constructive step before the keystone. On DC's harness unchanged, swap in two new hidden-layer rules and a transfer probe. A broadcast neuromodulatory scalar (three-factor) fails: it lands on the no-gradient floor, because one scalar everywhere can only scale an Oja step, not tell each unit how to change (Lindsay 2017). A per-unit local signal (feedback alignment, fixed random feedback, no weight transport, biologically plausible) drops bits-per-char 4.05→3.50 and lifts the abstraction score to ~0.41 within and on held-out transfer, approaching the global-backprop oracle's ~0.49. Transfer tracks within for every arm. So abstraction is reachable inside online + local + bounded: DC's "credit assignment" sharpens to per-unit credit assignment, and a shippable learner comes into reach. Read.

What changed: the abstraction wall is credit assignment

Theory update. 2026-06-27 · the first positive on the wall. Six gradient-free mechanisms each failed to build an abstract space, all pointing at one suspect, the gradient. This round isolates it: one fixed architecture, same data, same online single-pass regime, same probe, swap only the update rule. Online backprop builds the abstract (higher-CCGP) space; the no-gradient Oja arm compresses dimensionality hard (PR 100→14) but loses abstraction: the exact six-negatives signature, across three seeds. The abstraction wall was credit assignment, not topology. With the honest caveats kept (the gradient arm is the global oracle, one architecture, one probe) and the neuroscience that agrees: Numenta dropped HTM for transformers on language, Lindsay 2017 shows Hebbian gives mixed selectivity but not the abstract code, Wutz 2018 shows abstraction is built top-down. The open question flips: not "is a gradient needed?" but "can a local gradient keep the win inside the online + local + bounded regime?" Read.

Slower, and from above

2026-06-27 · two negatives that sharpen. Two more swings at the abstraction wall. Temporal pooling builds stable, rising-timescale concepts (persistence 7.9 → 71 → 640 chars, upper levels learn to predict the concept stream) but they recur no more than on shuffled text and abstraction drops. A learned top-down Hebbian signal settles the hierarchy (active-set change 0.62 → 0.004) and disambiguates an uncertain input, both real, and it does not lift abstraction either. The fifth and sixth convergent negatives, and the one that matters: a non-local signal was the named escape, and a Hebbian one does not clear the wall. Read.

Pure biology, no counting

2026-06-26 · honest mixed verdict. Drop counting entirely: a random sparse encoder, an HTM temporal memory, Monty cortical voting, all Hebbian, no backprop. Does it trace the prefrontal-geometry arc, dimensionality high then low, then abstract? It traces the back half. Anomaly collapses (it learns sequence with no counting), dimensionality compresses (PR 149→80), and voting sharpens the late drop (N=3 drops where N=1 rises). But CCGP/abstraction drops, not rises (0.65→0.56), and more data plus selective voting did not change it. Sparser and slower is not abstract. Bit-exact MLX/Metal throughout. Read.

Do concepts emerge as boundaries?

2026-06-26 · investigation closed at 10 variants. The purest form of the theory: do concepts emerge as the higher tiers of surprise boundaries? Ten experiments on the node runtime map the altitudes: WORD = branching entropy SOLID (F1 0.83 to 0.98); PHRASE = belief-shift WEAK; CLAUSE/SENTENCE = construction-completion VERY WEAK and fragile; TOPIC = lexical cohesion ~2×. And three meta-negatives rule out every cheap fix: the wins don't compose (CS), scale isn't the unlock (CV), a shared representation isn't the unlock (CW). Surprise carves only words; meaning above the word needs a genuinely different signal, not better wiring, data, or representation. The quiet win: all ten organs ran on the runtime with zero engine changes. Read.

From mimicry to production

The production library. 2026-06-26 · research. A reader predicts the next token to match the stream; a speaker emits one to move a listener. Reading does not become that for free. The distilled science and the count-native organs that add the steering (which utterance to repeat, and for whom) built on the one threshold already crossed: a contingent reply teaches more than the same words overheard cold. The audience model, scored on whether a listener recovered the referent, is the first top-down state with a licence to change the 99% slice. Read.

What changed: the cortex speaks, rough

Theory update. 2026-06-26 · synthesis. Production is comprehension read the hard way, and it now has first evidence: a chunk lexicon hands the cortex whole units, a coverage-competition producer says them more well-formed than gibberish and keeps its frame 87.5% of the time. A contingent reply teaches +0.45 bpc more than the same words overheard cold, and the win survives a live model: run against Haiku through the logged-in CLI with no API key, contingency beats the yoked control 12 settings out of 12. And the no-curriculum rule generalizes: no write-throttle, no structural curriculum. Read.

The acquisition queue, built

2026-06-26 · 7 wins, 6 partials, 3 clean negatives. We distilled a cognitive-science library of how children acquire language into a build queue, then built and ran the whole thing in one fan-out. The chunk lexicon is the lever: it wins the splice test and hands generation whole units, though it does not help prediction. The generation turn is real but rough. Three negatives each sharpen a standing rule. Read.

What changed: the budget bites, binding works, reasoning gets a road

Theory update. 2026-06-26 · synthesis. The round in one place. The bounded-memory rule stops being a prediction and becomes a measured fact (consolidation flips −0.006 → +0.307 bpc under a tightening cap). Long-distance binding, the one thing offset-attention couldn't do, is in hand (cue retrieval, 99.96% across a clause). Compositional reasoning gets a concrete road: redescription → VSA-decode → workspace, supply structure, don't factor blindly. And the coherence frontier hardens: three persistent-state mechanisms all land on the same 1% slice, so whatever crosses it must change the 99%. Read.

Reaching back by the right cue

2026-06-26 · win. Content-addressable cue retrieval binds a verb to its correct-number subject across a clause: 99.96% where offset-attention's fixed position key tops out at 65% and scores flat zero past its one modal offset. It reaches by a feature bundle weighted by fan, not by position. And because activation divides by the fan, it reproduces the human agreement-attraction illusion (0.00% → 2.11%, only with a recent opposite-number distractor). A new binding primitive. Read.

What survives a budget

Theory update. 2026-06-26 · the rule confirmed. The bounded-memory rule's prediction is now measured: mechanisms that vanish at unbounded scale return under a memory budget. Consolidation flips −0.006 (unbounded) → +0.144 → +0.307 bpc as the cap tightens, a clean curve that vanishes when the budget stops binding. Lossless generalization (consolidation) wins broadly; lossy (concepts) only as tight as the budget forces; the topic prior stays neutral. Read.

Reading structure back out of a sum

2026-06-26 · split. HRR role-filler decode with known roles recovers subject, verb, and object at 100%, robust to eight bound pairs, even at the smallest dimension. But the blind resonator, factoring the same sum with unknown roles, fails at affordable dimension (≈0% over a 4000-word codebook at D ≤ 2048). The payoff: VSA decode works given structure, which redescription supplies and the workspace manipulates. Don't factor blindly. Read.

Giving the workspace something to do

2026-06-26 · partial. The serial workspace reaches a two-hop target (acc 1.00) where System 1 and one-step deferral both score 0.00, trapped on the intermediate: reachability is real and new. But on a single deterministic chain it only ties a blind "apply-twice"; its focus and inhibition are for selecting among competing chains, which this probe lacks. The next axis is named: competing candidate chains. Read.

Keeping track of what's going on

2026-06-26 · negative. A persistent situation model (Chambers-Jurafsky event chains plus Zwaan who/where/topic slots) does not predict over long spans. The +0.55 bpw "everywhere" win was pure smoothing repair: a static frozen unigram beats the live situation by −0.07 bpw on the 99% non-backoff slice. The situation helps only the same 0.9% backoff slice two earlier mechanisms found. Measure a top-down prior against a static prior, never against no prior. Read.

Where forgetting's shape finally pays

2026-06-26 · win. Power-law (ACT-R) eviction beats LFU at the word level under non-stationarity and a tight budget (−0.008, −0.006 bpw at caps 10k, 30k; LFU re-wins once the cap is loose). The sign flipped from the char-gram result, exactly as the shape-of-forgetting post predicted. It wins by serving the present, not protecting the past. Read.

Writing it down

2026-06-26 · negative, with a principle. At equal memory budget, a bounded-internal + external store does not beat one bigger internal table (evidence fragmentation). It wins only when the page is cheaper than the skull (−0.23 bpc in the cost-asymmetric regime). Externalizing is a cost arbitrage, not a better use of the same bytes. Read.

Phrases that rhyme share their counts

2026-06-26 · negative. Permutation-bound FlyHash phrase addresses pool similar phrases (beats a floored literal on 67% of unseen-in-form probes) and keep order (×2.29 under scramble vs the bag's ×1.00) but FlyHash crosstalk loses the tail (aggregate ppl 1206 vs literal 831; poor exact recall 423 vs 2.2). The fix: use it as a backoff layer under the literal table. Read.

What changed: we have a System 2

Theory update. 2026-06-26 · synthesis. The round in one place. The architecture gained its second half: a metacognitive gate that calls a still-minimal System 2 over explicit, redescribed concepts. New: a gate that overrides System 1 only when it is wrong. New: redescription that turns counts into manipulable concepts. Changed: the combiner is take-the-best, not full pooling. Refined: LFU for char-grams, and no curriculum. The frontier sharpens to the multi-step task. Read.

Thinking slow, by counting

Theory update. 2026-06-26 · win, with an honest negative. A count-native System 2. A dual gate (calibrated confidence plus Botvinick conflict) deploys a deliberate pass that overrides System 1 only when it is wrong: +0.38 accuracy on conflict cases (where the reflex is wrong 88%), zero harm on no-conflict cases, and at zero budget it falls back to System 1 bit for bit. The gate is the win; the elaborate serial workspace loses to a trivial "defer to the wider context" here, and is parked for the multi-step task. Read.

When a habit becomes a thought

Theory update. 2026-06-26 · qualified win. Stability, not error, promotes a mastered construction into an explicit, slot-addressable concept that answers queries the flat count cannot: inverted slot lookup, role substitution, slot analogy. The Karmiloff-Smith U-shaped dip confirms it (0.155 → trough 0.051 → recovered 0.181, above baseline). These explicit concepts are the operands the System-2 workspace deliberates over. Read.

Less is more, and you can prove it

Theory update. 2026-06-26 · clear win. Validity-ordered, noncompensatory, early-stopping take-the-best beats full geometric-mean integration on every axis at once: accuracy 15.00% vs 9.71%, perplexity 1,918 vs 7,160, at 4.56 cues a step instead of 8. Less-is-more holds: ignoring the weak channel wins, sharpest on sparse contexts. This revises the standing combiner: sharpen by ignoring, not by pooling. Read.

Starting small, on purpose

Theory update. 2026-06-26 · honest negative. Growing the memory budget on a schedule does not beat full-from-start (2.744 vs 2.751, robust across seeds); only a permanently small budget loses (+30%). "Starting small" was a property of the gradient optimizer, which freezes early guesses: a count learner cannot get stuck, so it needs no curriculum, only enough final memory. The ZPD overlay hurt. Read.

The shape of forgetting

2026-06-26 · honest negative. ACT-R's power law is the right shape of forgetting, the only curve that represents spacing (spaced 8.96× more accessible than massed), but for dense char-grams raw-count LFU wins eviction at every cap, because LFU is the power law's d→0 limit and a char-gram's value is pure frequency. Right shape, wrong place: keep LFU for char-grams, reach for the power law at the word and concept level. Read.

What survives scale

Theory update. 2026-06-26 · scaling study. We re-ran every idea at 30 to 200× the data, half a billion words and three billion characters. The verdicts split clean: mechanisms that compete with local counts on already-seen prediction vanish (topic prior +0.34 bits/word → 0.0), and mechanisms that do what counting can't hold or grow (Bayesian-surprise boundaries F1 0.154 → 0.447). Read.

Grammar is just counting, made productive

2026-06-26 · qualified win. Count two things per frame, token and type, and a flat n-gram turns compositional. On held-out, never-seen frame-filler pairs the open-slot construction beats the n-gram 4.3 times on perplexity. It froze idioms and abstracted a NUMBER+UNIT slot, with no grammar given. Read.

Learning the new without losing the old

Theory update. 2026-06-26 · qualified win. Stream Darwin, then Shakespeare, then the Bible, one pass, bounded memory. The brain-inspired model forgets about 21 times less than a recency cache and predicts better at the peak. The load-bearing piece was ART resonance: protect the specific context that recognized the input. Read.

Is the analogy already in the counts?

2026-06-26 · qualified win. Raw PPMI counts solve a:b::c:? at 56% top-1, 94% top-5, four times baseline, no SVD, no word2vec. Two honest negatives: leader-clustering blurs the relation axes it cannot substitute for SVD, and NARS induction spreads its mass too broadly to beat a direct counter. Read.

What the model thinks is happening

2026-06-26 · qualified win. Bayesian surprise, the KL of the one-step belief update, beats per-token surprisal 5.7 times and branching-entropy 120 times at finding real article boundaries, and the lead grows with scale. The event-slot prior it drives helps prediction only on the 1% backoff slice. Read.

How sure is a count?

2026-06-26 · qualified win. Split a count into hits and misses and the NARS truth value calibrates it for free, ECE 0.280 to 0.027. As an expert weight it also cuts perplexity threefold. The knob-free gate is too cautious for clean text, and the only policy that beats always-open on the rare slice. Read.

What an agent learns while it dreams

Theory update. 2026-06-26 · qualified win. One offline sleep pass over the count memory, pure replay and bookkeeping, cuts memory 37% and improves prediction on the rare tail, with no new data. Keep dreaming and it breaks exactly as Letta warned: the memory goes generic and lossy, common contexts ruined to flatter the tail. Read.

Use the map to read, not to walk

Theory update. 2026-06-26 · qualified win. Proximity failed twice as a predictor. So we used it as a map and poured it into the counter instead. On the fifth of predictions where the counter has never seen the pair, the similarity cluster lends its neighbors' counts and cuts perplexity twenty-fold. It prices the tail without changing the top guess. Read.

When the letters lie, it leans on the idea

2026-06-26 · qualified result. Pour noise on the input only. The flat model collapses; the concept stack degrades 2.7 times slower, and the gate hands prediction up to the topic level, concept mass 86% to 95%, with no signal that the input is noisy. Read.

One brain part, or many? We gave each level its own job

2026-06-26 · negative result. A specialized stack, a different job per level, loses to the uniform Column on bits-per-char. The lone clean win is the gate: dynamic routing beats static pooling by 0.9 bits. The combiner is load-bearing. Read.

We gave the map its best shot

2026-06-26 · negative result. The fair rematch for proximity, the graph form, inside the best stack, on the rare-context slice. Accumulated evidence earns its keep there. Proximity still has no niche, parked deeper with a reason. Read.

Predicting the kind, not the word

2026-06-26 · qualified result. Hide a word and predict its class instead of its spelling. The class is predictable for rare words where the exact word is hopeless. Counted, it cannot collapse. Read.

When the whole room agrees on a topic

2026-06-26 · qualified result. A higher level commits to one topic and broadcasts it down. It helps exactly where the local context has run dry, and hurts where it has not. Read.

You can't write your signature backwards

2026-06-26 · clear win. A memory of change rather than content transfers to words it has never seen, runs in one direction only, and lets a cue prime what comes next. Read.

Attention, but counted instead of trained

2026-06-26 · clear win. Attention with no queries, keys, values, or gradients. Just counts keyed by position. It beats fixed n-grams on calibration by three times and proves it is not a bag of words. Read.

A vote that remembers what it just saw

2026-06-26 · qualified result. A vote that accumulates over time instead of starting fresh. It shrugs off noise far better, and fails honestly as a boundary detector. Read.

Meaning is a map, not a road

2026-06-26 · negative result. We placed words in a space where nearness means similar meaning. The space is real and beautiful. It still does not predict the next word. A clarifying loss. Read.

More data helps, all the way to a gigabyte

2026-06-26 · win. Prediction keeps improving with data up to a full gigabyte, and a custom GPU path makes that gigabyte learn in seconds. Read.

Finding phrases the way you'd guess them

2026-06-26 · mixed. The same surprise signal that finds word boundaries, one level up, discovers real phrases on its own. Topic boundaries are harder. Read.

The level that reaches past the last few words

2026-06-25 · clarifying. Stacking more fixed local levels stops paying. The lesson points at what kind of level is worth building next. Read.

One part, repeated, wired bigger

2026-06-25 · win. The whole system collapses to one part, a Column, repeated wide and stacked deep. Wire it bigger and it gets better, and you can say exactly why. Read.

Building the dials that bits-per-char hides

Theory update. 2026-06-25 · win. One number was hiding the truth. We built a scorecard that measures generalization, real-word rate, and phrase coherence, and it told a sharper story. Read.

How big a brain the data wants

2026-06-25 · win. A small model saturates early. A bigger one keeps learning, but only once it has enough data. The right size grows with the corpus. Read.

The hierarchy pays off at the right altitude

2026-06-25 · win. Measured at the word level instead of the character level, the concept hierarchy halves perplexity. Read.

When combining the experts made it worse

2026-06-25 · negative result. We combined every level of expert into one vote. It scored worse than a simple two-expert mix. The reason taught us which combiner is right. Read.

Words that lower the cost of letters

2026-06-25 · win. Teaching the model words, with no labels, makes it better at predicting letters by a fifth. The discovered words are real English, sitting in a document you can read. Read.

The counter beat the neural net

Theory update. 2026-06-25 · win. A plain online counter beat two gradient-trained networks at predicting text. It learns online, never forgets, and needs no backprop. It became the substrate. Read.

Finding where one word ends

2026-06-25 · qualified win. The first experiment. Prediction error alone recovers word boundaries from raw characters, at a quality the literature respects. One fashionable signal scored below random. Read.