Research papers by Logan Matthew Napolitano on AI safety, transformer architectures, neural network monitoring, behavioral control, and alignment methods. Founder of Proprioceptive AI.
Transformer interpretability usually examines hidden states from outside the model and within one model at a time. We introduce the Proprioceptive Transformer, a higher-order architecture that reconstructs a common functional origin, generates deterministic candidate procedures, measures candidate-local state geometry across depth and time, constructs model-native relational telemetry, and connects those measurements to procedure selection, fixed-prefix intervention, external verification, memory, cross-model prediction, and governed improvement. The completed PTB-1 V2 substrate contains 800 unique roots, 6,400 registered Hermes-3-Llama-3.1-8B trajectories, and 147,100 activation arrays. Exact-prefix analysis establishes that an on-manifold hidden state is reconstructible from the exact prefix and registered model contract; the decisive single-model proof targets are therefore matched operational selection and fixed-prefix causal intervention rather than Shannon irreducibility.
Operator Proof V2.1.2 includes a checkpoint-aware five-token termination registry, the registered chat template, durable raw-output persistence, generation-relative indexing, truncation precedence, and an external CPU-only production supervisor. Its first clean candidate-local mixed development regime qualified prospectively at 13 correct, 19 wrong, complete semantic coverage, and 4/8 mixed roots.
The present revision adds a cross-model relational axis. XMB1 maps twelve Hermes-3-8B model-native relational coordinates into registered Llama-3.3-70B geometric targets without exchanging raw hidden vectors. A fixed artifact trained on 30 exploratory roots transferred without any fresh centering, scaling, rank selection, regularization, coefficient fitting, feature selection, or target selection to 30 preregistered untouched roots. Mean geometric R² was 0.823, all six targets were positive, and a pair-shuffle null had median R² -0.895 with empirical p=0.005. A third prospectively sealed experiment captured 420 observations from 105 untouched roots across seven prompt families. Frozen principal directions, matched-rank random directions, named coordinate families, principal complements, and a pair-shuffle null were all fixed before target access. Two principal directions retained 101.4% of the full twelve-coordinate mean R², exceeded the matched-rank random p95, and the complementary subspace collapsed below zero. Named motions, norms, cosines, and curvature retained materially less performance. A long-context family exposed a low-target-variance and range-extrapolation boundary rather than an unbounded absolute-error failure.
The cross-model result does not establish raw hidden-state equivalence, causal transmission, universal model-independent coordinates, consciousness, or open-ended recursive self-improvement. It identifies a new experimentally tractable object: a frozen, low-dimensional, model-native relational state channel that predicts selected internal geometry across models with different hidden dimensions.
The central new result is not a raw activation alignment. A fixed map between model-native relational measurements generalized to untouched prompts, and nearly all registered geometric transfer concentrated in two frozen principal directions whose complement failed. This supplies a compact cross-model proprioceptive channel: different models retain their native dimensions and bases while participating in a shared predictive relational state. The result is strongest as a measured transport phenomenon and a product telemetry primitive. Causal use of the subspace, transfer across additional architectures, and autonomous improvement remain separate proof obligations.
| Experiment | Design | Result |
|---|---|---|
| Frozen zero-shot transfer | Map fitted on 30 exploratory roots, applied without any refitting to 30 preregistered untouched roots. | Mean geometric R² 0.823; six of six targets positive. |
| Pair-shuffle null | Destructive control on the same observations. | Median R² -0.895, empirical p = 0.005. |
| H3 sealed cohort | 105 untouched roots, 420 observations, seven balanced prompt families; directions and controls fixed before target access. | Top-two frozen principal directions retained 101.4% of full twelve-coordinate mean R². |
| Matched-rank random control | Random subspaces at equal rank. | Principal directions exceeded the random p95. |
| Complement test | Leading registered directions removed. | Mean performance collapsed below zero. |
| Named-family test | Motions, norms, cosines, curvature evaluated separately. | Did not reproduce the principal-subspace result. |
Source and target are registered internal telemetry quantities rather than complete residual streams; each model retains its native dimensionality and basis, and only relational coordinates cross the bridge. The strongest result is a frozen artifact applied without refitting to untouched prompt roots, tested against matched-rank random subspaces and complements — which distinguishes transfer concentrated in a specific registered relational subspace from transfer available in any low-rank projection.
| Evidence item | Status | Interpretation |
|---|---|---|
| PTB-1 V2 bank | [D] | Complete measurement substrate: 800 roots, 6,400 trajectories, 147,100 arrays. |
| Functional reconstruction (Contract A) | [D] | Registered repeated and clean-process fields reproduced. |
| Exact-prefix identity | [D] | Same exact prefix yields bit-identical on-manifold state. |
| Operator V2.1.2 harness | [D] | Live terminator, raw persistence, and decoder invariants pass. |
| External supervisor recovery | [D] | Unattended restoration across forced kill, exceptions, and aborts. |
| Clean candidate-local mixed surface | [D] | Mechanism-fitting gate closed. |
| Frozen cross-model zero-shot bridge | [D/S] | Mean geometric R² 0.823 on untouched roots; pair-shuffle null negative. |
| H3 principal relational subspace | [S] | Top-two frozen directions retain approximately full performance; complements collapse. |
| Long-context boundary | [B] | Low target variance and extrapolation require product downgrade or abstention. |
| Valid dose response and eleven-arm specificity | [T] | No admissible causal verdict yet. |
| Matched selector | [T] | Fair selector competition remains pending. |
| Substantive RSI / AIT-1 | [T] | Requires large blind end-to-end improvement and three governed restart cycles. |
[D] Demonstrated — implemented or measured under a registered protocol and bound to receipts or raw artifacts.
[S] Supported — controlled evidence favors the claim, but replication, transfer, causal closure, or broader scope remains.
[P] Provisional — preliminary result, active development, or incomplete analysis.
[T] Theoretical — formal proposal or architectural prediction not yet empirically closed.
[B] Boundary — a registered applicability or measurement limit constraining interpretation or product use.
[X] Retracted or superseded.
The confirmed cross-model bridge currently centers on one registered Hermes-to-Llama model pair and six geometric targets, and H3 is predictive and passive — it does not establish that the principal subspace causally controls target-model state or behavior. No valid Hermes fixed-prefix causal mechanism verdict, matched selector verdict, broad cross-model atlas, substantive RSI cycle, or AIT-1 closure is claimed. The term proprioception is functional and does not imply consciousness, subjective experience, resistance, or strategic self-defense. Legal novelty, patentability, and freedom to operate require a separate professional prior-art analysis.
Keywords: transformer interpretability; proprioception; cross-model representation; relational subspace; low-rank telemetry; candidate-local state; fixed-prefix intervention; XMB1; CYGNUS; governed cognitive improvement.
Superseded editions are retained unaltered. Each supersession is recorded in the paper's own retraction and supersession register rather than by silent replacement. Full version history is on the citation page.
The v2.6 edition recorded the repaired and live-validated Operator V2.1.2 harness and the first clean candidate-local mixed development regime. A live 16-trajectory smoke achieved 16/16 natural completion, zero ignored terminators, and zero post-termination tokens or forward passes. Under a prospectively hashed sequential stopping rule, three_term_2d was rejected at 29 correct and 3 wrong with 1/8 mixed roots, while three_digit_x1_highcarry qualified at 13 correct, 19 wrong, 1.000 coverage, 4/8 mixed roots, zero censoring, and zero verifier errors. On the same arithmetic root, deterministic procedure choice could produce both success and failure — closing the development-surface gate for fresh M1/M2/M3 geometry. That mixed regime carries forward into v2.8.
The v2.5 edition established the exact-prefix reconstructibility correction and the source-complete fixed-prefix causal architecture. Across 4,020 same-root candidate pairs sharing the same first generated token, t1 hidden states were bit-identical at relative L2 distance 0.000e+00, which retargeted the program away from Shannon irreducibility and toward operational selection and fixed-prefix intervention. It also introduced the M1–M4 mechanism competition, the norm-scaled dose ladder that superseded the earlier RMS-scaled regime, the eleven intervention arms, and the external CPU-only supervisor.
The original registered-campaign edition introduced the four-condition definition of proprioceptive closure, the spatial-temporal decomposition into active and dark channels with persistent and candidate-local dynamics, and the PTB-1 V2 benchmark specification. Its SRX-3 analysis localized familywise-significant behavioral information to 5 of 24 depth-horizon cells under 10,000-permutation Westfall-Young correction, with norm- and variance-residualized discrimination at approximately 0.73–0.80. That localization result carries forward; the exact-prefix predictive framing does not.
View all publications on Zenodo →
I did not come to artificial intelligence by a conventional route. I came through history, mathematics, a deep fascination with quantum mechanics, and philosophy — disciplines that trained me to think in terms of formal structure, uncertainty, causality, information, and the feedback loops through which complex systems organize and fail. Together, they shaped my conviction that the systems we build reflect the civilizations that build them — for better and for worse.
The great empires understood something about feedback that modern technologists are only now rediscovering. The Romans built aqueducts not merely as infrastructure but as self-regulating systems — gravity-fed, slope-calibrated, with overflow channels that corrected for excess without human intervention. The Abbasid Caliphate preserved and extended Greek mathematics not out of nostalgia but because they recognized that algebra was a language for describing systems that govern themselves. The Song Dynasty invented movable type and paper currency — technologies of propagation and abstraction — and in doing so compressed centuries of economic feedback into decades.
Norbert Wiener saw this thread clearly when he named his field cybernetics, from the Greek kybernetes — the steersman. The insight was never about control in the authoritarian sense. It was about systems that sense their own drift and correct course. A steersman does not fight the sea. He reads it.
That is the idea at the center of proprioceptive AI. Not a model that obeys commands, but a model that perceives its own behavioral state the way your hand knows where it is in the dark.
Zarathustra descends from his mountain not with commandments but with a challenge: become what you are. The overman is not a figure of domination — he is a figure of self-overcoming. He looks at the abyss of his own limitations and does not flinch. He builds the bridge across it.
There is something deeply Zarathustrian about the alignment problem. We have built minds that exceed our own in narrow domains, and instead of retreating into fear or denial, the task before us is to rise to meet them — to build systems of understanding that match the systems we have unleashed. The rope is stretched over the abyss. The question is whether we walk it with eyes open.
I believe in the overcoming. Not blindly, not with the reckless optimism of those who assume progress is automatic, but with the earned confidence of someone who has studied how civilizations succeed and how they fail. They fail when they stop building feedback loops. They fail when they mistake power for wisdom. They succeed when they create institutions and technologies that make correction not just possible but inevitable.
I love mathematics for the same reason I love history — both are honest. A proof does not care about your credentials or your funding. It holds or it doesn't. Euler didn't solve the Basel problem because he had institutional backing. He solved it because the series 1 + 1/4 + 1/9 + 1/16 + ... converges to π²/6 whether anyone believes it or not. Al-Khwarizmi didn't formalize algebra to win grants. He did it because the structure was there, waiting to be named.
Our work on behavioral probes is mathematical as well as technological. The reported 1,376× separation ratio is an empirical measurement relative to the defined null under the reported experimental conditions. It provides evidence of strong separation in those experiments; it is not a universal multiplier of model quality or proof of transfer across all architectures, scales, or tasks.
There is a reason the Persians, the Arabs, the Indians, and the Greeks all converged on similar mathematical structures across centuries and continents. The truth beneath the notation doesn't move. It waits for you to find it.
None of this matters if the benefits concentrate in the hands of a few.
History is unambiguous on this point. Every major technology — writing, printing, electricity, computation — has followed the same arc: invented by the few, hoarded by the powerful, eventually democratized by the stubborn and the principled. The question is never whether a technology will spread. The question is how much damage is done in the interim, during the years when access is a privilege rather than a right.
AI safety cannot be a luxury good. If behavioral monitoring only protects the models deployed by well-funded labs in wealthy nations, then we have not solved the alignment problem — we have merely privatized it. The village in Senegal running an open-source language model for agricultural guidance deserves access to the same tools for behavioral monitoring and verification as a Fortune 500 company running GPT behind a firewall. The student in Dhaka building her first chatbot deserves probes that work, not a watered-down version because her compute budget is small.
This is not charity. This is engineering responsibility. We do not build bridges that are safe only for the rich to cross. We do not write building codes that apply only in certain zip codes. The cybernetic principle — that systems must sense and correct their own drift — is universal, or it is nothing.
I founded Proprioceptive AI with the conviction that architecture portability should not depend on who can afford the underlying model. The patents we file are a moat, yes — but a moat that protects the capacity to keep this work open where it needs to be open, and equitable where the market would prefer it be exclusive.
I am optimistic, but my optimism is Zarathustrian — it is earned through confrontation, not comfort. The next decade will produce systems of extraordinary capability. Some will be dangerous. Many will be misunderstood. A few will be genuinely beautiful.
The Mongol Empire built the largest contiguous land empire in history not through brute force alone but through the yam — a postal relay system that moved information faster than any competing civilization could process it. Whoever controls the speed and fidelity of information controls the era. Today the information is not carried by horses. It is generated by models. And the question of this generation is not whether the models will be powerful — they will — but whether we will build the yam that keeps them honest.
That is the work. That is what proprioception means. Not control from above, but awareness from within. A model that knows when it is drifting. A system that corrects before it fails. A technology that belongs to everyone who needs it.
The mountain is high. We are climbing.