[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"project-92514":3},{"id":4,"name":5,"fullName":6,"owner":7,"repo":5,"description":8,"homepage":9,"htmlUrl":9,"language":10,"languages":9,"totalLinesOfCode":9,"stars":11,"forks":12,"watchers":13,"openIssues":14,"contributorsCount":14,"subscribersCount":14,"size":14,"stars1d":14,"stars7d":14,"stars30d":15,"stars90d":14,"forks30d":14,"starsTrendScore":14,"compositeScore":16,"rankGlobal":9,"rankLanguage":9,"license":17,"archived":18,"fork":18,"defaultBranch":19,"hasWiki":20,"hasPages":20,"topics":21,"createdAt":9,"pushedAt":9,"updatedAt":29,"readmeContent":30,"aiSummary":31,"trendingCount":14,"starSnapshotCount":14,"syncStatus":32,"lastSyncTime":33,"discoverSource":34},92514,"Subtext","ninjahawk\u002FSubtext","ninjahawk","To know what models don't say out loud.",null,"HTML",185,19,1,0,58,49.7,"Other",false,"main",true,[22,23,24,25,26,27,28],"interpretability","llm","mechanistic-interpretability","pytorch","qwen","transformers","visualization","2026-07-22 04:02:06","\u003Cdiv align=\"center\">\n\n# Subtext\n\n*A real-time instrument for observing the verbal workspace of a language model\u003Cbr>as it reads, reasons, and speaks.*\n\n![Subtext demo](media\u002Fdemo.gif)\n\n[![Python](https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FPython-3.11+-3776AB?logo=python&logoColor=white)](https:\u002F\u002Fwww.python.org\u002F)\n[![PyTorch](https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FPyTorch-CUDA-EE4C2C?logo=pytorch&logoColor=white)](https:\u002F\u002Fpytorch.org\u002F)\n[![Model](https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FModel-Qwen3.5--4B-6cb8e0)](https:\u002F\u002Fhuggingface.co\u002FQwen\u002FQwen3.5-4B)\n[![Lens](https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FLens-qwen--n1000-e8a13c)](https:\u002F\u002Fhuggingface.co\u002Fneuronpedia\u002Fjacobian-lens)\n[![License](https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FLicense-Apache_2.0-green)](LICENSE)\n\n**[🌐 Watch a live replay in your browser](https:\u002F\u002Fninjahawk.github.io\u002FSubtext\u002F)** · **[▶ Demo video](media\u002Fdemo.mp4)** · **[📄 The paper](https:\u002F\u002Ftransformer-circuits.pub\u002F2026\u002Fworkspace\u002Findex.html)** · **[🔬 Reference implementation](https:\u002F\u002Fgithub.com\u002Fanthropics\u002Fjacobian-lens)**\n\n\u003C\u002Fdiv>\n\n---\n\n## Getting started & staying tuned with us.\n\nStar us, and you will receive all release notifications from GitHub without any delay!\n\n\u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fninjahawk\u002FSubtext\u002Fstargazers\">\n \u003Cpicture>\n   \u003C!-- Chart is regenerated daily by .github\u002Fworkflows\u002Fstar-history.yml -->\n   \u003Csource media=\"(prefers-color-scheme: dark)\" srcset=\"media\u002Fstar-history-dark.svg\" \u002F>\n   \u003Csource media=\"(prefers-color-scheme: light)\" srcset=\"media\u002Fstar-history.svg\" \u002F>\n   \u003Cimg alt=\"Star History Chart\" src=\"media\u002Fstar-history.svg\" \u002F>\n \u003C\u002Fpicture>\n\u003C\u002Fa>\n\n---\n\n## Overview\n\nRecent work from Anthropic identified a small set of internal representations in\nlanguage models — the *J-space* — that behaves like a global workspace: its\ncontents can be verbally reported by the model, deliberately modulated, and are\ncausally used for multi-step reasoning, while the surrounding majority of neural\nactivity remains inaccessible to report. The identification tool is the\nJacobian lens, which transports a residual-stream activation at any layer into\nthe final-layer basis and decodes it through the model's own unembedding,\nanswering: *which vocabulary words is this internal state disposed to produce,\nnow or later?*\n\nSubtext applies that method continuously during live conversation with a local\nmodel. On every token — both while the model ingests the user's message and\nwhile it generates its reply — the lens is read at nine depths and the result\nis rendered as it happens. The intermediate steps of the model's computation\nbecome directly watchable: verdicts form during reading, several tokens before\nany output; planned words hold at high strength while unrelated tokens are\nbeing emitted; two-hop questions surface their unspoken middle term.\n\nSubtext differs from the interactive readouts already available (e.g. the\nNeuronpedia demo) in that it is conversational and continuous: it renders the\nlens during a live chat, includes the reading phase over the user's message,\nstreams at generation speed via a KV cache, and pairs the canvas with a\nper-token ledger and per-word inspector. Sessions can be exported and replayed\nin any browser without a GPU.\n\n## What the lens shows that the output does not\n\nThe value of the instrument is the gap between the model's internal state and\nits visible text. Three moments from the demo session:\n\n**1. The verdict precedes the reply.** Zero tokens of output exist; the model\nis still reading `Is this correct? 12 + 5 = 1`. The workspace already holds\n*math*, *addition*, *arithmetic*, *modulo* — the phase indicator is amber\n(reading).\n\n![Reading phase: thoughts before any output](media\u002Fstill_reading.png)\n\n**2. The judgment is formed, then verbalized.** As the reply begins (\"No, that\nis **not…\"), *incorrect* dominates the workspace at high strength, with\n*equation*, *calculation*, *statement* co-active — the conclusion is\ninternally settled several tokens before the words \"not correct\" appear.\n\n![The verdict forming](media\u002Fstill_verdict.png)\n\n**3. Plans are held while other words are being said.** Mid-explanation, the\nworkspace holds *modulo*, *bitwise*, *system*, *numbers* — the technical\ncaveat the model is about to raise — while the current output token is\nunrelated.\n\n![Planning ahead of speech](media\u002Fstill_planning.png)\n\nThese reproduce, on an open 4B model on consumer hardware, the reporting and\nplanning phenomena described in the paper (which used Claude-scale models),\nincluding the two-hop signature: *Italy* at layer 20 and *euros* at layer 26\non the country-shaped-like-a-boot question, before generation begins.\n\n## Reading the display\n\n- **Each rendered word is a lens readout, not model output.** It indicates an\n  internal activation disposing the model toward that word.\n- **Vertical position corresponds to layer.** Early layers (perception) are at\n  the top; readouts approach the bottom rail as they approach emission.\n- **Size and opacity encode absolute readout strength.** The display applies a\n  fixed monotone mapping from lens probability; weak readouts are rendered\n  weak. Amber marks readouts taken while reading the user; blue while\n  generating.\n- **Hover** shows a word's per-layer activation profile; **click** opens an\n  inspector with peak strength, mean depth, and strength history.\n- The right panel records everything the canvas curates: the conversation, a\n  live ranking of currently-active readouts, and a per-token ledger.\n\n## Timeline, words, and trace\n\nLive playback is fast; nothing is lost. Every frame of the current response is\nkept, so the whole display can be paused and re-inspected.\n\n- **Scrub the response.** A transport bar under the canvas (step \u002F play \u002F\n  scrubber \u002F speed) seeks to any token; the canvas, top-of-mind ranking,\n  partial reply, and stats reconstruct to that exact moment. Click any ledger\n  row to jump to it. During generation the view rides the live edge — scrub\n  back freely, then hit **live** to catch up. Replays are scrubbable the same\n  way.\n- **The words tab** (beside the ledger) aggregates every word the lens read\n  out during the response — how many tokens it was active, its peak layer and\n  strength — ranked by presence. Click a word to jump to its peak moment and\n  load it in the trace view.\n- **The trace view** (cloud \u002F trace, top left) plots one word's readout\n  strength across layers × tokens: the x-axis is shared with the scrubber,\n  amber while reading, blue while generating. It shows a concept climbing the\n  stack — and igniting across it just before being spoken — structure the\n  instantaneous cloud cannot show. Click anywhere on it to seek.\n\n## Method\n\n```\nbrowser (single HTML file)  ⇐ websocket ⇐  server.py\n    Qwen3.5-4B (bf16, HF transformers, KV cache)\n    pre-fitted Jacobian lens: neuronpedia\u002Fjacobian-lens, revision qwen-n1000\n    per token: residual hooks at 9 layers → J_l transport → unembed\n             → full-vocabulary softmax → word-start top-k → frame\n```\n\nEach exchange has two phases. A single prefill pass covers the user's message,\nwith lens readouts taken at every position (the *reading* phase); generation\nthen proceeds token-by-token with a KV cache, reading the lens at the newest\nposition each step (the *thinking* phase). The lens adds a per-layer\nmatrix-vector product and an unembedding per token, so streaming runs at the\nmodel's native generation speed.\n\n**Display filtering.** Raw lens top-k contains punctuation and BPE\ncontinuation fragments (e.g. *itude*, from *cert‑itude*), which are not\nmeaningful as readouts. Display is restricted to word-initial vocabulary\ntokens, following the reference implementation's `mask_display` with a\nstricter word-start criterion. Probabilities are computed over the full\nvocabulary before any filtering, so filtering affects legibility only, never\nthe readout itself.\n\n## Validation\n\n`verify_accuracy.py` compares this implementation's live path (forward hooks,\nKV cache enabled) against the reference `JacobianLens.apply()` on identical\ninputs. Across 4 layers × 3 positions on the walkthrough prompt, top-5\nreadouts match exactly, with cosine similarity ≥ 0.99998 between logit\nvectors, and reproduce the expected two-hop intermediates. The audit can be\nre-run at any time with the server stopped.\n\n## Setup\n\nRequirements: an NVIDIA GPU with ~10 GB of VRAM (and a CUDA build of PyTorch),\nor an Apple Silicon Mac with 16 GB+ unified memory (PyTorch ≥ 2.3, macOS 14+;\nruns on the `mps` device), plus Python 3.11+. Without either, the server falls\nback to CPU (slow, but usable for smoke tests). The device is picked\nautomatically at startup. First launch downloads the model and lens (~9 GB\ntotal) and builds a display-token mask (~1 minute, cached).\n\n```bash\ngit clone https:\u002F\u002Fgithub.com\u002Fninjahawk\u002FSubtext\ncd Subtext\npip install -r requirements.txt\npython server.py\n# → http:\u002F\u002Flocalhost:8765\n```\n\nOn Windows, run `python -u -X utf8 server.py`, or use `start.bat`.\n\n**Other models.** The server is configured for Qwen3.5-4B because Neuronpedia\npublishes a pre-fitted lens for it (a 27B lens is also published, for larger\nGPUs — edit `MODEL_NAME`\u002F`LENS_FILE` in `server.py`). Any HuggingFace decoder\ncan be used by fitting your own lens with `jlens.fit()`; ~100 prompts produces\na usable lens, and fitting a 4B-scale model takes on the order of an hour on a\nsingle consumer GPU. See the [reference repo](https:\u002F\u002Fgithub.com\u002Fanthropics\u002Fjacobian-lens)\nfor details.\n\n**Replays.** The ⤓ session button exports the current conversation — every\nlens frame included — as a JSON file. Open the app with `?replay=\u003Cfile-url>`\nto play one back with live pacing, no GPU required; that is exactly what the\n[hosted demo](https:\u002F\u002Fninjahawk.github.io\u002FSubtext\u002F) is.\n\n## Limitations\n\nThe instrument inherits the method's limitations. The lens reads only concepts\nthat correspond to single vocabulary tokens; multi-token concepts are invisible\nor fragmentary. It approximately captures the workspace identified in the\npaper, not the entirety of the model's internal state, and layers below the\nfitted range are not observed. Interpretation should also respect the paper's\nown framing: workspace readouts demonstrate functional availability of\ninformation for report and reasoning; they do not demonstrate subjective\nexperience.\n\n## Acknowledgements\n\nThe method and reference implementation are by Anthropic\n([jacobian-lens](https:\u002F\u002Fgithub.com\u002Fanthropics\u002Fjacobian-lens), Apache 2.0).\nPre-fitted lens weights are published by\n[Neuronpedia](https:\u002F\u002Fhuggingface.co\u002Fneuronpedia\u002Fjacobian-lens). The model is\n[Qwen3.5-4B](https:\u002F\u002Fhuggingface.co\u002FQwen\u002FQwen3.5-4B). Subtext is an\nindependent project and is not affiliated with Anthropic.\n\nLicensed under Apache 2.0.\n","Subtext 是一个用于实时观测大语言模型内部推理过程的可视化工具。它基于 Jacobian lens 方法，在模型处理用户输入和生成回复的每个 token 步骤中，持续将各层残差流激活映射至最终词表空间，以可视化模型“正在思考但尚未说出”的词汇倾向（即 J-space 工作区状态）。项目支持本地运行、实时流式渲染，技术上依赖 PyTorch 和 Transformers 库，预集成 Qwen3.5-4B 模型及专用 lens。适用于 LLM 可解释性研究、模型行为调试、教学演示及机制分析等需要细粒度追踪模型内部动态的场景。",2,"2026-07-09 02:30:12","CREATED_QUERY"]