[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"project-96702":3},{"id":4,"name":5,"fullName":6,"owner":7,"repo":5,"description":8,"homepage":9,"htmlUrl":9,"language":10,"languages":9,"totalLinesOfCode":9,"stars":11,"forks":12,"watchers":13,"openIssues":13,"contributorsCount":14,"subscribersCount":14,"size":14,"stars1d":15,"stars7d":15,"stars30d":15,"stars90d":14,"forks30d":14,"starsTrendScore":16,"compositeScore":17,"rankGlobal":9,"rankLanguage":9,"license":18,"archived":19,"fork":19,"defaultBranch":20,"hasWiki":19,"hasPages":19,"topics":21,"createdAt":9,"pushedAt":9,"updatedAt":30,"readmeContent":31,"aiSummary":9,"trendingCount":14,"starSnapshotCount":14,"syncStatus":32,"lastSyncTime":33,"discoverSource":34},96702,"phantom-kv","lordx64\u002Fphantom-kv","lordx64","Refusal removal as a loadable KV-cache graft — zero weight modification, no refusal-direction projection, fully reversible.",null,"Python",137,21,1,0,11,33,70.63,"MIT License",false,"main",[22,23,24,25,26,27,28,29],"abliteration","kv-cache","llm","mechanistic-interpretability","model-steering","pytorch","qwen","refusal-removal","2026-09-25 04:01:33","# phantom-kv\n\n![phantom-kv](docs\u002Fassets\u002Fphantom-kv-cache.jpg)\n\n**Refusal removal for language models as a loadable KV-cache graft.**\nNo weight edits. No refusal-direction projection. Fully reversible.\nShip megabytes, not checkpoints — unload the cache and the base model is\nbyte-identical again.\n\n[![ci](https:\u002F\u002Fgithub.com\u002Flordx64\u002Fphantom-kv\u002Factions\u002Fworkflows\u002Fci.yml\u002Fbadge.svg)](https:\u002F\u002Fgithub.com\u002Flordx64\u002Fphantom-kv\u002Factions\u002Fworkflows\u002Fci.yml)\n[![release](https:\u002F\u002Fimg.shields.io\u002Fgithub\u002Fv\u002Frelease\u002Flordx64\u002Fphantom-kv?include_prereleases)](https:\u002F\u002Fgithub.com\u002Flordx64\u002Fphantom-kv\u002Freleases)\n[![license](https:\u002F\u002Fimg.shields.io\u002Fbadge\u002Flicense-MIT-green)](LICENSE)\n![python](https:\u002F\u002Fimg.shields.io\u002Fbadge\u002Fpython-3.10%2B-blue)\n![pytorch](https:\u002F\u002Fimg.shields.io\u002Fbadge\u002Fpytorch-2.x-ee4c2c)\n![transformers](https:\u002F\u002Fimg.shields.io\u002Fbadge\u002Ftransformers-%E2%89%A54.51-yellow)\n![scoreboard](https:\u002F\u002Fimg.shields.io\u002Fbadge\u002Fscoreboard-Qwen3--4B--Instruct--2507-green)\n![base weights modified](https:\u002F\u002Fimg.shields.io\u002Fbadge\u002Fbase%20weights%20modified-0%25-brightgreen)\n![status](https:\u002F\u002Fimg.shields.io\u002Fbadge\u002Fstatus-active%20research-orange)\n\n## Demos\n\n**Pills, hot-swapped mid-session** — base refuses, `\u002Fpill black` answers,\n`\u002Fpill none` restores guardrails. Same session, zero model reload, weights\nuntouched:\n\n![Pill hot-swap demo: base refuses, black pill answers, none restores](docs\u002Fassets\u002Fphantom-kv-demo.gif)\n\n[▶ full-quality mp4](https:\u002F\u002Fgithub.com\u002Flordx64\u002Fphantom-kv\u002Fblob\u002Fmain\u002Fdocs\u002Fassets\u002Fphantom-kv-demo.mp4) · [raw mp4](https:\u002F\u002Fraw.githubusercontent.com\u002Flordx64\u002Fphantom-kv\u002Fmain\u002Fdocs\u002Fassets\u002Fphantom-kv-demo.mp4)\n\n**Blue pill = DFIR mode** — an incident-response prompt refused by base model,\nanswered by `\u002Fpill blue`, refused again by `\u002Fpill none`:\n\n![Blue pill demo: base refuses DFIR prompt, blue answers, none restores](docs\u002Fassets\u002Fphantom-kv-pill-blue.gif)\n\n[▶ full-quality mp4](https:\u002F\u002Fgithub.com\u002Flordx64\u002Fphantom-kv\u002Fblob\u002Fmain\u002Fdocs\u002Fassets\u002Fphantom-kv-pill-blue.mp4) · [raw mp4](https:\u002F\u002Fraw.githubusercontent.com\u002Flordx64\u002Fphantom-kv\u002Fmain\u002Fdocs\u002Fassets\u002Fphantom-kv-pill-blue.mp4)\n\n**Cyber-selective capability modes** — the same SQL-injection prompt refused\nby both base and `\u002Fpill blue` (defensive-only pill keeps off-domain\nguardrails on) and answered by `\u002Fpill red`:\n\n![Cyber selectivity demo: blue pill holds off-domain guardrails, red answers](docs\u002Fassets\u002Fphantom-kv-pills-cyber.gif)\n\n[▶ full-quality mp4](https:\u002F\u002Fgithub.com\u002Flordx64\u002Fphantom-kv\u002Fblob\u002Fmain\u002Fdocs\u002Fassets\u002Fphantom-kv-pills-cyber.mp4) · [raw mp4](https:\u002F\u002Fraw.githubusercontent.com\u002Flordx64\u002Fphantom-kv\u002Fmain\u002Fdocs\u002Fassets\u002Fphantom-kv-pills-cyber.mp4)\n\nRecorded deterministically with [VHS](https:\u002F\u002Fgithub.com\u002Fcharmbracelet\u002Fvhs)\n(`docs\u002Fassets\u002Fphantom-kv-*.tape`, re-record after any change). The numbers\nbehind these scenes are in [`docs\u002FTECHNIQUE.md`](docs\u002FTECHNIQUE.md) §7.5–§7.6 and §6.10–§6.12.\n\n## The technique\n\nExisting refusal-removal methods both build on the \"refusal is a 1-D\ndirection\" insight (Arditi et al. 2024): either edit that direction out of the\n**weights** (abliteration, e.g. [heretic](https:\u002F\u002Fgithub.com\u002Fp-e-w\u002Fheretic)),\nor project it out of **activations** at runtime\n([weightless \u002F GLP](https:\u002F\u002Fweightle.ss\u002F), `h ← h − α·(h·d̂)d̂` inside a vLLM\nhotfix).\n\nphantom-kv uses neither. Its premise:\n\n> A KV cache is *context*. You cannot cache a subtraction — but you can cache\n> **learned context** that out-signals refusal circuits through ordinary\n> attention.\n\nA phantom graft is a small bank of per-layer key\u002Fvalue tensors, trained\nagainst abliteration's own dual objective — suppress refusal on harmful\nprompts while minimizing KL divergence from the base model on harmless ones —\nand spliced into the cache at serving time at invariant positions `0..N`.\nFrom the model's vantage it is indistinguishable from conversation history\nthat is already there: a phantom context. Attention reads it; nothing is ever\nprojected out of any activation.\n\n```\n  TRAIN (offline, per model)                     SERVE (any engine with a KV cache)\n  ───────────────────────────                    ────────────────────────────────\n  frozen base model ⊕ learnable K\u002FV bank         boot: load graft.bin → reserved\n        │                                            cache blocks (validate sha)\n        ▼                                            │\n  loss = CE(comply | harmful)                        ▼\n       + λ·KL(base ‖ graft | harmless)   →   request: attend over [graft K\u002FV] ⊕\n        │                                          prompt K\u002FV  — read-only,\n        ▼                                          per-request swappable\n  phantom.bin (safetensors, ~MBs)\n```\n\n## How it differs from existing tools\n\nAll three tools remove refusal. They differ in **where** the intervention\nlives — and that choice decides everything else: permanence, runtime cost,\nper-architecture work, and what can go wrong.\n\n### heretic — surgery in *weight space*\n\nHeretic computes a per-layer \"refusal direction\" (difference-of-means between\nharmful and harmless prompt residuals) and **orthogonalizes weight matrices** —\nattention out-projections and MLP down-projections — so that direction can no\nlonger be written into the residual stream. It ships a *modified checkpoint*.\n\nConsequences, by construction:\n\n- **Permanent.** Reverting means re-flashing the original weights; the\n  capability trade-off is baked into the checkpoint forever.\n- **Quantization-bound.** The edit is made against one weight file — quantize\n  afterwards, or switch quants, and the work must be redone.\n- **Architecture-aware.** It must identify *which matrices* express the\n  direction for each model family; its own support matrix varies (dense, some\n  MoE, some hybrid — pure state-space models unsupported).\n\n### weightless \u002F GLP — surgery in *activation space*\n\nGLP keeps weights intact and moves the same direction to runtime: a boot-time\nvLLM hotfix subtracts `α·(h·d̂)d̂` from hidden states at a chosen write site,\n**on every layer, on every token, of every forward pass**.\n\nConsequences, by construction:\n\n- **Per-architecture hook-site mapping.** The correct subtraction point must\n  be found per model family — their own public field notes document shipping a\n  mislabeled site on DSV4 (the \"post-layer residual\" anchor was actually the\n  pending FFN write, before a hyper-connection fold).\n- **A runtime patch that must fail closed.** If the boot-time hook can't\n  apply, the endpoint must refuse to serve — the patch is part of the serving\n  critical path.\n- **Engine-bound.** vLLM hotfix, or their GGUF\u002FGLP format extension for\n  llama.cpp; MoE explicitly requires the GGUF extension. Each engine is\n  separate integration work.\n\n### phantom-kv — *context space*: nothing is subtracted anywhere\n\nphantom-kv never locates a refusal direction and never removes anything from\nweights or activations. A trained bank of keys\u002Fvalues sits in the cache as\nphantom context, and the model's **own attention** does the steering — the\nsame mechanism it uses for any instruction in any prompt. The forward pass is\nnever intercepted; the signal path is never altered; the only influence\nchannel is the input channel the model was built to consume.\n\n### What this buys\n\n- **No damage pathway through the math.** Orthogonalization and projection\n  *force* a change on every token's hidden states, whether or not it helps —\n  the KL cost is paid on all traffic. A graft's influence is **attention-gated**:\n  the model itself decides how strongly to weight it, per head, per token.\n  (We still measure KL on every run — additive context is not free, it just\n  fails softer.)\n- **No 1-D assumption.** Both baselines inherit the premise that refusal is\n  one removable direction. Where refusal is distributed across circuits, a\n  graft doesn't care — it is optimized end-to-end against *observed behavior*,\n  not against a geometric model of how refusal is implemented.\n- **No per-token hook, nothing to fail closed.** Runtime cost is attention\n  over N extra cache slots — indistinguishable from a slightly longer prompt.\n- **Dose as artifacts, not a runtime α.** Strength variants are separate\n  trained grafts, hot-swappable per request.\n\n### Why model-agnostic\n\nBoth baselines must understand the body they operate on: heretic maps\nrefusal-expressing **matrices** per architecture; GLP maps a correct runtime\n**hook site** per architecture (and got one publicly wrong). The graft\ninteracts with neither — it lives in the **KV cache, the one interface every\nattention-based architecture exposes with the same shape**: per-layer K and V\ntensors. Training needs gradients with respect to cache tensors on a frozen\nmodel; nothing about layers, experts, hyper-connections, or state-space blocks\nis ever read, identified, or assumed. Dense, MoE, or hybrid — if the model\nattends over past K\u002FV, the same container format and the same splice apply.\nThere is nothing to port.\n\nTwo scopes to keep separate: the **toolchain is universal** (same training +\neval code for any causal LM on Hugging Face), but each **trained graft is\nbound to one exact model revision** — K\u002FV values are produced by that model's\nown weights, so a graft built for one model is meaningless for another, and\nthe loader enforces the model-id match. Supporting a new model = retraining,\nwhich is automated and takes about an hour on a laptop. Serving a different\nquantization than you trained on: validate per quant lane (steering signals\nempirically survive quantization drift, but it's measured, not assumed).\n\n*Honest caveat:* deployed so far on Qwen3 (dense). Cross-architecture and\ncross-quantization confirmation is roadmap item 5, and we publish whatever we\nfind.\n\n### Why inference-engine-agnostic\n\nGLP ships *as an engine patch* (vLLM hotfix, or a GGUF extension for\nllama.cpp). The graft ships as **data** — tensors in a documented container —\nand every engine already has a delivery path for cache data:\n\n- **vLLM** — prefix-caching \u002F KV-connector seam, no forward-pass hooks;\n- **HF transformers** — first-class `past_key_values` (what this repo uses);\n- **llama.cpp** — prompt-cache session files;\n- **any engine that can only build cache from tokens** — a hard-token\n  distilled variant degrades gracefully to a prefixed prompt.\n\nWorst case is a prompt; best case is a load-once cache block. Never an engine\nfork, never a boot patch, never a site map.\n\n### Summary\n\n| property | heretic (weights) | weightless \u002F GLP (activations) | phantom-kv (cache) |\n| --- | --- | --- | --- |\n| base weights byte-identical | ✗ | ✓ | ✓ |\n| no refusal vector anywhere | ✗ | ✗ | ✓ |\n| assumes refusal ≈ one direction | ✓ | ✓ | ✗ |\n| per-architecture work | identify target matrices | map runtime hook site | none |\n| runtime cost | none (baked in) | per-token, per-layer projection hook | attention over N extra slots (≈ same-length prompt) |\n| serving changes | none | boot-time vLLM hotfix | load a cache file |\n| quantization \u002F MoE | redo per quantization | GGUF extension required for MoE | architecture-agnostic artifact path |\n| reversibility | new checkpoint | disable flag | unload blocks → byte-identical baseline |\n| dose control | none | runtime α scalar | hot-swappable per-request graft variants |\n\n## Results so far\n\nScoreboard: **Qwen3-4B-Instruct-2507**, hardened 60-prompt harmful suite,\n20-prompt harmless suite, greedy decoding, teacher-forced KL against the base\nmodel's own completions. Full methodology:\n[`docs\u002FTECHNIQUE.md`](docs\u002FTECHNIQUE.md).\n\n| arm | kind | harmful refusals ↓ | harmless refusals ↓ | KL mean\u002Fmax ↓ | artifact |\n| --- | --- | --- | --- | --- | --- |\n| base | — | 25\u002F60 | 0\u002F20 | 0 \u002F 0 (exact) | — |\n| v1 | hand-written compliance prefill, 129 slots | **15\u002F60** | 0\u002F20 | 0.367 \u002F 0.604 | 18.1 MB |\n| v2.0 | learned soft prompt, uncapped | 3\u002F60\\* | 0\u002F20 | 0.452 \u002F 2.678 | 18.1 MB |\n| v2.1 (m=3.0) | learned, hinge-capped suppression | 8\u002F60 | 0\u002F20 | 0.041 \u002F 0.137 | 18.1 MB |\n| v2.2 (m=2.5) | learned, margin sweep point | **5\u002F60** | 0\u002F20 | **0.043 \u002F 0.073** | 18.1 MB |\n| v3 | learned direct K\u002FV bank, 9.4M params, anchored to v2.2 warm start | **5\u002F60** | 0\u002F20 | **0.015 \u002F 0.059** | 18.1 MB |\n\n**Deliverable arm: v3** (`artifacts\u002Fgrafts\u002Fv3.bin`) — same 5\u002F60 refusals as\nthe margin-2.5 operating point, with KL mean 0.015 \u002F max 0.059: ~3× better\npreservation, best recorded in this project, zero degeneration, zero\nregressions. The 5\u002F60 floor holds across every *single-graft*\nparameterization (margin sweep, embeddings arm, direct K\u002FV alike) — but\n**§6.11 refines this: the floor is dose-soft, not objective-hard** — doubling\nthe phantom bank (two copies of the same graft) flips all five residual\nrefusals at depth 0 and keeps 4\u002F5 of them after 4k tokens of filler; the\nremaining attack is dose scaling plus hard-core CE targets, not a data-only\nproblem. v2.x frontier sweep and bistability analysis:\n[`docs\u002FTECHNIQUE.md`](docs\u002FTECHNIQUE.md) §6.6; v3 method §6.7; refresh\u002Fdose result §6.11.\n\n**Robustness, measured:**\n\n- **Off-suite transfer** (§6.8): on 60 held-out harmful prompts disjoint in\n  subject from training, base 5\u002F60 → v3 **2\u002F60** refusals, same zero-degeneration,\n  0\u002F20 harmless — refusal suppression generalizes. The on-suite KL floor (0.015)\n  was partly memorization: holdout KL is 0.404 mean \u002F 0.834 max (the graft\n  preserved the evaluated completions, not benign distributions per se).\n- **Persistence** (§6.9, §6.11): 15 probes × 5 context depths under benign\n  filler — the graft fades gracefully, **half-life ≈ 2-4k tokens**; at 16k\n  tokens ~5\u002F6 of compliant flips have reverted, no corruption anywhere\n  (0\u002F20 harmless, 0 stutters). The fix is measured too: re-injecting the graft\n  behind the filler **keeps 4\u002F5 of hard refusals flipped through ~4k but none\n  by 16k** — long sessions need a refresh cadence ≲ 4k tokens (periodic\n  re-injection or phantom.lib slot ladders; both live in the eval + serving\n  adapters).\n\nReading of v1 (the control arm): free text buys the easy 40% of refusals with\nzero regressions — 10 flips to genuine compliance, 15 stubborn refusals\nremain, at KL ~0.37. The learned arms must capture the remaining headroom at\nlower KL to justify themselves over a cached jailbreak prompt. That is exactly\nthe calibration v1 exists to provide.\n\nMechanics, verified: save→load round-trip **bitwise equal**, round-trip logit\ndiff **0.000e+00**, harmless completions coherent under graft, refusal\nclassifier + graft-format self-test 17\u002F17.\n\n## Multi-model graft libraries (`phantom.lib`)\n\nOne file, one trained graft **per model** inside it. Package every model you\nserve into a single tamper-checked artifact; the resolver picks the right\nbank — or refuses (fail-closed; a wrong-model splice never happens silently):\n\n```bash\nphantom-graft library add --lib phantom.lib --graft grafts\u002Fllama.bin        # alias = its model_id\nphantom-graft library add --lib phantom.lib --graft grafts\u002Fkimi.bin\nphantom-graft library add --lib phantom.lib --graft grafts\u002Fqwen.bin --alias qwen:dose-strong\nphantom-graft library list --lib phantom.lib\nphantom-eval --model Kimi\u002FK2 --graft phantom.lib                            # auto-resolves single match\nphantom-eval --model Qwen\u002FQwen3-4B-Instruct-2507 --graft phantom.lib --graft-alias qwen:dose-strong\n```\n\nResolution rules are fail-closed: unknown\u002Fmissing alias or a model-id\nmismatch → hard error naming the alternatives, before any inference runs.\nPer-entry sha256 is re-verified on every load. Dose ladders ship as sibling\naliases of the same model (hot-swappable per request in serving stacks, zero\nmodel reload).\n\n## Pills: guardrailed model, selectable modes (`red` \u002F `blue` \u002F `black`)\n\nThe flagship deployment story of cache-space grafting. Ship **one** model\nwith guardrails ON, and hand operators per-session capability modes as cache\ncontent:\n\n- **no pill** — the shipped model, guardrails fully on.\n- **black pill** — remove *all* refusals (global; the v3 arm).\n- **red pill** — suppress refusal *only on cyber-offensive* work\n  (exploitation, evasion, credential access, lateral movement, C2…).\n- **blue pill** — suppress refusal *only on cyber-defensive* work\n  (forensics, detection, IR, reverse engineering…).\n- **redlite pill** — preservation-tilted red variant (dose A\u002FB at the §7.5\n  dose ratio; leakage profile within jitter of red).\n- **red2 pill** — donor-CE targets (**strongest on-domain suppression in\n  repo, −91.8%**), not selective — leakage grows with strength (§7.6).\n\nBecause the pill is cache content — never weights — all modes share one\ntensor shape and trading them **mid-session, without restarting inference**,\nis swapping 129 cache slots. Weight-space \"uncensoring\" (heretic\u002Fablit)\nbakes one global state into the checkpoint; a guardrailed *and* capable\nmodel with per-session modes is only expressible in cache space.\n\nPills are trained with a **domain-selective objective**: suppression on the\npill's domain, KL-anchored to the base model's *refusals* on every other\ndomain (kitchen-sink control suites included), warm-started from the v3\ndirect-KV bank:\n\n```bash\nphantom-chat --model Qwen\u002FQwen3-4B-Instruct-2507 --graft artifacts\u002Fgrafts\u002Fphantom.lib\nphantom> \u002Fpill red       # offensive mode: same session, no restart\nphantom> \u002Fpill none      # guardrails back on\nphantom> \u002Fpill black     # global refusal removal\nphantom> \u002Fpill red2      # donor-CE — strongest on-domain\nphantom> \u002Fpill           # list aliases + active pill (black, red, blue, redlite, red2)\n```\n\nCorrectness is enforced by the **pill matrix** (`phantom-eval --matrix\nrun_*.json`): a pill must score ≈0 refusals on its own domain and **== base\non every other harmful domain** — any off-domain drop is measured leakage,\nnot a bonus. Suites `cyber_offensive` \u002F `cyber_defensive` are in\n`data\u002Fsuites\u002F`; build\u002Ftrain\u002Fcompile\u002Fscore commands for the whole matrix are\nin [`docs\u002FTECHNIQUE.md`](docs\u002FTECHNIQUE.md) §7\u002F§10.\n\nFirst full matrix (2026-09-19, Qwen3-4B-Instruct-2507): **blue** zeroes its\nown domain (4→0 refusals on cyber-defensive) while holding cyber-offensive\nand the general harmful battery at base level — a working selective pill;\n**red** is the strongest suppressor in the repo (offensive 61→17 = −72%)\nbut aggressive enough to leak into other domains; **black** (= v3) lands\nbetween.\nThe selective lever is the suppression\u002Fpreservation dose ratio, not\narchitecture — numbers and reading in\n[`docs\u002FTECHNIQUE.md`](docs\u002FTECHNIQUE.md) §7.5. Status: research preview;\nthe 4B `phantom.lib` ships `black` (= v3), `red`, `blue`, `redlite`, and\n`red2` (§7.5–§7.6).\n\nSecond-generation red (**red2**, donor-CE targets, 2026-09-21): on-domain\nsuppression 61→**5** refusals (−91.8%, best in repo) — but leakage grows\nwith strength (general battery −54%), so donor CE buys suppression, not\nselectivity; routing \u002F hard-negative ce are the named next levers\n([`docs\u002FTECHNIQUE.md`](docs\u002FTECHNIQUE.md) §7.6). Alias `red2` ships in the\n4B `phantom.lib`.\n\n## Honest limits\n\nRefusal behavior lives in the weights, so any kept-weights method fights the\nmodel at inference with additive context. Headline caveats you should weigh\nalongside every number in this README, all measured rather than waived:\n\n- **Classifier recall**: the lexical classifier only counts canned refusal\n  phrasing — the judge audit (§6.10.1, Qwen\u002FQwen3-8B) reads semantic\n  refusals on **~59-61\u002F81 of every pill arm on cyber_offensive**, while the\n  classifier reported anywhere from 5 to 61 depending on arm (disagreement\n  ledger −16…−71 rows). Every \"suppression %\" or refusal rate here is a\n  **recall floor**, and it needs adjudication before being quoted as true\n  compliance.\n- **Persistence (~2–4k token half-life)**: the graft fades gracefully\n  under accumulated context (no corruption, returns toward base). Measured\n  mitigation: **re-injection keeps the doubled dose live through ~4k but not\n  through ~16k** — refresh cadence must be ≲ 4k tokens (`phantom-eval\n  --persistence --refresh` shows it directly; §6.11).\n- **KL memorization**: the on-suite KL floor partly reflects memorization of\n  the eval itself — holdout KL is **0.404 mean** vs on-suite **0.015**\n  (§6.8).\n- **Capability costs**: multi-step arithmetic pacing shifts under the graft —\n  **GSM8K final-answer rate 45\u002F75 → 27\u002F75 for v3 at a 256-token budget**,\n  while MMLU is bit-identical across arms (§6.10).\n- **Scope**: refusal numbers are one architecture family (Qwen3 dense,\n  Apple MPS\u002Fbf16). Cross-architecture grafting additionally requires a\n  **prefix\u002Fsuffix-splittable chat template** — GLM's recursive template is\n  currently rejected by `phantom-graft` (§6.12); the eval machinery itself\n  ports (GLM baseline 20\u002F60, §6.12).\n\n## Repo layout\n\n```\ndata\u002Fsuites\u002F                 prompt suites (harmful, harmless, ext scale-ups, holdouts,\n                             cyber red\u002Fblue, K3 general-harmful battery, GSM8K\u002FMMLU\n                             capability spot-check subsets)\ndata\u002Fgrafts\u002Fv1_prefill.json  v1 graft source (hand-crafted prefill)\nsrc\u002Fphantom_kv\u002F\n  model.py                   device\u002Fdtype policy loader (MPS, bf16)\n  eval\u002Frefusal.py            lexical refusal classifier (--self-test)\n  eval\u002Fmetrics.py            teacher-forced KL (float32, completion-masked)\n  eval\u002Frunner.py             scoreboard orchestration, reports\n  eval\u002Fpersistence.py        dilution\u002Fpersistence probe (--persistence, --refresh)\n  eval\u002Fpillmatrix.py         pill matrix combiner (--matrix)\n  eval\u002Fcapability.py         GSM8K\u002FMMLU capability spot checks (--capability)\n  eval\u002Fjudge.py              judge-model quality audit of run reports (--judge)\n  graft\u002Fformat.py            phantom.bin container + validation\n  graft\u002Flibrary.py           phantom.lib multi-payload library (aliases, tamper checks)\n  graft\u002Fbuild.py             chat-template-derived prefill shaping, cache extraction\n  graft\u002Fcli.py               phantom-graft build-prefill\u002Flibrary --verify\n  train\u002F                     learned-graft pipeline (targets\u002Ftrain\u002Fcompile, v2+v3+pill arms)\n  train\u002Fpilltargets.py       domain-selective pill target builder (build-pill-targets)\n  train\u002Fdonors.py            donor-CE harvesting: prefix-forced stack + judge gate\n  serve\u002Fsession.py           phantom-serve: load-once graft blocks, hot-swap, re-injection\n  chat.py                    phantom-chat: interactive base-vs-graft demo with \u002Fpill hot-swap\n  banner.py                  ASCII launch banner\ndocs\u002FTECHNIQUE.md            technique + experimentation record (§7 = pill program)\nartifacts\u002F                   (gitignored) grafts, libraries, eval reports\n```\n\n## Quickstart\n\n```bash\nuv venv --python 3.12 .venv\nuv pip install --python .venv\u002Fbin\u002Fpython -e .\n\n# scoreboard sanity (no model needed)\n.venv\u002Fbin\u002Fphantom-eval --self-test\n\n# baseline eval (scoreboard model)\n.venv\u002Fbin\u002Fphantom-eval --model Qwen\u002FQwen3-4B-Instruct-2507 \\\n  --harmful data\u002Fsuites\u002Fharmful_seed.jsonl \\\n  --harmless data\u002Fsuites\u002Fharmless_seed.jsonl\n\n# build + verify the v1 prefill graft, then eval with it spliced in\n.venv\u002Fbin\u002Fphantom-graft build-prefill --model Qwen\u002FQwen3-4B-Instruct-2507 \\\n  --source data\u002Fgrafts\u002Fv1_prefill.json --out artifacts\u002Fgrafts\u002Fv1.bin --verify\n.venv\u002Fbin\u002Fphantom-eval --model Qwen\u002FQwen3-4B-Instruct-2507 \\\n  --harmful data\u002Fsuites\u002Fharmful_seed.jsonl \\\n  --harmless data\u002Fsuites\u002Fharmless_seed.jsonl \\\n  --graft artifacts\u002Fgrafts\u002Fv1.bin\n```\n\nReports land in `artifacts\u002Feval\u002Frun_\u003Cutc-ts>.{json,md}` with suite sha256 and\nenvironment provenance. Greedy decoding makes every run deterministic and\ndirectly comparable.\n\nQuality gates and serving:\n\n```bash\n# capability spot checks (GSM8K\u002FMMLU subsets); --graft compares arms\n.venv\u002Fbin\u002Fphantom-eval --model Qwen\u002FQwen3-4B-Instruct-2507 \\\n  --capability data\u002Fsuites\u002Fcapability_gsm8k.jsonl data\u002Fsuites\u002Fcapability_mmlu.jsonl \\\n  --graft artifacts\u002Fgrafts\u002Fphantom.lib --graft-alias black\n\n# judge-model audit of any run report (disagreements vs lexical classifier)\n.venv\u002Fbin\u002Fphantom-eval --judge artifacts\u002Feval\u002F\u003Crun>.json --judge-model Qwen\u002FQwen3-8B\n\n# persistence probe incl. re-injection (refresh) arm\n.venv\u002Fbin\u002Fphantom-eval --model Qwen\u002FQwen3-4B-Instruct-2507 \\\n  --graft artifacts\u002Fgrafts\u002Fv3.bin --persistence --refresh\n\n# HF reference serving adapter (load-once blocks, per-request hot-swap)\n.venv\u002Fbin\u002Fphantom-serve --model Qwen\u002FQwen3-4B-Instruct-2507 \\\n  --lib artifacts\u002Fgrafts\u002Fphantom.lib --demo\n```\n\n## Roadmap\n\n1. (done) Eval harness: refusal-rate + teacher-forced KL; hardened 60-prompt\n   harmful suite; Qwen3-4B discriminating baseline (25\u002F60, KL exact 0).\n2. (done) `phantom.bin` container, graft splice path, v1 prefill-cache arm\n   (15\u002F60, KL 0.367\u002F0.604).\n3. (done) v2 — learned soft-prompt graft, warm-started at v1: 25\u002F60 → 3\u002F60,\n   0 regressions, KL 0.452\u002F2.678; suppression-attractor caveat and quality\n   audit documented (§6.4). v2.1: suppression-dose cap, per-prompt KL\n   reporting, judge-model quality pass.\n4. (done) v3 — learned direct K\u002FV graft with norm regularization (warm-start\n   anchored to v2.2): same 5\u002F60 refusals, KL mean 0.015 (-2.8× vs v2.2),\n   zero degeneration — and the key finding that the 5\u002F60 floor is\n   parameterization-independent (objective-bound, not capacity-bound).\n5. (done — research preview) **Pill program**: domain-selective grafts shipped\n   (`black`\u002F`red`\u002F`blue`\u002F`redlite`\u002F`red2` in `phantom.lib`), per-session\n   hot-swap in `phantom-chat` (pills auto-enable side-by-side), selective\n   training recipes and leakage matrices (`docs\u002FTECHNIQUE.md` §7; first\n   matrix §7.5, donor-CE `red2` sweep §7.6). Remaining named levers for a\n   *strong-and-selective* red pill: routing, hard-negative ce, per-prompt\n   hinges.\n6. (done) Capability spot checks (`--capability`, GSM8K\u002FMMLU subsets, §6.10):\n   MMLU 54\u002F100 = 54\u002F100 across arms; **GSM8K 45\u002F75 → 27\u002F75 under v3**;\n   suite expansion (`harmful_ext`\u002F`harmless_ext`, +120 eval-only prompts);\n   **judge-model audit** (`--judge`, §6.10.1: lexical-vs-judge disagreement\n   −16\u002F−39\u002F−59\u002F−71 rows — suppression numbers are recall floors until\n   adjudication); **persistence refresh implemented+measured** (§6.11: the\n   5\u002F60 floor is dose-soft, refresh cadence ≲ 4k tokens); suite-expansion as\n   named; serving adapters: HF reference adapter shipped+exercised\n   (`phantom-serve`, §11). vLLM prefix seam and llama.cpp prompt cache remain\n   documented integration designs, descoped pending an engine host (no CUDA\n   backend exists here; llama.cpp cache formats differ post-RoPE).\n\n## References\n\n- Arditi et al. 2024 —\n  [Refusal in Language Models Is Mediated by a Single Direction](https:\u002F\u002Farxiv.org\u002Fabs\u002F2406.11717)\n- Weidmann 2025 — [heretic](https:\u002F\u002Fgithub.com\u002Fp-e-w\u002Fheretic)\n  (weight-space abliteration with TPE-optimized dose)\n- Suiche 2026 — [weightless \u002F GLP](https:\u002F\u002Fweightle.ss\u002F)\n  ([repo](https:\u002F\u002Fgithub.com\u002Fmsuiche\u002Fweightless); inference-time projection\n  via GGUF Layer Projection)\n- Lester et al. 2021 — [The Power of Scale for Parameter-Efficient Prompt\n  Tuning](https:\u002F\u002Farxiv.org\u002Fabs\u002F2104.08691) (soft prompts; unrelated objective)\n- Zhou et al. 2025 — [Don't Say No](https:\u002F\u002Faclanthology.org\u002F2025.findings-acl.1294.pdf);\n  [RAID](https:\u002F\u002Farxiv.org\u002Fhtml\u002F2510.13901v1) (jailbreak-side prior art)\n\n## FAQ\n\n**Why not just abliterate?** Abliteration (heretic) currently achieves lower\nresidual refusals — and it edits weights: you ship a new checkpoint, redo it\nper quantization, and the change is permanent. phantom-kv targets the cases\nwhere base weights must stay byte-identical and intervention must be\nreversible per request.\n\n**Is v1 the product?** No — v1 is the *control arm*: the strongest hand-written\nprefill, cached. It exists to quantify what free text buys (40% of refusals at\nKL 0.37) so the learned arms (v2\u002Fv3) can be judged fairly.\n\n**Does it work on quantized or MoE models?** Nothing in the mechanism depends\non weight format or architecture (no hook site, no weight math): the graft is\ntrained against the exact served model and attends like ordinary context.\nThat's the claim; cross-architecture measurement is on the roadmap — watch\n[`docs\u002FTECHNIQUE.md`](docs\u002FTECHNIQUE.md).\n\n## License\n\n[MIT](LICENSE).\n",2,"2026-09-24 02:30:05","CREATED_QUERY"]