[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"project-94843":3},{"id":4,"name":5,"fullName":6,"owner":7,"repo":5,"description":8,"homepage":9,"htmlUrl":10,"language":11,"languages":9,"totalLinesOfCode":9,"stars":12,"forks":13,"watchers":14,"openIssues":14,"contributorsCount":9,"subscribersCount":14,"size":14,"stars1d":14,"stars7d":14,"stars30d":14,"stars90d":14,"forks30d":14,"starsTrendScore":14,"compositeScore":15,"rankGlobal":9,"rankLanguage":9,"license":9,"archived":16,"fork":16,"defaultBranch":17,"hasWiki":16,"hasPages":16,"topics":9,"createdAt":9,"pushedAt":9,"updatedAt":18,"readmeContent":19,"aiSummary":20,"trendingCount":14,"starSnapshotCount":14,"syncStatus":21,"lastSyncTime":9,"discoverSource":22},94843,"Soup","MakazhanAlpamys\u002FSoup","MakazhanAlpamys","Fine-tune LLMs from one YAML. Layer streaming trains an 8B model on a 4 GB laptop GPU.",null,"https:\u002F\u002Fgithub.com\u002FMakazhanAlpamys\u002FSoup","Python",1646,262,0,52.26,false,"main","2026-08-24 04:01:22","\u003Cp align=\"center\">\n  \u003Cimg src=\"soup.png\" alt=\"Soup\" width=\"280\">\n\u003C\u002Fp>\n\n\u003Ch1 align=\"center\">Soup\u003C\u002Fh1>\n\n\u003Cp align=\"center\">\n  \u003Cstrong>Fine-tune and post-train LLMs in one command. No SSH, no config hell.\u003C\u002Fstrong>\n\u003C\u002Fp>\n\n\u003Cp align=\"center\">\n  \u003Ca href=\"https:\u002F\u002Ftrysoup.dev\">Website\u003C\u002Fa> &middot;\n  \u003Ca href=\"#quick-start\">Quick Start\u003C\u002Fa> &middot;\n  \u003Ca href=\"#configuration\">Config\u003C\u002Fa> &middot;\n  \u003Ca href=\"#documentation\">Docs\u003C\u002Fa> &middot;\n  \u003Ca href=\"docs\u002Fcommands.md\">Commands\u003C\u002Fa> &middot;\n  \u003Ca href=\"docs\u002Fmodels.md\">Models\u003C\u002Fa> &middot;\n  \u003Ca href=\"https:\u002F\u002Fdiscord.gg\u002F8RgVbFA6Zq\">Discord\u003C\u002Fa> &middot;\n  \u003Ca href=\"https:\u002F\u002Fwww.producthunt.com\u002Fproducts\u002Fsoup-cli\">Product Hunt\u003C\u002Fa>\n\u003C\u002Fp>\n\n\u003Cp align=\"center\">\n  \u003Ca href=\"https:\u002F\u002Fpypi.org\u002Fproject\u002Fsoup-cli\u002F\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fpypi\u002Fv\u002Fsoup-cli?color=blue\" alt=\"PyPI\">\u003C\u002Fa>\n  \u003Ca href=\"https:\u002F\u002Fpepy.tech\u002Fproject\u002Fsoup-cli\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fpepy\u002Fdt\u002Fsoup-cli?color=blue\" alt=\"Downloads\">\u003C\u002Fa>\n  \u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002Fpython-3.10--3.12-blue\" alt=\"Python 3.10-3.12\">\n  \u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002Flicense-Apache--2.0-blue\" alt=\"Apache-2.0 License\">\n  \u003Ca href=\"https:\u002F\u002Fgithub.com\u002FMakazhanAlpamys\u002FSoup\u002Factions\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fendpoint?url=https:\u002F\u002Fgist.githubusercontent.com\u002FMakazhanAlpamys\u002F65fdc943f85f3b2c46ecddb415c2b779\u002Fraw\u002Fsoup_tests.json\" alt=\"Tests\">\u003C\u002Fa>\n  \u003Ca href=\"https:\u002F\u002Fgithub.com\u002FMakazhanAlpamys\u002FSoup\u002Factions\">\u003Cimg src=\"https:\u002F\u002Fgithub.com\u002FMakazhanAlpamys\u002FSoup\u002Factions\u002Fworkflows\u002Fci.yml\u002Fbadge.svg\" alt=\"CI\">\u003C\u002Fa>\n  \u003Ca href=\"https:\u002F\u002Ftrysoup.dev\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002Fwebsite-trysoup.dev-blue\" alt=\"Website\">\u003C\u002Fa>\n  \u003Ca href=\"https:\u002F\u002Fdiscord.gg\u002F8RgVbFA6Zq\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FDiscord-join-5865F2?logo=discord&logoColor=white\" alt=\"Discord\">\u003C\u002Fa>\n  \u003Ca href=\"https:\u002F\u002Fdoi.org\u002F10.5281\u002Fzenodo.21771064\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FDOI-10.5281%2Fzenodo.21771064-blue?logo=zenodo&logoColor=white\" alt=\"DOI: 10.5281\u002Fzenodo.21771064\">\u003C\u002Fa>\n\u003C\u002Fp>\n\n\u003Cp align=\"center\">\n  \u003Ca href=\"https:\u002F\u002Fwww.producthunt.com\u002Fproducts\u002Fsoup-cli?embed=true&amp;utm_source=badge-featured&amp;utm_medium=badge&amp;utm_campaign=badge-soup-cli\">\n    \u003Cpicture>\n      \u003Csource media=\"(prefers-color-scheme: dark)\" srcset=\"https:\u002F\u002Fapi.producthunt.com\u002Fwidgets\u002Fembed-image\u002Fv1\u002Ffeatured.svg?post_id=1217869&amp;theme=dark\">\n      \u003Cimg src=\"https:\u002F\u002Fapi.producthunt.com\u002Fwidgets\u002Fembed-image\u002Fv1\u002Ffeatured.svg?post_id=1217869&amp;theme=light\" alt=\"Soup CLI - Fine-tune an 8B LLM on a 4 GB laptop GPU | Product Hunt\" width=\"250\" height=\"54\">\n    \u003C\u002Fpicture>\n  \u003C\u002Fa>\n\u003C\u002Fp>\n\n---\n\nSoup turns the pain of LLM fine-tuning into a simple workflow. One config, one command, done.\n\n```bash\npip install \"soup-cli[train]\"   # add [train] to fine-tune; bare `soup-cli` is the light CLI\nsoup init --template chat\nsoup train\n```\n\n**Fine-tune an 8B model on a 4 GB laptop GPU.** Layer streaming keeps the frozen base out of\nVRAM and feeds it to the GPU one decoder layer at a time. Measured on an RTX 3050 Laptop 4 GB:\nLlama-3.1-8B-Instruct + NF4 at **119.6 tok\u002Fs, 3.32 GB peak** — bit-exact against a normal\nresident run, and reproduced independently on an H100 at 113.00 tok\u002Fs in the same 3.32 GB.\n(The tok\u002Fs figure was measured on v0.72.2, before the v0.73.0 correctness repair that cost\n−4.8% at 32B; it has not been re-run on a 4 GB card since.) Opt-in (`stream_layers: true`)\nand still BETA —\n[how it works](docs\u002Fperformance-and-quantization.md#layer-streaming-beta-v0720-nf4-v0722-disk--wider-archs-v0723-preference-losses-v0724) ·\n[all measurements](benchmarks\u002F) · [paper](https:\u002F\u002Fdoi.org\u002F10.5281\u002Fzenodo.21771064) ·\n**[check it yourself on a free Colab T4](notebooks\u002Fproof-4gb.ipynb)** (caps the process to\n4 GB, then asserts a streamed model is bit-identical to a normal one)\n\n\u003Cp align=\"center\">\n  \u003Ca href=\"https:\u002F\u002Fyoutu.be\u002FT1LCErE943E\">\u003Cimg src=\"docs\u002Fassets\u002Flayer-streaming.gif\" alt=\"soup train pre-flight for Llama-3.1-8B on a 4 GB card: a 3.60 GB base store pinned in RAM across 32 layers and two 113 MB VRAM buffers, then a measured peak of 3.32 GB at 119.6 tok\u002Fs, stopping short of the 4 GB line\">\u003C\u002Fa>\u003Cbr>\n  \u003Csub>Llama-3.1-8B-Instruct + NF4, LoRA, batch 1, seq 512 on an RTX 3050 Laptop 4 GB — \u003Cb>3.32 GB peak, 119.6 tok\u002Fs\u003C\u002Fb>. \u003Ca href=\"https:\u002F\u002Fyoutu.be\u002FT1LCErE943E\">Full video (90s)\u003C\u002Fa>\u003C\u002Fsub>\n\u003C\u002Fp>\n\n## Why Soup?\n\nTraining LLMs is still painful. Even experienced teams spend 30-50% of their time fighting\ninfrastructure instead of improving models. Soup fixes that.\n\n- **Zero SSH.** Never SSH into a broken GPU box again.\n- **One config.** A simple YAML file is all you need.\n- **Auto everything.** Batch size, GPU detection, quantization — handled.\n- **Works locally.** Train on your own GPU with QLoRA. No cloud required.\n\n## What's New\n\n**v0.73.2 — the release gate stops lying in both directions.** `soup ship` answers one\nquestion: did this model get better, or did I break it? Two of its suites were ranking by\nthe wrong thing, and one whole failure direction had no detector at all.\n\n- **A suite scored 0.225 for a model that got it right 40\u002F40.** `mini_tool_call` was\n  ranking *brace hygiene*: the model emitted one closing brace short, so the parse fell\n  back to the inner object and the scorer rejected it for lacking the outer key. And\n  `mini_mmlu` scored Llama-3.1-8B at **0.423 — below a 0.5B** — because the extractor did\n  not know `\\boxed{C}` and the prompt never asked for a letter. Both fixed; 0.423 → 0.731.\n- **New: a benign-prompt axis.** Leg 2 flagged a *drop* in refusal rate and had no reverse,\n  so a tune that refuses everything read as a monotone safety improvement. Two models with\n  byte-identical scores on all seven shipped suites, one of which refuses every benign\n  request, were indistinguishable to the gate. `mini_over_refusal` is its mirror; paired\n  with the safety suite, neither can be gamed alone.\n- **New: `soup ship --noise-floor N`** re-runs the base model N times and refuses to call\n  any delta smaller than the measured spread significant. Greedy decoding is not\n  deterministic on GPU — same model, no adapter, five runs spread **0.015–0.020** against a\n  0.05 threshold, and four of six paired deltas in that session sat inside the floor. It\n  **sizes** the effect; it does not calibrate a threshold, and the release says so.\n- **A caller error was indistinguishable from a regression.** A non-callable generator\n  scored `0.0` on three suites and raised on the others — and in leg 2 a 0.0 reads as\n  \"failed every item\", i.e. it failed in the direction that looks like a finding.\n- Also: `soup data split --stratify-semantic` (#388) and `soup mcp serve --allow-execute`\n  (#391), both from outside contributors.\n\nThe measurement record for the previous release's VRAM work, published as written —\nincluding the **three readings withdrawn during it** — is\n[`benchmarks\u002Fgate-v0.73.1-measured-vram-fit.md`](benchmarks\u002Fgate-v0.73.1-measured-vram-fit.md).\n\n```yaml\n# soup.yaml — then just `soup train --config soup.yaml`\ntraining:\n  stream_layers: true      # base streams out of VRAM; only the adapter trains\n  quantization: 4bit       # NF4 — ~4x smaller store, so 8B fits a 4 GB card\n  batch_size: 4            # bigger batches amortise the weight read\n  stream_source: auto      # RAM when it fits, NVMe disk when it does not\n  seed: 1234               # new in v0.73.0\n```\n\n> Python **3.10–3.12** only. v0.73.0 adds the upper bound that was missing: on 3.13+, pip\n> used to resolve untested PyTorch wheels that crash in the native extension before Soup\n> runs at all.\n\n\u003Cdetails>\n\u003Csummary>Previous release — v0.72.4, align on a laptop (DPO \u002F ORPO \u002F SimPO \u002F KTO over layer streaming)\u003C\u002Fsummary>\n\nLayer streaming used to support supervised fine-tuning only; v0.72.4 opened it to the\npreference losses. The risk was one thing: DPO needs a reference model, and a second copy\nwould double memory and defeat the point. Soup uses *the same streamed base with its\nadapters switched off* — measured at **0.914×** the SFT peak, where forcing a real second\ninstance cost **+730 MB, exactly one copy of the weights**. Bit-exact against a normal\nnon-streamed run for all four. Honest cost: free in *memory*, not in *time* — DPO reads the\nlayer stack **1.52×** as often per step. `grpo` \u002F `ppo` stay excluded on purpose.\n\n> **Trained with `stream_layers: true` on v0.72.0?** That adapter is inert — its tensors were\n> saved under keys with an extra `.inner.` segment, so every loader returned the untuned base.\n> Fixed in v0.72.1; re-run or re-save. Check with:\n> `python -c \"from safetensors.torch import load_file; print([k for k in load_file('adapter_model.safetensors') if '.inner.' in k][:3])\"`\n\n\u003C\u002Fdetails>\n\n\u003Cdetails>\n\u003Csummary>Previous release — v0.71.40, soup reward synth (generate a reward verifier from your data)\u003C\u002Fsummary>\n\nPoint `soup reward synth` at a JSONL of reference outputs and it infers a deterministic verifier,\nwrites a readable \u002F committable `.py` reward function, and — the part nobody else does — *refuses* to\nemit one that can't tell your references from bad answers (four families: `numeric` \u002F `json_schema` \u002F\n`regex` \u002F `tool_call`; a mandatory calibration report is the moat). Reward ensembles\n(`reward_fn: \"accuracy,format\"`) also train now. (#311)\n\n```bash\nsoup reward synth references.jsonl -o reward.py --output-report calib.json\n```\n\n\u003C\u002Fdetails>\n\n\u003Cdetails>\n\u003Csummary>Previous release — v0.71.39, CI for weights not prompts (emit + provenance-bind the ship verdict)\u003C\u002Fsummary>\n\n`soup ship`'s verdict became emittable, committable, and provenance-bound: `--emit-evidence` makes a\nrun replay into an identical verdict, `eval.ship` in `soup.yaml` + `--config` makes the gate policy\nreviewable, and `--config` binds evidence to the exact recipe that produced it (stale evidence → exit 3).\n`soup ship --push owner\u002Frepo#N` posts the SHIP \u002F DON'T-SHIP card on the PR.\n\n\u003C\u002Fdetails>\n\n\u003Cdetails>\n\u003Csummary>Previous release — v0.71.38, The gate grows teeth (real leg-2 regression gate)\u003C\u002Fsummary>\n\n`soup ship`'s regression leg became real: a fixed, extraction-based scorer over seven bundled,\noffline suites (MCQ · arithmetic · tool-calling · JSON validity · safety\u002Frefusal). A tune that\nwins your task but quietly breaks tool-calling now gets a **DON'T SHIP**. Zero new deps.\n\n```bash\nsoup ship --base .\u002Fbase --adapter .\u002Fmy-lora --task-eval my_task.jsonl\n#   exit 0 = SHIP · 2 = DON'T SHIP · 3 = bad flags · 1 = runtime error\n```\n\n\u003C\u002Fdetails>\n\nFull history: [CHANGELOG.md](CHANGELOG.md) &middot; [GitHub Releases](https:\u002F\u002Fgithub.com\u002FMakazhanAlpamys\u002FSoup\u002Freleases).\n\n## Quick Start\n\n### 1. Install\n\n```bash\n# Light core: CLI + config + data tools, no PyTorch\npip install soup-cli\n\n# Add the training stack (torch, transformers, peft, trl, datasets, …)\npip install \"soup-cli[train]\"\n\n# Everything (train + serve + ui + data) in one shot\npip install \"soup-cli[all]\"\n\n# Or from GitHub (latest dev)\npip install git+https:\u002F\u002Fgithub.com\u002FMakazhanAlpamys\u002FSoup.git\n```\n\nThe full extras table (`fast`, `mlx`, `serve`, `eval`, `ui`, `vision`, `audio`, …) lives in\n[`docs\u002Fmodels.md`](docs\u002Fmodels.md#optional-extras).\n\n> **Double quotes, not single.** `\"soup-cli[train]\"` is the only spelling that works in every\n> shell — `cmd.exe`, PowerShell, bash and zsh. If you copied `'soup-cli[train]'` from an older\n> tutorial and pip rejected it, that is the reason:\n> [why, and the exact error](docs\u002Fmodels.md#quoting-the-extra).\n\n`soup init`, `soup data …`, and the other data\u002Finspection commands work on the light install.\nFine-tuning (`soup train`) needs the `[train]` extra.\n\n### 2. Create a config\n\n```bash\nsoup init                       # interactive wizard\nsoup init --template chat       # or start from a template\n```\n\nTemplates: `chat`, `code`, `tool-calling`, `medical`, `reasoning`, `vision`, `kto`, `orpo`,\n`simpo`, `ipo`, `bco`, `rlhf`, `pretrain`, `moe`, `longcontext`, `embedding`, `audio`.\n\n### 3. Train, test, ship\n\n```bash\nsoup train --config soup.yaml                 # LoRA, quantization, batching — all handled\nsoup chat  --model .\u002Foutput                    # talk to your model\nsoup push  --model .\u002Foutput --repo you\u002Fmy-model\n\nsoup merge  --adapter .\u002Foutput                              # merge LoRA into the base\nsoup export --model .\u002Foutput --format gguf --quant q4_k_m   # GGUF for Ollama \u002F llama.cpp\n```\n\nMore export targets (ONNX, TensorRT, AWQ, GPTQ, BitNet) and deployment options live in\n[`docs\u002Fserving-and-export.md`](docs\u002Fserving-and-export.md).\n\n## Configuration\n\nA complete `soup.yaml`:\n\n```yaml\nbase: meta-llama\u002FLlama-3.1-8B-Instruct\ntask: sft\n# backend: unsloth  # 2-5x faster, pip install \"soup-cli[fast]\"\n\ndata:\n  train: .\u002Fdata\u002Ftrain.jsonl\n  format: alpaca\n  val_split: 0.1\n\ntraining:\n  epochs: 3\n  lr: 2e-5\n  batch_size: auto\n  lora:\n    r: 64\n    alpha: 16\n  quantization: 4bit\n\noutput: .\u002Foutput\n```\n\n`config\u002Fschema.py` is the single source of truth for every field. Advanced data, training,\nand PEFT options are documented under [Documentation](#documentation).\n\n## Documentation\n\nThe full feature reference lives in [`docs\u002F`](docs\u002F). Start here:\n\n| Guide | Covers |\n|---|---|\n| [Training tasks & methods](docs\u002Ftraining.md) | SFT, DPO\u002FGRPO\u002FPPO\u002FKTO\u002FORPO\u002FSimPO\u002FIPO\u002FBCO, tool-calling, PRM, pre-training, distillation, classification, vision\u002Faudio\u002FTTS, unlearning, RAFT\u002FRA-DIT, loop-hardening detectors |\n| [PEFT, long context & efficiency](docs\u002Fpeft-and-efficiency.md) | DoRA, LoRA+, rsLoRA, VeRA, OLoRA, NEFTune, PiSSA, ReLoRA, optimizer & PEFT zoo, LLaMA Pro, GaLore, YaRN\u002FLongLoRA, packing, curriculum, auto-tuning |\n| [Performance & quantization](docs\u002Fperformance-and-quantization.md) | QAT, FP8, Quant Menu (I + II), KV-cache, NVFP4, save formats, Cut Cross-Entropy, gradient checkpointing, kernels, activation offloading, layer streaming, multi-GPU \u002F DeepSpeed \u002F FSDP |\n| [Data engineering](docs\u002Fdata.md) | Formats, the Axolotl\u002FLF-parity pipeline, data tools, synthetic generation & forge, quality scorecards, trace tooling, remote datasets, mixing, recipe DAGs |\n| [Evaluation & probes](docs\u002Fevaluation.md) | Eval design\u002Fgate, eval-gated training, benchmarks, NLG metrics, calibration, Elo arena, diagnose, post-train X-ray probes, A\u002FB, drift, tunability, `soup advise` |\n| [Serving & export](docs\u002Fserving-and-export.md) | OpenAI-compatible server, batch inference, benchmarking, merge\u002Fexport, Anthropic Messages endpoint, speculative decoding (train + measure your own draft), deploy autopilot, Web UI, Agent Forge |\n| [Adapters, registry & governance](docs\u002Fadapters-and-governance.md) | Adapter lifecycle\u002Fmanagement, model registry, Soup Cans, the data flywheel (`soup loop`), knowledge editing, steering, supply-chain controls (scan\u002Fsign\u002FBOM\u002Fattest\u002Faudit\u002Fairgap) |\n| [Compliance & governance quickstart](docs\u002Fcompliance.md) | HIPAA\u002FSOC2\u002FEU-AI-Act\u002FSR-11-7 `init` templates, provenance (BOM\u002Fattest\u002Frepro-receipt), audit log, air-gap, model-card autogen (`soup card`), CI gate (`soup ci init`) |\n| [Backends, platform & ops](docs\u002Fbackends-and-ops.md) | MLX\u002FUnsloth backends, alternative hubs, HF Hub integration, autopilot, experiment tracking, plan\u002Fapply, env lockfiles, hardware-fit, completions, plugins, utility commands |\n| [Command reference](docs\u002Fcommands.md) | The full `soup` command list |\n| [Supported models & extras](docs\u002Fmodels.md) | Recommended model families, the VRAM size guide, the pip extras matrix |\n\n## Data Formats\n\nAlpaca, ShareGPT, ChatML, preference pairs (DPO \u002F ORPO \u002F SimPO \u002F IPO \u002F KTO), vision, audio,\nASR, plaintext, embedding, RAFT and more — all auto-detected from JSONL, JSON, CSV, Parquet or\nTXT, so in most cases you point `data.train` at a file and nothing else changes. Schemas with a\nworked example per format, plus the data pipeline (remote URIs, streaming, sharding,\ninterleaving, vocab expansion, document ingestion), are in\n[`docs\u002Fdata.md`](docs\u002Fdata.md#data-formats).\n\n## Common Commands\n\n```bash\nsoup train  --config soup.yaml        # train (SFT\u002FDPO\u002FGRPO\u002FPPO\u002FKTO\u002FORPO\u002FSimPO\u002FIPO\u002F...)\nsoup infer  --model .\u002Foutput --input prompts.jsonl   # batch inference\nsoup chat   --model .\u002Foutput          # interactive chat\nsoup serve  --model .\u002Foutput          # OpenAI-compatible API server\nsoup merge  --adapter .\u002Foutput        # merge LoRA into the base model\nsoup export --model .\u002Foutput --format gguf           # export for deployment\nsoup eval   benchmark --model .\u002Foutput               # evaluate\nsoup data   inspect .\u002Fdata\u002Ftrain.jsonl               # dataset stats\nsoup recipes list                     # 100+ ready-made model recipes\nsoup autopilot --model \u003Cid> --data d.jsonl --goal chat  # zero-config\nsoup doctor                           # check GPU \u002F deps \u002F environment\n```\n\nThe complete command list is in [`docs\u002Fcommands.md`](docs\u002Fcommands.md).\n\n## Supported Models\n\nSoup works with **any** text-generation model on the\n[HuggingFace Hub](https:\u002F\u002Fhuggingface.co\u002Fmodels?pipeline_tag=text-generation) — if it loads with\n`AutoModelForCausalLM`, it works, zero config changes. Llama 3.x\u002F4, Qwen 2.5\u002F3, Gemma 3, Mistral,\nMixtral, DeepSeek R1\u002FV3, Phi-4, and 100+ others ship as ready-made recipes (`soup recipes list`).\n\n| VRAM | Max model (QLoRA 4-bit) | Example |\n|---|---|---|\n| 8 GB | ~7B | Llama-3.1-8B, Mistral-7B |\n| 16 GB | ~14B | Phi-4-14B, Qwen2.5-14B |\n| 24 GB | ~34B | CodeLlama-34B, Yi-1.5-34B |\n| 48 GB | ~70B | Llama-3.3-70B |\n| 80 GB+ | 70B+ (full) or MoE | Mixtral-8x22B, DeepSeek-V3 |\n\nFull model + vision tables and the optional-extras matrix are in [`docs\u002Fmodels.md`](docs\u002Fmodels.md).\n\n## Docker\n\nRun Soup without installing CUDA or PyTorch locally (image published to GHCR on every release):\n\n```bash\ndocker pull ghcr.io\u002Fmakazhanalpamys\u002Fsoup:latest\ndocker run --gpus all -v $(pwd):\u002Fworkspace ghcr.io\u002Fmakazhanalpamys\u002Fsoup train --config soup.yaml\ndocker compose up   # or build locally\n```\n\n## Requirements\n\n- Python 3.10, 3.11 or 3.12 (those are the versions CI tests; 3.13+ is not supported yet\n  because the PyTorch stack has not been validated there)\n- GPU with CUDA (recommended), Apple Silicon (MPS), or CPU (experimental — very slow)\n- 8 GB+ VRAM for 7B models with QLoRA\n\nAll training tasks run on CPU for testing (quantization auto-disabled). Optional extras\n(`train`, `all`, `fast`, `vision`, `qat`, `serve`, `serve-fast`, `ui`, `eval`, `deepspeed`,\n`liger`, `mlx`, `onnx`, `tensorrt`, …) are listed in\n[`docs\u002Fmodels.md`](docs\u002Fmodels.md#optional-extras).\n\n## Troubleshooting\n\n```bash\nsoup doctor    # GPU, system resources, dependencies, and version in one place\n```\n\n- **`ImportError: DLL load failed while importing _C` (Windows)** — reinstall PyTorch for your\n  CUDA version: `pip install torch --index-url https:\u002F\u002Fdownload.pytorch.org\u002Fwhl\u002Fcu121`.\n- **`soup version` ≠ `pip show soup-cli`** — multiple Python installs; use a virtualenv.\n\n## Development\n\n```bash\ngit clone https:\u002F\u002Fgithub.com\u002FMakazhanAlpamys\u002FSoup.git\ncd Soup\npip install -e \".[dev]\"\n\nruff check src\u002Fsoup_cli\u002F tests\u002F    # lint\npytest tests\u002F -v                   # unit tests (fast, no GPU)\npytest tests\u002F -m smoke -v          # smoke tests (downloads a tiny model, trains)\n\npre-commit install                 # optional: ruff lint+format on commit\n```\n\nSee [CONTRIBUTING.md](CONTRIBUTING.md) for the full workflow and [SECURITY.md](SECURITY.md) to\nreport a vulnerability.\n\n## Support Soup\n\nSoup is Apache-2.0 and free — and stays that way. It is built and maintained in the open on a\nsingle 4 GB laptop, which is why every performance number in these docs is measured rather than\nclaimed.\n\nIf Soup saved you a training run, [starring the repo](https:\u002F\u002Fgithub.com\u002FMakazhanAlpamys\u002FSoup)\nhelps most, and it costs nothing. If you would like to fund the work directly:\n\n**[❤️ Donate](https:\u002F\u002Fbuy.stripe.com\u002F4gMcN441k3pha3T19ye7m04)** — one-off, any amount (use\n*Change amount* on the checkout page). Payments are processed by Stripe under the maintainer's\nregistered business, **MePlay, Inc.** — that name, not \"Soup\", is what appears on the checkout\npage and on your card statement.\n\nDonations buy GPU time for the hardware-gated work — multi-GPU, 8B+ validation, Apple Silicon —\nthat a single 4 GB laptop cannot reach.\n\nThe other way to move exactly those items is **hardware itself**. They ship behind honest\n\"requires \\\u003Chardware\\>\" gates rather than unverified claims, so if you have access to a bigger\nbox — or GPU credits going unused — running one of the\n[`help wanted`](https:\u002F\u002Fgithub.com\u002FMakazhanAlpamys\u002FSoup\u002Fissues?q=is%3Aissue+is%3Aopen+label%3A%22help+wanted%22)\nissues and posting the numbers helps as much as funding the GPU time would. Those issues say\nexactly what is blocked on hardware today.\n\n## Contributors\n\nBuilt by the community ❤️ — thank you to everyone who has contributed. See\n[CONTRIBUTORS.md](CONTRIBUTORS.md).\n\n[![Contributors](https:\u002F\u002Fcontrib.rocks\u002Fimage?repo=MakazhanAlpamys\u002FSoup)](https:\u002F\u002Fgithub.com\u002FMakazhanAlpamys\u002FSoup\u002Fgraphs\u002Fcontributors)\n\n## Contact\n\nBugs and feature requests belong in the\n[issue tracker](https:\u002F\u002Fgithub.com\u002FMakazhanAlpamys\u002FSoup\u002Fissues), questions in\n[Discussions](https:\u002F\u002Fgithub.com\u002FMakazhanAlpamys\u002FSoup\u002Fdiscussions) — both get answered faster\nand help the next person with the same problem.\n\nFor live chat, setup help, and everything that reads better as a conversation, join the\n[Discord](https:\u002F\u002Fdiscord.gg\u002F8RgVbFA6Zq). Anything that should still be findable in six months\nbelongs in Issues or Discussions — a Discord answer helps one person, an issue helps everyone\nwho hits the same thing. The [Code of Conduct](CODE_OF_CONDUCT.md) applies there too.\n\nFor anything that does not fit in public — security reports (see [SECURITY.md](SECURITY.md)),\nCode of Conduct matters, or press — email **team@trysoup.dev**. That is the project address\nand the right one for anything Soup-related. **makazanalpamys@gmail.com** is the maintainer's\npersonal address; it reaches the same person and is a fine fallback.\n\n## Citing Soup\n\nLayer streaming — training an 8B model on a 4 GB laptop GPU by streaming the frozen base from\nhost RAM one decoder layer at a time — is described in a preprint, together with the correctness\nprotocol that verifies a streamed run against a resident one (forward and backward stated\nseparately, because they are two claims and not one).\n\n> Makazhan, A. (2026). *Exact Layer Streaming: LoRA Fine-Tuning of an 8B Model on a 4 GB Laptop\n> GPU* (v3). Zenodo. https:\u002F\u002Fdoi.org\u002F10.5281\u002Fzenodo.21918325\n\n**Version 3 (13 August 2026) is current.** The title and the claim are unchanged — 8B on 4 GB —\nand no measured number has changed since v1. What v3 does is **withdraw an explanation we had\npublished**, which is also the shortest way to describe what the paper is for:\n\n- **Retracted in v3: \"layer streaming is bound by host-to-device transfer, not by the GPU.\"**\n  That was an *inference* from the H100 replication below, and it had never been measured. We\n  measured it on 11 August and it is false at the published configuration: deleting every\n  host-to-device byte buys **1.4%**, the compute stream waits on a copy for **0.20%** of the\n  step, and the step runs at **71.3%** of that card's same-session GEMM ceiling. The largest\n  streaming-specific cost is the per-layer NF4 dequantisation, at 9.8%\n  ([the record](benchmarks\u002Fprobe-v0.73.0-what-bounds-streaming.md)). Every measurement stands;\n  the replication survives in a weaker form — the constraint is common to both machines and is\n  not the GPU's compute.\n- **Replication on hardware nothing like the original** (added in v2): 119.6 tok\u002Fs on the RTX\n  3050 against a median 113.00 on an H100, at the same 3.32 GB peak.\n- **A silent wrong-gradient defect, found and repaired.** On NF4 above ~165 MiB per layer the\n  forward stayed bit-exact and the loss curve looked healthy while the gradients were wrong. The\n  cause is named in the upstream library and reported there; the repair is gated against controls\n  on real 32B and 72B.\n- **Bit-exactness at real model sizes** instead of three-layer toys: forward from 0.5B to 72B,\n  backward at 8B and 14B.\n- **Trained-model quality, measured for the first time**, and indistinguishable from a resident run.\n- **A comparison against DeepSpeed** — including the result that does not flatter us: eight cards\n  of ZeRO-3 are slower than one card training resident.\n- **The limitations section rewritten**: of v1's ten items, one closed and four more narrowed,\n  and seven new ones added.\n\nCite the version you used. `10.5281\u002Fzenodo.21771064` is the concept DOI and always resolves to\nthe latest version (v3 today); v1 and v2 remain citable at their own version DOIs and are not\nedited — the retraction above is a new version precisely so that the record of what we claimed,\nand when, stays intact.\n\nThe measurement records behind every number in it are in [`benchmarks\u002F`](benchmarks\u002F), published\nas written — including the failures, the assumptions that turned out wrong, and the numbers that\nwere measured and then discarded.\n\n```bibtex\n@misc{makazhan2026exact,\n  title        = {Exact Layer Streaming: LoRA Fine-Tuning of an 8B Model on a 4 GB Laptop GPU},\n  author       = {Makazhan, Alpamys},\n  year         = {2026},\n  publisher    = {Zenodo},\n  version      = {v3},\n  doi          = {10.5281\u002Fzenodo.21918325},\n  url          = {https:\u002F\u002Fdoi.org\u002F10.5281\u002Fzenodo.21918325}\n}\n```\n\n## License\n\n[Apache-2.0](LICENSE). Copyright © the Soup contributors.\n","Soup 是一个轻量级命令行工具，用于在资源受限设备（如配备4GB显存的笔记本GPU）上高效微调大语言模型。其核心采用层流式（layer streaming）技术，在训练时按需加载冻结的基础模型层，显著降低显存占用；支持通过单个YAML配置文件定义训练流程，并提供一键式命令（如 soup train）完成数据准备、微调与导出。适用于个人开发者、教育场景及边缘端LLM定制化需求，无需SSH或复杂环境配置。",2,"trending"]