[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"project-93601":3},{"id":4,"name":5,"fullName":6,"owner":7,"repo":5,"description":8,"homepage":9,"htmlUrl":9,"language":10,"languages":9,"totalLinesOfCode":9,"stars":11,"forks":12,"watchers":13,"openIssues":14,"contributorsCount":15,"subscribersCount":15,"size":15,"stars1d":16,"stars7d":16,"stars30d":16,"stars90d":15,"forks30d":15,"starsTrendScore":17,"compositeScore":18,"rankGlobal":9,"rankLanguage":9,"license":19,"archived":20,"fork":20,"defaultBranch":21,"hasWiki":22,"hasPages":20,"topics":23,"createdAt":9,"pushedAt":9,"updatedAt":32,"readmeContent":33,"aiSummary":9,"trendingCount":15,"starSnapshotCount":15,"syncStatus":34,"lastSyncTime":35,"discoverSource":36},93601,"lora-speedrun","Saivineeth147\u002Flora-speedrun","Saivineeth147","Speedrunning LoRA fine-tuning: frozen task, frozen hardware, public wall-clock leaderboard. modded-nanogpt for fine-tuning.",null,"Python",134,7,3,6,0,19,57,74.11,"MIT License",false,"main",true,[24,25,26,27,28,29,30,31],"benchmark","fine-tuning","leaderboard","llm","lora","peft","qlora","speedrun","2026-07-22 04:02:09","# LoRA Speedrun 🏁\n\n[![CI](https:\u002F\u002Fgithub.com\u002FSaivineeth147\u002Flora-speedrun\u002Factions\u002Fworkflows\u002Fci.yml\u002Fbadge.svg)](https:\u002F\u002Fgithub.com\u002FSaivineeth147\u002Flora-speedrun\u002Factions\u002Fworkflows\u002Fci.yml)\n[![License: MIT](https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FLicense-MIT-yellow.svg)](.\u002FLICENSE)\n[![PRs Welcome](https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FPRs-welcome-brightgreen.svg)](.\u002FCONTRIBUTING.md)\n\n**How fast can you LoRA-fine-tune [Qwen2.5-1.5B](https:\u002F\u002Fhuggingface.co\u002FQwen\u002FQwen2.5-1.5B) to ≥ 57% on GSM8K — on a single L40S?**\n\nThis is [modded-nanogpt](https:\u002F\u002Fgithub.com\u002FKellerJordan\u002Fmodded-nanogpt) for fine-tuning: a\nfrozen task, frozen hardware, and a public leaderboard of wall-clock records. Every record is\nindependently re-run 3× with fresh seeds on identical hardware before it counts.\n\n**Attempting and verifying are free**: official timing runs on a Modal L40S sandbox, and\nModal's free monthly compute credits cover full runs — so anyone can compete, and anyone\ncan re-verify any record with one command.\n\n\u003Cp align=\"center\">\n  \u003Cimg src=\"docs\u002Fleaderboard.png\" alt=\"LoRA Speedrun leaderboard — 11:57 baseline to 1:44 current record (−86%) across four verified records\" width=\"760\">\n\u003C\u002Fp>\n\n## Leaderboard\n\n\u003C!-- LEADERBOARD:START — auto-generated by scripts\u002Fupdate_leaderboard.py, do not edit by hand -->\n### Track 1 — GSM8K · Qwen2.5-1.5B · target ≥ 57.0% · 1× L40S\n\n**Current record: 1m 44s** by [@stared](https:\u002F\u002Fgithub.com\u002Fstared) — Shortest-4k data pruning, 1 aggressive-LR epoch, \u003C\u003C...>> annotations stripped, custom GPU-resident packed loop, chunked completion-only CE (no full logits).\n\n| # | Date | Author | Train time | GSM8K\u002FEM | Δ | Technique |\n|---|------|--------|-----------|----------|---|-----------|\n| 0 | 2026-07-18 | [@Saivineeth147](https:\u002F\u002Fgithub.com\u002FSaivineeth147) | 11m 57s | 59.4% | — | Baseline: plain LoRA r=16 on all linear layers, 3 epochs, cosine LR. No tricks. ([report](records\u002Fverifications\u002F000-baseline.md)) |\n| 1 | 2026-07-18 | [@Saivineeth147](https:\u002F\u002Fgithub.com\u002FSaivineeth147) | 6m 05s | 61.1% | −49% | Sequence packing + completion-only loss masking, 2 epochs. Same LoRA config as #0; ~2x faster at higher accuracy. ([report](records\u002Fverifications\u002F001-saivineeth147.md)) |\n| 2 | 2026-07-20 | [@stared](https:\u002F\u002Fgithub.com\u002Fstared) | 1m 53s | 59.6% | −69% | One 4e-4 epoch over 3,000 examples with 2x optimizer updates (integrated-LR), 512-token packs. First outside record. ([report](records\u002Fverifications\u002F002-pmigdal.md)) |\n| 3 | 2026-07-20 | [@stared](https:\u002F\u002Fgithub.com\u002Fstared) | 1m 44s | 60.3% | −8% | Shortest-4k data pruning, 1 aggressive-LR epoch, \u003C\u003C...>> annotations stripped, custom GPU-resident packed loop, chunked completion-only CE (no full logits). ([report](records\u002Fverifications\u002F003-pmigdal.md)) |\n\n### Track 2 — SQuAD v1.1 · SmolLM2-1.7B · target ≥ 75.5% · 1× L40S\n\n**Current record: 11m 08s** by [@Saivineeth147](https:\u002F\u002Fgithub.com\u002FSaivineeth147) — Track 2 baseline: plain LoRA r=16 on SQuAD, first 20k examples, 1 epoch, full-sequence loss. No tricks.\n\n| # | Date | Author | Train time | GSM8K\u002FEM | Δ | Technique |\n|---|------|--------|-----------|----------|---|-----------|\n| 0 | 2026-07-20 | [@Saivineeth147](https:\u002F\u002Fgithub.com\u002FSaivineeth147) | 11m 08s | 77.5% | — | Track 2 baseline: plain LoRA r=16 on SQuAD, first 20k examples, 1 epoch, full-sequence loss. No tricks. ([report](records\u002Fverifications\u002Ft2-000-baseline.md)) |\n\n\u003C!-- LEADERBOARD:END -->\n\nFull history with verification reports: [records\u002FRECORDS.md](.\u002Frecords\u002FRECORDS.md) ·\nLive leaderboard: [huggingface.co\u002Fspaces\u002Fvineeth98\u002Flora-speedrun](https:\u002F\u002Fhuggingface.co\u002Fspaces\u002Fvineeth98\u002Flora-speedrun)\n\n## The tracks\n\nTwo frozen tracks, deliberately different model families and task types — so a technique\nonly proves general by winning on both. Same hardware, caps, and verification everywhere.\n\n| | Track 1 | Track 2 |\n|---|---|---|\n| **Base model** | `Qwen\u002FQwen2.5-1.5B` | `HuggingFaceTB\u002FSmolLM2-1.7B` |\n| **Task** | GSM8K math → ≥ **57.0%** exact-match | SQuAD v1.1 QA → ≥ **75.5% EM** |\n| **Training data** | GSM8K `train` split only | SQuAD `train` split only |\n| **Metric** | **Training wall-clock. Lower wins.** | same |\n| **Hardware** | 1× L40S (48 GB), Modal sandbox | same |\n| **Constraint** | adapter-only, ≤ 30M trainable params | same |\n\nMachine-readable specs: [`spec.yaml`](.\u002Fspec.yaml) · [`spec-t2.yaml`](.\u002Fspec-t2.yaml). Full rules: [TASK.md](.\u002FTASK.md).\n\nYou control everything else: LoRA rank and placement, quantization, learning-rate schedules,\nsequence packing, data subset selection and ordering, custom kernels, when to stop. Train on\n1,000 well-chosen examples for 90 seconds if you can make it clear the bar.\n\n## Why this exists\n\nI fine-tune small models on a budget, and I couldn't tell which speedup claims were real.\nEvery technique — DoRA, rsLoRA, Unsloth, packing tricks — reports numbers on different\nmodels, data, and hardware. In practice the claims are unfalsifiable.\n\nThe only format I've seen actually settle arguments like that is a frozen task with a\nreferee. That's what nanoGPT's speedrun did for pretraining optimizers (Muon came out of\nit). This is the same arena for fine-tuning.\n\nWhat it builds toward: a verified public record of which training tricks pay for\nthemselves in wall-clock and which don't survive replication. Every record has to explain\nits mechanism in its write-up, so the leaderboard doubles as a lab notebook — the\nrejected-variants sections are often as useful as the records. And because tricks can\noverfit one setup, records live on tracks with different model families and task types:\na technique only proves general by transferring.\n\n## Quickstart\n\n**Official-hardware run (free).** Make a [Modal](https:\u002F\u002Fmodal.com) account, then:\n\n```bash\ngit clone https:\u002F\u002Fgithub.com\u002FSaivineeth147\u002Flora-speedrun && cd lora-speedrun\npip install modal pyyaml && modal setup      # one-time browser auth\n\npython harness\u002Fmodal_verify.py --prefetch    # one-time: cache model + data in a volume\n\n# one timed, evaluated attempt of the baseline on the exact spec hardware:\npython harness\u002Fmodal_verify.py --submission submissions\u002F000-baseline --runs 1\n\n# full record-style verification (3 fresh seeds, all must pass):\npython harness\u002Fmodal_verify.py --submission submissions\u002F000-baseline --runs 3\n```\n\n**Local iteration (optional).** Any 24 GB+ card runs the baseline for fast experimenting —\n`bash scripts\u002Fsetup_gpu.sh`, then `python harness\u002Frun_submission.py submissions\u002F000-baseline --runs 1`.\nLocal times aren't official; the leaderboard clock is the Modal L40S.\n\nThen copy `submissions\u002FTEMPLATE\u002F`, make it faster, and open a PR. See [CONTRIBUTING.md](.\u002FCONTRIBUTING.md).\n\n## How records get verified\n\n\u003Cp align=\"center\">\n  \u003Cimg src=\"docs\u002Fverification-replay.gif\" alt=\"Terminal replay of a real record verification: training, integrity + adapter audit, GSM8K eval, and the 3-seed verdict\" width=\"760\">\n\u003C\u002Fp>\n\n\u003Cp align=\"center\">\u003Cem>A real verification, replayed from the logs (seed 463953844, time-compressed): train → integrity + adapter audit → eval → 3-seed verdict.\u003C\u002Fem>\u003C\u002Fp>\n\n1. You open a PR with your training script, config, notes, and self-reported numbers.\n2. CI statically validates it, and an automated Claude security screen reviews the diff\n   (exfiltration attempts, network use, harness tampering, test-set contact) and posts\n   its findings publicly.\n3. A maintainer reviews the code, then comments `\u002Fverify` — which re-runs your submission\n   **3× with fresh seeds** in a network-blocked Modal sandbox on the spec L40S. All 3 runs\n   must clear the target; official time is the mean.\n4. The harness audits the adapter param count and re-verifies model\u002Fdata content hashes\n   (anti-tampering), and the verification report is posted on the PR and committed to\n   [records\u002Fverifications\u002F](.\u002Frecords\u002Fverifications\u002F) with the accept\u002Freject reasoning.\n\nFull protocol, rubric, and threat model: [JUDGING.md](.\u002FJUDGING.md) · [SECURITY.md](.\u002FSECURITY.md).\n\n## Ideas nobody has claimed yet\n\n_Taken so far: sequence packing + completion-only masking (record #1)._\n\n1-epoch aggressive-LR schedules · data pruning (train on the hardest 2k examples?) ·\nblock-diagonal\u002Fvarlen packing attention · QLoRA NF4 vs bf16 tradeoff · rsLoRA \u002F DoRA \u002F\nPiSSA init · LoRA+ (asymmetric LR for A\u002FB) · NEFTune noise · curriculum ordering ·\nrank\u002Fplacement search (MLP-only vs attention-only) · `torch.compile` · Unsloth kernels ·\nLiger kernels · fused cross-entropy · smarter warmup for short runs\n\nClaim one, beat 6m 05s, get your name on the board.\n\n## FAQ\n\n**Is this LoRa the radio protocol?** No — LoRA ([Low-Rank Adaptation](https:\u002F\u002Fen.wikipedia.org\u002Fwiki\u002FLoRA_(machine_learning))),\nthe standard cheap way to fine-tune a language model: train a small adapter on top of a\nfrozen model. The competition: everyone fine-tunes the same model to the same score on\nthe same GPU, and the fastest verified training run holds the record.\n\n**Won't techniques overfit to one model + one task?** That's exactly why there are two\ntracks with different model families and task types — and more will follow the same\nfreeze-and-calibrate protocol. A trick that only wins on one track is a record, but the\ntechniques worth trusting are the ones that transfer. The track system makes that an\nempirical question instead of an argument.\n\n**Why Qwen2.5-1.5B? Isn't it pretrained on math?** Probably, like every modern base model.\nIt doesn't matter: the target is an *anchor*, not a claim about mathematical discovery. The\nrace is the interesting part — same reason nanoGPT speedrunning targets an arbitrary val loss.\n(Track 2 uses a different family, SmolLM2, partly for this reason.)\n\n**Why wall-clock instead of FLOPs or steps?** Because wall-clock is what you pay for, and it\nforces kernels, data loading, and algorithms to compete in the same currency. Same rule as\nmodded-nanogpt.\n\n**Why an L40S on Modal instead of a 4090 or H100?** Three reasons. It's one consistent\ndatacenter SKU, so times are actually comparable (rented consumer cards vary host-to-host).\nIt's free to use via Modal's monthly credits, so competing and re-verifying costs nothing.\nAnd submissions are strangers' code — Modal sandboxes run them network-blocked and\nsecretless. (The L40S is the same AD102 silicon as the 4090, so consumer-GPU tricks transfer.)\n\n**Can I train on other data \u002F distill from a bigger model?** No. GSM8K train split only,\nno teacher models, no synthetic data. See [TASK.md](.\u002FTASK.md) for the full banned list.\n\n**Multiple GPUs?** No. One L40S. That's the point.\n\n## License\n\nMIT. Records, reports, and write-ups are as public as the code.\n",2,"2026-07-21 02:30:09","CREATED_QUERY"]