[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"project-9687":3},{"id":4,"name":5,"fullName":6,"owner":7,"repo":5,"description":8,"homepage":9,"htmlUrl":10,"language":11,"languages":10,"totalLinesOfCode":10,"stars":12,"forks":13,"watchers":14,"openIssues":15,"contributorsCount":16,"subscribersCount":16,"size":16,"stars1d":17,"stars7d":18,"stars30d":19,"stars90d":16,"forks30d":16,"starsTrendScore":20,"compositeScore":21,"rankGlobal":10,"rankLanguage":10,"license":22,"archived":23,"fork":23,"defaultBranch":24,"hasWiki":23,"hasPages":23,"topics":25,"createdAt":10,"pushedAt":10,"updatedAt":36,"readmeContent":37,"aiSummary":38,"trendingCount":16,"starSnapshotCount":16,"syncStatus":39,"lastSyncTime":40,"discoverSource":41},9687,"LoRA","microsoft\u002FLoRA","microsoft","Code for loralib, an implementation of \"LoRA: Low-Rank Adaptation of Large Language Models\"","https:\u002F\u002Farxiv.org\u002Fabs\u002F2106.09685",null,"Python",13590,907,77,80,0,1,12,79,8,43.87,"MIT License",false,"main",[26,27,28,29,30,31,32,33,34,35],"adaptation","deberta","deep-learning","gpt-2","gpt-3","language-model","lora","low-rank","pytorch","roberta","2026-06-12 02:02:11","# LoRA: Low-Rank Adaptation of Large Language Models\n\nThis repo contains the source code of the Python package `loralib` and several examples of how to integrate it with PyTorch models, such as those in Hugging Face.\nWe only support PyTorch for now.\nSee our paper for a detailed description of LoRA.\n\n**LoRA: Low-Rank Adaptation of Large Language Models** \u003Cbr>\n*Edward J. Hu\\*, Yelong Shen\\*, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen* \u003Cbr>\nPaper: https:\u002F\u002Farxiv.org\u002Fabs\u002F2106.09685 \u003Cbr>\nVideo explainer: https:\u002F\u002Fwww.youtube.com\u002Fwatch?v=DhRoTONcyZE \u003Cbr>\n\n*Update 2\u002F2023: LoRA is now supported by the [State-of-the-art Parameter-Efficient Fine-Tuning (PEFT)](https:\u002F\u002Fgithub.com\u002Fhuggingface\u002Fpeft) library by Hugging Face.*\n\nLoRA reduces the number of trainable parameters by learning pairs of rank-decompostion matrices while freezing the original weights.\nThis vastly reduces the storage requirement for large language models adapted to specific tasks and enables efficient task-switching during deployment all without introducing inference latency.\nLoRA also outperforms several other adaptation methods including adapter, prefix-tuning, and fine-tuning.\n\nWe obtain result comparable or superior to full finetuning on the GLUE benchmark using [RoBERTa (Liu et al., 2019)](https:\u002F\u002Farxiv.org\u002Fabs\u002F1907.11692) base and large and [DeBERTa (He et al., 2020)](https:\u002F\u002Farxiv.org\u002Fabs\u002F2006.03654) XXL 1.5B, while only training and storing a fraction of the parameters. Click the numbers below to download the RoBERTa and DeBERTa LoRA checkpoints.\n\n|   |         | RoBERTa base \u003Cbr> Fine-tune  |  RoBERTa base \u003Cbr> LoRA  | DeBERTa XXL \u003Cbr> Fine-tune | DeBERTa XXL \u003Cbr> LoRA  |\n|---|-------------------------|----------------|--------------------------|-----------------|-----------------|\n|   | # of Trainable Params.  | 125M | 0.8M | 1.5B | 4.7M     |\n|   | MNLI (m-Acc\u002Fmm-Acc)     | \u003Cb>87.6\u003C\u002Fb> | [\u003Cb>87.5\u003C\u002Fb>±.3\u002F86.9±.3](https:\u002F\u002Fgithub.com\u002Fmicrosoft\u002FLoRA\u002Freleases\u002Fdownload\u002FRoBERTa-base\u002Froberta_base_lora_mnli.bin) |91.7\u002F\u003Cb>91.9\u003C\u002Fb>| [\u003Cb>91.9\u003C\u002Fb>±.1\u002F\u003Cb>91.9\u003C\u002Fb>±.2](https:\u002F\u002Fgithub.com\u002Fmicrosoft\u002FLoRA\u002Freleases\u002Fdownload\u002FDeBERTa\u002Fdeberta_v2_xxlarge_lora_mnli.bin)       |\n|   | SST2 (Acc)              | 94.8 | [\u003Cb>95.1\u003C\u002Fb>±.2](https:\u002F\u002Fgithub.com\u002Fmicrosoft\u002FLoRA\u002Freleases\u002Fdownload\u002FRoBERTa-base\u002Froberta_base_lora_sst2.bin) | \u003Cb>97.2\u003C\u002Fb>    | [96.9±.2](https:\u002F\u002Fgithub.com\u002Fmicrosoft\u002FLoRA\u002Freleases\u002Fdownload\u002FDeBERTa\u002Fdeberta_v2_xxlarge_lora_sst2.bin)                    |\n|   | MRPC (Acc)              | \u003Cb>90.2\u003C\u002Fb> | [\u003Cb>89.7\u003C\u002Fb>±.7](https:\u002F\u002Fgithub.com\u002Fmicrosoft\u002FLoRA\u002Freleases\u002Fdownload\u002FRoBERTa-base\u002Froberta_base_lora_mrpc.bin) | 92.0           | [\u003Cb>92.6\u003C\u002Fb>±.6](https:\u002F\u002Fgithub.com\u002Fmicrosoft\u002FLoRA\u002Freleases\u002Fdownload\u002FDeBERTa\u002Fdeberta_v2_xxlarge_lora_mrpc.bin)             |\n|   | CoLA (Matthew's Corr)   | \u003Cb>63.6\u003C\u002Fb> | [\u003Cb>63.4\u003C\u002Fb>±1.2](https:\u002F\u002Fgithub.com\u002Fmicrosoft\u002FLoRA\u002Freleases\u002Fdownload\u002FRoBERTa-base\u002Froberta_base_lora_cola.bin) | \u003Cb>72.0\u003C\u002Fb>    | [\u003Cb>72.4\u003C\u002Fb>±1.1](https:\u002F\u002Fgithub.com\u002Fmicrosoft\u002FLoRA\u002Freleases\u002Fdownload\u002FDeBERTa\u002Fdeberta_v2_xxlarge_lora_cola.bin)           |\n|   | QNLI (Acc)              | 92.8 | [\u003Cb>93.3\u003C\u002Fb>±.3](https:\u002F\u002Fgithub.com\u002Fmicrosoft\u002FLoRA\u002Freleases\u002Fdownload\u002FRoBERTa-base\u002Froberta_base_lora_qnli.bin) | \u003Cb>96.0\u003C\u002Fb>    | [\u003Cb>96.0\u003C\u002Fb>±.1](https:\u002F\u002Fgithub.com\u002Fmicrosoft\u002FLoRA\u002Freleases\u002Fdownload\u002FDeBERTa\u002Fdeberta_v2_xxlarge_lora_qnli.bin)            |\n|   | QQP (Acc)               | \u003Cb>91.9\u003C\u002Fb> | [90.8±.1](https:\u002F\u002Fgithub.com\u002Fmicrosoft\u002FLoRA\u002Freleases\u002Fdownload\u002FRoBERTa-base\u002Froberta_base_lora_qqp.bin) | 92.7           | [\u003Cb>92.9\u003C\u002Fb>±.1](https:\u002F\u002Fgithub.com\u002Fmicrosoft\u002FLoRA\u002Freleases\u002Fdownload\u002FDeBERTa\u002Fdeberta_v2_xxlarge_lora_qqp.bin)           |\n|   | RTE (Acc)               | 78.7 | [\u003Cb>86.6\u003C\u002Fb>±.7](https:\u002F\u002Fgithub.com\u002Fmicrosoft\u002FLoRA\u002Freleases\u002Fdownload\u002FRoBERTa-base\u002Froberta_base_lora_rte.bin) | 93.9           | [\u003Cb>94.9\u003C\u002Fb>±.4](https:\u002F\u002Fgithub.com\u002Fmicrosoft\u002FLoRA\u002Freleases\u002Fdownload\u002FDeBERTa\u002Fdeberta_v2_xxlarge_lora_rte.bin)           |\n|   | STSB (Pearson\u002FSpearman Corr) | 91.2 | [\u003Cb>91.5\u003C\u002Fb>±.2\u002F\u003Cb>91.3\u003C\u002Fb>±.2](https:\u002F\u002Fgithub.com\u002Fmicrosoft\u002FLoRA\u002Freleases\u002Fdownload\u002FRoBERTa-base\u002Froberta_base_lora_stsb.bin) |\u003Cb>92.9\u003C\u002Fb>\u002F92.6| [\u003Cb>93.0\u003C\u002Fb>±.2\u002F\u003Cb>92.9\u003C\u002Fb>±.3](https:\u002F\u002Fgithub.com\u002Fmicrosoft\u002FLoRA\u002Freleases\u002Fdownload\u002FDeBERTa\u002Fdeberta_v2_xxlarge_lora_stsb.bin)      |\n|   | Average  | 86.40 | \u003Cb>87.24\u003C\u002Fb> | 91.06 | \u003Cb>91.32\u003C\u002Fb> |\n\n\u003Ci>Note: You still need the original pre-trained checkpoint from [Hugging Face](https:\u002F\u002Fhuggingface.co\u002F) to use the LoRA checkpoints.\u003C\u002Fi>\n\nFine-tuning numbers are taken from [Liu et al. (2019)](https:\u002F\u002Farxiv.org\u002Fabs\u002F1907.11692) and [He et al. (2020)](https:\u002F\u002Farxiv.org\u002Fabs\u002F2006.03654).  We include confidence intervals on results from our experiments. Please follow the instructions in `examples\u002FNLU\u002F` to reproduce our results.\n\nOn GPT-2, LoRA compares favorably to both full finetuning and other efficient tuning methods, such as [adapter (Houlsby et al., 2019)](https:\u002F\u002Farxiv.org\u002Fabs\u002F1902.00751) and [prefix tuning (Li and Liang, 2021)](https:\u002F\u002Farxiv.org\u002Fabs\u002F2101.00190). We evaluated on E2E NLG Challenge, DART, and WebNLG:\n\n|   | Method              | # of Trainable Params | E2E (BLEU)   | DART (BLEU)  | WebNLG (BLEU-U\u002FS\u002FA)            |\n|---|---------------------|-----------------------|--------------|--------------|--------------------------------|\n|   | GPT-2 M (Fine-Tune) | 354.92M               | 68.2         | 46.0         | 30.4\u002F\u003Cb>63.2\u003C\u002Fb>\u002F47.6          |\n|   | GPT-2 M (Adapter)   | 0.37M                 | 66.3         | 42.4         | 45.1\u002F54.5\u002F50.2                 |\n|   | GPT-2 M (Prefix)    | 0.35M                 | 69.7         | 45.7         | 44.1\u002F63.1\u002F54.4                 |\n|   | GPT-2 M (LoRA)      | 0.35M                 |\u003Cb>70.4\u003C\u002Fb>±.1|\u003Cb>47.1\u003C\u002Fb>±.2| \u003Cb>46.7\u003C\u002Fb>±.4\u002F62.1±.2\u002F\u003Cb>55.3\u003C\u002Fb>±.2 |\n|   | GPT-2 L (Fine-Tune) | 774.03M               | 68.5         | 46.5         | 41.7\u002F\u003Cb>64.6\u003C\u002Fb>\u002F54.2          |\n|   | GPT-2 L (Adapter)   | 0.88M                 | 69.1±.1      | 45.7±.1      | \u003Cb>49.8\u003C\u002Fb>±.0\u002F61.1±.0\u002F56.0±.0 |\n|   | GPT-2 L (Prefix)    | 0.77M                 | 70.3         | 46.5         | 47.0\u002F64.2\u002F56.4                 |\n|   | GPT-2 L (LoRA)      | 0.77M                 |\u003Cb>70.4\u003C\u002Fb>±.1|\u003Cb>47.5\u003C\u002Fb>±.1| 48.4±.3\u002F\u003Cb>64.0\u003C\u002Fb>±.3\u002F\u003Cb>57.0\u003C\u002Fb>±.1 |\n\nNon-LoRA baselines, except for adapter on GPT-2 large, are taken from [Li and Liang (2021)](https:\u002F\u002Farxiv.org\u002Fabs\u002F2101.00190). We include confidence intervals on results from our experiments.\n\nDownload the GPT-2 LoRA checkpoints:\n * [GPT-2 Medium E2E](https:\u002F\u002Fgithub.com\u002Fmicrosoft\u002FLoRA\u002Freleases\u002Fdownload\u002FGPT-2\u002Fgpt2_md_lora_e2e.pt) (1.5 MB)\n * [GPT-2 Medium DART](https:\u002F\u002Fgithub.com\u002Fmicrosoft\u002FLoRA\u002Freleases\u002Fdownload\u002FGPT-2\u002Fgpt2_md_lora_dart.pt) (1.5 MB)\n * [GPT-2 Medium WebNLG](https:\u002F\u002Fgithub.com\u002Fmicrosoft\u002FLoRA\u002Freleases\u002Fdownload\u002FGPT-2\u002Fgpt2_md_lora_webnlg.pt) (1.5 MB)\n * [GPT-2 Large E2E](https:\u002F\u002Fgithub.com\u002Fmicrosoft\u002FLoRA\u002Freleases\u002Fdownload\u002FGPT-2\u002Fgpt2_lg_lora_e2e.pt) (2.3 MB)\n * [GPT-2 Large DART](https:\u002F\u002Fgithub.com\u002Fmicrosoft\u002FLoRA\u002Freleases\u002Fdownload\u002FGPT-2\u002Fgpt2_lg_lora_dart.pt) (2.3 MB)\n * [GPT-2 Large WebNLG](https:\u002F\u002Fgithub.com\u002Fmicrosoft\u002FLoRA\u002Freleases\u002Fdownload\u002FGPT-2\u002Fgpt2_lg_lora_webnlg.pt) (2.3 MB)\n\nPlease follow the instructions in `examples\u002FNLG\u002F` to reproduce our result.\n## Repository Overview\n\n\u003Ci>(The initial release of this repo has been archived in the branch \"snapshot-9-15-2021\")\u003C\u002Fi>\n\nThere are several directories in this repo:\n* [loralib\u002F](loralib) contains the source code for the package `loralib`, which needs to be installed to run the examples we provide;\n* [examples\u002FNLG\u002F](examples\u002FNLG) contains an example implementation of LoRA in GPT-2 using our package, which can be used to reproduce the result in our paper;\n* [examples\u002FNLU\u002F](examples\u002FNLU) contains an example implementation of LoRA in RoBERTa and DeBERTa using our package, which produces competitive results on the GLUE benchmark;\n* See how we use `loralib` in [GPT-2](examples\u002FNLG\u002Fsrc\u002Fmodel.py), [RoBERTa](examples\u002FNLU\u002Fsrc\u002Ftransformers\u002Fmodels\u002Froberta\u002Fmodeling_roberta.py), and [DeBERTa v2](examples\u002FNLU\u002Fsrc\u002Ftransformers\u002Fmodels\u002Fdeberta_v2\u002Fmodeling_deberta_v2.py)\n\n## Quickstart\n\n 1. Installing `loralib` is simply\n ```bash\n pip install loralib\n # Alternatively\n # pip install git+https:\u002F\u002Fgithub.com\u002Fmicrosoft\u002FLoRA\n ```\n\n 2. You can choose to adapt some layers by replacing them with counterparts implemented in `loralib`. We only support `nn.Linear`, `nn.Embedding`, and `nn.Conv2d` for now. We also support a `MergedLinear` for cases where a single `nn.Linear` represents more than one layers, such as in some implementations of the attention `qkv` projection (see Additional Notes for more).\n ```python\n # ===== Before =====\n # layer = nn.Linear(in_features, out_features)\n\n # ===== After ======\n import loralib as lora\n # Add a pair of low-rank adaptation matrices with rank r=16\n layer = lora.Linear(in_features, out_features, r=16)\n ```\n\n 3. Before the training loop begins, mark only LoRA parameters as trainable.\n ```python\n import loralib as lora\n model = BigModel()\n # This sets requires_grad to False for all parameters without the string \"lora_\" in their names\n lora.mark_only_lora_as_trainable(model)\n # Training loop\n for batch in dataloader:\n    ...\n ```\n 4. When saving a checkpoint, generate a `state_dict` that only contains LoRA parameters.\n ```python\n # ===== Before =====\n # torch.save(model.state_dict(), checkpoint_path)\n # ===== After =====\n torch.save(lora.lora_state_dict(model), checkpoint_path)\n ```\n 5. When loading a checkpoint using `load_state_dict`, be sure to set `strict=False`.\n ```python\n # Load the pretrained checkpoint first\n model.load_state_dict(torch.load('ckpt_pretrained.pt'), strict=False)\n # Then load the LoRA checkpoint\n model.load_state_dict(torch.load('ckpt_lora.pt'), strict=False)\n ```\n\n#### Now training can proceed as usual.\n\n## Additional Notes\n\n1. While we focus on a simple yet effect setup, namely adapting only the `q` and `v` projection in a Transformer, in our examples, LoRA can be apply to any subsets of pre-trained weights. We encourage you to explore different configurations, such as adapting the embedding layer by replacing `nn.Embedding` with `lora.Embedding` and\u002For adapting the MLP layers. It's very likely that the optimal configuration varies for different model architectures and tasks.\n\n2. Some Transformer implementation uses a single `nn.Linear` for the projection matrices for query, key, and value. If one wishes to constrain the rank of the updates to the individual matrices, one has to either break it up into three separate matrices or use `lora.MergedLinear`. Make sure to modify the checkpoint accordingly if you choose to break up the layer.\n```python\n# ===== Before =====\n# qkv_proj = nn.Linear(d_model, 3*d_model)\n# ===== After =====\n# Break it up (remember to modify the pretrained checkpoint accordingly)\nq_proj = lora.Linear(d_model, d_model, r=8)\nk_proj = nn.Linear(d_model, d_model)\nv_proj = lora.Linear(d_model, d_model, r=8)\n# Alternatively, use lora.MergedLinear (recommended)\nqkv_proj = lora.MergedLinear(d_model, 3*d_model, r=8, enable_lora=[True, False, True])\n```\n3. Training bias vectors in tandem with LoRA might be a cost-efficient way to squeeze out extra task performance (if you tune the learning rate carefully). While we did not study its effect thoroughly in our paper, we make it easy to try in `lora`. You can mark some biases as trainable by passing \"all\" or \"lora_only\" to `bias=` when calling `mark_only_lora_as_trainable`. Remember to pass the corresponding `bias=` argument to `lora_state_dict` when saving a checkpoint.\n```python\n# ===== Before =====\n# lora.mark_only_lora_as_trainable(model) # Not training any bias vectors\n# ===== After =====\n# Training all bias vectors associated with modules we apply LoRA to \nlora.mark_only_lora_as_trainable(model, bias='lora_only')\n# Alternatively, we can train *all* bias vectors in the model, including LayerNorm biases\nlora.mark_only_lora_as_trainable(model, bias='all')\n# When saving a checkpoint, use the same bias= ('all' or 'lora_only')\ntorch.save(lora.lora_state_dict(model, bias='all'), checkpoint_path)\n```\n4. Calling `model.eval()` will trigger the merging of LoRA parameters with the corresponding pretrained ones, which eliminates additional latency for subsequent forward passes. Calling `model.train()` again will undo the merge. This can be disabled by passing `merge_weights=False` to LoRA layers.\n\n## Contact\nPlease contact us or post an issue if you have any questions.\n\nFor questions related to the package `loralib`:\n* Edward Hu (edward@edwardjhu.com)\n* Phillip Wallis (phwallis@microsoft.com)\n* Weizhu Chen (wzchen@microsoft.com)\n\nThe GPT-2 example:\n* Phillip Wallis (phwallis@microsoft.com)\n* Yelong Shen (yeshe@microsoft.com)\n\nThe RoBERTa\u002FDeBERTa example:\n* Lu Wang (luw@microsoft.com)\n\n## Acknowledgements\nWe thank in alphabetical order Jianfeng Gao, Jade Huang, Jiayuan Huang, Lisa Xiang Li, Xiaodong Liu, Yabin Liu, Benjamin Van Durme, Luis Vargas, Haoran Wei, Peter Welinder, and Greg Yang for providing valuable feedback.\n\n## Citation\n```BibTeX\n@inproceedings{\nhu2022lora,\ntitle={Lo{RA}: Low-Rank Adaptation of Large Language Models},\nauthor={Edward J Hu and Yelong Shen and Phillip Wallis and Zeyuan Allen-Zhu and Yuanzhi Li and Shean Wang and Lu Wang and Weizhu Chen},\nbooktitle={International Conference on Learning Representations},\nyear={2022},\nurl={https:\u002F\u002Fopenreview.net\u002Fforum?id=nZeVKeeFYf9}\n}\n```\n\n## Contributing\n\nThis project welcomes contributions and suggestions.  Most contributions require you to agree to a\nContributor License Agreement (CLA) declaring that you have the right to, and actually do, grant us\nthe rights to use your contribution. For details, visit https:\u002F\u002Fcla.opensource.microsoft.com.\n\nWhen you submit a pull request, a CLA bot will automatically determine whether you need to provide\na CLA and decorate the PR appropriately (e.g., status check, comment). Simply follow the instructions\nprovided by the bot. You will only need to do this once across all repos using our CLA.\n\nThis project has adopted the [Microsoft Open Source Code of Conduct](https:\u002F\u002Fopensource.microsoft.com\u002Fcodeofconduct\u002F).\nFor more information see the [Code of Conduct FAQ](https:\u002F\u002Fopensource.microsoft.com\u002Fcodeofconduct\u002Ffaq\u002F) or\ncontact [opencode@microsoft.com](mailto:opencode@microsoft.com) with any additional questions or comments.\n","LoRA是一个用于大型语言模型低秩适应的Python库。它通过学习一对秩分解矩阵来减少可训练参数的数量，同时冻结原始权重，从而大幅降低存储需求，并在不引入推理延迟的情况下实现高效的任务切换。该技术特别适用于需要对预训练模型进行特定任务微调但又受限于计算资源或希望保持模型轻量化的场景。基于PyTorch框架开发，支持包括RoBERTa和DeBERTa在内的多种模型，并已被Hugging Face整合进其PEFT库中。实验表明，在GLUE基准测试上，LoRA仅需训练少量参数即可达到甚至超越全量微调的效果。",2,"2026-06-11 03:24:11","top_topic"]