QLoRA vs. LoRA: An Empirical Study of Performance–Efficiency Trade-offs in Fine-Tuning Mistral‑7B on Instruction‑Following Data
Abstract
Parameter‑Efficient Fine‑Tuning (PEFT) approaches have emerged as a key strategy for efficiently adapting Large Language Models (LLMs) in resource‑constrained environments. Within this family of methods, LoRA and its quantized extension QLoRA provide effective alternatives by substantially decreasing the number of trainable parameters required for adaptation. In this study, a systematic comparison of LoRA and QLoRA is presented using a controlled experimental setup with identical base models and training data, with all experiments repeated across three random seeds and statistical significance assessed via paired t‑tests. Both approaches are evaluated across multiple dimensions, including generation quality (ROUGE‑L, Exact Match, token‑level F1, and the semantic metric BERTScore), training efficiency, and GPU memory consumption. Baseline (untrained) model performance is reported to quantify the gain from fine‑tuning, and loss curves are provided to confirm convergence. The results show that QLoRA not only reduces memory usage by approximately 70% but also achieves superior performance compared to LoRA, with improvements in token‑level F1 and BERTScore F1 reaching statistical significance (p < 0.05), while other lexical metrics showed consistent but non‑significant improvements. Furthermore, the performance–efficiency trade‑offs are analyzed, and it is demonstrated that QLoRA offers a more favorable balance between computational cost and model quality. These findings highlight QLoRA as a highly effective and practical fine‑tuning strategy for resource‑constrained environments, supported by a reproducible, statistically robust evaluation pipeline.
Keywords
References
- Ben Zaken, E., Goldberg, Y., & Ravfogel, S. (2022). BitFit: Simple Parameter-efficient Fine-tuning for Transformer-based Masked Language-models. In: S. Muresan, P. Nakov, & A. Villavicencio (Eds.), Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Vol. 2, pp. 1–9), (22-27 May 2022), Dublin, Ireland. https://doi.org/10.18653/v1/2022.acl-short.1
- Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., … Amodei, D. (2020). Language models are few-shot learners. In: H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, H. Lin (Eds.), Proceedings of the 34th International Conference on Neural Information Processing Systems (NeurIPS 2020) (pp. 1877–1901), (6-12 December 2020), Vancouver BC Canada. https://doi.org/10.48550/arXiv.2005.14165
- Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tai, Y., Fedus, W., Li, Y., Wang, X., Dehghani, M., Brahma, S., Webson, A., Gu, S. S., Dai, Z., Suzgun, M., Chen, X., Chowdhery, A., Castro-Ros, A., Pellat, M., Robinson, K., … Wei, J. (2024). Scaling Instruction-Finetuned Language Models. Journal of Machine Learning Research, 25(1), 3381–3433. https://doi.org/10.48550/arXiv.2210.11416
- Dettmers, T., Lewis, M., Belkada, Y., & Zettlemoyer, L. (2022). LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale. In: S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, & A. Oh (Eds.), Proceedings of the 36th International Conference on Neural Information Processing Systems (NeurIPS 2022) (pp. 30318 - 30332), (28 November - 9 December 2022), New Orleans LA USA. https://doi.org/10.48550/arXiv.2208.07339
- Dettmers, T., Pagnoni, A., Holtzman, A., & Zettlemoyer, L. (2023). QLoRA: Efficient Finetuning of Quantized LLMs. In: A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, & S. Levine (Eds.), Proceedings of the 37th International Conference on Neural Information Processing Systems (NeurIPS 2023), (pp. 10088–10115), (10-16 December 2023), New Orleans LA USA. https://doi.org/10.48550/arXiv.2305.14314
- Ding, N., Qin, Y., Yang, G., Wei, F., Yang, Z., Su, Y., Hu, S., Chen, Y., Chan, C.-M., Chen, W., Yi, J., Zhao, W., Wang, X., Liu, Z., Zheng, H.-T., Chen, J., Liu, Y., Tang, J., Li, J., & Sun, M. (2023). Parameter-efficient fine-tuning of large-scale pre-trained language models. Nature Machine Intelligence, 5(3), 220–235. https://doi.org/10.1038/s42256-023-00626-4
- Frantar, E., Ashkboos, S., Hoefler, T., & Alistarh, D. (2023). OPTQ: Accurate Quantization for Generative Pre-trained Transformers. In: Proceedings of the 11th International Conference on Learning Representations (ICLR 2023) (pp. 1–16), (1-5 May 2023), Kigali Rwanda.
- Guan, Z. K., & Wang, Y. (2025). Evolving Paradigms in Task-Based Search and Learning: A Comparative Analysis of Traditional Search Engine with LLM-Enhanced Conversational Search System. https://doi.org/10.48550/arXiv.2512.00313
Details
Primary Language
English
Subjects
Deep Learning, Natural Language Processing
Journal Section
Research Article
Authors
Early Pub Date
September 23, 2026
Publication Date
September 30, 2026
Submission Date
April 24, 2026
Acceptance Date
July 16, 2026
Published in Issue
Year 2026 Volume: 13 Number: 3