Evaluating Trade-offs of Quantized LLMs for Requirements and Test Alignment

Loading...
Thumbnail Image

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

Large Language Models (LLMs) have shown impressive capabilities in various domains due to their ability to process and interpret natural language. Meanwhile, as software systems continue to expand in size and complexity, the number of associated artifacts (e.g, Requirements and Test Cases) grows as well, leading to challenges in aligning REST (Requirement Engineering and System Test) efforts. There have been earlier initiatives in using LLMs for REST alignment, as it is a widely used measure for software quality assurance. However, the costs associated with model deployment and execution have limited their feasibility. There is a need for a faster yet reasonable solution to cope with the rate at which software artifacts keep growing. In this paper, we investigate whether quantized LLMs can serve as a viable alternative, given their smaller size and less demanding hardware requirements. We choose Mistral, a widely used open-weight LLM, and assess it in conjunction with three different quantization techniques: AWQ, GPTQ, and AQLM— comparing these four versions of Mistral against each other. The experiment is performed with four requirement specification datasets encompassing 433 Requirements and 408 Tests in total. We offer insights into the feasibility of adopting quantized LLMs for REST alignment, highlighting the efficacy, efficiency, and trade-offs of adopting such models, along with a actionable guidance for practitioners. Index Terms—Large Language Models, REST, Traceability, Quantization, Software Testing

Description

Keywords

Citation

ISBN

Articles

Department

Defence location

Endorsement

Review

Supplemented By

Referenced By