Closing the Semantic Gap: Automated Test Suite Management using Multi-Agent Systems - A Design Science Research Study on Autonomous Test Suite Augmentation

Loading...
Thumbnail Image

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

Requirement-test links can show that two artifacts are related, but they do not show whether a test checks the behavior stated in a requirement. A linked test may exercise the relevant feature while omitting the expected result or assertion needed to verify it. This thesis studies these cases as possible gaps in test evidence and presents Lanterns, a multi-agent research artifact for diagnosing such gaps and generating candidate tests. The study follows Design Science Research and evaluates Lanterns on EBT and LibEST datasets. The implementation separates traceability recovery from assertion level analysis because a related test is not necessarily a verifying test. The assertion report records structured gap entries, which are passed to test generation together with traceability and project evidence. This staged design links the diagnostic work in RQ1 with the gap-directed generation examined in RQ2, while keeping the evidence and decisions from each stage available for review. Traceability predictions were compared with the dataset reference links, and assertion predictions were compared with a manually constructed ground truth that was checked through external review. Across two EBT runs, median recall was 93.14% for traceability and 97.50% for assertion analysis, with median precision of 57.93% and 54.87%, respectively. For LibEST, the median recall was 60.52% for traceability and 68.81% for assertion analysis, while median precision was 40.00% and 35.63%. Most values were close between the two runs, although EBT assertion precision showed greater variation. The results show that Lanterns recovered many positive links, especially in EBT, but its false positives and missed links still require human review. For RQ2, selected LibEST gap reports were used to generate candidate tests. Three cases were examined qualitatively to determine how closely each candidate addressed its target gap. The assessment showed that gap-directed generation can target specific missing behaviors, although partial output and hallucinations remain risks that require a technical review before adoption. The thesis contributes an inspectable workflow that connects traceability, verification analysis, and gap-directed test generation. The results support Lanterns as an aid for reviewing possible gaps in test evidence, rather than as a source of final coverage decisions or validated tests.

Description

Keywords

software testing, semantic gaps, requirements traceability, assertion analysis, test generation, multi-agent systems, large language models, design science research.

Citation

ISBN

Articles

Department

Defence location

Collections

Endorsement

Review

Supplemented By

Referenced By