Debiasing LLM Annotations in Low-Sample Regimes: A Benchmarking Study

Loading...
Thumbnail Image

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

Recent research has shown an increasing trend toward automating data annotation using LLMs. The main drawback of this approach is that LLM annotations suffer from systematic bias. Consequently, debiasing methods such as DSL and PPI++ have been developed to correct for this bias. These methods have not yet been evaluated in low-sample regimes where annotation budgets are most constrained. This thesis is a benchmarking study, evaluating these debiasing methods on four binary tasks, using five different LLMs for annotation, three downstream models, and varying annotation budgets between n = 20 to 200 labels. The main finding is that debiasing consistently improves model estimation errors even at budgets as low as n = 20. The benefit is largest where the expert-only θ† baseline has the highest uncertainty. Downstream model selection is a critical factor affecting the efficacy of debiasing methods: class prevalence and LPM yield the strongest and most stable improvements, while logistic regression- the natural choice for binary outcomes- introduces a persistent bias floor due to the L2 regularization required in the low-sample regime. These results provide insights to help practitioners working with tight annotation budgets, and introduce annotation informativeness as a potential pre-screening diagnostic for whether debiasing is likely to succeed.

Description

Keywords

Large language models, LLMs, data annotation, debiasing, low-sample regime, benchmarking

Citation

ISBN

Articles

Department

Defence location

Collections

Endorsement

Review

Supplemented By

Referenced By