Is That Polite? How Linguistic Features Shape Large Language Model Judgments in Chinese Politeness Classification
Date
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
Politeness, as an important pragmatic phenomenon, plays a key role in human communication. However, its automated recognition is still challenging in the field of Natural Language Processing (NLP), especially in the Chinese context, which is characterized by high context-dependency and distinct cultural features. This study takes the Chinese Mandarin politeness classification as its research focus, and systematically compares the performance of several methodologies, including traditional machine learning, handcrafted feature-based methods, pre-trained language models (BERT), and Large Language Models (LLMs). The research centers on two primary questions: first, the performance of LLMs in the Mandarin politeness classification task; and second, the specific linguistic features that drive the models’ politeness judgements. Methodologically, this thesis provides a comparative analysis of classification performance across various models. From an interpretability perspective, it employs feature importance analysis and feature ablation experiments to explore the role of handcrafted features in traditional models. Furthermore, feature grouping analysis and perturbation experiments are designed to examine the sensitivity of LLMs to different linguistic feature. Experimental results show that BERT model achieves the best overall performance, while LLMs show significant performance gain under few-shot settings. Feature analysis further found that LLMs rely more on surface-level formal features, such as exclamation marks and strong-tone expressions in politeness judgments, while showing limited understanding of pragmatic strategies such as hedging and indirectness. The study suggests that current models still have shortcomings in Chinese pragmatic understanding, particularly when handling deep pragmatic strategies. This thesis systematically analyzes the Chinese Mandarin politeness classification task from the perspective of model performance and feature interpretability, providing a new perspective for understanding the behavioral patterns of LLMs in pragmatic tasks.