Text analysis for email multi label classification

Paniskaki, Kyriaki
Harsha Kadam, Sanjit
Göteborgs universitet/Institutionen för data- och informationsteknikswe
University of Gothenburg/Department of Computer Science and Engineeringeng
2020-07-08T11:30:41Z
2020-07-08T11:30:41Z
2020-07-08
This master’s thesis studies a multi label text classification task on a small data set of bilingual, English and Swedish, short texts (emails). Specifically, the size of the data set is 5800 emails and those emails are distributed among 107 classes with the special case that the majority of the emails includes the two languages at the same time. For handling this task different models have been employed: Support Vector Machines (SVM), Gated Recurrent Units (GRU), Convolution Neural Network (CNN), Quasi Recurrent Neural Network (QRNN) and Transformers. The experiments demonstrate that in terms of weighted averaged F1 score, the SVM outperforms the other models with a score of 0.96 followed by the CNN with 0.89 and the QRNN with 0.80.sv
http://hdl.handle.net/2077/65588
engsv
CSE 20-14sv
Technology
natural language processingsv
machine learningsv
multi label text classificationsv
deep neural networkssv
bilingual textssv
emailssv
short textssv
Text analysis for email multi label classificationsv
Text analysis for email multi label classificationsv
text
Student essay
H2

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
gupea_2077_65588_1.pdf
Size:
2.03 MB
Format:
Adobe Portable Document Format
Description:
Master thesis

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
876 B
Format:
Item-specific license agreed upon to submission
Description:

Collections