Romance-MoLE: Rich Cousins and Their Benefits An Approach to Low-Resource Language Modelling

Loading...
Thumbnail Image

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

Integrating low-resource languages into Large Language Models (LLMs) is fundamentally limited by persistent data scarcity and the “Curse of Multilinguality,” where dominant source languages overwrite minority language representations and cause structural “translationese”. To preserve the endangered Languedocien dialect of Occitan without catastrophic forgetting, this thesis presents Romance-MoLE (Mixture of LoRA Experts), a parameter-efficient framework applying targeted cross-lingual transfer from Occitan’s phylogenetic “rich cousins”: Catalan and French. Using the pre-trained Llama-3.1- Carballo architecture, the pipeline applies Trans-Tokenisation to repair subword fragmentation and utilises curated synthetic data to enforce standardised dialectal norms. Independent Low-Rank Adaptation (LoRA) experts extracted for French and Catalan are fused via TIES-Merging to construct a stable baseline initialisation. Finally, a Hierarchical Mixture of Ranked Adapters routing strategy (HMoRA) applies token-level gating in shallow layers for localised morphological precision and sequence-level pooling in deep layers to eliminate mid-sentence code-switching. Romance-MoLE yields a +14.57 chrF++ improvement in Occitan translation quality over a standard fine-tuning baseline, while mitigating catastrophic forgetting in Catalan and allowing the model to follow French instructions with a significant reduction in structural code-switching.

Description

Keywords

cross-lingual transfer learning, mixture of LoRA experts (MoLE), hierarchical mixture of LoRA experts (HMoRA), trans-tokenisation, occitan, parameterefficient fine-tuning (PEFT)

Citation

ISBN

Articles

Department

Defence location

Endorsement

Review

Supplemented By

Referenced By