<?xml version="1.0" encoding="UTF-8"?><?xml-stylesheet type="text/xsl" href="static/style.xsl"?><OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd"><responseDate>2026-09-19T21:31:46Z</responseDate><request verb="GetRecord" identifier="oai:gupea.ub.gu.se:2077/21418" metadataPrefix="dim">https://gupea.ub.gu.se/server/oai/request</request><GetRecord><record><header><identifier>oai:gupea.ub.gu.se:2077/21418</identifier><datestamp>2015-07-09T02:06:18Z</datestamp><setSpec>com_2077_18218</setSpec><setSpec>com_2077_4716</setSpec><setSpec>com_2077_10556</setSpec><setSpec>col_2077_18219</setSpec><setSpec>col_2077_10557</setSpec></header><metadata><dim:dim xmlns:dim="http://www.dspace.org/xmlns/dspace/dim" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:doc="http://www.lyncode.com/xoai" xsi:schemaLocation="http://www.dspace.org/xmlns/dspace/dim http://www.dspace.org/schema/dim.xsd">
   <dim:field mdschema="dc" element="contributor" qualifier="author">Hammarström, Harald</dim:field>
   <dim:field mdschema="dc" element="date" qualifier="accessioned">2009-11-16T10:20:28Z</dim:field>
   <dim:field mdschema="dc" element="date" qualifier="available">2009-11-16T10:20:28Z</dim:field>
   <dim:field mdschema="dc" element="date" qualifier="issued">2009-11-16T10:20:28Z</dim:field>
   <dim:field mdschema="dc" element="identifier" qualifier="isbn">978-91-628-7942-6</dim:field>
   <dim:field mdschema="dc" element="identifier" qualifier="uri">http://hdl.handle.net/2077/21418</dim:field>
   <dim:field mdschema="dc" element="description" qualifier="abstract" lang="en">This thesis presents work in two areas; Language Technology and Linguistic&#xd;
Typology.&#xd;
In the field of Language Technology, a specific problem is addressed: Can a&#xd;
computer extract a description of word conjugation in a natural language using&#xd;
only written text in the language? The problem is often referred to as Unsupervised&#xd;
Learning of Morphology and has a variety of applications, including&#xd;
Machine Translation, Document Categorization and Information Retrieval. The&#xd;
problem is also relevant for linguistic theory. We give a comprehensive survey&#xd;
of work done so far on the problem and then describe a new approach to the&#xd;
problem as well as a number of applications. The idea is that concatenative&#xd;
affixation, i.e., how stems and affixes are stringed together to form words, can,&#xd;
with some success, be modelled simplistically. Essentially, words consist of highfrequency&#xd;
strings (“affixes”) attached to low-frequency strings (“stems”), e.g.,&#xd;
as in the English play-ing. Case studies show how this naive model can be used&#xd;
for stemming, language identification and bootstrapping language description.&#xd;
There are around 7 000 languages in the world, exhibiting a bewildering&#xd;
structural diversity. Linguistic Typology is the subfield of linguistics that aíms&#xd;
to understand this diversity. Many of the languages in the world today are&#xd;
spoken only by relatively small groups of people and are threatened by extinction&#xd;
and it is therefore a priority to record them. Language documentation, is and&#xd;
has been, an extremely decentralised activity, carried out not only by linguists,&#xd;
but also missionaries, travellers, anthropologists etc foremostly throughout the&#xd;
past 200 years. There is no central record of which and how many languages have&#xd;
been described. To meet the priority, we have attempted to list those languages&#xd;
which are the most poorly described which do not belong to a language family&#xd;
where some other languages is decently described – a task requiring both analysis&#xd;
and diligence. Next, the thesis includes typological work on one of the more&#xd;
tractable aspects of language structure, namely numeral systems, i.e., normed&#xd;
expressions used to denote exact quantities. In one of the first surveys to cover&#xd;
the whole world, we look at rare number bases among numeral systems. One&#xd;
major rarity is base-6-36 systems which are only attested in South/Southwest&#xd;
New Guinea and we make a special inquiry into its emergence.&#xd;
Traditionally, linguists have had headaches over what counts as a language&#xd;
as opposed to a dialect, and have therefore been reluctant to give counts of the&#xd;
number of languages in a given area. One chapter of the present thesis shows&#xd;
that, contrary to popular belief, there is an intuitively sound way to count&#xd;
languages (as opposed to dialects). The only requirement is that, for each pair&#xd;
of varieties, we are told whether they are mutually intelligible or not.</dim:field>
   <dim:field mdschema="dc" element="language" qualifier="iso" lang="en">eng</dim:field>
   <dim:field mdschema="dc" element="relation" qualifier="haspart" lang="en">Hammarström, H. (2005). A New Algorithm for Unsupervised Induction&#xd;
of Concatenative Morphology In Yli-Jyrä, A., Karttunen, L., and&#xd;
Karhumäki, J., editors, Finite State Methods in Natural Language Processing:&#xd;
5th International Workshop, FSMNLP 2005, Helsinki, Finland,&#xd;
September 1-2, 2005. Revised Papers, volume 4002 of Lecture Notes in&#xd;
Computer Science, pages 288–289. Springer-Verlag, Berlin.</dim:field>
   <dim:field mdschema="dc" element="relation" qualifier="haspart" lang="en">Hammarström, H. (2006a). A naive theory of morphology and an algorithm&#xd;
for extraction. In Wicentowski, R. and Kondrak, G., editors,&#xd;
SIGPHON 2006: Eighth Meeting of the Proceedings of the ACL Special Interest&#xd;
Group on Computational Phonology, 8 June 2006, New York City,&#xd;
USA, pages 79–88. Association for Computational Linguistics.</dim:field>
   <dim:field mdschema="dc" element="relation" qualifier="haspart" lang="en">Hammarström, H. (2006b). Poor man’s stemming: Unsupervised recognition&#xd;
of same-stem words. In Ng, H. T., Leong, M.-K., Kan, M.-Y., and Ji,&#xd;
D., editors, Information Retrieval Technology: Proceedings of the Third&#xd;
Publicatons and Contributions 5&#xd;
Asia Information retrieval Symposium, AIRS 2006, Singapore, October&#xd;
2006, volume 4182 of Lecture Notes in Computer Science, pages 323–337.&#xd;
Springer-Verlag, Berlin.</dim:field>
   <dim:field mdschema="dc" element="relation" qualifier="haspart" lang="en">Hammarström, H. (2007a). A fine-grained model for language identification.&#xd;
In Proceedings of iNEWS-07 Workshop at SIGIR 2007, 23-27 July&#xd;
2007, Amsterdam, pages 14–20. ACM.</dim:field>
   <dim:field mdschema="dc" element="relation" qualifier="haspart" lang="en">Hammarström, H. (2007b). A survey and classification of methods for&#xd;
(mostly) unsupervised learning of morphology. In NODALIDA 2007,&#xd;
the 16th Nordic Conference of Computational Linguistics, Tartu, Estonia,&#xd;
25-26 May 2007. NEALT.</dim:field>
   <dim:field mdschema="dc" element="relation" qualifier="haspart" lang="en">Hammarström, H., Thornell, C., Petzell, M., and Westerlund, T. (2008).&#xd;
Bootstrapping language description: The case of Mpiemo (Bantu A, Central&#xd;
African Republic). In Proceedings of LREC-2008, pages 3350–3354.&#xd;
European Language Resources Association (ELRA).</dim:field>
   <dim:field mdschema="dc" element="relation" qualifier="haspart" lang="en">Hammarström, H. (2009a). Poor man’s word-segmentation: Unsupervised&#xd;
morphological analysis for indonesian. In Proceedings of the Third&#xd;
International Workshop on Malay and Indonesian Language Engineering&#xd;
(MALINDO). Singapore: ACL.</dim:field>
   <dim:field mdschema="dc" element="relation" qualifier="haspart" lang="en">Hammarström, H. (2009b). A Survey of Computational Morphological&#xd;
Resources for Low-Density Languages Submitted.</dim:field>
   <dim:field mdschema="dc" element="relation" qualifier="haspart" lang="en">Forsberg, M., Hammarström, H., and Ranta, A. (2006). Lexicon extraction&#xd;
from raw text data. In Salakoski, T., Ginter, F., Pyysalo, S., and&#xd;
Pahikkala, T., editors, Advances in Natural Language Processing: Proceedings&#xd;
of the 5th International Conference, FinTAL 2006 Turku, Finland,&#xd;
August 23-25, 2006, volume 4139 of Lecture Notes in Computer Science,&#xd;
pages 488–499. Springer-Verlag, Berlin.</dim:field>
   <dim:field mdschema="dc" element="relation" qualifier="haspart" lang="en">Hammarström, H. (2008a). Automatic annotation of bibliographical references&#xd;
with target language. In Proceedings of MMIES-2: Wokshop&#xd;
on Multi-source, Multilingual Information Extraction and Summarization,&#xd;
pages 57–64. ACL.</dim:field>
   <dim:field mdschema="dc" element="relation" qualifier="haspart" lang="en">Hammarström, H. (2008b). Counting languages in dialect continua using&#xd;
the criterion of mutual intelligibility. Journal of Quantitative Linguistics,&#xd;
15(1):34–45.</dim:field>
   <dim:field mdschema="dc" element="relation" qualifier="haspart" lang="en">Hammarström, H. (2009c). Whence the Kanum base-6 numeral system?&#xd;
Linguistic Typology, 13(2):305–319.&#xd;
m. Hammarström, H. (2009d [to appear]). Rarities in numeral systems. In&#xd;
Wohlgemuth, J. and Cysouw, M., editors, Rara &amp;amp; Rarissima: Collecting&#xd;
and interpreting unusual characteristics of human languages, Empirical&#xd;
Approaches to Language Typology, pages 7–55. Mouton de Gruyter.</dim:field>
   <dim:field mdschema="dc" element="relation" qualifier="haspart" lang="en">Hammarström, H. (2009e). The Status of the Least Documented Language&#xd;
Families in the World Submitted.</dim:field>
   <dim:field mdschema="dc" element="subject" lang="en">Computational Linguistics</dim:field>
   <dim:field mdschema="dc" element="subject" lang="en">Language typology</dim:field>
   <dim:field mdschema="dc" element="title" lang="en">Unsupervised Learning of Morphology and the Languages of the World</dim:field>
   <dim:field mdschema="dc" element="type">Text</dim:field>
   <dim:field mdschema="dc" element="type" qualifier="svep">Doctoral thesis</dim:field>
   <dim:field mdschema="dc" element="type" qualifier="degree" lang="en">Doctor of Engineering</dim:field>
   <dim:field mdschema="dc" element="gup" qualifier="mail" lang="en">harald@bombo.se</dim:field>
   <dim:field mdschema="dc" element="gup" qualifier="origin" lang="en">Göteborgs universitet. IT-fakulteten</dim:field>
   <dim:field mdschema="dc" element="gup" qualifier="department" lang="en">Department of Computer Science and Engineering ; Institutionen för data- och informationsteknik</dim:field>
   <dim:field mdschema="dc" element="gup" qualifier="defenceplace" lang="en">10:15 in room HB1, Hörsalsväagen 8</dim:field>
   <dim:field mdschema="dc" element="gup" qualifier="defencedate">2009-12-11</dim:field>
   <dim:field mdschema="dc" element="gup" qualifier="dissdb-fakultet">ITF</dim:field>
   <dim:field mdschema="dc" element="citation" qualifier="doi">ITF</dim:field>
   <dim:field mdschema="others" element="access-status">open.access</dim:field>
</dim:dim>
</metadata></record></GetRecord></OAI-PMH>