Lupa

Show document Help

A- | A+ | Print
Title:UzbekPOS : a multi-domain dataset for Uzbek part-of-speech tagging
Authors:ID Sharipov, Maksud (Author)
ID Kuriyozov, Elmurod (Author)
ID Vičič, Jernej (Author)
Files:.pdf RAZ_Sharipov_Maksud_2026.pdf (787,49 KB)
MD5: 0CCF384767E7ADC9CF5239718D55BA63
 
URL https://www.sciencedirect.com/science/article/pii/S2352340926001939?via%3Dihub
 
Language:English
Work type:Article
Typology:1.01 - Original Scientific Article
Organization:FAMNIT - Faculty of Mathematics, Science and Information Technologies
Abstract:In this paper, we introduce UzbekPOS — a part-of-speech (POS) tagged dataset manually annotated for the Uzbek language, designed for natural language processing, artificial intelligence models, and corpus linguistics applications. This tagged corpus is currently the largest publicly available POS- tagged corpus for the Uzbek language. The dataset comprises sentences drawn from a diverse range of Uzbek text sources, including literature, news outlets, science, education, and public speaking, to reflect linguistic and topical diversity. The sentences are tokenized and annotated by professional annotators, utilizing a finely grained POS tagset which integrates standard Universal Dependencies with additional labels that are specific to the morphological and syntactic features of the Uzbek language, comprising 16 tags in total. The UzbekPOS contains almost 4.5K sentences and more than 53K token/tag pairs, with each annotation cross-verified by at least two annotators for highest reliability. It also comes with both raw (txt) and generally accepted formats of distribution (TSV, JSON), as well as the universal POS-tagging format (conllu). This resource is one of the first and the largest openly published POS-tagged dataset for Uzbek, an under-resourced and morphologically complex Turkic language. This dataset can also act as a key foundation for training POS taggers, as a test set for machine learning models, and as a source for linguistic studies. The resource also bears the reusability potential for tasks of related kinds, such as morphological analysis, syntactic parsing, and transfer learning across languages of the Turkic family. Furthermore, this dataset can serve as seed material for creating similar corpora of POS for other Turkic languages and can help conduct cross-linguistic analyses and tool building.
Keywords:POS tagging, Uzbek language, morphological annotation, natural language processing
Publication version:Version of Record
Publication date:28.02.2026
Year of publishing:2026
Number of pages:str. 1-10
Numbering:Vol. 66, article 112640
PID:20.500.12556/RUP-23571 This link opens in a new window
UDC:004.65:811.5
ISSN on article:2352-3409
DOI:10.1016/j.dib.2026.112640 This link opens in a new window
COBISS.SI-ID:270694915 This link opens in a new window
Publication date in RUP:08.09.2026
Views:61
Downloads:2
Metadata:XML DC-XML DC-RDF
:
Copy citation
  
Average score:(0 votes)
Your score:Voting is allowed only for logged in users.
Share:Bookmark and Share


Hover the mouse pointer over a document title to show the abstract or click on the title to get all document metadata.

Record is a part of a journal

Title:Data in brief
Publisher:Elsevier
ISSN:2352-3409
COBISS.SI-ID:32117977 This link opens in a new window

Licences

License:CC BY 4.0, Creative Commons Attribution 4.0 International
Link:http://creativecommons.org/licenses/by/4.0/
Description:This is the standard Creative Commons license that gives others maximum freedom to do what they want with the work as long as they credit the author.

Secondary language

Language:Slovenian
Abstract:V tem članku predstavljamo UzbekPOS – nabor podatkov, označen z vrstami govora (POS), ki je ročno označen za uzbeški jezik in je zasnovan za obdelavo naravnega jezika, modele umetne inteligence in aplikacije korpusnega jezikoslovja. Ta označeni korpus je trenutno največji javno dostopni korpus z oznakami POS za uzbeški jezik. Nabor podatkov vsebuje stavke, vzete iz raznolikih uzbeških besedilnih virov, vključno z literaturo, novicami, znanostjo, izobraževanjem in javnim nastopanjem, da bi odražali jezikovno in tematsko raznolikost. Stavke žetonizirajo in označujejo profesionalni anotatorji z uporabo natančno zrnatega nabora oznak POS, ki združuje standardne univerzalne odvisnosti z dodatnimi oznakami, značilnimi za morfološke in sintaktične značilnosti uzbeškega jezika, skupaj pa obsega 16 oznak. UzbekPOS vsebuje skoraj 4,5 tisoč stavkov in več kot 53 tisoč parov žetonov/oznak, pri čemer vsako opombo navzkrižno preverita vsaj dva anotatorja za najvišjo zanesljivost. Na voljo je tudi v surovi (txt) in splošno sprejeti obliki distribucije (TSV, JSON) ter v univerzalni obliki označevanja POS (conllu). Ta vir je eden prvih in največjih odprto objavljenih naborov podatkov z oznakami POS za uzbeščino, turški jezik s premalo viri in morfološko kompleksnimi viri. Ta nabor podatkov lahko služi tudi kot ključna osnova za usposabljanje označevalcev POS, kot testni nabor za modele strojnega učenja in kot vir za jezikoslovne študije. Vir ima tudi potencial ponovne uporabe za naloge sorodnih vrst, kot so morfološka analiza, sintaktično razčlenjevanje in prenos učenja med jeziki turške družine. Poleg tega lahko ta nabor podatkov služi kot začetni material za ustvarjanje podobnih korpusov POS za druge turške jezike in lahko pomaga pri izvajanju medjezikovnih analiz in gradnji orodij.
Keywords:označevanje POS, uzbeški jezik, morfološka anotacija, obdelava naravnega jezika


Comments

Leave comment

You must log in to leave a comment.

Comments (0)
0 - 0 / 0
 
There are no comments!

Back
Logos of partners University of Maribor University of Ljubljana University of Primorska University of Nova Gorica