1. Dataset of Uzbek base words : extraction and data analysis based on the school corpusKhabibulla Madatov, Surayyo Khajibaeva, Jernej Vičič, 2026, izvirni znanstveni članek Opis: The article presents a dataset of Uzbek base words extracted from a purposefully prepared corpus using the Synonym Thesaurus Support method. This method identifies base words for each school-grade by analysing a large text corpus comprising 142 textbooks intended for school education in Uzbekistan. The definition of the base word used in this article and in the proposed dataset is a word within a synonymic series that: - is the most widely used. - is distinguished by semantic clarity and stability. - has stylistic neutrality. Based on the proposed approach, school textbooks were analysed by dividing them into Primary (school grades 1 - 4), Basic Secondary (school grades 5 - 9), and Secondary (school grades 10 - 11) blocks. Base words that stand out from the general corpus were identified for each school-grade. This method extracted new base words not found in previous school grades and specific to the observed grade. The main idea of the method is to extract base words from the lemma sset of each school-grade using a corpus of synonyms. This allows analysing the level of lexical complexity and class-specific vocabulary richness of texts intended for schoolchildren. The final results are lists of base words specifically extracted from primary (school-grades 1 - 4), basic secondary (school-grades 5 - 9), and secondary (school-grades 10 - 11) school texts; 17,599,48,203, and 20,491 base words, respectively. Ključne besede: school corpus, base word, basic vocabulary, Uzbek language Objavljeno v RUP: 20.05.2026; Ogledov: 352; Prenosov: 10
Celotno besedilo (2,42 MB) Gradivo ima več datotek! Več... |
2. TF-IDF-based classification of Uzbek educational textsKhabibulla Madatov, Sapura Sattarova, Jernej Vičič, 2025, izvirni znanstveni članek Opis: This paper presents a baseline study on automatic Uzbek text classification. Uzbek is a morphologically rich and low-resource language, which makes reliable preprocessing and evaluation challenging. The approach integrates Term Frequency–Inverse Document Frequency (TF–IDF) representation with three conventional methods: linear regression (LR), k-Nearest Neighbors (k-NN), and cosine similarity (CS, implemented as a 1-NN retrieval model). The objective is to categorize school learning materials by grade level (grades 5–11) to support improved alignment between curricular texts and students’ intellectual development. A balanced dataset of Uzbek school textbooks across different subjects was constructed, preprocessed with standard NLP tools, and converted into TF–IDF vectors. Experimental results on the internal test set of 70 files show that LR achieved 92.9% accuracy (precision = 0.94, recall = 0.93, F1 = 0.93), while CS performed comparably with 91.4% accuracy (precision = 0.92, recall = 0.91, F1 = 0.92). In contrast, k-NN obtained only 28.6% accuracy, confirming its weakness in high-dimensional sparse feature spaces. External evaluation on seven Uzbek literary works further demonstrated that LR and CS yielded consistent and interpretable grade-level mappings, whereas k-NN results were unstable. Overall, the findings establish reliable baselines for Uzbek educational text classification and highlight the potential of extending beyond lexical overlap toward semantically richer models in future work. Ključne besede: Uzbek language, text classification, low-resource languages, TF-IDF, cosine similarity, linear regression, k-Nearest Neighbors Objavljeno v RUP: 17.10.2025; Ogledov: 940; Prenosov: 8
Celotno besedilo (286,87 KB) Gradivo ima več datotek! Več... |
3. Dataset of vocabulary in Uzbek primary education : extraction and analysis in case of the school corpusKhabibulla Madatov, Sapura Sattarova, Jernej Vičič, 2025, izvirni znanstveni članek Opis: The main goal of this research work is to determine the number of new words that a primary school pupil should know/acquire during each academic year. To accomplish this, we have created two datasets. The first dataset was compiled based on the "Explanatory Vocabulary of the Uzbek Language" (EDUL). The second dataset was created from 35 primary school textbooks for grades 1-4 approved by the Ministry of Preschool and School Education of the Republic of Uzbekistan, and it was named the "Uzbek Primary School Corpus" (UPSC) by authors. Using the "Comparative Lemma Extraction Method" (CLEM) proposed by the authors of the article, a vocabulary for grades 1-4 was created, and the problem of determining the number of new words (disregarding word forms as Uzbek is a morphologically rich language) that primary school pupils should learn each academic year was solved. Ključne besede: Uzbek language, primary school, corpus construction, natural language processing (NLP), comparative Lemma extraction method Objavljeno v RUP: 08.08.2025; Ogledov: 1138; Prenosov: 10
Celotno besedilo (342,87 KB) Gradivo ima več datotek! Več... |
4. Dataset of Uzbek verbs with formation and suffixesMaksud Sharipov, Jernej Vičič, 2025, drugi znanstveni članki Opis: The main goal of this work is to create a dataset of Uzbek language verbs. This dataset stores information about which words verbs are derived from and with which affixes. The affixes are classified into distinct categories. With the help of this dataset, it is possible to determine from which parts of speech each Uzbek verb is derived and with which affixes. It also plays a key role in identifying verbs in Uzbek language texts and developing rule-based models for their analysis. Additionally, this dataset plays a key role in building various artificial intelligence models for the morphological and syntactic analysis of Uzbek language texts. Verbs play a crucial role in learning any language; therefore, students in schools and higher education institutions can also use this dataset during the learning process. The obtained dataset serves as a valuable resource for researchers and practitioners interested in Uzbek language processing tasks. Ključne besede: verb phrase, Uzbek language, Uzbek web corpus, verb form, verb affixes Objavljeno v RUP: 02.06.2025; Ogledov: 1353; Prenosov: 14
Celotno besedilo (281,07 KB) Gradivo ima več datotek! Več... |
5. |