1. UzbekPOS : a multi-domain dataset for Uzbek part-of-speech taggingMaksud Sharipov, Elmurod Kuriyozov, Jernej Vičič, 2026, izvirni znanstveni članek Opis: In this paper, we introduce UzbekPOS — a part-of-speech (POS) tagged dataset manually annotated for the Uzbek language, designed for natural language processing, artificial intelligence models, and corpus linguistics applications. This tagged corpus is currently the largest publicly available POS- tagged corpus for the Uzbek language. The dataset comprises sentences drawn from a diverse range of Uzbek text sources, including literature, news outlets, science, education, and public speaking, to reflect linguistic and topical diversity. The sentences are tokenized and annotated by professional annotators, utilizing a finely grained POS tagset which integrates standard Universal Dependencies with additional labels that are specific to the morphological and syntactic features of the Uzbek language, comprising 16 tags in total. The UzbekPOS contains almost 4.5K sentences and more than 53K token/tag pairs, with each annotation cross-verified by at least two annotators for highest reliability. It also comes with both raw (txt) and generally accepted formats of distribution (TSV, JSON), as well as the universal POS-tagging format (conllu). This resource is one of the first and the largest openly published POS-tagged dataset for Uzbek, an under-resourced and morphologically complex Turkic language. This dataset can also act as a key foundation for training POS taggers, as a test set for machine learning models, and as a source for linguistic studies. The resource also bears the reusability potential for tasks of related kinds, such as morphological analysis, syntactic parsing, and transfer learning across languages of the Turkic family. Furthermore, this dataset can serve as seed material for creating similar corpora of POS for other Turkic languages and can help conduct cross-linguistic analyses and tool building. Ključne besede: POS tagging, Uzbek language, morphological annotation, natural language processing Objavljeno v RUP: 08.09.2026; Ogledov: 84; Prenosov: 2
Celotno besedilo (787,49 KB) Gradivo ima več datotek! Več... |
2. Digital signal processing tools for radar-based human-computer interactionNuwan Attygalle, Matjaž Kljun, Klen Čopič Pucihar, 2026, samostojni znanstveni sestavek ali poglavje v monografski publikaciji Opis: In recent years, miniature radar-on-chip sensors have been explored for HCI by both academia and industry. This is driven by the availability of affordable radar hardware and advances in signal processing and machine learning. However, comparative evaluation of radar-based gesture interaction systems is challanging and rear. One important dimension for comparison is the set of radar signal representa- tions derived from raw voltage data. These representations commonly include range- Doppler, range-azimuth-angle, range-elevation-angle, point cloud and In-phase and Quadrature (IQ) radar cube formats. However, existing studies often restrict com- parative analysis to a single signal representation type, typically focusing on gesture recognition algorithms or minor variations within Digital Signal Processing (DSP) pipelines. To promote comparative evaluation of radar signal representations, this chapter develops an open-source application designed to facilitate efficient and reli- able dataset preparation of various radar signal representations. The application sup- ports visualisation of generated signal representations and includes a command-line interface for batch processing, thereby streamlining the dataset preparation workflow. Ključne besede: digital signal processing, radar, human-computer interaction Objavljeno v RUP: 18.05.2026; Ogledov: 531; Prenosov: 8
Celotno besedilo (41,69 MB) Gradivo ima več datotek! Več... |
3. Dataset of sentiment tagged language resources for Macedonian languageSofija Kochovska, Jernej Vičič, Branko Kavšek, 2026, izvirni znanstveni članek Opis: Macedonian is a South Slavic language spoken by about 2 million people, primarily in North Macedonia and among diaspora communities worldwide. It’s known for a few distinctive features. Most notably, it uses definite articles attached to the end of nouns, for example, kniga (a book) becomes knigata (the book). Furthermore, it doesn’t use grammatical cases, which makes its grammar relatively straightforward compared to other Slavic languages. The dataset comprises two lists of sentiment annotated words that present the core of the Macedonian sentiment-annotated lexicon, a list of the stopwords, and a list of Affirmative and non-Affirmative words (AnAwords) composed mostly of intensifiers and diminishers, and a list of polarity shifters. The main usage of the presented materials is in rule-based sentiment analysis, but the usage of some of the lists can be much broader. Ključne besede: Macedonian language, sentiment analysis, sentiment lexicon, sentiment analys, rule-based methods, natural language processing, low-resource languages, AnA words, stopwords, intensifiers, diminishers, polarity shifters Objavljeno v RUP: 20.01.2026; Ogledov: 810; Prenosov: 5
Celotno besedilo (251,79 KB) Gradivo ima več datotek! Več... |
4. Dataset of vocabulary in Uzbek primary education : extraction and analysis in case of the school corpusKhabibulla Madatov, Sapura Sattarova, Jernej Vičič, 2025, izvirni znanstveni članek Opis: The main goal of this research work is to determine the number of new words that a primary school pupil should know/acquire during each academic year. To accomplish this, we have created two datasets. The first dataset was compiled based on the "Explanatory Vocabulary of the Uzbek Language" (EDUL). The second dataset was created from 35 primary school textbooks for grades 1-4 approved by the Ministry of Preschool and School Education of the Republic of Uzbekistan, and it was named the "Uzbek Primary School Corpus" (UPSC) by authors. Using the "Comparative Lemma Extraction Method" (CLEM) proposed by the authors of the article, a vocabulary for grades 1-4 was created, and the problem of determining the number of new words (disregarding word forms as Uzbek is a morphologically rich language) that primary school pupils should learn each academic year was solved. Ključne besede: Uzbek language, primary school, corpus construction, natural language processing (NLP), comparative Lemma extraction method Objavljeno v RUP: 08.08.2025; Ogledov: 1329; Prenosov: 11
Celotno besedilo (342,87 KB) Gradivo ima več datotek! Več... |
5. |
6. |
7. Parallelizing an algorithm to find the maximal clique on interval graphs on graphical processing unitsChristian Trefftz, Andrés Santamaría-Galvis, Roberto Cruz Rodes, 2014, objavljeni znanstveni prispevek na konferenci Ključne besede: graph theory, graphics processing units, parallel algorithms, CUDA, Thrust library, interval graphs, maximal clique Objavljeno v RUP: 18.10.2021; Ogledov: 3315; Prenosov: 30
Povezava na celotno besedilo |
8. |
9. Viscoelastic properties of thermo-hydro-mechanically treated beech (Fagus sylvatica L.) determined using dynamic mechanical analysisAndreja Kutnar, Jane O'Dell, Christopher Hunt, Charles R. Frihart, Frederick A. Kamke, Matthew Schwarzkopf, 2020, izvirni znanstveni članek Ključne besede: Fagus sylvatica, Viscoelastic vroperties, thermo-hydro-mechanical processing, THM Objavljeno v RUP: 02.12.2020; Ogledov: 3704; Prenosov: 37
Povezava na celotno besedilo |
10. |