I'm a PhD student in Computer Science at the Ukrainian Catholic University (UCU), working on developing
unsupervised and semi-supervised methods to obtain high-quality machine
learning models using
minimal data.
My goal is to bring state-of-the-art performance to mid- and low-resource
languages across different modalities: text, speech, and vision.
This is important for languages that don't have the same amount of data available as for English,
bringing those communities equal economic opportunities.
Here is a list of my research projects, exploring how to shorten the gap between high-resource and
low-resource languages in natural language processing.
I annotated part of the Ukrainian examples for the Last Translation Benchmark, a collection of challenging translation examples created and reviewed by people. Each example includes verification rules that identify specific translation failures, supporting more reliable evaluation of machine translation systems.
Data-Efficient Adaptation of Multilingual LLMs to Ukrainian
This is a paper for Lapa LLM. We presented a reproducible approach to adapting multilingual LLMs to Ukrainian using tokenizer adaptation, data quality filtering transferred through translation, and instruction data generation. Applied to Gemma-3-12B, the approach improved Ukrainian benchmark performance while requiring 1.5 times fewer tokens for the same text. We released the models, datasets, classifiers, and code to support adaptation to other languages.
Bridging Applied Experience and Research Contexts in Ukrainian NLP Education
Yurii Paniv, Viktoriia Makovska
Seventh Workshop on Teaching Natural Language Processing (TeachNLP 2026), 2026
๐ paper
/
code
We described our open undergraduate NLP course at Ukrainian Catholic University, taught in Ukrainian with publicly available slides, notebooks, recordings, and assignments. The paper shares how we adapted English-centric materials to the Ukrainian context, supported students with different technical backgrounds, and combined theory, practical projects, and ethics in the curriculum.
BabyBabelLM: A Multilingual Benchmark of Developmentally Plausible Training Data
Jaap Jumelet, Abdellah Fourtassi, Akari Haga, Bastian Bunzeck, Bhargav Shandilya, Diana Galvan-Sosa, Faiz Ghifari Haznitrama, Francesca Padovani, Francois Meyer, Hai Hu, Julen Etxaniz, Laurent Prevot, Linyang He, Marรญa Grandury, Mila Marcheva, Negar Foroutan, Nikitas Theodoropoulos, Pouya Sadeghi, Siyuan Song, Suchir Salhan, Susana Zhou, Yurii Paniv, Ziyin Zhang, Arianna Bisazza, Alex Warstadt, Leshem Choshen
19th Conference of the European Chapter of the Association for Computational Linguistics (EACL 2026), Long Papers, 2026
๐ paper
/
code
/
๐ datasets
/
๐ค model
/
๐ website
We introduced BabyBabelLM, a collection of pretraining datasets in 45 languages designed to approximate the language exposure children receive while acquiring their native language. The project targets the equivalent of 100 million English words per language and includes evaluation suites and baseline models to support multilingual language modeling and cognitive research.
Isolating LLM Performance Gains in Pre-training versus Instruction-tuning for Mid-resource Languages: The Ukrainian Benchmark Study
Yurii Paniv 15th International Conference on Recent Advances in Natural Language Processing (RANLP 2025), 2025
๐ paper
/
code
/
๐ datasets
I compared base and instruction-tuned language models on Ukrainian summarization, question answering, and translation tasks, and introduced LongFlores for evaluating paragraph-level translation. In these experiments, base models given a few task examples performed better than their instruction-tuned counterparts evaluated without examples, pointing to opportunities to improve Ukrainian instruction tuning.
Context-Aware Lexical Stress Prediction and Phonemization for Ukrainian TTS Systems
Anastasiia Senyk, Mykhailo Lukianchuk, Valentyna Robeiko, Yurii Paniv Fourth Ukrainian Natural Language Processing Workshop (UNLP 2025), 2025
๐ paper
/
code
/
๐ datasets
/
๐ค model
We developed a rule-based Ukrainian phonemizer and a sentence-level lexical stress prediction model for text-to-speech systems, alongside a benchmark with annotated stress patterns. The phonemizer achieved a 1.23% word error rate on a manually constructed pronunciation dataset, while the stress prediction pipeline outperformed existing neural approaches.
UAlign: LLM Alignment Benchmark for the Ukrainian Language
We introduced UAlign, a benchmark for evaluating LLM alignment in Ukrainian, with 3,682 moral judgment scenarios and 1,700 ethical situations in parallel Ukrainian and English versions. Our evaluation of six models found a gap between their alignment performance in the two languages, highlighting the need for Ukrainian evaluation resources.
Sovereign Large Language Models: Advantages, Strategy and Regulations
I was a PM for this policy research paper. This report analyzes key trends, challenges, risks, and opportunities associated with the development of Large Language Models (LLMs) globally. It examines national experiences in developing LLMs and assesses the feasibility of investment in this sector. Additionally, the report explores strategies for implementing, regulating, and financing AI projects at the state level. We analyzed more than 18 macroregions, covering about 85% of the worldโs population.
Benchmarking Multimodal Models for Ukrainian Language Understanding Across Academic and Cultural Domains
Yurii Paniv, Artur Kiulian, Dmytro Chaplynskyi, Mykola Khandoga, Anton Polishko, Tetiana Bas, Guillermo Gabrielli
Fourth Ukrainian Natural Language Processing Workshop (UNLP 2025), 2024
๐ paper
/
code
/
๐ MMZNO
/
๐ UACuisine
/
๐ Multi30K-UK
Here we benchmarked most of the existing multimodal LLMs for Ukrainian language for academic performance and cultural understanding to establish an understanding what models perform best for Ukrainian in multi-modal scenario. We gathered a national exam dataset for this purpose.
Setting up the Data Printer with Improved English to Ukrainian Machine Translation
Yurii Paniv, Dmytro Chaplynskyi, Nikita Trynus, Volodymyr Kyrylov
Third Ukrainian Natural Language Processing Workshop (UNLP) @ LREC-COLING 2024, 2024
๐ arxiv
/
code
/
๐ค model
/
๐น๏ธ demo
Here we trained a new model for English to Ukrainian machine translation, utilizing unsupervised data selection to reach SOTA performance on English-Ukrainian translation (FLORES test set). High-quality translations should help in the development of the Ukrainian language processing tools and resources.
Unsupervised Data Validation Methods for Efficient Model Training
The updated version of my research proposal with motivation why and how we need to focus on bringing state-of-the-art performance to low-resource languages and domains. Old version of proposal can be found here.
๐ค Open Source Projects
These are projects that I have contributed to or created, that led to my research work.
First Ukrainian instruction-tuned language models and datasets. This is a first publicly-known exploration of extracting Ukrainian-language capabilities in Large Language Models.
Experiment to bring natural-sounding Text-to-Speech to low-resource Crimean Tatar language with just 2 hours of audio data, validating a recipe for other low-resource languages. Showcased on national TV in Ukraine! Sample:
Speech-to-Text training scripts for Ukrainian WER 12,22%. Outdated by now, there are a lot of better models here.
โ๏ธ Industry experience
Overall I have 10+ years of industry experience, spanning recommendation systems and data
analysis for retail, automatic microchip quality assurance, financial domain and serving and processing
vast
amounts of geospatial data. More details and up-to-date information could be found on my
LinkedIn.
๐ค Other Projects
These include talks, teaching, community projects, podcast appearances, and unpublished research work.
UNLP 2026 Keynote: Addressing the Ukrainian NLP Gap Using Agents
I gave a keynote at the Fifth Ukrainian Natural Language Processing Conference (UNLP 2026) about our attempt to use agents to automate research and close the NLP gap between Ukrainian and English.
EuroHPC Lecture: Accessing Free Compute for Training LLMs
I gave a lecture at AI & BigData Online Day 2026 about how Ukrainian businesses can access free computing resources for training large language models through EuroHPC, with practical guidance on where and how to apply.
I was named one of Forbes Ukraineโs 20 AI leaders of Ukraine for 2025. The profile highlighted my work leading the development of Lapa, an open Ukrainian language model, and our efforts to improve its Ukrainian language capabilities and efficiency. Read the magazine profile.
I joined Roman Kyslyi on the AI HOUSE Podcast to discuss how we built Lapa LLM, from collecting data and training a Ukrainian tokenizer to evaluating the model. We also discussed filtering propaganda from the training corpus, reasoning, multimodality, and plans for the modelโs development and distribution.
I co-created and taught the undergraduate Natural Language Processing course at Ukrainian Catholic University in Fall 2025 alongside Viktoriia Makovska and our teaching team. Delivered in Ukrainian, the course covers classical NLP, neural models, transformers, LLMs, fine-tuning, retrieval, and responsible AI. Lecture slides, practical notebooks, assignments, and video recordings are publicly available.
I was a guest lecturer at UCU Audio Processing Course, talking about Speaker Diarization.
Advisor for several BSc and MSc students
Advisor
2025-02-16
Iโm currently an advisor for several BSc and MSc students at Ukrainian Catholic University, helping them with their research projects in the field of NLP, processing text and speech data in particular.
I was one of the authors and lecturers for GenAI course at Ukrainian Catholic University in Fall 2024. Alongside my colleagues Nazarii Drushchak and Igor Babin we introduced students to generative NLP and CV concepts, tools and latest research. Course syllabus is available here.