Yurii Paniv

I'm a PhD student in Computer Science at the Ukrainian Catholic University (UCU), working on developing unsupervised and semi-supervised methods to obtain high-quality machine learning models using minimal data.
My goal is to bring state-of-the-art performance to mid- and low-resource languages across different modalities: text, speech, and vision.

This is important for languages that don't have the same amount of data available as for English, bringing those communities equal economic opportunities.

I lead the development of Lapa LLM, an open Ukrainian language model. I was named one of Forbes Ukraine's 20 AI leaders of Ukraine for 2025.

profile photo

๐Ÿช„ Research


Here is a list of my research projects, exploring how to shorten the gap between high-resource and low-resource languages in natural language processing.

Last Translation Benchmark contributor interface for comparing translations and defining verification rules

Last Translation Benchmark


Vilรฉm Zouhar et al., including Yurii Paniv (dataset contributor)
arXiv, 2026
๐Ÿ“– arxiv / code / ๐Ÿ“Š datasets / ๐ŸŒ website

I annotated part of the Ukrainian examples for the Last Translation Benchmark, a collection of challenging translation examples created and reviewed by people. Each example includes verification rules that identify specific translation failures, supporting more reliable evaluation of machine translation systems.

Benchmark results comparing Lapa pretrained and instruction-tuned models with other Ukrainian language models

Data-Efficient Adaptation of Multilingual LLMs to Ukrainian


Yurii Paniv, Bohdan Didenko, Mykola Haltiuk, Vladyslav Humennyy, Andrian Kravchenko, Roman Kyslyi, Viktoriia Makovska, Artem Orlovskyi, Bohdan Ruban, Maksym-Yurii Rudko, Anastasiia Senyk, Nazarii Drushchak, Dmytro Chaplynskyi, Mariana Romanyshyn
Fifth Ukrainian Natural Language Processing Conference (UNLP 2026), 2026
๐Ÿ“– paper / code / ๐Ÿ“Š datasets / ๐Ÿค— model

This is a paper for Lapa LLM. We presented a reproducible approach to adapting multilingual LLMs to Ukrainian using tokenizer adaptation, data quality filtering transferred through translation, and instruction data generation. Applied to Gemma-3-12B, the approach improved Ukrainian benchmark performance while requiring 1.5 times fewer tokens for the same text. We released the models, datasets, classifiers, and code to support adaptation to other languages.

NLP course progression from foundations and classical NLP through neural methods, applied tasks, modern NLP, and production and ethics

Bridging Applied Experience and Research Contexts in Ukrainian NLP Education


Yurii Paniv, Viktoriia Makovska
Seventh Workshop on Teaching Natural Language Processing (TeachNLP 2026), 2026
๐Ÿ“– paper / code

We described our open undergraduate NLP course at Ukrainian Catholic University, taught in Ukrainian with publicly available slides, notebooks, recordings, and assignments. The paper shares how we adapted English-centric materials to the Ukrainian context, supported students with different technical backgrounds, and combined theory, practical projects, and ethics in the curriculum.

BabyBabelLM training data distribution by source category across languages and three data tiers

BabyBabelLM: A Multilingual Benchmark of Developmentally Plausible Training Data


Jaap Jumelet, Abdellah Fourtassi, Akari Haga, Bastian Bunzeck, Bhargav Shandilya, Diana Galvan-Sosa, Faiz Ghifari Haznitrama, Francesca Padovani, Francois Meyer, Hai Hu, Julen Etxaniz, Laurent Prevot, Linyang He, Marรญa Grandury, Mila Marcheva, Negar Foroutan, Nikitas Theodoropoulos, Pouya Sadeghi, Siyuan Song, Suchir Salhan, Susana Zhou, Yurii Paniv, Ziyin Zhang, Arianna Bisazza, Alex Warstadt, Leshem Choshen
19th Conference of the European Chapter of the Association for Computational Linguistics (EACL 2026), Long Papers, 2026
๐Ÿ“– paper / code / ๐Ÿ“Š datasets / ๐Ÿค— model / ๐ŸŒ website

We introduced BabyBabelLM, a collection of pretraining datasets in 45 languages designed to approximate the language exposure children receive while acquiring their native language. The project targets the equivalent of 100 million English words per language and includes evaluation suites and baseline models to support multilingual language modeling and cognitive research.

Radar chart comparing Gemma 2 base, instruction-tuned, and MamayLM models on Ukrainian benchmarks

Isolating LLM Performance Gains in Pre-training versus Instruction-tuning for Mid-resource Languages: The Ukrainian Benchmark Study


Yurii Paniv
15th International Conference on Recent Advances in Natural Language Processing (RANLP 2025), 2025
๐Ÿ“– paper / code / ๐Ÿ“Š datasets

I compared base and instruction-tuned language models on Ukrainian summarization, question answering, and translation tasks, and introduced LongFlores for evaluating paragraph-level translation. In these experiments, base models given a few task examples performed better than their instruction-tuned counterparts evaluated without examples, pointing to opportunities to improve Ukrainian instruction tuning.

Five steps of the Ukrainian phonemization pipeline applied to the word for sixty

Context-Aware Lexical Stress Prediction and Phonemization for Ukrainian TTS Systems


Anastasiia Senyk, Mykhailo Lukianchuk, Valentyna Robeiko, Yurii Paniv
Fourth Ukrainian Natural Language Processing Workshop (UNLP 2025), 2025
๐Ÿ“– paper / code / ๐Ÿ“Š datasets / ๐Ÿค— model

We developed a rule-based Ukrainian phonemizer and a sentence-level lexical stress prediction model for text-to-speech systems, alongside a benchmark with annotated stress patterns. The phonemizer achieved a 1.23% word error rate on a manually constructed pronunciation dataset, while the stress prediction pipeline outperformed existing neural approaches.

UAlign benchmark development pipeline: dataset selection, filtration, translation, and linguistic refinements

UAlign: LLM Alignment Benchmark for the Ukrainian Language


Andrian Kravchenko, Yurii Paniv, Nazarii Drushchak
Fourth Ukrainian Natural Language Processing Workshop (UNLP 2025), 2025
๐Ÿ“– paper / code / ๐Ÿ“Š datasets

We introduced UAlign, a benchmark for evaluating LLM alignment in Ukrainian, with 3,682 moral judgment scenarios and 1,700 ethical situations in parallel Ukrainian and English versions. Our evaluation of six models found a gap between their alignment performance in the two languages, highlighting the need for Ukrainian evaluation resources.

project image

Sovereign Large Language Models: Advantages, Strategy and Regulations


Mykhailo Bondarenko, Sviatoslav Lushnei, Yurii Paniv, Oleksii Molchanovsky, Mariana Romanyshyn, Yurii Filipchuk, Artur Kiulian
arxiv, 2025
๐Ÿ“– arxiv

I was a PM for this policy research paper. This report analyzes key trends, challenges, risks, and opportunities associated with the development of Large Language Models (LLMs) globally. It examines national experiences in developing LLMs and assesses the feasibility of investment in this sector. Additionally, the report explores strategies for implementing, regulating, and financing AI projects at the state level. We analyzed more than 18 macroregions, covering about 85% of the worldโ€™s population.

project image

Benchmarking Multimodal Models for Ukrainian Language Understanding Across Academic and Cultural Domains


Yurii Paniv, Artur Kiulian, Dmytro Chaplynskyi, Mykola Khandoga, Anton Polishko, Tetiana Bas, Guillermo Gabrielli
Fourth Ukrainian Natural Language Processing Workshop (UNLP 2025), 2024
๐Ÿ“– paper / code / ๐Ÿ“Š MMZNO / ๐Ÿ“Š UACuisine / ๐Ÿ“Š Multi30K-UK

Here we benchmarked most of the existing multimodal LLMs for Ukrainian language for academic performance and cultural understanding to establish an understanding what models perform best for Ukrainian in multi-modal scenario. We gathered a national exam dataset for this purpose.

project image

Setting up the Data Printer with Improved English to Ukrainian Machine Translation


Yurii Paniv, Dmytro Chaplynskyi, Nikita Trynus, Volodymyr Kyrylov
Third Ukrainian Natural Language Processing Workshop (UNLP) @ LREC-COLING 2024, 2024
๐Ÿ“– arxiv / code / ๐Ÿค— model / ๐Ÿ•น๏ธ demo

Here we trained a new model for English to Ukrainian machine translation, utilizing unsupervised data selection to reach SOTA performance on English-Ukrainian translation (FLORES test set). High-quality translations should help in the development of the Ukrainian language processing tools and resources.

project image

Unsupervised Data Validation Methods for Efficient Model Training


Yurii Paniv
arxiv, 2023
๐Ÿ“– arxiv

The updated version of my research proposal with motivation why and how we need to focus on bringing state-of-the-art performance to low-resource languages and domains. Old version of proposal can be found here.


๐Ÿค— Open Source Projects


These are projects that I have contributed to or created, that led to my research work.

project image

UAlpaca


GitHub
2023-03-28
code / ๐Ÿค— model / ๐Ÿ•น๏ธ demo

First Ukrainian instruction-tuned language models and datasets. This is a first publicly-known exploration of extracting Ukrainian-language capabilities in Large Language Models.

project image

Crimean Tatar Text-to-Speech


GitHub
2022-10-24
code / ๐Ÿค— model / ๐Ÿ•น๏ธ demo

Experiment to bring natural-sounding Text-to-Speech to low-resource Crimean Tatar language with just 2 hours of audio data, validating a recipe for other low-resource languages. Showcased on national TV in Ukraine!
Sample:

project image

Ukrainian Question and Answering with BERT


GitHub
2022-05-30
code / ๐Ÿค— model

Extractive Ukrainian Question Answering models with BERT, useful for analyzing Ukrainian text.

project image

Ukrainian Text-to-Speech


GitHub
2021-10-01
code / ๐Ÿค— model / ๐Ÿ•น๏ธ demo

The most popular open source Ukrainian Text-to-Speech model with completely open and MIT-licensed stack.
Sample:

project image

Ukrainian Speech-to-Text


GitHub
2020-08-10
code / ๐Ÿค— model / ๐Ÿ•น๏ธ demo

Speech-to-Text training scripts for Ukrainian WER 12,22%. Outdated by now, there are a lot of better models here.




โš™๏ธ Industry experience


Overall I have 10+ years of industry experience, spanning recommendation systems and data analysis for retail, automatic microchip quality assurance, financial domain and serving and processing vast amounts of geospatial data. More details and up-to-date information could be found on my LinkedIn.

๐Ÿค Other Projects


These include talks, teaching, community projects, podcast appearances, and unpublished research work.

Slide from Yurii Paniv's UNLP 2026 keynote on automating Ukrainian NLP research

UNLP 2026 Keynote: Addressing the Ukrainian NLP Gap Using Agents


Speaker
2026-05-30
๐Ÿ“บ video

I gave a keynote at the Fifth Ukrainian Natural Language Processing Conference (UNLP 2026) about our attempt to use agents to automate research and close the NLP gap between Ukrainian and English.

Yurii Paniv's talk announcement for AI & BigData Online Day 2026

EuroHPC Lecture: Accessing Free Compute for Training LLMs


Lecturer
2026-04-04
๐Ÿ“บ video / slides

I gave a lecture at AI & BigData Online Day 2026 about how Ukrainian businesses can access free computing resources for training large language models through EuroHPC, with practical guidance on where and how to apply.

Forbes Ukraine magazine profile of Yurii Paniv in its 2025 list of 20 AI leaders

Forbes Ukraine: 20 AI Leaders of Ukraine 2025


Recognition
2025-12-01
๐ŸŒ website

I was named one of Forbes Ukraineโ€™s 20 AI leaders of Ukraine for 2025. The profile highlighted my work leading the development of Lapa, an open Ukrainian language model, and our efforts to improve its Ukrainian language capabilities and efficiency. Read the magazine profile.

AI HOUSE Podcast episode 47 featuring Yurii Paniv and Roman Kyslyi discussing Lapa LLM

AI HOUSE Podcast: Building the Ukrainian Lapa LLM


Podcast
2025-11-07
๐Ÿ“บ video

I joined Roman Kyslyi on the AI HOUSE Podcast to discuss how we built Lapa LLM, from collecting data and training a Ukrainian tokenizer to evaluating the model. We also discussed filtering propaganda from the training corpus, reasoning, multimodality, and plans for the modelโ€™s development and distribution.

UCU NLP course progression from classical NLP to neural methods, applications, and production and ethics

NLP Course Lecturer at UCU


Lecturer
2025-09-01
๐ŸŒ website / ๐Ÿ“บ video

I co-created and taught the undergraduate Natural Language Processing course at Ukrainian Catholic University in Fall 2025 alongside Viktoriia Makovska and our teaching team. Delivered in Ukrainian, the course covers classical NLP, neural models, transformers, LLMs, fine-tuning, retrieval, and responsible AI. Lecture slides, practical notebooks, assignments, and video recordings are publicly available.

project image

Audio Processing Course Lecturer at UCU


Lecturer
2025-03-24
๐ŸŒ website / ๐Ÿ“บ video

I was a guest lecturer at UCU Audio Processing Course, talking about Speaker Diarization.

project image

Advisor for several BSc and MSc students


Advisor
2025-02-16

Iโ€™m currently an advisor for several BSc and MSc students at Ukrainian Catholic University, helping them with their research projects in the field of NLP, processing text and speech data in particular.

project image

GenAI Course Lecturer at UCU


Lecturer
2024-10-25
๐ŸŒ website

I was one of the authors and lecturers for GenAI course at Ukrainian Catholic University in Fall 2024. Alongside my colleagues Nazarii Drushchak and Igor Babin we introduced students to generative NLP and CV concepts, tools and latest research. Course syllabus is available here.

project image

'Text-to-Speech Speedrun' Lecture


Lecturer
2024-03-25
๐Ÿ“บ video

Was a public speaker on 2 events in the same month with the same topic: ยซText-to-Speech speedrunยป.

project image

ContribuLing 2022 (Wikimedia Foundation conference) Speaker


Lecturer
2022-04-20

Topic: ยซRecognition and synthesis of Ukrainian languageยป


Design and source code from Jon Barron's website