Conference Program
Join Zoom Meeting: https://us02web.zoom.us/j/85467881508?pwd=vwip0po9HYZ7LUwaSXnMqRlBAHu7NG.1
Passcode: 688765
Join Zoom Meeting: https://us02web.zoom.us/j/85467881508?pwd=vwip0po9HYZ7LUwaSXnMqRlBAHu7NG.1
Passcode: 688765
10:00-10:15 Opening Remarks
10:15-11:15 Morning Session: Translation & Evaluation
Chair: Roman Kyslyi
| 10:15-10:35 | SimIdioms: A Corpus and Benchmark for Ukrainian Idiom Translation
Yaryna Petruniv, Iuliia Makogon and Roman Kyslyi |
| 10:35-10:55 | Semantic Fidelity Versus Literary Quality: A Construct Validity Study of Neural Machine Translation Metrics
Dmytro Chaplynskyi, Ivan Kulynych, Maria Shvedova and Lesia Ivashkevych |
| 10:55-11:15 | Professional Translators Versus Quality Estimation Models: Reliability and Agreement in English-Ukrainian Translation Evaluation
Dmytro Chaplynskyi, Kyrylo Zakharov and Lesia Ivashkevych |
11:15-11:45 Morning Coffee Break
11:45-13:05 Morning Session: Core NLP & Linguistic Resources
Chair: Oleksii Ignatenko
| 11:45-12:05 | Data-Efficient Adaptation of Multilingual LLMs to Ukrainian
Yurii Paniv, Bohdan Didenko, Mykola Haltiuk, Vladyslav Humennyy, Andrian Kravchenko, Roman Kyslyi, Viktoriia Makovska, Artem Orlovskyi, Bohdan Ruban, Maksym-Yurii Rudko, Anastasiia Senyk, Nazarii Drushchak, Dmytro Chaplynskyi and Mariana Romanyshyn |
| 12:05-12:25 | Dictionary-Based Speculative Decoding for Non-Latin-Script Languages
Oleksiy Syvokon |
| 12:25-12:45 | How Far Can Prompting Go for Minimal-Edit Ukrainian Grammatical Error Correction?
Kateryna Karpo and Artem Chernodub |
| 12:45-13:05 | Entropy of Ukrainian
Anton Lavreniuk, Mykyta Mudryi and Markiian Chaklosh |
13:05-14:15 Lunch
14:15-15:15 Keynote: Anna Rogers. What’s Next for NLP after LLMs?
15:15-16:15 Afternoon Session: Language Proficiency & Paraphrasing
Chair: Roman Kyslyi
| 15:15-15:35 | Toward a Gold-Standard Benchmark for Evaluating Ukrainian Language Proficiency in LLMs
Svitlana Galeshchuk, Yuliia Maksymiuk, Yuliia Chernobrov, Oleksandra Antoniv, Nina Stankevych, Nataliia Faryna and Oksana Popkova |
| 15:35-15:55 | Automated CEFR-Level Assessment for Ukrainian Texts
Olha Kanishcheva and Mikhail Kopotev |
| 15:55-16:15 | Mining Native Ukrainian Paraphrases: A Multi-Source Comparison
Vladyslav Fesenko, Hanna Dydyk-Meush and Volodymyr Mudryi |
16:15-16:45 Afternoon Coffee Break
16:45-18:00 Social Session: Informal discussion of Day 1 papers
18:00-18:10 End of Day 1
10:00-10:10 Day 2 Welcome
10:10-11:10 Morning Session: Speech, Multimodality & OCR
Chair: Oleksii Ignatenko
| 10:10-10:30 | Scaling ASR for Hutsul Dialect: Multi-Speaker Data Collection, Enhanced Transcription and Cross-Speaker Evaluation
Artem Orlovskyi, Zakhar Guzii, Bohdan Onyshchenko, Roman Kyslyi and Pavlo Khomenko |
| 10:30-10:50 | UkrSL: Towards a Ukrainian Continuous Sign Language Dataset
Oleksandr Sobetskyi, Maryna Kosse, Roman Kyslyi and Angelina Savchenko |
| 10:50-11:10 | Digitizing Old Ukrainian Texts: A Prompt-Based OCR Pipeline and Evaluation Dataset
Dmytro Chaplynskyi and Hanna Dydyk-Meush |
11:10-11:40 Morning Coffee Break
11:40-13:00 Morning Session: Applications, Systems & Tools
Chair: Roman Kyslyi
| 11:40-12:00 | Improving Domain-Specific Translation from English into Ukrainian with Retrieval-Augmented Generation
Anton Shpigunov |
| 12:00-12:20 | UAReviews: A Multi-Task Ukrainian Dataset for Emotion and Intent Classification
Roman Kyslyi, Ihor Pysmennyi and Denys Mykhailov |
| 12:20-12:40 | Graph-Based Detection of Disinformation Narrative Diffusion between Russian and Ukrainian Telegram Channels
Yuliia Vistak, Viktoriia Makovska, Vera Schmitt and Veronika Solopova |
| 12:40-13:00 | A Two-Axis Framework for Analyzing Ukrainian Dialogues
Artem Korotenko and Roman Kyslyi |
13:00-14:15 Lunch
14:15-14:35 Afternoon Session: Applications, Systems & Tools (continued)
Chair: Oleksii Ignatenko
| 14:15-14:35 | Belief Propagation in LLM World Models: Measuring Strategic Information Bias with Prediction Markets
Mykola Khandoga, Yevhen Kostiuk, Anton Polishko, Yuri Filipchuk, Kostiantyn Kozlov, Dmytro Zamriy and Artur Kiulian |
14:35-15:35 Keynote: Yurii Paniv. On Automating Research.
15:35-16:55 Afternoon Session: Shared Task on Multi-Domain Document Understanding
Chair: Roman Kyslyi
| 15:35-15:55 | The UNLP 2026 Shared Task on Multi-Domain Document Understanding
Volodymyr Sydorskyi, Nataliia Romanyshyn, Roman Kyslyi and Olena Nahorna |
| 15:55-16:15 | RAG Pipeline Strategies for Ukrainian Multi-Domain Document Understanding Task
Mykola Nosenko and Pavlo Kilko |
| 16:15-16:35 | An End-to-End Ukrainian RAG for Local Deployment: Optimized Hybrid Search and Lightweight Generation
Mykola Trokhymovych, Yana Oliinyk and Nazarii Nyzhnyk |
| 16:35-16:55 | Qwen Goes Brrr: Off-the-Shelf RAG for Ukrainian Multi-Domain Document Understanding
Anton Bazdyrev, Ivan Bashtovyi, Ivan Havlytskyi, Oleksandr Kharytonov and Artur Khodakovskyi |
17:00-18:00 Panel Discussion: Oleksii Molchanovskyi, Olena Andriienko, and Roman Kyslyi.
Ethics of AI Usage.
18:00-18:15 Closing Words
18:15 Afterparty on campus
Illia Strelnykov, Data Scientist at YouScan, Ukraine

Topic: Leveraging User Feedback to Improve Your Models
While academic research provides a strong foundation for model development, the ultimate goal is to deploy these models in real-world applications, where they interact with actual users. This talk addresses the critical challenge of effectively leveraging user feedback to enhance model performance in practical scenarios. We’ll explore ways to incorporate the highly valuable — yet inherently noisy — user-provided data into model training and fine-tuning pipelines. First, we’ll cover methods for collecting user feedback and the challenges involved in processing it, including issues like bias and conflicting information. Then we will examine various solutions for tackling these challenges and how to use refined feedback for model improvement.
Sebastian Ruder, Research Scientist at Meta, Germany

Topic: Multilinguality in Llama 4 and Beyond
Abstract: Multilingual LLMs have become so powerful that they can be used in real-world conversations in a variety of applications. While this presents many opportunities, it also poses challenges associated with the complexity of natural language. In this talk, I will seek to connect academic research to real-world challenges of multilingual conversational AI. I will first provide an overview of multilinguality in Llama 4, highlighting the importance of evaluation. I will then discuss what it takes to bridge the gap between academic and real-world evaluations. Finally, I will discuss how we can develop models that are useful to speakers in their local context, across the globe and for the Ukrainian language.
Kateryna Burovova, ML Engineer at LetsData

Kateryna specializes in AI-powered solutions for detecting and combating harmful information operations, leveraging NLP and computational social science to create threat detection pipelines that analyze content semantics, user behavior patterns, network dynamics, and other contextual signals.
Nataliia Romanyshyn, AI Specialist at Texty.org.ua

Nataliia focuses on the detection and analysis of Russian disinformation. Her expertise includes natural language processing, specifically topic modeling, named entity recognition, large language models, and multilingual NLP. She plays a key role in developing analytical frameworks that transform complex textual data into actionable insights aimed at uncovering disinformation mechanisms.
Yaroslav Peliushenko, Head of Analytics at Osavul

Yaroslav is the Head of Analysis at Osavul, a technology company developing AI-powered solutions for deep intelligence and countering information threats. At UNLP, he will share insights into how their team analyses and structures information, the frameworks they use, and the thinking behind their approach.
Yuliia Dukach, PhD, Data Journalist and Head of Disinformation Investigations at OpenMinds

Yuliia is an expert in disinformation research with over five years of experience in investigative data journalism. Yuliia applies advanced skills in Python and machine learning to analyze computational propaganda and online misinformation.
Vasyl Starko, Ukrainian Catholic University, Ukraine

Andriy Rysin, Independent researcher, USA

Topic: BRUK Team’s Resources for Ukrainian Corpus Creation
The talk will focus on the key resources and tools developed by the BRUK team for the automatic processing of Ukrainian texts, especially for building Ukrainian corpora. The resources include:
* BRUK (Ukrainian Brown Corpus, a projected one-million-word POS gold standard)
* VESUM (A Large Electronic Dictionary of Ukrainian, over 420,000 lemmas and counting, for POS tagging)
* USL (Ukrainian Semantic Lexicon for semantic tagging).
The tools come in the form of the NLP_UK suite for Ukrainian text tokenization, lemmatization, POS tagging, and cleaning. The application of NLP_UK to build multiple iterations of the GRAC corpus will be discussed.