Závěrečná práce: Bc. Jan Franěk: Transformer-Based Word Hyphenation
Diplomová práce
Transformer-Based Word Hyphenation
Anotace
Tato práce prozkoumává aplikaci neuronových sítí založených na architektuře transformerů na úkon identifikace dělících bodů slov přirozeného jazyka, zaměřujíc se zejména na český a německý jazyk. S rostoucími možnosti strojového učení v oblasti zpracování přirozeného jazyka jsme se soustředili na možnosti využití transformerů jako efektivní, kompaktní a obecné alternativy k současným nasazeným řešením …více
Abstract
This thesis investigates the application of transformer-based neural network architectures for the task of end-of-line hyphenation, focusing on the Czech and German languages. With the increasing capabilities of machine learning in natural language processing, we explored whether transformer models can serve as an effective, compact, and generalizable alternative to the current state-of-the-art approaches …více
Zadání práce
The student will find the minimal architecture of the transformer comparable to the current solution (minimal patterns generated by Patgen with optimizations suggested in Ondřej Sojka's thesis). He will design a series of architectural experiments and compare their precision, recall, generalization, effectiveness, and efficiency. Then, in the minimal architecture, he will describe how the cascade of context-dependent exceptions is handled in the transformer without memorization to achieve both high accuracy and generalization (measured by cross-validation).
22. 5. 2025 09:09, doc. RNDr. Petr Sojka, Ph.D., učo 2378
Práce na příbuzné téma
Seznam prací, které mají shodná klíčová slova.
-
Judy
Ing. Jakub Máca -
Automatic Detection of Fake News
Mgr. Martin Bažík -
Extrakce argumentace založená na znalostech
Mgr. Jiří Procházka -
Structured Information Extraction from Pharmaceutical Records
Mgr. Michaela Bamburová -
Automated Solution and Difficulty Prediction for Elementary School Mathematics
Mgr. Milan Horký -
Efficient Hyphenation Pattern Generation of Slavic Languages
Bc. František Hůlka -
Utilisation of language representations for Information Retrieval
Ing. Petr Mička -
Automatic Classification of Legal Documents
Mgr. Jiří Mauritz




