Automation and Information TechnologiesDepartment of Automated Systems for DataDepartment of Data Analysis and ProgrammingDmukhtasibovich -Doctor of Physical and MathematicalInstitute of ManagementKazan Federal UniversityKazan National Research TechnologicalNusratullovich -Doctor of TechnicalProfessorUNIVERSITY INSTITUTE OF COMPUTATIONALMay 5, 2026arXiv:2605.03799

Natural Language Processing: A Comprehensive Practical Guide from Tokenisation to RLHF

Mullosharaf K. Arabov

AI Summary

This paper presents a hands-on practicum covering the modern NLP pipeline, from tokenization to RLHF, emphasizing reproducibility and open-weight models. It features twelve sessions combining theory with implementation, evaluation metrics, and assessment criteria, all applied to a single evolving corpus. The practicum incorporates original research on low-resource languages like Tajik and Tatar, providing resources and benchmarks for adapting NLP techniques to data-scarce environments.

Key Contribution

Learn to build and evaluate your own NLP pipeline, from tokenization to RLHF, using open-weight models and reproducible research practices.

Abstract

This preprint presents a systematic, research-oriented practicum that guides the reader through the entire modern NLP pipeline: from tokenisation and vectorisation to fine-tuning of large language models, retrieval-augmented generation, and reinforcement learning from human feedback. Twelve hands-on sessions combine concise theory with detailed implementation plans, formalised evaluation metrics, and transparent assessment criteria. The work is not a conventional textbook: it is designed as a reproducible research artefact where every session requires publishing code, models, and reports in public repositories. All experiments are conducted on a single evolving corpus, and the work advocates open-weight models over commercial APIs, with special attention to the Hugging Face ecosystem. The material is enriched by original research on low-resource languages, incorporating linguistic resources for Tajik and Tatar (subword tokenisers, embeddings, lexical databases, and transliteration benchmarks), demonstrating how modern NLP can be adapted to data-scarce environments. Designed for senior undergraduates, graduate students, and practising developers seeking to implement, compare, and deploy methods from classical ML to state-of-the-art LLM-based systems.

Natural Language Processing Recommendation & Information Retrieval RLHF & Preference Learning

Citation Metrics

Citations0

Influential citations0

References0

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

Natural Language Processing: A Comprehensive Practical Guide from Tokenisation to RLHF

Related Papers