Skip to main content
Skip header

Natural Language Processing

Type of study Follow-up Master
Language of instruction Czech
Code 460-4168/01
Abbreviation ZPJ
Course title Natural Language Processing
Credits 5
Coordinating department Department of Computer Science
Course coordinator doc. Mgr. Jiří Dvorský, Ph.D.

Subject syllabus

- Fundamentals of text processing
Introduction to NLP. Pattern matching and regular expressions for searching patterns in text. Text tokenization. Stemming and lemmatization. Sentence segmentation. Edit distance, measuring similarity between words.

- Statistical language models (N-grams)
N-grams: unigrams, bigrams, trigrams. Calculating sequence probabilities. Perplexity.

- Text classification – Naive Bayes
The principle of the Naive Bayes classifier. Training the model on data. Evaluation metrics: Precision, Recall, F1 score. Examples: sentiment analysis, spam detection.

- Text classification – Logistic regression
Linear and logistic regression. Representation of text using features. Calculation of weights using gradient descent. Prevention of overfitting. Multinomial logistic regression.

- Vector model and Word Embeddings
Traditional vector model: TF-IDF. Word Embeddings: Word2vec, Skip-gram, Continuous bag of words. Properties of embeddings. Bias in data and models.

- Introduction to neural networks for NLP
Perceptron. Feedforward neural networks. Activation functions (ReLU, sigmoid). Network training. Use for text classification.

- Recurrent Neural Networks (RNN and LSTM)
The principle of recurrent neural networks (RNN) and their internal state (memory). The problem of vanishing and exploding gradients. LSTM (Long Short-Term Memory) and GRU cells. Bidirectional and layered RNN architecture.

- Encoder-Decoder Model and Attention Mechanism
Encoder-Decoder architecture. Attention mechanism.

- Transformer Architecture
Self-Attention. Multi-Head Attention. Positional Encoding. Transformer block structure (encoder and decoder).

- Large Language Models (LLM)
GPT (Generative Pre-trained Transformer). Pre-training principle. Text generation and sampling techniques.

- Bidirectional BERT Models and Fine Tuning
BERT (Bidirectional Encoder Representations from Transformers). Masked Language Modeling (MLM). Fine tuning: adapting a pre-trained model to specific tasks (classification, entity recognition - NER).

- The future of NLP
Prompting: Zero-shot, one-shot, and few-shot learning. Chain-of-Thought prompting for complex tasks. RLHF (Reinforcement Learning from Human Feedback).

Literature

1. Jurafsky, Dan, and James H. Martin. Speech and Language Processing. 3rd ed. draft, Stanford University, 2025, web.stanford.edu/~jurafsky/slp3/. Accessed 30 Aug. 2025.
2. Tunstall, Lewis, et al. Natural Language Processing with Transformers. O'Reilly Media, 2022.
3. Manning, Christopher D., and Hinrich Schütze. Foundations of Statistical Natural Language Processing. MIT Press, 1999.
4. Goodfellow, Ian, et al. Deep Learning. MIT Press, 2016.

Advised literature

1. Manning, C. D.; Raghavan, P. & Schutze, H.: Introduction to Information Retrieval, Cambridge University Press, 2008. Dostupné z https://nlp.stanford.edu/IR-book/
2. Witten I. H., Moffat A., Bell T. C.: Managing Gigabytes (2nd ed.): Compressing and Indexing Documents and Images, Morgan Kaufmann Publishers Inc., 1999, ISBN 1-55860-570-3