Skip to main content

BIM 432 - Natural Language Processing

Faculty of Engineering and Natural Sciences · Computer Engineering · Undergraduate

ECTS: 5 T+P+L: 2+0+1 Departmental Elective
Coordinator: Dr. Öğr. Üyesi Şengül BAYRAK

Course Objective

The aim of this course is to teach the basic concepts and techniques in the field of Natural Language Processing (NLP), to develop skills in working with text data, and to enable students to implement various NLP applications with Python and related libraries. In addition, the course aims to enable students to practice language models, sentiment analysis, text summarization, machine translation and chatbot development, and to develop solutions to real-world NLP problems.

Course Content

The definition of NLP, its uses and the structure of the language (morphology, syntax, semantics) will be explained. Basic usage of Python and popular NLP libraries (NLTK, spaCy, Hugging Face) will be demonstrated. Students will learn how to process text data using tokenization, stopwords extraction, lemmatization and stemming. Students should practice text preprocessing steps with Python. Bag of Words (BoW) and TF-IDF methods are used to convert text to numeric data. Discover relationships between words with Word2Vec, GloVe and FastText. Logistic regression and Naive Bayes classifiers are used for sentiment analysis. RNN, LSTM, GRU and Transformer architectures are introduced. The basics of language models such as BERT, GPT, T5 and their use with Hugging Face are explained. Turkish NLP challenges and preprocessing operations on Turkish texts using Zemberek library. Students practice on Turkish texts. Noun entity recognition (NER) and word sense disambiguation (WSD) methods are taught. NER model is developed with spaCy to recognize entities in texts. Machine translation methods (rule-based, statistical, neural) and Transformer-based models are taught. English-Turkish translation is practiced with MarianMT. Abstractive and extractive summarization methods and TextRank algorithm are explained. Students summarize Turkish news texts. Rule-based and learning-based chatbots are introduced. University information chatbot is developed using Rasa and DialoGPT. Knowledge extraction and Transformer based question-answer systems are taught. Question and answer applications with BERT. Big language models (GPT-4, LLaMA) and ethical issues in NLP (bias, misinformation) are discussed. Text generation is done using GPT-4 API. Challenges in NLP projects and model evaluation metrics (Perplexity, BLEU, ROUGE, F1-score) are taught. Students perform model optimization on the projects. Students present their projects and demonstrate what they have learned.

Required Resources

Daniel Jurafsky, and James H. Martin, "Speech and Language Processing", Third Edition, Prentice Hall, 2020.

Recommended Resources

1- Indurkhya, N., & Damerau, F. J. (2010). Handbook of natural language processing. Chapman and Hall/CRC.

2- Manning, C., & Schutze, H. (1999). Foundations of statistical natural language processing. MIT press.

Explanations

1. Quiz and/or homework will be related to the topics of the week. 1 homework/quiz will be given.

2. For the application part, students are expected to develop projects in accordance with the content of the course and the projects developed by the students are expected to be presented by the students in the 14th and 15th weeks.

Rules

1. Attendance: According to the regulations, if a student fails to attend 30% of the total class time, he/she will be marked absent (DZ) and will fail the course.

2. Academic honesty: In cases of unethical behavior such as plagiarism, copying, etc., you will be subject to disciplinary proceedings and penal sanctions in accordance with IZU's relevant regulations.

Course Learning Outcomes

  1. Explain the fundamental concepts of Natural Language Processing and text processing workflows.
  2. Apply text preprocessing techniques such as tokenization, normalization, stemming, and lemmatization.
  3. Develop machine learning and deep learning models for NLP tasks such as text classification and sentiment analysis.
  4. Evaluate and interpret the performance of NLP models using metrics such as accuracy, precision, recall, and F1-score.
  5. Design and implement an NLP project on real-world text data, and analyze and report the results in a technical manner.
  6. Develops engineering solutions for natural language processing and machine learning systems while taking into account ethical, privacy, and social implications.

Core Area Distribution

(46) Mathematics and Statistics%30 (48) Computing%40 (52) Engineering and Engineering Trades%30

Teaching Methods

ExpressionQuestion-AnswerDiscussionExercise and PracticeGroup StudyBrain StormingExperiment - Test / Lab/ Workshop / Field PracticeSelf studyProblem Solving

Assessment & Evaluation

HomeworkPerformance Assignment ( Lab / Workshop / Field Work / Seminar / Presentation / Completion Study / ThesisProject / DesignTesting (Essay / Tests: True-Falls, multiple-choice, short answer, matching)

ECTS / Workload

ActivityQuantityDuration (h)Total Workload
Course Duration (Including Exam Week)16348
Out of Class Study Period000
Midterm122
Quiz000
Assignment177
Practice11010
Final122

Course Schedule

WeekSubjectPreparation
1Introduction to Natural Language Processing and the Basic Conceptslecture notes
2Text Preprocessinglecture notes
3Numeric Representations - TF- IDF and Word Embeddingslecture notes
4Sentiment Analysis and Text Classificationlecture notes
5Language Models and Introduction to Transformerslecture notes
6Turkish Natural Language ProcessingLecture notes
7Named Entity Recognition (NER) and Word Sense Disambiguation (WSD)lecture notes
8MidtermMidterm
9Machine Translation and Language Modelinglecture notes
10Automatic Text Summarizationlecture notes
11Chatbot and Conversation Systemslecture notes
12Knowledge Extraction and Question and Answer Systemslecture notes
13Large Language Models (LLM) and Ethical Issueslecture notes
14Project presentations-
15Project presentations-
16FinalFinal