Lecture 2. Tokenization and word counts

Материал из Wiki - Факультет компьютерных наук
Версия от 20:44, 22 августа 2015; Polidson (обсуждение | вклад) (Новая страница: «== How many words? == == Zipf's law == == Heaps' law == == Why tokenization is difficult? == == Rule-based tokenization == == Sentence segmentation == == N…»)
(разн.) ← Предыдущая версия | Текущая версия (разн.) | Следующая версия → (разн.)
Перейти к навигации Перейти к поиску

How many words?

Zipf's law

Heaps' law

Why tokenization is difficult?

Rule-based tokenization

Sentence segmentation

Natural Language Toolkit (NLTK)

Learning to tokenize

Exercise 1.1 Word counts

Lemmatization (Normalization)

Stemming

Exercise 1.2 Word counts (continued)

Exercise 1.3 Do we need all words?