Lecture 2. Tokenization and word counts: различия между версиями
Перейти к навигации
Перейти к поиску
Polidson (обсуждение | вклад) Новая страница: «== How many words? == == Zipf's law == == Heaps' law == == Why tokenization is difficult? == == Rule-based tokenization == == Sentence segmentation == == N…» |
Polidson (обсуждение | вклад) |
||
| Строка 1: | Строка 1: | ||
== How many words? == | == How many words? == | ||
"The rain in Spain stays mainly in the plain." | |||
9 '''tokens''': The, rain, in, Spain, stays, mainly, in, the, plain | |||
7 (or 8) '''types''': T = the rain, in, Spain, stays, mainly, plain | |||
=== Type and token === | |||
''Type'' is an element of the vocabulary. | |||
''Token'' is an instance of that type in the text. | |||
N = number of tokens; | |||
V - vocabulary (i.e. all types); | |||
|V| = size of vocabulary (i.e. number of types). | |||
How are N and |V| related? | |||
== Zipf's law == | == Zipf's law == | ||
Версия от 20:49, 22 августа 2015
How many words?
"The rain in Spain stays mainly in the plain." 9 tokens: The, rain, in, Spain, stays, mainly, in, the, plain 7 (or 8) types: T = the rain, in, Spain, stays, mainly, plain
Type and token
Type is an element of the vocabulary.
Token is an instance of that type in the text.
N = number of tokens;
V - vocabulary (i.e. all types);
|V| = size of vocabulary (i.e. number of types).
How are N and |V| related?