COCT 書面語語料庫2017
Metadata for COCT 書面語語料庫2017
Corpus title COCT 書面語語料庫2017
CQPweb's short handles for this corpus yl2017v2 / YL2017V2
Total number of texts in corpus 159,051
Total word tokens in all corpus texts 303,766,044
Word types in the corpus 2,522,973
Standardised type:token ratio (1,000-token basis) 0.4359 types per token
Non-standardised type:token ratio 0.0083 types per token
Text metadata and word-level annotation
The database stores the following information for each text in the corpus: 圖書分類
出版年份
The primary classification of texts is based on: 圖書分類
Words in this corpus are annotated with: pos
The primary word-level annotation scheme is: No primary word-level annotation scheme has been set