COCT 口語語料庫 (2026彙整)
Metadata for COCT 口語語料庫 (2026彙整)
Corpus title COCT 口語語料庫 (2026彙整)
CQPweb's short handles for this corpus da_spoken / DA_SPOKEN
Total number of texts in corpus 12,627
Total word tokens in all corpus texts 44,125,306
Word types in the corpus 296,603
Standardised type:token ratio (1,000-token basis) 0.3716 types per token
Non-standardised type:token ratio 0.0067 types per token
Text metadata and word-level annotation
The database stores the following information for each text in the corpus: 語料採購年份
The primary classification of texts is based on: 語料採購年份
Words in this corpus are annotated with: Part of Speech
The primary word-level annotation scheme is: Part of Speech