COCT 書面語語料庫2019
Analyse corpus

This page contains controls for advanced corpus analysis functions. WARNING: currently under development. You have been warned.

Click here for the experimental Lexical Growth Curve tool

Select analysis
Choose an option for corpus analysis:
Saved feature matrices
Name Subcorpus Object unit N features N objects Date created Actions
You have no saved features matrices.
Design feature matrix for multivariate analysis
Explanation of the use of feature matrices goes here. Note it can be used for PCA, cluster analysis or factor analysis.
Select unit of analysis
Choose a unit of analysis (for factoring, clustering, etc.) At the moment, the only choice is "text".
Define object labelling method
All data objects (e.g. texts) in a feature matrix need to have a label.
Choose one of the methods opposite for creation of object labels.
Select texts
Select a subcorpus or the full corpus.
Only the texts in the subcorpus you select will be included in the feature matrix.
Select features (from saved queries)
Use the tickboxes below to select the saved queries you want to include as features.
Use? Name No. of hits Date Discount?
(100% = No discount)
 
You don't have any saved queries.
 
Select features (based on query permutation)
This is for features whose value can only be deduced by mathemtaical manipulation of more than one saved query. Typical example: where a feature is equal to (search for soemthing ) minus (search for something elsE)
Use? Operand # 1 Op Operand # 2 Discount?
(100% = No discount)
 
You don't have any saved queries.
 
Select additional features
Use? Description

Allow extra features to be added that are not queries. The list of these is:

  • Standardised type-token ratio (by increments of 400/1,000/2,000)
  • Average word length (optionally limited to regular forms)

Other statistical features can be defined via the saved-query feature function: e.g. lexical density.


Every token, including punctuation, is included.

Excludes tokens that are punctuation, formulae, or otherwise unlikely to be “real” words.
Enter a name for this new feature matrix: