An n-gram based approach to the automatic classification of schoolchildren’s writing
DOI:
https://doi.org/10.35869/vial.v0i16.93Keywords:
writing, n-grams, primary school, morphological categories, automatic classificationAbstract
This article focuses on the analysis of schoolchildren’s writing (throughout the whole primary school period) using sets of morphological labels (n-grams). We analyzed the sets of bigrams and trigrams from a group of literary texts written by Catalan schoolchildren in order to identify which bigrams and trigrams can help discriminate between texts from the three cycles into which the Spanish primary education system is divided: lower cycle (6- and 7-year-olds), middle cycle (8- and 9-year- olds) and upper cycle (10- and 11-year-olds). The results obtained are close to 70% of correct classifications (77.5% bigrams and 68.6% trigrams), making this technique useful for automatic document classification by age.
Downloads
Downloads
Published
Issue
Section
License
Revistas_UVigo es el portal de publicación en acceso abierto de las revistas de la Universidade de Vigo. La puesta a disposición y comunicación pública de las obras en el portal se efectúa bajo licencias Creative Commons (CC).
Para cuestiones de responsabilidades, propiedad intelectual y protección de datos consulte el aviso legal de la Universidade de Vigo.