Volver a resultados
Ficha bibliográfica · Consulta y acceso
Artículo

Segmenting corpora of texts

Tony Berber Sardinha · Pontifícia Universidade Católica de São Paulo - PUC-SP · 2002

Acceso abierto disponible
Lectura rápida. Revisá los datos básicos del recurso y luego accedé al contenido desde el botón principal. En esta ficha solo se muestra la información necesaria para identificar la obra, citarla y abrirla.

Acceso al recurso

Entrá al contenido desde la opción principal o elegí otra fuente disponible.

DOAJ DOAJ Articles
Entrar por DOAJ
Acceso principal

Acceso abierto disponible

Recurso identificado como acceso abierto, sin confirmar automáticamente si es texto completo directo.
Abrir recurso

Resumen

Descripción general del contenido del recurso.

The aim of the research presented here is to report on a corpus-based method for discourse analysis that is based on the notion of segmentation, or the division of texts into cohesive portions. For the purposes of this investigation, a segment is defined as a contiguous portion of written text consisting of at least two sentences. The segmentation procedure developed for the study is called LSM (link set median), which is based on the identification of lexical repetition in text. The data analysed in this investigation were three corpora of 100 texts each. Each corpus was composed of texts of one particular genre: research articles, annual business reports, and encyclopaedia entries. The total number of words in the three corpora was 1,262,710 words. The segments inserted in the texts by the LSM procedure were compared to the internal section divisions in the texts. Afterwards, the results obtained through the LSM procedure were then compared to segmentation carried out at random. The results indicated that the LSM procedure worked better than random, suggesting that lexical repetition accounts in part for the way texts are segmented into sections.

Cómo citar

Elegí el formato que necesitás y copiá la referencia al portapapeles.

APA 7

Sardinha, T. B. (2002). Segmenting corpora of texts. https://doi.org/10.1590/S0102-44502002000200004

MLA

Sardinha, Tony Berber. "Segmenting corpora of texts." 2002. https://doi.org/10.1590/S0102-44502002000200004.

Chicago

Sardinha, Tony Berber. 2002. "Segmenting corpora of texts.". https://doi.org/10.1590/S0102-44502002000200004.

Harvard

Sardinha, T. B. 2002, Segmenting corpora of texts, Pontifícia Universidade Católica de São Paulo - PUC-SP, available at: https://doi.org/10.1590/S0102-44502002000200004 [Accessed 5 Aug. 2026].

Compartir e imprimir

Guardá la ficha, copiá su enlace permanente o imprimila como PDF.

Exportar referencia

Si usás un gestor bibliográfico, podés exportar el registro en los formatos más comunes.

Detalles del recurso

Información bibliográfica útil para confirmar que se trata del material correcto.

Título
Segmenting corpora of texts
Autor / colaboradores
Tony Berber Sardinha
Editorial
Pontifícia Universidade Católica de São Paulo - PUC-SP
Año de publicación
2002
ISSN
1678-460X
ISSN
1678-460X
Idioma
Inglés

Materias

Explorá otros recursos relacionados a partir de estas materias.

Copiado