MARSEC: A Machine-Readable Spoken English Corpus
- 1 December 1993
- journal article
- research article
- Published by Cambridge University Press (CUP) in Journal of the International Phonetic Association
- Vol. 23 (2) , 47-54
- https://doi.org/10.1017/s0025100300004849
Abstract
The purpose of this paper is to describe a new version of the Spoken English Corpus which will be of interest to phoneticians and other speech scientists. The Spoken English Corpus is a well-known collection of spoken-language texts that was collected and transcribed in the 1980's in a joint project involving IBM UK and the University of Lancaster (Alderson and Knowles forthcoming, Knowles and Taylor 1988). One valuable aspect of it is that the recorded material on which it was based is fairly freely available and the recording quality is generally good. At the time when the recordings were made, the idea of storing all the recorded material in digital form suitable for computer processing was of limited practicality. Although storage on digital tape was certainly feasible, this did not provide rapid computer access. The arrival of optical disk technology, with the possibility of storing very large amounts of digital data on a compact disk at relatively low cost, has brought about a revolution in ideas on database construction and use. It seemed to us that the recordings of the Spoken English Corpus (hereafter SEC) should now be converted into a form which would enable the user to gain access to the acoustic signal without the laborious business of winding through large amounts of tape. Once this was done, we should be able not only to listen to the recordings in a very convenient way, but also to carry out many automatic analyses of the material by computer.Keywords
This publication has 4 references indexed in Scilit:
- The role of context in the automatic recognition of stressed syllablesPublished by International Speech Communication Association ,1993
- The machine-readable Spoken English CorpusPublished by Brill ,1993
- From text to waveform: converting the Lancaster/IBM Spoken English Corpus into a speech databasePublished by Brill ,1993
- Prosodic labelling: The problem of tone group boundariesPublished by Walter de Gruyter GmbH ,1991