Large Corpus of Czech Parliament Plenary Hearings

Jonáš Kratochvil, Peter Polák, Ondřej Bojar


Abstract
We present a large corpus of Czech parliament plenary sessions. The corpus consists of approximately 1200 hours of speech data and corresponding text transcriptions. The whole corpus has been segmented to short audio segments making it suitable for both training and evaluation of automatic speech recognition (ASR) systems. The source language of the corpus is Czech, which makes it a valuable resource for future research as only a few public datasets are available in the Czech language. We complement the data release with experiments of two baseline ASR systems trained on the presented data: the more traditional approach implemented in the Kaldi ASRtoolkit which combines hidden Markov models and deep neural networks (NN) and a modern ASR architecture implemented in Jaspertoolkit which uses deep NNs in an end-to-end fashion.
Anthology ID:
2020.lrec-1.781
Volume:
Proceedings of the Twelfth Language Resources and Evaluation Conference
Month:
May
Year:
2020
Address:
Marseille, France
Editors:
Nicoletta Calzolari, Frédéric Béchet, Philippe Blache, Khalid Choukri, Christopher Cieri, Thierry Declerck, Sara Goggi, Hitoshi Isahara, Bente Maegaard, Joseph Mariani, Hélène Mazo, Asuncion Moreno, Jan Odijk, Stelios Piperidis
Venue:
LREC
SIG:
Publisher:
European Language Resources Association
Note:
Pages:
6363–6367
Language:
English
URL:
https://aclanthology.org/2020.lrec-1.781
DOI:
Bibkey:
Cite (ACL):
Jonáš Kratochvil, Peter Polák, and Ondřej Bojar. 2020. Large Corpus of Czech Parliament Plenary Hearings. In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 6363–6367, Marseille, France. European Language Resources Association.
Cite (Informal):
Large Corpus of Czech Parliament Plenary Hearings (Kratochvil et al., LREC 2020)
Copy Citation:
PDF:
https://aclanthology.org/2020.lrec-1.781.pdf