Abstract
Danish language technology has been hindered by a lack of broad-coverage corpora at the scale modern NLP prefers. This paper describes the Danish Gigaword Corpus, the result of a focused effort to provide a diverse and freely-available one billion word corpus of Danish text. The Danish Gigaword corpus covers a wide array of time periods, domains, speakers’ socio-economic status, and Danish dialects.
| Original language | English |
|---|---|
| Title of host publication | Proceedings of the 23rd Nordic Conference on Computational Linguistics (NoDaLiDa) |
| Number of pages | 9 |
| Publisher | Linköping University Electronic Press |
| Publication date | 2021 |
| Pages | 413-421 |
| Publication status | Published - 2021 |
Cite this
- APA
- Standard
- Harvard
- Vancouver
- Author
- BIBTEX
- RIS