Te Hiku Media te reo Māori corpus
Te Hiku MediaTe Hiku Media's te reo Māori corpus is a labeled speech collection that the iwi broadcaster uses to build Māori speech recognition. It combines decades of archived radio interviews with native speakers, many now deceased, and sentences read aloud by the public through the Kōrero Māori site. Te Hiku Media keeps the data under Māori stewardship.
Openness
2 high confidence- license
- Kaitiakitanga-License(Requires permission to access or use and forbids commercial use without explicit grant.)
- access
- manual(access by request to Te Hiku Media, which prioritizes requests from Indigenous communities)
- dataset_card
- no(No card
Te Hiku Media holds the recordings under its own stewardship license and grants access on request, giving priority to Indigenous communities. The license bars commercial use unless Te Hiku Media grants it explicitly, case by case, so commercial use stays with the community rather than being open to anyone.
- https://blog.papareo.nz/whisper-is-another-case-study-in-colonisation/ recorded 2026-09-26
"Our labelled te reo Māori speech corpus is upwards of 500 hours"; "We treat data as taonga".
- https://raw.githubusercontent.com/TeHikuMedia/Kaitiakitanga-License/master/LICENSE.md recorded 2026-09-26
"You must contact us and seek permission to access, use, contribute towards, or modify code in this repository"; "You may not use code ... for commercial purposes unless we explicitly grant you the right".
- https://raw.githubusercontent.com/TeHikuMedia/Kaitiakitanga-License/master/README.md recorded 2026-09-26
"You need permission from Te Reo Irirangi o Te Hiku o Te Ika (Te Hiku Media) if you would like to use or adapt this license."
- https://tehiku.nz/te-hiku-tech/te-hiku-dev-korero/25141/data-sovereignty-and-the-kaitiakitanga-license recorded 2026-09-26
"We will prioritise requests from Indigenous communities and individuals/organisations that are committed to supporting Indigenous communities".
Adoption
not assessedAccess to the recordings is granted on request by Te Hiku Media rather than through a public download, and nothing it publishes counts the requests or uses, so there is no usage figure to read.
- https://tehiku.nz/te-hiku-tech/te-hiku-dev-korero/25141/data-sovereignty-and-the-kaitiakitanga-license recorded 2026-09-24
"We will prioritise requests from Indigenous communities and individuals/organisations that are committed to supporting Indigenous communities"
Capability
2 medium confidenceThe corpus holds archival speech from native speakers and trains Te Hiku Media's own speech recognizers, including a fine-tuned Whisper model. It is described only in the owner's blog posts, with no paper or card, which places it below BibleTTS, whose recordings and alignment are set out in a paper.
- https://blog.papareo.nz/whisper-is-another-case-study-in-colonisation/ recorded 2026-09-24
"Our labelled te reo Māori speech corpus is upwards of 500 hours, but we decided to fine tune Whisper on about 125 hours of labelled speech"; examples "from ... Legacy Collection of interviews recorded for radio decades ago".
- https://tehiku.nz/te-hiku-tech/korero-maori/ recorded 2026-09-24
"Visit koreromaori.com and read te reo Māori sentences so we can teach machines te reo Māori."
Verified 2026-09-24