【24h】

Formosa Speech Recognition Challenge 2020 and Taiwanese Across Taiwan Corpus

机译:Formosa语音识别挑战2020和台湾跨台湾语料库

获取原文

摘要

Taiwanese (a.k.a. Taiwanese Hokkien, Hoklo, Taigi, Southern Min or Min-Nan) is an endangered language, because the domination of Mandarin, the number of Taiwanese speakers continues to drop, especially among the youth generations. In addressing this problem, a Taiwanese speech-enabled human-computer interface for supporting people's daily life is essential. Therefore, a Formosa Speech in the Wild (FSW) project was established to collect a large-scale Taiwanese speech across Taiwan (TAT) corpus to boost the development of Taiwanese speech recognition (TSR). A Formosa Speech Recognition Challenge 2020 (FSR-2020) was also hosted to promote the corpus as well as to evaluate the performance of state-of-the-art TSR systems. This paper briefly introduces TAT corpus and FSR-2020 challenge, presents the provided data profile, evaluation plan and reports experimental baseline results. A subset of TAT corpus, TAT-Vol1, is given away for free for all participants (non-commercial license), and its corresponding Kaldi baseline recipes have been published online. Experimental results have showed that the combination of TAT corpus and the baseline recipes is a good resource pack for TSR research and development.
机译:台湾人(A.K.A.台湾北海道,北海,太极拳,南部分钟或闽南)是一种濒临灭绝的语言,因为普通话的统治,台湾扬声器的数量继续下降,特别是青年世代。在解决这个问题时,为支持人们日常生活的一个台湾支持的人类计算机界面至关重要。因此,建立了野外(FSW)项目中的福尔科斯言论,以收集台湾(TAT)语料库的大规模台湾演讲,以提高台湾语音识别(TSR)的发展。 OFFOOSA语音识别挑战2020(FSR-2020)也被托管,以促进语料库,并评估最先进的TSR系统的性能。本文简要介绍了TAT语料库和FSR-2020挑战,提出了提供的数据配置文件,评估计划和报告实验基线结果。对于所有参与者(非商业许可),免费对TAT-Vol1提供TAT-Vol1的一个小词组,其相应的Kaldi基准配方在线发布。实验结果表明,TAT语料库和基线配方的组合是TSR研究和开发的良好资源包。

相似文献

  • 外文文献
  • 中文文献
  • 专利
获取原文

客服邮箱:kefu@zhangqiaokeyan.com

京公网安备:11010802029741号 ICP备案号:京ICP备15016152号-6 六维联合信息科技 (北京) 有限公司©版权所有
  • 客服微信

  • 服务号