Listening while speaking: Speech chain by deep learning

机译：边听边说：深度学习的语音链

获取原文

获取原文并翻译 | 示例

页面导航

摘要
著录项
相似文献
相关主题

摘要

Despite the close relationship between speech perception and production, research in automatic speech recognition (ASR) and text-to-speech synthesis (TTS) has progressed more or less independently without exerting much mutual influence on each other. In human communication, on the other hand, a closed-loop speech chain mechanism with auditory feedback from the speaker's mouth to her ear is crucial. In this paper, we take a step further and develop a closed-loop speech chain model based on deep learning. The sequence-to-sequence model in close-loop architecture allows us to train our model on the concatenation of both labeled and unlabeled data. While ASR transcribes the unlabeled speech features, TTS attempts to reconstruct the original speech waveform based on the text from ASR. In the opposite direction, ASR also attempts to reconstruct the original text transcription given the synthesized speech. To the best of our knowledge, this is the first deep learning model that integrates human speech perception and production behaviors. Our experimental results show that the proposed approach significantly improved the performance more than separate systems that were only trained with labeled data.

机译：尽管语音感知和产生之间有着密切的关系，但自动语音识别（ASR）和文本语音合成（TTS）的研究或多或少地独立进行，彼此之间没有太多相互影响。另一方面，在人类交流中，具有从说话者的嘴到她的耳朵的听觉反馈的闭环语音链机制至关重要。在本文中，我们将进一步采取措施，并基于深度学习开发闭环语音链模型。闭环体系结构中的序列到序列模型使我们能够在标记数据和未标记数据的串联上训练我们的模型。当ASR转录未标记的语音特征时，TTS会尝试根据ASR的文本来重建原始语音波形。在相反的方向上，ASR还尝试在给定合成语音的情况下重建原始文本转录。据我们所知，这是第一个将人类语音感知和生产行为整合在一起的深度学习模型。我们的实验结果表明，与仅使用标记数据进行训练的单独系统相比，所提出的方法可显着提高性能。

著录项

来源
《2017 IEEE Automatic Speech Recognition and Understanding Workshop》|2017年|301-308|共8页
会议地点 Okinawa(JP)
作者
Andros Tjandra; Sakriani Sakti; Satoshi Nakamura;
展开▼
作者单位

Graduate School of Information Science, Nara Institute of Science and Technology, Japan;

Graduate School of Information Science, Nara Institute of Science and Technology, Japan;

Graduate School of Information Science, Nara Institute of Science and Technology, Japan;

展开▼
会议组织
原文格式 PDF
正文语种 eng
中图分类
关键词
Speech; Speech processing; Decoding; Hidden Markov models; Machine learning; Spectrogram; Data models;

机译：语音;语音处理;解码;隐马尔可夫模型;机器学习;声谱图;数据模型;;

相似文献

外文文献
中文文献
专利

1. SpeakerBeam: A New Deep Learning Technology for Extracting Speech of a Target Speaker Based on the Speaker’s Voice Characteristics [J] . Marc Delcroix, Katerina Zmolikova, Keisuke Kinoshita, NTT Technical Review . 2018,第11期

机译：SpeakerBeam：一种新的深度学习技术，用于根据说话者的语音特征提取目标说话者的语音
2. An Efficient Deep Learning Based Method for Speech Assessment of Mandarin-Speaking Aphasic Patients [J] . Mahmoud Seedahmed S., Kumar Akshay, Tang Yiting, Biomedical and Health Informatics, IEEE Journal of . 2020,第11期

机译：一种高效的基于深度学习的讲话评估方法，讲话者的失性患者
3. Developing AI that Pays Attention to Who You Want to Listen to: Deep-learning-based Selective Hearing with SpeakerBeam [J] . Marc Delcroix, Tsubasa Ochiai, Hiroshi Sato, NTT Technical Review . 2021,第9期

机译：开发一个人注意到要听取谁的人：基于深度学习的选择性听力与扬声器
4. Listening while speaking: Speech chain by deep learning [C] . Andros Tjandra, Sakriani Sakti, Satoshi Nakamura IEEE Workshop on Automatic Speech Recognition and Understanding . 2017

机译：在说话时听：深度学习的言论
5. Deep learning for speech classification and speaker recognition [D] . Saleem, Muhammad Muneeb. 2014

机译：深度学习用于语音分类和说话人识别
6. Speech Audiometry at Home: Automated Listening Tests via Smart Speakers With Normal-Hearing and Hearing-Impaired Listeners [O] . Jasper Ooster, Melanie Krueger, Jörg-Hendrik Bach, 2020

机译：主页言语听力测量：通过智能扬声器自动聆听测试具有正常听力和听力受损的听众
7. Listening while Speaking: Speech Chain by Deep Learning [O] . Tjandra, Andros, Sakti, Sakriani, Nakamura, Satoshi 2017

机译：口语聆听：深度学习的语音链

Listening while speaking: Speech chain by deep learning

摘要

著录项

相似文献

相关主题

期刊订阅