Select Language

公开数据集

MAGICDATA 汉语普通话朗读语料数据库(测试数据集)

MAGICDATA 汉语普通话朗读语料数据库(测试数据集)

Scene:

Music Analysis

Data Type:

Audio
所需积分:10 去赚积分?
  • 565浏览
  • 14下载
  • 1点赞
  • 收藏
  • 分享

贡献者查看主页

小小程序员

致力于人工智能业务的研究、数据集处理。

Data Preview ? 2.2G

    Data Structure ?

    *数据结构实际以真实数据为准

    MAGICDATA Mandarin Chinese Read Speech Corpus was developed by MAGIC DATA Technology Co., Ltd. and freely published for non-commercial use.

    The contents and the corresponding descriptions of the corpus include:


    • The corpus contains 755 hours of speech data, which is  mostly mobile recorded data.

    • 1080 speakers from different accent areas in China are  invited to participate in the recording.

    • The sentence transcription accuracy is higher than 98%.

    • Recordings are conducted in a quiet indoor environment.

    • The database is divided into training set, validation set, and testing  set in a ratio of 51: 1: 2.

    • Detail information such as speech data coding and speaker information is  preserved in the metadata file.

    • The domain of recording texts is diversified, including interactive  Q&A, music search, SNS messages, home command and control, etc.

    • Segmented transcripts are also provided.

    The corpus aims to support researchers in speech recognition, machine translation, speaker recognition, and other speech-related fields. Therefore, the corpus is totally free for academic use.

    The corpus is a subset of a much bigger data ( 10566.9 hours Chinese Mandarin Speech Corpus ) set which was recorded in the same environment. Please feel free to contact us via business@magicdatatech.com for more details.

    Citation

    Please cite the corpus as "Magic Data Technology Co., Ltd., "http://www.imagicdatatech.com/index.php/home/dataopensource/data_info/id/101", 05/2019".

    about us

    Magic Data Technology Co., Ltd. (referred to as Magic Data) was established in 2016. Through our higher-expertise and higher-precision data services, Magic Data has quickly grown into one of the foremost companies in artificial intelligence industry. We strive to provide the most efficient and highest quality one-stop data services for customers in the fields of speech recognition, intelligent imaging and Natural Language Understanding (NLU). Our services include data scheme design, data collection, data annotation/transcription, etc.

    Contact


    • Tel:  (+86) 10-82527250

    • Email:  business@magicdatatech.com

    • http://www.imagicdatatech.com


    External URL: http://www.imagicdatatech.com/index.php/home/dataopensource/data_info/id/101    Full description from the company website


    0相关评论
    ×

    帕依提提提温馨提示

    该数据集正在整理中,为您准备了其他渠道,请您使用

    注:部分数据正在处理中,未能直接提供下载,还请大家理解和支持。