Filters
Search
Product type
Language
Country
Year of Collection

Korean (South Korea) scripted microphone

More info
Common Use CasesASR, Virtual Assistant, Chatbot
Dataset IDKOR_ASR001
TypeAudio
Unit20 hours
LanguageKorean
CountrySouth Korea

Latin American Spanish Inverse text normalisation

More info
Common Use CasesASR, Language Modelling, Closed Captioning
Dataset IDSPA_ITN001
TypeText
Unit3795 test cases
LanguageSpanish
CountryN/A

Mandarin Chinese (China) scripted microphone

More info
Common Use CasesASR, Virtual Assistant, Chatbot
Dataset IDMAC_ASR002
TypeAudio
Unit26 hours
LanguageMandarin Chinese
CountryChina

Mandarin Chinese Inverse text normalisation

More info
Common Use CasesASR, Language Modelling, Closed Captioning
Dataset IDCMN_ITN001
TypeText
Unit4230 test cases
LanguageMandarin Chinese
CountryN/A

Polish (Poland) scripted microphone

More info
Common Use CasesASR, Virtual Assistant, Chatbot
Dataset IDPOL_ASR001
TypeAudio
Unit25 hours
LanguagePolish
CountryPoland

Polish (Poland) scripted smartphone

More info
Common Use CasesASR, Virtual Assistant, Chatbot
Dataset IDPOL_ASR002_CN
TypeAudio
Unit293 hours
LanguagePolish
CountryPoland

Get Started with Off-the-Shelf AI Training Datasets

Appen’s extensive catalog of off-the-shelf (OTS) datasets spans multiple data types and industries, providing comprehensive coverage for various AI applications. These datasets are crafted to the highest standards of quality and accuracy, ensuring reliable training data for AI models.

Talk to an expert