For the complete documentation index, see llms.txt. This page is also available as Markdown.

💡Which languages and regions do Thordata's datasets cover?

Coverage spans major global languages plus a strong focus on low-resource and regional languages — particularly Southeast Asian languages that are poorly served elsewhere.

Verified language coverage includes: speech corpora in Indonesian (3,200 hours), Thai (800 hours), Malay (500 hours), Vietnamese, accented English (Thai, Indonesian, and Malaysian accents), and Japanese conversational audio; OCR image datasets in 11 languages — Chinese, English, Japanese, Korean, French, German, Spanish, Italian, Portuguese, Russian, and Indonesian; and text datasets covering global platform content from YouTube, TikTok, and X. For languages or regional variants not listed, custom collection can often be arranged — low-resource and long-tail languages are a deliberate specialty, not an afterthought.

Last updated