> For the complete documentation index, see [llms.txt](https://doc.thordata.com/doc/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://doc.thordata.com/doc/faq/product-problem/multimodal-dataset/which-languages-and-regions-do-thordatas-datasets-cover.md).

# Which languages and regions do Thordata's datasets cover?

Coverage spans major global languages plus a strong focus on low-resource and regional languages — particularly Southeast Asian languages that are poorly served elsewhere.

Verified language coverage includes: speech corpora in Indonesian (3,200 hours), Thai (800 hours), Malay (500 hours), Vietnamese, accented English (Thai, Indonesian, and Malaysian accents), and Japanese conversational audio; OCR image datasets in 11 languages — Chinese, English, Japanese, Korean, French, German, Spanish, Italian, Portuguese, Russian, and Indonesian; and text datasets covering global platform content from YouTube, TikTok, and X. For languages or regional variants not listed, custom collection can often be arranged — low-resource and long-tail languages are a deliberate specialty, not an afterthought.
