> For the complete documentation index, see [llms.txt](https://doc.thordata.com/doc/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://doc.thordata.com/doc/faq/product-problem/multimodal-dataset/what-types-of-multimodal-datasets-does-thordata-offer.md).

# What types of multimodal datasets does Thordata offer?

Thordata offers multimodal datasets across seven major modalities: video, audio, image, text, question banks, embodied AI (Ego), and 3D data — covering everything from off-the-shelf collections to fully customized builds.

Representative datasets include: low-resource ASR speech corpora with precisely aligned transcripts (3,200 hours of Indonesian, 800 hours of Thai, 500 hours of Malay, plus Vietnamese and accented English); 20 billion YouTube comments; 90 billion TikTok metadata records; 1 billion X (formerly Twitter) posts; OCR image datasets in 11 languages including Chinese, English, Japanese, Korean, and major European languages; a professional exam question bank of 9 million questions; first-person Ego videos spanning home, office, and factory scenes (90,000+ hours); and 3D scanner datasets.&#x20;
