> For the complete documentation index, see [llms.txt](https://doc.thordata.com/doc/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://doc.thordata.com/doc/faq/product-problem/multimodal-dataset/which-ai-applications-are-thordatas-datasets-used-for.md).

# Which AI applications are Thordata's datasets used for?

Thordata's datasets power applications across speech recognition, LLM training and evaluation, computer vision, e-commerce intelligence, content moderation, and embodied AI robotics.

Concrete examples: the low-resource ASR corpora (Indonesian, Thai, Malay, Vietnamese, accented English) train speech recognition for cross-border customer service and localization; the 20-billion-entry YouTube comment and 90-billion-record TikTok metadata collections feed recommendation, trend analysis, and sentiment models; 11-language OCR image datasets support document digitization; the 9-million-question professional exam bank serves LLM evaluation and fine-tuning; and first-person Ego videos covering home, office, and factory scenes train embodied AI and VLA models. If your application falls outside these areas, a custom collection can be scoped around your exact scenario.
