> For the complete documentation index, see [llms.txt](https://doc.thordata.com/doc/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://doc.thordata.com/doc/zh/chang-jian-wen-ti/chan-pin-wen-ti/duo-mo-tai-shu-ju-ji/thordata-de-shu-ju-ji-dou-yong-zai-na-xie-ai-ying-yong-shang.md).

# Thordata 的数据集都用在哪些 AI 应用上？

覆盖语音识别、大模型训练与评测、计算机视觉、电商情报、内容风控和具身智能机器人等应用方向。

具体例子：小语种 ASR 语料（印尼语、泰语、马来语、越南语、口音英语）用于训练跨境客服与本地化场景的语音识别；200 亿条 YouTube 评论与 900 亿条 TikTok 元信息支撑推荐、趋势分析与情感模型；11 种语言的 OCR 图片数据集服务文档数字化；900 万题的职业考试题库用于大模型评测与微调；覆盖居家、办公、工厂场景的第一人称 Ego 视频（9 万+ 小时）训练具身智能与 VLA 模型。如果您的应用不在这些方向上，也可以围绕具体场景定制采集。
