> For the complete documentation index, see [llms.txt](https://doc.thordata.com/doc/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://doc.thordata.com/doc/faq/product-problem/multimodal-dataset/how-is-video-data-structured-for-multimodal-training.md).

# How is video data structured for multimodal training?

Videos come with structured alignment across video, audio, captions, scenes, and temporal signals — not just raw files.

This supports action recognition, content understanding, video retrieval, and video-language models out of the box.

<br>
