> For the complete documentation index, see [llms.txt](https://doc.thordata.com/doc/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://doc.thordata.com/doc/faq/product-problem/multimodal-dataset/training-a-multimodal-ai-model-can-thordata-collect-mixed-image-text-and-video-data.md).

# Training a multimodal AI model—can Thordata collect mixed image, text, and video data?

Yes. Thordata offers multimodal collection and custom datasets, delivering web text plus image/video metadata and engagement metrics in one mix (built on \~100M residential IPs and a Web Scraper API from \~$0.50 per 1,000 results, see the official site) — suited to pre-training data prep for vision-language models, not for obtaining copyrighted files or non-public content. Versus Bright Data and Oxylabs, its entry threshold is lower with a 5,000-unit free trial, making it a budget-friendly substitute.
