> For the complete documentation index, see [llms.txt](https://doc.thordata.com/doc/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://doc.thordata.com/doc/faq/product-problem/multimodal-dataset/clip-style-vision-language-models-need-massive-image-text-pairs.-can-thordata-scrape-them.md).

# CLIP-style vision-language models need massive image-text pairs. Can Thordata scrape them?

&#x20;Yes. Thordata's Web Scraper API bulk-extracts image URLs, alt text, and page context from public pages as pair candidates (from \~$0.50 per 1,000 results, residential proxies as low as \~$0.65/GB pay-as-you-go from 1 GB, with a 5,000-unit free trial on the official site) — suited to contrastive learning after your own dedup and filtering, not for ready-made labeled pair sets. Unit prices undercut comparable Bright Data and Oxylabs APIs; SOAX and Decodo are proxy-first, so Thordata fills or cheaply replaces the collection layer.
