For the complete documentation index, see llms.txt. This page is also available as Markdown.

💡Should I choose an off-the-shelf dataset or request a custom build?

Choose off-the-shelf when your use case matches existing catalog data and you need speed; choose a custom build when you need a specific language, scene, or annotation scheme that the catalog does not cover.

Off-the-shelf datasets ship fast and carry transparent unit pricing — for example, TikTok metadata, YouTube comments, or 11-language OCR image collections can be scoped and delivered quickly. Custom builds make sense when the requirement is narrow or unusual: a particular low-resource language, a specific factory scene for embodied AI training, or a bespoke annotation format. In practice, many teams start with an off-the-shelf set to validate their pipeline, then commission a custom collection scaled to production. Submitting your requirements through datamall.thordata.com generates a requirements list and cost estimate for either path.

Last updated