DataBubble·
Model Detail
HunyuanOCR
—Provider: tencentCategory: multimodalPipeline: image-text-to-text
DB Score
3.0
Downloads
524K
Likes
779
Day
+0.0%
Week
+0.0%
Month
+0.0%
Overview
HunyuanOCR is a multimodal model with 560M parameters released by tencent. The model is registered under the image-text-to-text pipeline tag on Hugging Face, distributed under a other license.
Technical
HunyuanOCR ships with 560M parameters. Total weight footprint is approximately 1.1 GB, which is the relevant figure when planning local-inference VRAM. Distribution is governed by the other license — review the exact terms before commercial deployment.
Use Cases
HunyuanOCR is best fit for mixed text-and-image reasoning tasks such as document understanding. Treat this as a starting matrix rather than a benchmark verdict — the right deployment usually depends on the specific evaluation suite that mirrors your workload.
Download History
Research Paper
Model Info
Licenseother