Model Detail
HunyuanOCR
▲ 24.1%HunyuanOCR is a multimodal model with 560M parameters released by tencent. The model is registered under the image-text-to-text pipeline tag on Hugging Face, distributed under a other license.
HunyuanOCR ships with 560M parameters. Total weight footprint is approximately 1.1 GB, which is the relevant figure when planning local-inference VRAM. Distribution is governed by the other license — review the exact terms before commercial deployment.
Downloads of HunyuanOCR have moved +24.1% over the past 24 hours. That is a slight downtrend, consistent with normal cooling as newer models compete for the same workloads. These numbers are signal, not guarantee — week-over-week download counts on Hugging Face also reflect mirror traffic, CI scrapes, and one-off benchmarking runs.
HunyuanOCR is best fit for mixed text-and-image reasoning tasks such as document understanding. Treat this as a starting matrix rather than a benchmark verdict — the right deployment usually depends on the specific evaluation suite that mirrors your workload.