Model Detail
Qianfan-OCR
—Qianfan-OCR is a multimodal model with 2.4B parameters released by baidu. The model is registered under the image-text-to-text pipeline tag on Hugging Face, distributed under the permissive apache-2.0 license.
Qianfan-OCR ships with 2.4B parameters. Total weight footprint is approximately 4.7 GB, which is the relevant figure when planning local-inference VRAM. The apache-2.0 license is permissive, allowing commercial deployment and derivative work without per-seat fees, though attribution requirements still apply.
Qianfan-OCR is best fit for mixed text-and-image reasoning tasks such as document understanding. Treat this as a starting matrix rather than a benchmark verdict — the right deployment usually depends on the specific evaluation suite that mirrors your workload.