Model Detail
Dolphin-Mistral-24B-Venice-Edition
—Dolphin-Mistral-24B-Venice-Edition is a large language model with 24B parameters released by dphn. The model is registered under the text-generation pipeline tag on Hugging Face, and supports text->text inputs, distributed under the permissive apache-2.0 license.
Dolphin-Mistral-24B-Venice-Edition is priced at $0.2/M input tokens and $0.9/M output tokens. Operationally the model offers a 128K-token context window, which matters when sizing it for prompt-heavy or latency-sensitive workloads. At this input rate the model sits in the commodity tier and is suitable for high-volume workloads where per-call cost dominates the decision.
Dolphin-Mistral-24B-Venice-Edition ships with 24B parameters. The published knowledge cutoff is 2024-04-30, so newer events will not be reflected in zero-shot answers without retrieval. Total weight footprint is approximately 24.0 GB, which is the relevant figure when planning local-inference VRAM. The apache-2.0 license is permissive, allowing commercial deployment and derivative work without per-seat fees, though attribution requirements still apply.
Dolphin-Mistral-24B-Venice-Edition is best fit for general-purpose chat and instruction-following workloads, and high-volume batch jobs where per-call cost dominates the budget. Treat this as a starting matrix rather than a benchmark verdict — the right deployment usually depends on the specific evaluation suite that mirrors your workload.