Model Detail
NVIDIA-Nemotron-3-Super-120B-A12B-FP8
—NVIDIA-Nemotron-3-Super-120B-A12B-FP8 is a code generation model with 120B parameters released by NVIDIA. The model is registered under the text-generation pipeline tag on Hugging Face, and supports text->text inputs, distributed under a other license.
NVIDIA-Nemotron-3-Super-120B-A12B-FP8 is priced at $0.1/M input tokens and $0.5/M output tokens. Operationally the model offers a 1000K-token context window, which matters when sizing it for prompt-heavy or latency-sensitive workloads. At this input rate the model sits in the commodity tier and is suitable for high-volume workloads where per-call cost dominates the decision.
NVIDIA-Nemotron-3-Super-120B-A12B-FP8 ships with 120B parameters. Total weight footprint is approximately 123.6 GB, which is the relevant figure when planning local-inference VRAM. Distribution is governed by the other license — review the exact terms before commercial deployment.
NVIDIA-Nemotron-3-Super-120B-A12B-FP8 is best fit for code completion, repository-scale Q&A, and pair-programming integrations, high-volume batch jobs where per-call cost dominates the budget, and long-context tasks such as full-codebase analysis or book-length summarization (1000K tokens). It is a less obvious choice for one-shot generation of security-critical code without review. Treat this as a starting matrix rather than a benchmark verdict — the right deployment usually depends on the specific evaluation suite that mirrors your workload.
Fine-tune video and image models at scale with NVIDIA NeMo Automodel and 🤗 Diffusers
NVIDIA Nemotron 3 Embed Ranks #1 Overall on RTEB, Advancing Agentic Retrieval
Decomposing Runtime, Kernel, and Quantization Speedups via a Matched FP16 Intermediate: A Hardware-Conditioned Case Study on Four NVIDIA RTX A5000 GPUs
arXiv:2607.11368v1 Announce Type: cross Abstract: Reported serving speedups from quantized kernels typically bundle the weight format, the kernel, and the inference runtime into one number. We present an attribution study on four NVIDIA RTX A5000 GPUs, 24 GiB each, on a single host with NVLink-bridg
Paris-based AI voice startup Gradium raises $100M seed, backed by Nvidia
The company is using the cash to open an office in the Bay Area and compete for talent there, "strengthening its position at the heart of the world's leading AI ecosystem."
Nvidia is a victim of the compute marketplace it created
Having proven how valuable compute can be, the company finds itself at the center of a market everyone wants to be in — while simpler technologies and less interesting companies get rich on the sidelines.
Nvidia competitor Etched hits $5B valuation, $1B in sales for AI chip
Nvidia AI chip competitor Etched says it has already booked $1 billion under contract for the inference systems powered by its chip.