Model Detail
Inkling
—Inkling is a multimodal model with 476.2B parameters released by thinkingmachines. The model is registered under the image-text-to-text pipeline tag on Hugging Face, and supports text+image+audio->text inputs, distributed under the permissive apache-2.0 license.
Inkling is priced at $1/M input tokens and $4.05/M output tokens. Operationally the model offers a 1049K-token context window, which matters when sizing it for prompt-heavy or latency-sensitive workloads. Pricing in this range is the working middle of the API market — neither the cheapest nor the most expensive option per token, so cost-fit is usually a function of how much output you generate.
Inkling ships with 476.2B parameters. Total weight footprint is approximately 952.4 GB, which is the relevant figure when planning local-inference VRAM. The apache-2.0 license is permissive, allowing commercial deployment and derivative work without per-seat fees, though attribution requirements still apply.
Inkling is best fit for mixed text-and-image reasoning tasks such as document understanding, and long-context tasks such as full-codebase analysis or book-length summarization (1049K tokens). Treat this as a starting matrix rather than a benchmark verdict — the right deployment usually depends on the specific evaluation suite that mirrors your workload.
Thinking Machines amps up its bet against one-size-fits-all AI with its first open model, Inkling
It's the company's first public proof point after a year and a half spent building AI infrastructure largely out of public view.
Welcome Inkling by Thinking Machines
Hierarchical Multi-Fidelity Learning for Predicting Three-Dimensional Flame Wrinkling and Turbulent Burning Velocity
arXiv:2605.08232v1 Announce Type: new Abstract: High-fidelity experimental characterization of turbulent premixed flames remains limited by the cost and complexity of advanced diagnostics, particularly under elevated pressures and intense turbulence where measurements of coupled flame morphology and