·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
America needs to stop getting shocked by Chinese AI2h◆Advancing next-gen AI with materials science innovation2h◆Gritt exits stealth with $34 million for robots to build solar plants—then, everything else3h◆Capacity and Redundancy Trade-offs in Multi-Task Learning9h◆Predictive Training with Latent Imagination for Visual Quadruped Navigation9h◆Where Not to Learn: Prior-Aligned Training with Subset-based Attribution Constraints for Reliable Decision-Making9h◆Did We Actually Fix It? An Independent Adversarial Stress-Test of Post-Point-Adjustment Evaluation Metrics for Time-Series Anomaly Detection9h◆Supervised Reward Inference9h◆PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization9h◆Is Progressive Disclosure All You Need for Long-Context Agents?9h◆It Depends on the Dataset: When a Brain-Encoding Model's Predicted Responses Beat Their Visual Backbone for Video Memorability9h◆DMFNet: Dual-Backbone Multiscale Fusion Network for Urban Scene Classification9h◆Oracle Gap and Signal Fidelity: A Fixed-Pool Diagnostic for Test-Time Collaboration9h◆Time-Frequency Consistency Learning for Robust Speech Deepfake Detection9h◆Scientific reasoning does not reliably translate into scientific forecasting in frontier AI9h◆Language Triggers Hijack Language Circuits: A Mechanistic Analysis of Backdoor Behaviors in Large Language Models9h◆Kernel Regression with Tensor Trains and Hadamard Overparameterization9h◆AI-Augmented Human Resource Management? Insights from German companies9h◆Diagnosing Correctness Probes under Self-Judgement Confounding9h◆BLAD: A Historically Contextualized, Multilingual Dataset of Bangladeshi Legal Acts (1799 to 2025)9h◆America needs to stop getting shocked by Chinese AI2h◆Advancing next-gen AI with materials science innovation2h◆Gritt exits stealth with $34 million for robots to build solar plants—then, everything else3h◆Capacity and Redundancy Trade-offs in Multi-Task Learning9h◆Predictive Training with Latent Imagination for Visual Quadruped Navigation9h◆Where Not to Learn: Prior-Aligned Training with Subset-based Attribution Constraints for Reliable Decision-Making9h◆Did We Actually Fix It? An Independent Adversarial Stress-Test of Post-Point-Adjustment Evaluation Metrics for Time-Series Anomaly Detection9h◆Supervised Reward Inference9h◆PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization9h◆Is Progressive Disclosure All You Need for Long-Context Agents?9h◆It Depends on the Dataset: When a Brain-Encoding Model's Predicted Responses Beat Their Visual Backbone for Video Memorability9h◆DMFNet: Dual-Backbone Multiscale Fusion Network for Urban Scene Classification9h◆Oracle Gap and Signal Fidelity: A Fixed-Pool Diagnostic for Test-Time Collaboration9h◆Time-Frequency Consistency Learning for Robust Speech Deepfake Detection9h◆Scientific reasoning does not reliably translate into scientific forecasting in frontier AI9h◆Language Triggers Hijack Language Circuits: A Mechanistic Analysis of Backdoor Behaviors in Large Language Models9h◆Kernel Regression with Tensor Trains and Hadamard Overparameterization9h◆AI-Augmented Human Resource Management? Insights from German companies9h◆Diagnosing Correctness Probes under Self-Judgement Confounding9h◆BLAD: A Historically Contextualized, Multilingual Dataset of Bangladeshi Legal Acts (1799 to 2025)9h◆
DataBubble·

Model Detail

deepseek-ai logo

DeepSeek-V4-Flash

▲ 1.7%
Provider: DeepSeekCategory: llmPipeline: text-generation
DB Score
2.9
Downloads
3.0M
Likes
2K
Day
+1.7%
Week
+0.0%
Month
+0.0%
Overview

DeepSeek-V4-Flash is a large language model with 79.0B parameters released by DeepSeek. The model is registered under the text-generation pipeline tag on Hugging Face, and supports text->text inputs, distributed under the permissive mit license.

Pricing & Throughput

DeepSeek-V4-Flash is priced at $0.19/M input tokens and $0.51/M output tokens. Operationally the model offers a 1000K-token context window, which matters when sizing it for prompt-heavy or latency-sensitive workloads. At this input rate the model sits in the commodity tier and is suitable for high-volume workloads where per-call cost dominates the decision.

Technical

DeepSeek-V4-Flash ships with 79.0B parameters. Total weight footprint is approximately 158.1 GB, which is the relevant figure when planning local-inference VRAM. The mit license is permissive, allowing commercial deployment and derivative work without per-seat fees, though attribution requirements still apply.

Trending Signal

Downloads of DeepSeek-V4-Flash have moved +1.7% over the past 24 hours. That is a slight downtrend, consistent with normal cooling as newer models compete for the same workloads. These numbers are signal, not guarantee — week-over-week download counts on Hugging Face also reflect mirror traffic, CI scrapes, and one-off benchmarking runs.

Read about databubble_score →
Use Cases

DeepSeek-V4-Flash is best fit for general-purpose chat and instruction-following workloads, high-volume batch jobs where per-call cost dominates the budget, and long-context tasks such as full-codebase analysis or book-length summarization (1000K tokens). Treat this as a starting matrix rather than a benchmark verdict — the right deployment usually depends on the specific evaluation suite that mirrors your workload.

Download History
Pricing
Input ($/M tokens)
$0.19
Output ($/M tokens)
$0.51
Context Window
1000K
Research Paper
arXiv: 2606.19348→
Model Info
Licensemit
Modalitytext->text
Citations768 (69 influential)
Recent newsView all news →
Related News
arxiv9h ago

FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention

arXiv:2606.09079v3 Announce Type: replace-cross Abstract: Conventional LLMs keep the full KV cache loaded during decoding, causing a severe GPU memory bottleneck for ultra-long context serving. In this report, we propose \textbf{Lookahead Sparse Attention (LSA)}, a novel inference paradigm powered b

arxivneutral31d ago

DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence

arXiv:2606.19348v1 Announce Type: cross Abstract: We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models -- DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) -- both supporting a

arxiv41d ago

Instruction Finetuning DeepSeek-R1-8B Model Using LoRA and NEFTune

arXiv:2606.10392v1 Announce Type: new Abstract: Financial named-entity recognition (NER) is essential for translating unstructured financial reports and news into structured knowledge graphs. However, general-purpose large language models (LLMs) often misclassify financial entities or ignore domain-

arxiv56d ago

DeepSeekMath Meets Order Book: Group-Aware Policy Optimization for High-Frequency Directional Trading

arXiv:2605.25527v1 Announce Type: new Abstract: This paper studies reinforcement learning for high-frequency trading on limit order books by pairing an Order-Flow-based state model with policy-gradient methods. Instead of value-based RL techniques like tabular Q-learning, our approach deploys policy

arxiv56d ago

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models

arXiv:2506.18543v2 Announce Type: replace-cross Abstract: The rapid proliferation of Large Language Models (LLMs) has heightened concerns regarding their exposure to jailbreak attacks, which craft adversarial inputs designed to elicit unsafe content. Although proprietary models such as GPT-4 have be

arxivbullish33d ago

Attribution-Guided and Coverage-Maximized Pruning for Structural MoE Compression

arXiv:2606.18304v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) models scale compute efficiently, yet remain expensive to deploy due to their substantial memory footprint and inference overhead. Prior compression methods mainly operate at the expert level, either removing entire experts o

Related Models
deepseek-ai logo
DeepSeek-V3.2
DeepSeek · 11.2M downloads
deepseek-ai logo
DeepSeek-R1
DeepSeek · 8.9M downloads
google-bert logo
bert-base-uncased
google-bert · 69.6M downloads
sentence-transformers logo
paraphrase-multilingual-MiniLM-L12-v2
SBERT · 48.6M downloads
HomeModelsNews