·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
US threatens sanctions against Chinese AI models over IP theft41m◆Google launches a cheaper alternative to large AI security models like Mythos1h◆Music streamer Deezer says more than 50% of daily uploads are AI-generated2h◆Halliday’s latest smart glasses feature a much-improved display3h◆America needs to stop getting shocked by Chinese AI5h◆Advancing next-gen AI with materials science innovation5h◆Gritt exits stealth with $32 million for robots to build solar plants — then, everything else6h◆Capacity and Redundancy Trade-offs in Multi-Task Learning12h◆Predictive Training with Latent Imagination for Visual Quadruped Navigation12h◆Where Not to Learn: Prior-Aligned Training with Subset-based Attribution Constraints for Reliable Decision-Making12h◆Did We Actually Fix It? An Independent Adversarial Stress-Test of Post-Point-Adjustment Evaluation Metrics for Time-Series Anomaly Detection12h◆Supervised Reward Inference12h◆PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization12h◆Is Progressive Disclosure All You Need for Long-Context Agents?12h◆It Depends on the Dataset: When a Brain-Encoding Model's Predicted Responses Beat Their Visual Backbone for Video Memorability12h◆DMFNet: Dual-Backbone Multiscale Fusion Network for Urban Scene Classification12h◆Oracle Gap and Signal Fidelity: A Fixed-Pool Diagnostic for Test-Time Collaboration12h◆Scientific reasoning does not reliably translate into scientific forecasting in frontier AI12h◆Language Triggers Hijack Language Circuits: A Mechanistic Analysis of Backdoor Behaviors in Large Language Models12h◆AI-Augmented Human Resource Management? Insights from German companies12h◆US threatens sanctions against Chinese AI models over IP theft41m◆Google launches a cheaper alternative to large AI security models like Mythos1h◆Music streamer Deezer says more than 50% of daily uploads are AI-generated2h◆Halliday’s latest smart glasses feature a much-improved display3h◆America needs to stop getting shocked by Chinese AI5h◆Advancing next-gen AI with materials science innovation5h◆Gritt exits stealth with $32 million for robots to build solar plants — then, everything else6h◆Capacity and Redundancy Trade-offs in Multi-Task Learning12h◆Predictive Training with Latent Imagination for Visual Quadruped Navigation12h◆Where Not to Learn: Prior-Aligned Training with Subset-based Attribution Constraints for Reliable Decision-Making12h◆Did We Actually Fix It? An Independent Adversarial Stress-Test of Post-Point-Adjustment Evaluation Metrics for Time-Series Anomaly Detection12h◆Supervised Reward Inference12h◆PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization12h◆Is Progressive Disclosure All You Need for Long-Context Agents?12h◆It Depends on the Dataset: When a Brain-Encoding Model's Predicted Responses Beat Their Visual Backbone for Video Memorability12h◆DMFNet: Dual-Backbone Multiscale Fusion Network for Urban Scene Classification12h◆Oracle Gap and Signal Fidelity: A Fixed-Pool Diagnostic for Test-Time Collaboration12h◆Scientific reasoning does not reliably translate into scientific forecasting in frontier AI12h◆Language Triggers Hijack Language Circuits: A Mechanistic Analysis of Backdoor Behaviors in Large Language Models12h◆AI-Augmented Human Resource Management? Insights from German companies12h◆
News/model/deepseek-v4-flash-gguf

deepseek-v4-flash-gguf news

19 articles mentioning deepseek-v4-flash-gguf

arxiv12h ago

FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention

arXiv:2606.09079v3 Announce Type: replace-cross Abstract: Conventional LLMs keep the full KV cache loaded during decoding, causing a severe GPU memory bottleneck for ultra-long context serving. In this report, we propose \textbf{Lookahead Sparse Attention (LSA)}, a novel inference paradigm powered b

arxivJun 20

DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence

arXiv:2606.19348v1 Announce Type: cross Abstract: We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models -- DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) -- both supporting a

arxivJun 10

Instruction Finetuning DeepSeek-R1-8B Model Using LoRA and NEFTune

arXiv:2606.10392v1 Announce Type: new Abstract: Financial named-entity recognition (NER) is essential for translating unstructured financial reports and news into structured knowledge graphs. However, general-purpose large language models (LLMs) often misclassify financial entities or ignore domain-

arxivMay 26

DeepSeekMath Meets Order Book: Group-Aware Policy Optimization for High-Frequency Directional Trading

arXiv:2605.25527v1 Announce Type: new Abstract: This paper studies reinforcement learning for high-frequency trading on limit order books by pairing an Order-Flow-based state model with policy-gradient methods. Instead of value-based RL techniques like tabular Q-learning, our approach deploys policy

arxivMay 26

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models

arXiv:2506.18543v2 Announce Type: replace-cross Abstract: The rapid proliferation of Large Language Models (LLMs) has heightened concerns regarding their exposure to jailbreak attacks, which craft adversarial inputs designed to elicit unsafe content. Although proprietary models such as GPT-4 have be

arxivMay 22

RTPrune: Reading-Twice Inspired Token Pruning for Efficient DeepSeek-OCR Inference

arXiv:2605.00392v3 Announce Type: replace-cross Abstract: DeepSeek-OCR leverages visual-text compression to reduce long-text processing costs and accelerate inference, yet visual tokens remain prone to redundant textual and structural information. Moreover, current token pruning methods for conventi

techcrunchMay 6

DeepSeek could hit $45B valuation from its first investment round

The Chinese AI lab came to prominence in early 2025 after launching a large language model that trained on a fraction of the compute power and at a fraction of the cost of the big U.S. models like those from OpenAI and Anthropic.

mit-tech-reviewApr 24

Three reasons why DeepSeek’s new model matters

On April 24, Chinese AI firm DeepSeek released a preview of V4, its long-awaited new flagship model. The model can process much longer prompts than its last generation, thanks to a new design that helps it handle large amounts of text more efficiently. Like DeepSeek’s previous models, V4 is open sou

techcrunchApr 24

DeepSeek previews new AI model that ‘closes the gap’ with frontier models

DeepSeek says both models are more efficient and performant than DeepSeek V3.2 due to architectural improvements, and have almost "closed the gap" with current leading models, both open and closed, on reasoning benchmarks.

thevergeApr 24

China’s DeepSeek previews new AI model a year after jolting US rivals

Chinese AI company DeepSeek released a preview of its hotly anticipated next-generation AI model V4 on Friday, saying that the open-source model can compete with leading closed-source systems from US rivals including Anthropic, Google, and OpenAI. DeepSeek says V4 marks a major improvement over prio

huggingfaceApr 24

DeepSeek-V4: a million-token context that agents can actually use

arxivApr 22

Fine-tuning DeepSeek-OCR-2 for Molecular Structure Recognition

arXiv:2604.03476v2 Announce Type: replace-cross Abstract: Optical Chemical Structure Recognition (OCSR) is critical for converting 2D molecular diagrams from printed literature into machine-readable formats. While Vision-Language Models have shown promise in end-to-end OCR tasks, their direct applic

arxivMar 31

Can AI be a Teaching Partner? Evaluating ChatGPT, Gemini, and DeepSeek across Three Teaching Strategies

arXiv:2603.26673v1 Announce Type: cross Abstract: There are growing promises that Large Language Models (LLMs) can support students' learning by providing explanations, feedback, and guidance. However, despite their rapid adoption and widespread attention, there is still limited empirical evidence r

huggingfaceFeb 3

The Future of the Global Open-Source AI Ecosystem: From DeepSeek to AI+

huggingfaceJan 27

Architectural Choices in China's Open-Source AI Ecosystem: Building Beyond DeepSeek

huggingfaceJan 20

One Year Since the “DeepSeek Moment”

huggingfaceJan 31

Mini-R1: Reproduce Deepseek R1 „aha moment“ a RL tutorial

huggingfaceJan 30

How to deploy and fine-tune DeepSeek models on AWS

huggingfaceJan 28

Open-R1: a fully open reproduction of DeepSeek-R1

HomeModelsNews