·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning7h◆Subagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks7h◆Distribution-Consistent Inference for Dynamic Sparse Mixture-of-Experts7h◆In RAG We Trust? Measuring Robustness of Retrieval-Augmented Generation Under Document Poisoning7h◆FrontierChallenge: Evaluating Scientific Workflow Completion7h◆IBIB: A Protocol for Measuring Enterprise AI Systems by Serving Route, Not Model Identifier7h◆AgenticGen: Reward-Guided Agentic Video Generation for Advertising7h◆Strangers to Themselves: What Language Models Say About Themselves Is Generic7h◆Omni Interaction Agent Technical Report7h◆CoGReV: A Confidence-Gated Post-Hoc Non-Monotonic Belief Revision Framework for Phishing Website Classification7h◆Bit-Flip Attacks on Vision-Language-Action Models: Action-Decoding Architecture Shapes the Vulnerability7h◆Can Foundation Models Moderate Online Content? Evaluating Instruction- vs. Example-Driven Policy Operationalization7h◆RiLM: Parameter-Efficient Language Modeling via Geodesic Decoding7h◆Tracing Computation Density in LLMs7h◆Cultural Binding Heads in Language Models7h◆Direct Diversity Optimization for Diverse Successful Trajectories in Preference Post-Training7h◆Left-Branching Transformers Excel at Right-Branching Languages: Data Shapes Word Order Preferences in Language Models7h◆Palmyra x6 Technical Report: An Agentic, Tool-Use Model Post-Trained via Anchored Supervised Fine-Tuning7h◆'Ghaib in Translation' aka Unseen Harm: Measuring Cross-Script Safety Inconsistency with 'Missed-in-Urdu' Scores in LLM Hate Speech Detection7h◆DexterSQL: Deep Schema Exploration and Rule-based Correction for Text-to-SQL Generation7h◆Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning7h◆Subagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks7h◆Distribution-Consistent Inference for Dynamic Sparse Mixture-of-Experts7h◆In RAG We Trust? Measuring Robustness of Retrieval-Augmented Generation Under Document Poisoning7h◆FrontierChallenge: Evaluating Scientific Workflow Completion7h◆IBIB: A Protocol for Measuring Enterprise AI Systems by Serving Route, Not Model Identifier7h◆AgenticGen: Reward-Guided Agentic Video Generation for Advertising7h◆Strangers to Themselves: What Language Models Say About Themselves Is Generic7h◆Omni Interaction Agent Technical Report7h◆CoGReV: A Confidence-Gated Post-Hoc Non-Monotonic Belief Revision Framework for Phishing Website Classification7h◆Bit-Flip Attacks on Vision-Language-Action Models: Action-Decoding Architecture Shapes the Vulnerability7h◆Can Foundation Models Moderate Online Content? Evaluating Instruction- vs. Example-Driven Policy Operationalization7h◆RiLM: Parameter-Efficient Language Modeling via Geodesic Decoding7h◆Tracing Computation Density in LLMs7h◆Cultural Binding Heads in Language Models7h◆Direct Diversity Optimization for Diverse Successful Trajectories in Preference Post-Training7h◆Left-Branching Transformers Excel at Right-Branching Languages: Data Shapes Word Order Preferences in Language Models7h◆Palmyra x6 Technical Report: An Agentic, Tool-Use Model Post-Trained via Anchored Supervised Fine-Tuning7h◆'Ghaib in Translation' aka Unseen Harm: Measuring Cross-Script Safety Inconsistency with 'Missed-in-Urdu' Scores in LLM Hate Speech Detection7h◆DexterSQL: Deep Schema Exploration and Rule-based Correction for Text-to-SQL Generation7h◆
News/model/Kimi-K3

Kimi-K3 news

9 articles mentioning Kimi-K3

arxivJul 29

Kimi K3: Open Frontier Intelligence

arXiv:2607.24653v1 Announce Type: cross Abstract: We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is built on Kimi Delta Attention and Attention Residuals, which improve i

techcrunchJul 24

‘AI communism’, rogue models, and the why Kimi K3 spooked Wall Street

Chinese AI lab Moonshot’s open model Kimi went viral this week for reasons that had less to do with the model itself and more to do with how the U.S. AI industry reacted to it. Meanwhile, an unreleased OpenAI model wandered outside its test environment and ended up connected to a real security breac

techcrunchJul 23

Experts say exploiting Anthropic’s Fable isn’t how Kimi K3 got so good

"I don't think you get a model this strong and this quickly on the heels of Fable doing strictly distillation," one expert told TechCrunch.

techcrunchJul 18

Kimi: Threat or menace?

Chinese company Moonshot AI released a new version of its Kimi model this week, prompting concern about "full AI communism."

techcrunchJul 16

Moonshot’s upcoming Kimi 3 is expected to close the gap with Anthropic’s Opus 4.8

The FT reports Kimi K3 will be the largest open AI model from China, with a parameter count between 2 trillion and 3 trillion.

arxivApr 6

An Independent Safety Evaluation of Kimi K2.5

arXiv:2604.03121v1 Announce Type: cross Abstract: Kimi K2.5 is an open-weight LLM that rivals closed models across coding, multimodal, and agentic benchmarks, but was released without an accompanying safety evaluation. In this work, we conduct a preliminary safety assessment of Kimi K2.5 focusing on

huggingfaceAug 14

Kimina-Prover-RL

huggingfaceJul 10

Kimina-Prover: Applying Test-time RL Search on Large Formal Reasoning Models

arxivApr 6bearish

AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents

arXiv:2604.02947v1 Announce Type: new Abstract: Computer-use agents extend language models from text generation to persistent action over tools, files, and execution environments. Unlike chat systems, they maintain state across interactions and translate intermediate outputs into concrete actions. T

#safety#benchmark#autonomous agents
HomeModelsNews