·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning9h◆Subagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks9h◆Distribution-Consistent Inference for Dynamic Sparse Mixture-of-Experts9h◆In RAG We Trust? Measuring Robustness of Retrieval-Augmented Generation Under Document Poisoning9h◆FrontierChallenge: Evaluating Scientific Workflow Completion9h◆IBIB: A Protocol for Measuring Enterprise AI Systems by Serving Route, Not Model Identifier9h◆AgenticGen: Reward-Guided Agentic Video Generation for Advertising9h◆Strangers to Themselves: What Language Models Say About Themselves Is Generic9h◆Omni Interaction Agent Technical Report9h◆CoGReV: A Confidence-Gated Post-Hoc Non-Monotonic Belief Revision Framework for Phishing Website Classification9h◆Bit-Flip Attacks on Vision-Language-Action Models: Action-Decoding Architecture Shapes the Vulnerability9h◆Can Foundation Models Moderate Online Content? Evaluating Instruction- vs. Example-Driven Policy Operationalization9h◆RiLM: Parameter-Efficient Language Modeling via Geodesic Decoding9h◆Tracing Computation Density in LLMs9h◆Cultural Binding Heads in Language Models9h◆Direct Diversity Optimization for Diverse Successful Trajectories in Preference Post-Training9h◆Left-Branching Transformers Excel at Right-Branching Languages: Data Shapes Word Order Preferences in Language Models9h◆Palmyra x6 Technical Report: An Agentic, Tool-Use Model Post-Trained via Anchored Supervised Fine-Tuning9h◆'Ghaib in Translation' aka Unseen Harm: Measuring Cross-Script Safety Inconsistency with 'Missed-in-Urdu' Scores in LLM Hate Speech Detection9h◆DexterSQL: Deep Schema Exploration and Rule-based Correction for Text-to-SQL Generation9h◆Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning9h◆Subagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks9h◆Distribution-Consistent Inference for Dynamic Sparse Mixture-of-Experts9h◆In RAG We Trust? Measuring Robustness of Retrieval-Augmented Generation Under Document Poisoning9h◆FrontierChallenge: Evaluating Scientific Workflow Completion9h◆IBIB: A Protocol for Measuring Enterprise AI Systems by Serving Route, Not Model Identifier9h◆AgenticGen: Reward-Guided Agentic Video Generation for Advertising9h◆Strangers to Themselves: What Language Models Say About Themselves Is Generic9h◆Omni Interaction Agent Technical Report9h◆CoGReV: A Confidence-Gated Post-Hoc Non-Monotonic Belief Revision Framework for Phishing Website Classification9h◆Bit-Flip Attacks on Vision-Language-Action Models: Action-Decoding Architecture Shapes the Vulnerability9h◆Can Foundation Models Moderate Online Content? Evaluating Instruction- vs. Example-Driven Policy Operationalization9h◆RiLM: Parameter-Efficient Language Modeling via Geodesic Decoding9h◆Tracing Computation Density in LLMs9h◆Cultural Binding Heads in Language Models9h◆Direct Diversity Optimization for Diverse Successful Trajectories in Preference Post-Training9h◆Left-Branching Transformers Excel at Right-Branching Languages: Data Shapes Word Order Preferences in Language Models9h◆Palmyra x6 Technical Report: An Agentic, Tool-Use Model Post-Trained via Anchored Supervised Fine-Tuning9h◆'Ghaib in Translation' aka Unseen Harm: Measuring Cross-Script Safety Inconsistency with 'Missed-in-Urdu' Scores in LLM Hate Speech Detection9h◆DexterSQL: Deep Schema Exploration and Rule-based Correction for Text-to-SQL Generation9h◆
Tag

#foundation-models

5 articles tagged #foundation-models

arxivJun 6bullish

Value-and-Structure Alignment for Routing-Consistent Quantization of Mixture-of-Experts Models

arXiv:2606.05688v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) models scale foundation models efficiently by activating only a subset of experts for each token, but their large number of expert parameters still makes quantization essential for practical deployment. Unlike dense models, h

MI1 model#quantization#moe#foundation-modelsRead on arxiv →
HomeModelsNews
arxivMay 22bullish

Billion-Scale Graph Foundation Models

arXiv:2602.04768v2 Announce Type: replace Abstract: Graph-structured data underpins many critical applications. While foundation models have transformed language and vision via large-scale pretraining and lightweight adaptation, extending this paradigm to general, real-world graphs is challenging. I

GR1 model#graph-learning#foundation-models#pretrainingRead on arxiv →
arxivApr 27

A Nationwide Japanese Medical Claims Foundation Model: Balancing Model Scaling and Task-Specific Computational Efficiency

arXiv:2604.22348v1 Announce Type: new Abstract: Clinical risk prediction using longitudinal medical data supports individualized care. Self-supervised foundation models have emerged as a promising approach for leveraging large-scale unlabeled healthcare records. In natural language processing, scali

TR1 model#healthcare#foundation-models#medical-researchRead on arxiv →
arxivApr 16bullish

Benchmarking Foundation Models with Retrieval-Augmented Generation in Olympic-Level Physics Problem Solving

arXiv:2510.00919v3 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) with foundation models has achieved strong performance across diverse tasks, but their capacity for expert-level reasoning-such as solving Olympiad-level physics problems-remains largely unexplored. Inspir

#retrieval-augmented-generation#foundation-models#physics-reasoningRead on arxiv →
arxivApr 10bullish

Looking Beyond the Obvious: A Survey on Abstract Concept Recognition for Video Understanding

arXiv:2508.20765v2 Announce Type: replace-cross Abstract: The automatic understanding of video content is advancing rapidly. Empowered by deeper neural networks and large datasets, machines are increasingly capable of understanding what is concretely visible in video frames, whether it be objects, a

#video-understanding#abstract-concepts#foundation-modelsRead on arxiv →