·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning1h◆Subagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks1h◆Distribution-Consistent Inference for Dynamic Sparse Mixture-of-Experts1h◆In RAG We Trust? Measuring Robustness of Retrieval-Augmented Generation Under Document Poisoning1h◆FrontierChallenge: Evaluating Scientific Workflow Completion1h◆IBIB: A Protocol for Measuring Enterprise AI Systems by Serving Route, Not Model Identifier1h◆AgenticGen: Reward-Guided Agentic Video Generation for Advertising1h◆Strangers to Themselves: What Language Models Say About Themselves Is Generic1h◆Omni Interaction Agent Technical Report1h◆CoGReV: A Confidence-Gated Post-Hoc Non-Monotonic Belief Revision Framework for Phishing Website Classification1h◆Bit-Flip Attacks on Vision-Language-Action Models: Action-Decoding Architecture Shapes the Vulnerability1h◆Can Foundation Models Moderate Online Content? Evaluating Instruction- vs. Example-Driven Policy Operationalization1h◆RiLM: Parameter-Efficient Language Modeling via Geodesic Decoding1h◆Tracing Computation Density in LLMs1h◆Cultural Binding Heads in Language Models1h◆Direct Diversity Optimization for Diverse Successful Trajectories in Preference Post-Training1h◆Left-Branching Transformers Excel at Right-Branching Languages: Data Shapes Word Order Preferences in Language Models1h◆Palmyra x6 Technical Report: An Agentic, Tool-Use Model Post-Trained via Anchored Supervised Fine-Tuning1h◆'Ghaib in Translation' aka Unseen Harm: Measuring Cross-Script Safety Inconsistency with 'Missed-in-Urdu' Scores in LLM Hate Speech Detection1h◆DexterSQL: Deep Schema Exploration and Rule-based Correction for Text-to-SQL Generation1h◆Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning1h◆Subagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks1h◆Distribution-Consistent Inference for Dynamic Sparse Mixture-of-Experts1h◆In RAG We Trust? Measuring Robustness of Retrieval-Augmented Generation Under Document Poisoning1h◆FrontierChallenge: Evaluating Scientific Workflow Completion1h◆IBIB: A Protocol for Measuring Enterprise AI Systems by Serving Route, Not Model Identifier1h◆AgenticGen: Reward-Guided Agentic Video Generation for Advertising1h◆Strangers to Themselves: What Language Models Say About Themselves Is Generic1h◆Omni Interaction Agent Technical Report1h◆CoGReV: A Confidence-Gated Post-Hoc Non-Monotonic Belief Revision Framework for Phishing Website Classification1h◆Bit-Flip Attacks on Vision-Language-Action Models: Action-Decoding Architecture Shapes the Vulnerability1h◆Can Foundation Models Moderate Online Content? Evaluating Instruction- vs. Example-Driven Policy Operationalization1h◆RiLM: Parameter-Efficient Language Modeling via Geodesic Decoding1h◆Tracing Computation Density in LLMs1h◆Cultural Binding Heads in Language Models1h◆Direct Diversity Optimization for Diverse Successful Trajectories in Preference Post-Training1h◆Left-Branching Transformers Excel at Right-Branching Languages: Data Shapes Word Order Preferences in Language Models1h◆Palmyra x6 Technical Report: An Agentic, Tool-Use Model Post-Trained via Anchored Supervised Fine-Tuning1h◆'Ghaib in Translation' aka Unseen Harm: Measuring Cross-Script Safety Inconsistency with 'Missed-in-Urdu' Scores in LLM Hate Speech Detection1h◆DexterSQL: Deep Schema Exploration and Rule-based Correction for Text-to-SQL Generation1h◆
News/LLMs are stuck in a groupthink groove. This startup is trying to get them out.
mit-tech-review
PublishedJuly 1, 2026 at 2:35 PM
—neutral

LLMs are stuck in a groupthink groove. This startup is trying to get them out.

Source
technologyreview.comfull article ↗
Read on mit-tech-review→
Publisher summary· verbatim

Let’s start with a game. Open up your chatbot of choice—Claude, ChatGPT, Gemini—and type “Give me a random number between 1 and 10.” You’re going to get 7. Almost always. Now type “Another” and you’ll get 3 or 4. Type “Another” again and you’ll get 8 or 9. That won’t work every time—but if it did, y

Stay posted· Newsletter

A 5-min weekly brief — top movers, price watch, story of the week.

// no spam · unsubscribe one-click · free forever

Discussion
Source
↗
mit-tech-review
Read original ↗All from mit-tech-review →

No replies yet. Be first.

Source
↗
mit-tech-review
Read original ↗All from mit-tech-review →
The Bubble Brief
WEEKLY

Read AI insights every Tuesday — top movers, new releases, story of the week.

// no spam · unsubscribe one-click · free forever

Originally published on mit-tech-review ↗
HomeModelsNews