·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
Cursor makes its biggest India push yet ahead of SpaceX acquisition with localized pricing4h◆Reverso: Efficient Time Series Foundation Models for Zero-shot Forecasting5h◆Multinex: Lightweight Low-light Image Enhancement via Multi-prior Retinex5h◆Market Design for AI: Beyond the Copyright Binary5h◆Who Pays the Price? Stakeholder-Centric Prompt Injection Benchmarking for Real-world Web Agents5h◆TextRich: A Multi-Domain Benchmark for Detecting AI-Generated Text-Rich Images from GPT-Image-25h◆DiscoLoop: Looping Discrete Embeddings and Continuous Hidden States for Multi-hop Reasoning5h◆From World Models to World Action Models: A Concise Tutorial for Robotics5h◆QuantFlow: A Federated Mamba-Based Post-Transformer Foundation Model for Time-Series Forecasting5h◆Multi-Turn On-Policy Distillation with Prefix Replay5h◆Operational Proto-Introspection in Looped Language Models: Process-Quality Taps, Executable Branching, and the Readout-Control Boundary5h◆Skillware: A Software Ontology and Engineering Lifecycle for Persistent Behavioral Artifacts5h◆MedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation Agents5h◆PhantomFill: When the Form Demands an Answer, Language Models Invent One5h◆Error Certificates for KV-Cache Eviction via Randomized Design5h◆Explaining GAND: A Resource on Gender-Ambiguous Natural Data & Contrastive Attribution5h◆MioFFAn: an Annotation Software for Formula Formalization with LLM Automation Capabilities5h◆LA-RL: Label-Aware Self-Reflection for Reinforcement Learning in Information Extraction5h◆Mwando: Leveraging AI to Preserve and Teach shiKomori5h◆The JEPA Paradox in Language: The Geometry of Linguistic Alternatives5h◆Cursor makes its biggest India push yet ahead of SpaceX acquisition with localized pricing4h◆Reverso: Efficient Time Series Foundation Models for Zero-shot Forecasting5h◆Multinex: Lightweight Low-light Image Enhancement via Multi-prior Retinex5h◆Market Design for AI: Beyond the Copyright Binary5h◆Who Pays the Price? Stakeholder-Centric Prompt Injection Benchmarking for Real-world Web Agents5h◆TextRich: A Multi-Domain Benchmark for Detecting AI-Generated Text-Rich Images from GPT-Image-25h◆DiscoLoop: Looping Discrete Embeddings and Continuous Hidden States for Multi-hop Reasoning5h◆From World Models to World Action Models: A Concise Tutorial for Robotics5h◆QuantFlow: A Federated Mamba-Based Post-Transformer Foundation Model for Time-Series Forecasting5h◆Multi-Turn On-Policy Distillation with Prefix Replay5h◆Operational Proto-Introspection in Looped Language Models: Process-Quality Taps, Executable Branching, and the Readout-Control Boundary5h◆Skillware: A Software Ontology and Engineering Lifecycle for Persistent Behavioral Artifacts5h◆MedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation Agents5h◆PhantomFill: When the Form Demands an Answer, Language Models Invent One5h◆Error Certificates for KV-Cache Eviction via Randomized Design5h◆Explaining GAND: A Resource on Gender-Ambiguous Natural Data & Contrastive Attribution5h◆MioFFAn: an Annotation Software for Formula Formalization with LLM Automation Capabilities5h◆LA-RL: Label-Aware Self-Reflection for Reinforcement Learning in Information Extraction5h◆Mwando: Leveraging AI to Preserve and Teach shiKomori5h◆The JEPA Paradox in Language: The Geometry of Linguistic Alternatives5h◆
Tag

#generalization

5 articles tagged #generalization

arxivMay 16bullish

Beyond What to Select: A Plug-and-play Oscillatory Data-Volume Scheduling for Efficient Model Training

arXiv:2605.14773v1 Announce Type: cross Abstract: Data selection accelerates training by identifying representative training data while preserving model performance. However, existing methods mainly focus on designing sample-importance criteria, i.e., deciding what to select, while typically fixing

#optimization#machine-learning#efficiencyRead on arxiv →
arxivApr 29
HomeModelsNews

Out of Spuriousity: Improving Robustness to Spurious Correlations without Group Annotations

arXiv:2407.14974v2 Announce Type: replace-cross Abstract: Machine learning models are known to learn spurious correlations, i.e., features having strong relations with class labels but no causal relation. Relying on those correlations leads to poor performance in the data groups without these correl

#machine-learning#robustness#generalizationRead on arxiv →
arxivApr 6

Early-Warning Signals of Grokking via Loss-Landscape Geometry

arXiv:2602.16967v3 Announce Type: replace Abstract: Grokking -- the abrupt transition from memorization to generalization after prolonged training -- has been linked to confinement on low-dimensional execution manifolds in modular arithmetic. Whether this mechanism extends beyond arithmetic remains

TR1 model#machine-learning#generalization#arithmeticRead on arxiv →
arxivApr 6

Low-Dimensional and Transversely Curved Optimization Dynamics in Grokking

arXiv:2602.16746v3 Announce Type: replace Abstract: Grokking -- the delayed transition from memorization to generalization in small algorithmic tasks -- remains poorly understood. We present a geometric analysis of optimization dynamics in transformers trained on modular arithmetic. PCA of attention

TR1 model#machine-learning#optimization#generalizationRead on arxiv →
arxivApr 3

Semantic Interaction Information mediates compositional generalization in latent space

arXiv:2603.27134v2 Announce Type: replace Abstract: Are there still barriers to generalization once all relevant variables are known? We address this question via a framework that casts compositional generalization as a variational inference problem over latent variables with parametric interactions

REECFU4 models · +1#machine learning#generalization#reinforcement learningRead on arxiv →