·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning9h◆Subagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks9h◆Distribution-Consistent Inference for Dynamic Sparse Mixture-of-Experts9h◆In RAG We Trust? Measuring Robustness of Retrieval-Augmented Generation Under Document Poisoning9h◆FrontierChallenge: Evaluating Scientific Workflow Completion9h◆IBIB: A Protocol for Measuring Enterprise AI Systems by Serving Route, Not Model Identifier9h◆AgenticGen: Reward-Guided Agentic Video Generation for Advertising9h◆Strangers to Themselves: What Language Models Say About Themselves Is Generic9h◆Omni Interaction Agent Technical Report9h◆CoGReV: A Confidence-Gated Post-Hoc Non-Monotonic Belief Revision Framework for Phishing Website Classification9h◆Bit-Flip Attacks on Vision-Language-Action Models: Action-Decoding Architecture Shapes the Vulnerability9h◆Can Foundation Models Moderate Online Content? Evaluating Instruction- vs. Example-Driven Policy Operationalization9h◆RiLM: Parameter-Efficient Language Modeling via Geodesic Decoding9h◆Tracing Computation Density in LLMs9h◆Cultural Binding Heads in Language Models9h◆Direct Diversity Optimization for Diverse Successful Trajectories in Preference Post-Training9h◆Left-Branching Transformers Excel at Right-Branching Languages: Data Shapes Word Order Preferences in Language Models9h◆Palmyra x6 Technical Report: An Agentic, Tool-Use Model Post-Trained via Anchored Supervised Fine-Tuning9h◆'Ghaib in Translation' aka Unseen Harm: Measuring Cross-Script Safety Inconsistency with 'Missed-in-Urdu' Scores in LLM Hate Speech Detection9h◆DexterSQL: Deep Schema Exploration and Rule-based Correction for Text-to-SQL Generation9h◆Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning9h◆Subagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks9h◆Distribution-Consistent Inference for Dynamic Sparse Mixture-of-Experts9h◆In RAG We Trust? Measuring Robustness of Retrieval-Augmented Generation Under Document Poisoning9h◆FrontierChallenge: Evaluating Scientific Workflow Completion9h◆IBIB: A Protocol for Measuring Enterprise AI Systems by Serving Route, Not Model Identifier9h◆AgenticGen: Reward-Guided Agentic Video Generation for Advertising9h◆Strangers to Themselves: What Language Models Say About Themselves Is Generic9h◆Omni Interaction Agent Technical Report9h◆CoGReV: A Confidence-Gated Post-Hoc Non-Monotonic Belief Revision Framework for Phishing Website Classification9h◆Bit-Flip Attacks on Vision-Language-Action Models: Action-Decoding Architecture Shapes the Vulnerability9h◆Can Foundation Models Moderate Online Content? Evaluating Instruction- vs. Example-Driven Policy Operationalization9h◆RiLM: Parameter-Efficient Language Modeling via Geodesic Decoding9h◆Tracing Computation Density in LLMs9h◆Cultural Binding Heads in Language Models9h◆Direct Diversity Optimization for Diverse Successful Trajectories in Preference Post-Training9h◆Left-Branching Transformers Excel at Right-Branching Languages: Data Shapes Word Order Preferences in Language Models9h◆Palmyra x6 Technical Report: An Agentic, Tool-Use Model Post-Trained via Anchored Supervised Fine-Tuning9h◆'Ghaib in Translation' aka Unseen Harm: Measuring Cross-Script Safety Inconsistency with 'Missed-in-Urdu' Scores in LLM Hate Speech Detection9h◆DexterSQL: Deep Schema Exploration and Rule-based Correction for Text-to-SQL Generation9h◆
Tag

#representation-learning

10 articles tagged #representation-learning

arxivJul 30bullish

From Cellular Responses to Pharmacological Domains: Multimodal Zero-Shot Drug Representation Learning

arXiv:2607.25322v1 Announce Type: new Abstract: Multimodal drug discovery enables drug representation learning beyond chemical structure by incorporating cellular responses such as gene expression and cell morphology. However, direct fusion and instance-level contrastive alignment may mix mechanism-

#multimodal#drug-discovery#representation-learningRead on arxiv →
arxivJul 24
HomeModelsNews

Heat-Kernel Entropy Profiles and Geometric Effective Sample Size for Weighted Measures on Manifolds

arXiv:2607.06696v2 Announce Type: replace-cross Abstract: Weighted empirical measures on compact manifolds appear in importance sampling, particle approximations, posterior summaries, quadrature, and representation learning. Ordinary effective sample size and related weight summaries ignore the geom

#machine-learning#importance-sampling#representation-learningRead on arxiv →
arxivJul 10

From Performance to Viability: A Bootstrap Framework for Latent-Space Representation Learning in Adaptive Biological Systems

arXiv:2606.01374v2 Announce Type: replace Abstract: Observable performance is commonly used to characterize biological systems. In adaptive systems, however, similar performances may arise from distinct organizations, and configurations that appear comparable at a given time may follow different lon

#machine-learning#representation-learning#biological-systemsRead on arxiv →
arxivJul 2bullish

Identifying Latent Concepts and Structures for Generalized Category Discovery

arXiv:2607.00620v1 Announce Type: cross Abstract: Generalized Category Discovery (GCD) aims to recognize known classes while autonomously discovering novel ones in open-world settings. However, current approaches primarily focus on designing clustering objectives, often overlooking a critical bottle

CO1 model#computer-vision#representation-learning#open-world-recognitionRead on arxiv →
arxivJun 10bullish

Tractogram foundation model

arXiv:2606.09893v1 Announce Type: cross Abstract: Diffusion MRI (dMRI) tractography is the only noninvasive approach for mapping white-matter pathways in the living human brain. It represents each brain as a tractogram: a large, unordered set of three-dimensional streamlines that includes informatio

TR1 model#medical-imaging#representation-learning#diffusion-mriRead on arxiv →
arxivJun 5bullish

Toward Scalable and Valid Conditional Independence Testing with Spectral Representations

arXiv:2512.19510v2 Announce Type: replace Abstract: Conditional independence (CI) is central to causal inference, feature selection, and graphical modeling, yet it is untestable in many settings without additional assumptions. Existing CI tests often rely on restrictive structural conditions, limiti

#machine-learning#representation-learning#causal-inferenceRead on arxiv →
arxivMay 13bullish

MePo: Meta Post-Refinement for Rehearsal-Free General Continual Learning

arXiv:2602.07940v3 Announce Type: replace Abstract: To cope with uncertain changes of the external world, intelligent systems must continually learn from complex, evolving environments and respond in real time. This ability, collectively known as general continual learning (GCL), encapsulates practi

ME1 model#continual-learning#pretrained-models#meta-learningRead on arxiv →
arxivMay 11bullish

Don't Retrain, Align: Adapting Autoregressive LMs to Diffusion LMs via Representation Alignment

arXiv:2605.06885v1 Announce Type: cross Abstract: Diffusion language models (DLMs) have recently demonstrated capabilities that complement standard autoregressive (AR) models, particularly in non-sequential generation and bidirectional editing. Although recent work has shown that pretrained autoregr

DIAU2 models#diffusion#language-models#representation-learningRead on arxiv →
arxivMay 1

Robust Learning on Heterogeneous Graphs with Heterophily: A Graph Structure Learning Approach

arXiv:2604.27387v1 Announce Type: new Abstract: Heterogeneous graphs with heterophily have emerged as a powerful abstraction for modeling complex real-world systems, where nodes of different types and labels interact in diverse and often non-homophilous ways. Despite recent advances, robust represen

#heterogeneous-graphs#representation-learning#graph-structureRead on arxiv →
arxivApr 11

Lexical Tone is Hard to Quantize: Probing Discrete Speech Units in Mandarin and Yor\`ub\'a

arXiv:2604.07467v1 Announce Type: new Abstract: Discrete speech units (DSUs) are derived by quantising representations from models trained using self-supervised learning (SSL). They are a popular representation for a wide variety of spoken language tasks, including those where prosody matters. DSUs

#speech-recognition#representation-learning#quantisationRead on arxiv →