·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
The sameness problem behind those unappetizing AI-generated menus2h◆ExecRetrieval: Measuring the Functional-Correctness Gap in Code-Embedding Retrieval2h◆MasterControl Seventeen Every Time2h◆GrowPage: On-Demand KV Budgeting for Efficient LLM Reasoning Serving2h◆PPO-STGNN: A Proximal Policy Optimization Approach with Spatio-Temporal Graph Neural Networks for DAG Task Scheduling in Cloud-Edge-End Computing2h◆HalluPeer: A Taxonomy-driven Benchmark for Detecting Hallucinations in Scientific Peer Reviews2h◆The Attention Triangle in Audio-Video Models2h◆CauseCollab: Causal Unified and Modality-Agnostic Network for Heterogeneous Collaborative Perception2h◆Semantic Bayesian World Models2h◆Adapting to Evolving Requirements: Agentic AI for Retail Supply Chain Operations2h◆Bioinfoysis Technical Report2h◆STAIR (STructure Aware Information Retriever): A novel dataset and LLM based retriever for document structure augmentation2h◆Xiaomi-TabLDM: A Tabular Foundation Model Technical Report2h◆Inferring Affective Consciousness in an Artificial Agent: A Case Study2h◆Lose the Order, Keep the Hierarchy: Deordering HTN Plans2h◆Value-Preserving Architectures for Agentic AI Systems2h◆Speak for Me: Giving LLMs the Situational Awareness to Participate in a Meeting2h◆Towards Numerical TOHTN Planning with SMT-based HTN-SAT Encoding2h◆FiMI Banking: A Sovereign Model for Indian Retail Banking2h◆Privacy-Preserving Topology-Guided Safety for LLM-Based Multi-Agent Systems via Federated Graph Learning2h◆The sameness problem behind those unappetizing AI-generated menus2h◆ExecRetrieval: Measuring the Functional-Correctness Gap in Code-Embedding Retrieval2h◆MasterControl Seventeen Every Time2h◆GrowPage: On-Demand KV Budgeting for Efficient LLM Reasoning Serving2h◆PPO-STGNN: A Proximal Policy Optimization Approach with Spatio-Temporal Graph Neural Networks for DAG Task Scheduling in Cloud-Edge-End Computing2h◆HalluPeer: A Taxonomy-driven Benchmark for Detecting Hallucinations in Scientific Peer Reviews2h◆The Attention Triangle in Audio-Video Models2h◆CauseCollab: Causal Unified and Modality-Agnostic Network for Heterogeneous Collaborative Perception2h◆Semantic Bayesian World Models2h◆Adapting to Evolving Requirements: Agentic AI for Retail Supply Chain Operations2h◆Bioinfoysis Technical Report2h◆STAIR (STructure Aware Information Retriever): A novel dataset and LLM based retriever for document structure augmentation2h◆Xiaomi-TabLDM: A Tabular Foundation Model Technical Report2h◆Inferring Affective Consciousness in an Artificial Agent: A Case Study2h◆Lose the Order, Keep the Hierarchy: Deordering HTN Plans2h◆Value-Preserving Architectures for Agentic AI Systems2h◆Speak for Me: Giving LLMs the Situational Awareness to Participate in a Meeting2h◆Towards Numerical TOHTN Planning with SMT-based HTN-SAT Encoding2h◆FiMI Banking: A Sovereign Model for Indian Retail Banking2h◆Privacy-Preserving Topology-Guided Safety for LLM-Based Multi-Agent Systems via Federated Graph Learning2h◆
News/IHUBERT: Vector-Based Semantic Deduplication and Domain-Balanced Pretraining for Persian Resources
arxiv
PublishedJune 20, 2026 at 4:00 AM
—neutral

IHUBERT: Vector-Based Semantic Deduplication and Domain-Balanced Pretraining for Persian Resources

Source
arxiv.orgfull article ↗
Read on arxiv→
Publisher summary· verbatim

arXiv:2606.20089v1 Announce Type: cross Abstract: Persian pretrained language models (PLMs) are still limited by the scarcity of large-scale, high-quality pretraining corpora and by insufficient evaluation beyond standard classification and NER tasks. We present IHUBERT, a monolingual Persian PLM tr

Stay posted· Newsletter

A 5-min weekly brief — top movers, price watch, story of the week.

// no spam · unsubscribe one-click · free forever

Discussion
Source
↗
arxiv
Read original ↗All from arxiv →

No replies yet. Be first.

Source
↗
arxiv
Read original ↗All from arxiv →

Related coverage

More from ARXIV
arxivExecRetrieval: Measuring the Functional-Correctness Gap in Code-Embedding Retrieval2harxivMasterControl Seventeen Every Time2harxivGrowPage: On-Demand KV Budgeting for Efficient LLM Reasoning Serving2harxivPPO-STGNN: A Proximal Policy Optimization Approach with Spatio-Temporal Graph Neural Networks for DAG Task Scheduling in Cloud-Edge-End Computing2h
The Bubble Brief
WEEKLY

Read AI insights every Tuesday — top movers, new releases, story of the week.

// no spam · unsubscribe one-click · free forever

Originally published on arxiv ↗
HomeModelsNews