·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
US threatens sanctions against Chinese AI models over IP theft1h◆Google launches a cheaper alternative to large AI security models like Mythos1h◆Music streamer Deezer says more than 50% of daily uploads are AI-generated3h◆Halliday’s latest smart glasses feature a much-improved display3h◆America needs to stop getting shocked by Chinese AI5h◆Advancing next-gen AI with materials science innovation6h◆Gritt exits stealth with $32 million for robots to build solar plants — then, everything else6h◆Capacity and Redundancy Trade-offs in Multi-Task Learning12h◆Predictive Training with Latent Imagination for Visual Quadruped Navigation12h◆Where Not to Learn: Prior-Aligned Training with Subset-based Attribution Constraints for Reliable Decision-Making12h◆Did We Actually Fix It? An Independent Adversarial Stress-Test of Post-Point-Adjustment Evaluation Metrics for Time-Series Anomaly Detection12h◆Supervised Reward Inference12h◆PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization12h◆Is Progressive Disclosure All You Need for Long-Context Agents?12h◆It Depends on the Dataset: When a Brain-Encoding Model's Predicted Responses Beat Their Visual Backbone for Video Memorability12h◆DMFNet: Dual-Backbone Multiscale Fusion Network for Urban Scene Classification12h◆Oracle Gap and Signal Fidelity: A Fixed-Pool Diagnostic for Test-Time Collaboration12h◆Scientific reasoning does not reliably translate into scientific forecasting in frontier AI12h◆Language Triggers Hijack Language Circuits: A Mechanistic Analysis of Backdoor Behaviors in Large Language Models12h◆AI-Augmented Human Resource Management? Insights from German companies12h◆US threatens sanctions against Chinese AI models over IP theft1h◆Google launches a cheaper alternative to large AI security models like Mythos1h◆Music streamer Deezer says more than 50% of daily uploads are AI-generated3h◆Halliday’s latest smart glasses feature a much-improved display3h◆America needs to stop getting shocked by Chinese AI5h◆Advancing next-gen AI with materials science innovation6h◆Gritt exits stealth with $32 million for robots to build solar plants — then, everything else6h◆Capacity and Redundancy Trade-offs in Multi-Task Learning12h◆Predictive Training with Latent Imagination for Visual Quadruped Navigation12h◆Where Not to Learn: Prior-Aligned Training with Subset-based Attribution Constraints for Reliable Decision-Making12h◆Did We Actually Fix It? An Independent Adversarial Stress-Test of Post-Point-Adjustment Evaluation Metrics for Time-Series Anomaly Detection12h◆Supervised Reward Inference12h◆PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization12h◆Is Progressive Disclosure All You Need for Long-Context Agents?12h◆It Depends on the Dataset: When a Brain-Encoding Model's Predicted Responses Beat Their Visual Backbone for Video Memorability12h◆DMFNet: Dual-Backbone Multiscale Fusion Network for Urban Scene Classification12h◆Oracle Gap and Signal Fidelity: A Fixed-Pool Diagnostic for Test-Time Collaboration12h◆Scientific reasoning does not reliably translate into scientific forecasting in frontier AI12h◆Language Triggers Hijack Language Circuits: A Mechanistic Analysis of Backdoor Behaviors in Large Language Models12h◆AI-Augmented Human Resource Management? Insights from German companies12h◆
News/model/bert-base-uncased

bert-base-uncased news

47 articles mentioning bert-base-uncased

arxiv4d ago

Cross-Dataset Generalization in Urdu Fake News Detection: An Empirical Study with XLM-RoBERTa and a Length Confound Analysis

arXiv:2607.14131v1 Announce Type: new Abstract: Urdu fake news detection remains under-resourced despite Urdu being spoken by over 231 million people worldwide. While prior work has demonstrated strong in-domain performance on individual Urdu datasets, cross-dataset generalisation has received littl

arxiv5d ago

Translation as a Computationally Efficient Bridge: Feasibility of English BERT for Low-Resource Languages

arXiv:2607.12612v1 Announce Type: new Abstract: BERT models have revolutionised Natural Language Processing (NLP) through their ability to process unstructured text across diverse domains. However, developing high-quality BERT models for non-English languages remains challenging due to limited annot

arxivJul 14

Polarization Detection: A Hybrid Approach with AfroXLMR-Social and DeBERTa for Low- and High-Resource Settings

arXiv:2607.10312v1 Announce Type: cross Abstract: The rapid proliferation of online polarization threatens social cohesion, necessitating robust automated detection systems that operate effectively across diverse linguistic contexts. This paper presents our system description for the POLAR Shared Ta

arxivJul 14

Comparative Analysis of GAT and BERT for Human-Like Playtesting

arXiv:2607.11501v1 Announce Type: new Abstract: Accurately modeling and understanding player experience is crucial for designing engaging puzzle games. To achieve this, a common approach involves collecting diverse user data to train predictive playtesting models that mimic player behavior. However,

arxivJul 3

BamiBERT: A New BERT-based Language Model for Vietnamese

arXiv:2607.02259v1 Announce Type: new Abstract: In this paper, we introduce BamiBERT, a new BERT-based pre-trained language model for Vietnamese that addresses key limitations of PhoBERT -- the current de facto Vietnamese text encoder. Trained from scratch on a 129GB corpus of general-domain Vietnam

arxivJul 1

SpikeLogBERT: Energy-Efficient Log Parsing Using Spiking Transformer Networks

arXiv:2606.31781v1 Announce Type: cross Abstract: Log parsing is a fundamental step in automated log analysis, transforming raw system logs into structured event templates for downstream tasks such as anomaly detection and system monitoring. Existing log parsing methods range from rule-based and clu

arxivJun 30

BabyHuBERT: Multilingual Self-Supervised Learning for Segmenting Speakers in Child-Centered Long-Form Recordings

arXiv:2509.15001v3 Announce Type: replace-cross Abstract: Child-centered daylong recordings are essential for studying early language development, but existing speech models trained on clean adult data perform poorly due to acoustic and linguistic differences. We introduce BabyHuBERT, a self-supervi

arxivJun 30

Legal Domain Adaptation of Modern BERT Models

arXiv:2606.28538v1 Announce Type: new Abstract: We investigate domain adaptation of modern BERT models in the legal domain. We further pre-train ModernBERT on all US court opinions using the masked language modeling objective. Although ModernBERT has been trained on roughly 500x more data than origi

arxivJun 30

MauBERT: Universal Phonetic Inductive Biases for Few-Shot Acoustic Units Discovery

arXiv:2512.19612v2 Announce Type: replace Abstract: This paper introduces MauBERT, a multilingual extension of HuBERT that leverages articulatory features for robust cross-lingual phonetic representation learning. We continue HuBERT pre-training with supervision based on a phonetic-to-articulatory f

arxivJun 30

BERTomelo: Your Portuguese Encoder Best Friend

arXiv:2606.28999v1 Announce Type: cross Abstract: Encoders have become the state of the art for multiple NLP tasks, especially those requiring deep contextual understanding. While multilingual models offer broad coverage, dedicated monolingual encoders are essential for capturing the unique lexical

arxivJun 26bullish

RecallRisk-BERT: A Multi-Task Framework for Post-Report Medical Device Recall Triage

arXiv:2606.27174v1 Announce Type: new Abstract: Medical device recalls are a critical regulatory mechanism for protecting patient safety. The growing volume of FDA recall records presents challenges in post-report recall triage, severity assessment, and root-cause interpretation. Existing studies mo

#medical-device#recall-prediction#multi-task-learning
arxivJun 26

Comparing BERT Sentence-Pair Classification and Few-Shot LLM Prompting for Detecting Threat and Solution Framing in German Climate News

arXiv:2606.26489v1 Announce Type: new Abstract: News media play a central role in shaping public perceptions of climate change, and whether coverage emphasizes threats or solutions has measurable effects on audience engagement and policy support. Automated detection of these framing patterns at the

arxivJun 26

Quantum Maximum Likelihood Prediction via Hilbert Space Embeddings

arXiv:2602.18364v3 Announce Type: replace-cross Abstract: Maximum likelihood prediction (MLP) is a core task at the heart of modern large language models. Here, we study a quantum version of this task for a simplified data model consisting of independent and identically distributed samples, as a fir

arxivJun 25

Spam and Sentiment Detection in Arabic Tweets Using MARBERT Model

arXiv:2606.25495v1 Announce Type: new Abstract: Saudi Telecom Company (STC) is among the most popular companies in Saudi Arabia, with many customers. Yet, there is still a big room for improvement in users' satisfaction. Social media is the most robust platform to gauge users' satisfaction and deter

arxivJun 24

L3Cube-MahaPOS: A Marathi Part-of-Speech Tagging Dataset and BERT Models

arXiv:2606.24825v1 Announce Type: new Abstract: Part-of-Speech (POS) tagging is a foundational NLP task underpinning machine translation, information extraction, and syntactic parsing. Despite Marathi being spoken by over 83 million people and ranking among the top twenty most spoken languages world

arxivJun 20

IHUBERT: Vector-Based Semantic Deduplication and Domain-Balanced Pretraining for Persian Resources

arXiv:2606.20089v1 Announce Type: cross Abstract: Persian pretrained language models (PLMs) are still limited by the scarcity of large-scale, high-quality pretraining corpora and by insufficient evaluation beyond standard classification and NER tasks. We present IHUBERT, a monolingual Persian PLM tr

arxivJun 19

MiqraBERT: Regression-Based Sentence-BERT Finetuning for Biblical Hebrew Parallel Detection

arXiv:2606.19638v1 Announce Type: new Abstract: Textual reuse pervades the Hebrew Bible, yet the computational methods used to detect it still rest largely on lexical overlap, and they falter once a parallel involves paraphrase, lexical substitution, or syntactic reworking. This paper introduces Miq

arxivJun 18

Hilbert-Geo: Solving Solid Geometric Problems by Neural-Symbolic Reasoning

arXiv:2605.16385v3 Announce Type: replace-cross Abstract: Geometric problem solving, as a typical multimodal reasoning problem, has attracted much attention and made great progress recently, however most of works focus on plane geometry while usually fail in solid geometry due to 3D spatial diagrams

arxivJun 17

The Critical Role of Model Selection in Causal Inference: A Comparative Analysis of Classification Models within the InferBERT Framework for Pharmacovigilance

arXiv:2606.17113v1 Announce Type: cross Abstract: Distinguishing causal adverse drug events (ADEs) from spurious correlations remains a central challenge in pharmacovigilance. The InferBERT framework integrates transformer models with Do-calculus, but its success hinges on the underlying classificat

arxivJun 17

RooseBERT: A New Deal For Political Language Modelling

arXiv:2508.03250v4 Announce Type: replace-cross Abstract: The increasing amount of political debates and politics-related discussions calls for the definition of novel computational methods to automatically analyse such content with the final goal of lightening up political deliberation to citizens.

arxivJun 15

IntSeqBERT: Learning Arithmetic Structure in OEIS via Modulo-Spectrum Embeddings

arXiv:2603.05556v2 Announce Type: replace Abstract: Integer sequences in the OEIS span values from single-digit constants to astronomical factorials and exponentials, making prediction challenging for standard tokenised models that cannot handle out-of-vocabulary values or exploit periodic arithmeti

arxivJun 15

A Computational Audit of Demographic Association Encoding in ClinicalBERT Language Predictions

arXiv:2606.14460v1 Announce Type: new Abstract: Transformer-based clinical language models are increasingly integrated into high-stakes clinical decision support pipelines, yet the computational mechanisms through which demographic associations encoded in medical documentation propagate into model p

arxivJun 12

UR-BERT: Scaling Text Encoders for Massively Multilingual TTS Through Universal Romanization and Speech Token Prediction

arXiv:2606.11681v2 Announce Type: replace Abstract: We propose UR-BERT, a Romanized transcription-based text-to-speech (TTS) encoder for massively multilingual TTS systems. Conventional grapheme-to-phoneme (G2P)-based approaches are limited to around 100 languages due to the availability of reliable

arxivJun 12

MentalMARBERT: Domain-Adaptive Pre-training and Two-Stage Fine-Tuning for Arabic Mental Health Disorders Detection

arXiv:2606.12649v1 Announce Type: new Abstract: Detecting mental health disorders from Arabic social media text remains challenging due to dialectal variation, informal language, limited high-quality annotated resources, and severe class imbalance. While English mental health natural language proces

arxivJun 5

ColBERTSaR: Sparsified ColBERT Index via Product Quantization

arXiv:2606.05568v1 Announce Type: cross Abstract: While ColBERT is an effective neural retrieval architecture, it requires a heavy index structure to support candidate set retrieval based on approximated token embeddings, gathering and decompressing document token embeddings, and applying the MaxSim

arxivJun 3

KliniskVestBERT: BERT Model Specialised to Norwegian Clinical Texts

arXiv:2606.01904v2 Announce Type: replace-cross Abstract: The increasing application of Natural Language Processing (NLP) in healthcare demands language models specifically attuned to the complexities of clinical language. This work introduces KliniskVestBERT, a suite of three BERT-based encoder mod

arxivJun 3

The Word and the Way: Strategies for Domain-Specific BERT Pre-Training in German Medical NLP

arXiv:2606.03250v1 Announce Type: new Abstract: Digital healthcare generates vast amounts of clinical text that can support AI-assisted applications, yet German biomedical language models remain limited by older architectures or restricted training data. We present ChristBERT (Clinical- and Healthca

arxivJun 2

HalleluBERT: Let Every Token That Has Meaning Bear Its Weight

arXiv:2510.21372v2 Announce Type: replace Abstract: Transformer-based models have advanced NLP, yet Hebrew still lacks a RoBERTa encoder that is trained at scale and released in both base and large variants. We present HalleluBERT, a RoBERTa-based encoder family trained from scratch on 49.1~GB of de

arxivJun 2

Tiny Recursive Models for Solving the J2-Perturbed Lambert Problem

arXiv:2606.00895v1 Announce Type: cross Abstract: This paper presents a fast, recursive neural solver for the J2-perturbed Lambert problem based on Tiny Recursive Models (TRM), termed the TRM-Perturbed Lambert (TRM-PL) model. TRM is a weight-shared architecture whose effective capacity emerges from

arxivJun 2

PortBERT: Navigating the Depths of Portuguese Language Models

arXiv:2606.02100v1 Announce Type: new Abstract: Transformer models dominate modern NLP, but efficient, language-specific models remain scarce. In Portuguese, most focus on scale or accuracy, often neglecting training and deployment efficiency. In the present work, we introduce PortBERT, a family of

arxivJun 2

Construction of Historical Knowledge Graphs Based on BERT and Graph Neural Networks

arXiv:2606.01747v1 Announce Type: cross Abstract: Through digital humanities research and scale-up historical data analysis, a significant amount of traditional historical text is converted into structured knowledge graphs. This paper provides a high-level architecture that combines bidirectional en

arxivJun 2

EuroBERT: Scaling Multilingual Encoders for European Languages

arXiv:2503.05500v3 Announce Type: replace-cross Abstract: General-purpose multilingual vector representations, used in retrieval, regression and classification, are traditionally obtained from bidirectional encoder models. Despite their wide applicability, encoders have been recently overshadowed by

arxivJun 2

BERT4beam: Large AI Model Enabled Generalized Beamforming Optimization

arXiv:2509.11056v2 Announce Type: replace-cross Abstract: Artificial intelligence (AI) is anticipated to emerge as a pivotal enabler for the forthcoming sixth-generation (6G) wireless communication systems. However, current research efforts regarding large AI models for wireless communications prima

arxivJun 2

Cognitive-Linguistic Indicators of Depression in Online Communities: Analysed by DistilBERT and Holographic Reduced Representation

arXiv:2606.00026v1 Announce Type: new Abstract: This paper investigates whether combining cognitively grounded linguistic features with transformer-based embeddings improves automated detection of depression in online text. Using Beck's Cognitive Theory of Depression, the study extracts cognitive di

arxivJun 2

GottBERT: a pure German Language Model

arXiv:2012.02110v2 Announce Type: replace Abstract: Pre-trained language models have significantly advanced natural language processing (NLP), especially with the introduction of BERT and its optimized version, RoBERTa. While initial research focused on English, single-language models can be advanta

arxivJun 2

GeistBERT: Breathing Life into German NLP

arXiv:2506.11903v5 Announce Type: replace Abstract: Advances in transformer-based language models have highlighted the benefits of language-specific pre-training on high-quality corpora. In this context, German NLP stands to gain from updated architectures and modern datasets tailored to the linguis

arxivJun 2

SindBERT, the Sailor: Charting the Seas of Turkish NLP

arXiv:2510.21364v2 Announce Type: replace Abstract: Transformer models have revolutionized NLP, yet many morphologically rich languages remain underrepresented in large-scale pre-training efforts. With SindBERT, we set out to chart the seas of Turkish NLP, providing the first large-scale RoBERTa-bas

arxivJun 1

Separating Secrets from Placeholders: A Hybrid CNN-CodeBERT Framework for Three-Class Credential Leakage Detection

arXiv:2605.31520v1 Announce Type: cross Abstract: Credential leakage in public source code repositories poses a critical security threat, with over 23.8 million secrets exposed in 2024 alone. Existing detection tools suffer from high false-positive rates because rigid pattern matching and binary cla

arxivJun 1

Revisiting the Bertrand Paradox via Equilibrium Analysis of No-regret Learners

arXiv:2602.21620v2 Announce Type: replace-cross Abstract: We study the discrete Bertrand pricing game with a non-increasing demand function. The game has $n \ge 2$ players who simultaneously choose prices from the set $\{1/k, 2/k, \ldots, 1\}$, where $k\in\mathbb{N}$. The player who sets the lowest

arxivMay 29

Measure flow path recovery in Bayes Hilbert spaces

arXiv:2603.20329v2 Announce Type: replace-cross Abstract: We study the ill-posed problem of recovering a probability measure flow from finitely many moving localized sensors using a Bayes Hilbert framework. Relative to a fixed reference probability measure, a probability law is represented by its ce

arxivMay 28

ClinicalEncoder26AM: A Multlilingual Diagnosable ColBERT Model; Evidences from the MultiClinNER Shared Task

arXiv:2605.28521v1 Announce Type: new Abstract: ClinicalEncoder26AM is a multilingual Diagnosable ColBERT for clinical and biomedical texts, which aligns at multiple levels its token-level semantic with ClinicalMap25, a clinical latent space inspired by BioLORD-2023 and enriched with synthetic and a

arxivMay 27

DunbaaBERT: From Sacrifice to Semantics

arXiv:2605.26935v1 Announce Type: new Abstract: Large language models have achieved strong performance across many NLP tasks, yet Urdu remains comparatively underexplored due to limited resources and fragmented evaluation settings. To address this gap, we introduce DunbaaBERT, a family of Urdu RoBER

arxivMay 26

Fine-Tuning Over Architectural Complexity: Broad-Coverage PII Detection on PIIBench with DeBERTa

arXiv:2605.25816v1 Announce Type: cross Abstract: Personally identifiable information (PII) detection systems are frequently trained within narrow source or domain boundaries, limiting coverage when deployed on heterogeneous text. We study model fine-tuning on a corrected multi-source PIIBench prepa

arxivMay 26

Forgotten Words: Benchmarking NeoBERT for Dementia Detection in Low-Resource Conversational Filipino and English Speech

arXiv:2605.26007v1 Announce Type: new Abstract: Dementia detection from spontaneous speech offers a scalable approach to cognitive screening, yet NLP systems remain predominantly English-centric. This limitation is especially acute in the Philippines, where Filipino-English code-switching is pervasi

arxivMay 26

Complex Stochastic Gradient Descent and Directional Bias in Reproducing Kernel Hilbert Spaces

arXiv:2604.23017v2 Announce Type: replace Abstract: Stochastic Gradient Descent (SGD) is a known stochastic iterative method popular for large-scale convex optimization problems due to its simple implementation and scalability. Some objectives, such as those found in complex-valued neural networks,

arxivMay 25

A Comparative Evaluation of Structural Topic Models and BERTopic for Short, Open-Ended Survey Responses

arXiv:2605.23093v1 Announce Type: new Abstract: Topic modeling in applied psychology increasingly spans two methodological traditions: probabilistic bag-of-words models and newer embedding-based approaches. Yet many evaluations of these methods rely on longer and cleaner benchmark corpora, leaving l

arxivMay 25

A Fine-Tuned BERT Classifier for Personal-Letter Titles in Late-Ming and Early-Qing Collected Works

arXiv:2605.23103v1 Announce Type: cross Abstract: I present Lepton (Letter Prediction), a fine-tuned BERT classifier that predicts whether a title in a Classical Chinese wenji table of contents is a personal letter or a closely confusable preface (particularly the farewell-preface). Lepton fine-tune

HomeModelsNews