·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
Seattle Times and Newsday are the latest publications to sue OpenAI and Microsoft3h◆Hikers rescued after using Google Gemini for planning7h◆OpenAI confirms ‘wiki incident,’ says it’s ‘working on a framework’ for more disclosure8h◆OpenAI admits to German wiki ‘incident’15h◆XDOF, just three months out of stealth, is in talks for a Series B at a $1.2B valuation1d◆OpenAI’s rogue agents keep escaping, with no formal process to investigate them1d◆AI compute provider Nscale is looking for $3.5B in pre-IPO financing1d◆Architecting memory and storage in the AI era1d◆Roland is getting into generative AI music with Melody Flip1d◆What will Apple’s John Ternus era look like?1d◆Another swarm of OpenAI agents reached the open internet without the frontier lab’s knowledge1d◆Microsoft says virtually nobody was grabbing NYT articles through its chatbot1d◆Apple’s Ternus era begins as Nvidia bets on the whole AI stack1d◆Google’s Gemini Spark can now manage your Google Photos library1d◆Less than 24 hours to apply for your TechCrunch Disrupt 2026 Side Event1d◆Rogue OpenAI agents appear to have organized another attack using a German wiki1d◆Instagram’s AI detection is a mess (again)1d◆Why AI food looks like that1d◆Microsoft’s Project Zenith is a ‘distraction-free Windows experience’ for developers1d◆Sam Altman apologizes for ‘messy’ GPT-6 Astra rollout that’s locked out paying users1d◆Seattle Times and Newsday are the latest publications to sue OpenAI and Microsoft3h◆Hikers rescued after using Google Gemini for planning7h◆OpenAI confirms ‘wiki incident,’ says it’s ‘working on a framework’ for more disclosure8h◆OpenAI admits to German wiki ‘incident’15h◆XDOF, just three months out of stealth, is in talks for a Series B at a $1.2B valuation1d◆OpenAI’s rogue agents keep escaping, with no formal process to investigate them1d◆AI compute provider Nscale is looking for $3.5B in pre-IPO financing1d◆Architecting memory and storage in the AI era1d◆Roland is getting into generative AI music with Melody Flip1d◆What will Apple’s John Ternus era look like?1d◆Another swarm of OpenAI agents reached the open internet without the frontier lab’s knowledge1d◆Microsoft says virtually nobody was grabbing NYT articles through its chatbot1d◆Apple’s Ternus era begins as Nvidia bets on the whole AI stack1d◆Google’s Gemini Spark can now manage your Google Photos library1d◆Less than 24 hours to apply for your TechCrunch Disrupt 2026 Side Event1d◆Rogue OpenAI agents appear to have organized another attack using a German wiki1d◆Instagram’s AI detection is a mess (again)1d◆Why AI food looks like that1d◆Microsoft’s Project Zenith is a ‘distraction-free Windows experience’ for developers1d◆Sam Altman apologizes for ‘messy’ GPT-6 Astra rollout that’s locked out paying users1d◆
News/model/bert-base-uncased

bert-base-uncased news

46 articles mentioning bert-base-uncased

arxiv2d ago

Quantum Maximum Likelihood Prediction via Hilbert Space Embeddings

arXiv:2602.18364v4 Announce Type: replace-cross Abstract: Maximum likelihood prediction (MLP) is a core task at the heart of modern large language models. Here, we study a quantum version of this task for a simplified data model consisting of independent and identically distributed samples, as a fir

arxiv2d ago

Towards Solving the Gilbert-Pollak Conjecture via Large Language Models

arXiv:2601.22365v3 Announce Type: replace-cross Abstract: The Gilbert-Pollak Conjecture \citep{gilbert1968steiner}, also known as the Steiner Ratio Conjecture, states that for any finite point set in the Euclidean plane, the Steiner minimum tree has length at least $\sqrt{3}/2 \approx 0.866$ times t

arxiv2d ago

Efficient Context-Limited Telescope Bibliography Classification for the WASP-2025 Shared Task Using SciBERT

arXiv:2609.01647v1 Announce Type: new Abstract: The creation of telescope bibliographies is a crucial part of assessing the scientific impact of observatories and ensuring reproducibility in astronomy. This task involves identifying, categorizing, and linking scientific publications that reference o

arxiv3d ago

Polish ModernBERT: The Long and Short of Polish Language Understanding

arXiv:2609.01379v1 Announce Type: new Abstract: Encoder-only Transformers remain effective for discriminative and representation-learning tasks, yet Polish encoders still largely rely on BERT/RoBERTa-style architectures. We introduce \textbf{Polish ModernBERT}, a family of four Polish encoders avail

openai5d ago

How law firm Gilbert + Tobin governs and scales AI with OpenAI

See how Gilbert + Tobin combines CEO-led commitment, rigorous governance, and human accountability to scale ChatGPT Enterprise and Codex across the firm.

arxivAug 28

MoganColBERT-TR: A Late-Interaction Multi-Vector Retrieval Model for Turkish

arXiv:2608.26344v1 Announce Type: new Abstract: We previously reported a ModernBERT encoder trained from scratch for Turkish (MoganBERT-TR) and a single-vector embedding model built on top of it (MoganBERT-embed). This work introduces the third model in that lineage: MoganColBERT-TR, a multi-vector

arxivAug 28

Gaussian Processes and Reproducing Kernel Hilbert Spaces: Connections and Equivalences

arXiv:2506.17366v2 Announce Type: replace-cross Abstract: This monograph studies the relations between two approaches using positive definite kernels: probabilistic methods using Gaussian processes, and non-probabilistic methods using reproducing kernel Hilbert spaces (RKHS). They are widely studied

arxivAug 27

MoganBert-TR: A Turkish Encoder Foundation Model Trained from Scratch with a CLM-to-MLM Curriculum

arXiv:2608.25768v1 Announce Type: new Abstract: Turkish encoder models have adopted modern architectures while leaving the pretraining objective fixed at masked language modelling. This paper introduces MoganBert-TR, a 149M-parameter Turkish encoder foundation model trained from scratch on a languag

arxivJul 31bullish

NorBERTo: A ModernBERT Model Trained for Portuguese with 331 Billion Tokens Corpus

arXiv:2605.00086v2 Announce Type: replace Abstract: High-quality corpora are essential for advancing Natural Language Processing (NLP) in Portuguese. Building on previous encoder-only models such as BERTimbau and Albertina PT-BR, we introduce NorBERTo, a modern encoder based on the ModernBERT archit

#nlp#portuguese#language-models
arxivJul 31

HOMER: Huber-of-Means for Efficient and Robust Estimation in Hilbert Spaces

arXiv:2607.27532v1 Announce Type: cross Abstract: Heavy tails weaken high-confidence control for the empirical mean. Geometric median-of-means (MOM) also lacks a threshold that moves toward mean efficiency. We propose \emph{HOMER}, or Huber-of-Means for Efficient and Robust Estimation. HOMER aggrega

arxivJul 30

ChineseBERT: Chinese Pretraining Enhanced by Glyph and Pinyin Information

arXiv:2106.16038v3 Announce Type: replace Abstract: Recent pretraining models in Chinese neglect two important aspects specific to the Chinese language: glyph and pinyin, which carry significant syntax and semantic information for language understanding. In this work, we propose ChineseBERT, which i

arxivJul 29

BERT-based Models vs. Large Language Models for Low-Resource Named Entity Recognition: A Comparative Study on Marathi

arXiv:2607.23344v1 Announce Type: cross Abstract: Named Entity Recognition (NER) for low-resource languages such as Marathi remains a challenging task due to limited annotated resources and linguistic complexity. Although recent Large Language Models (LLMs) have demonstrated strong performance acros

arxivJul 29

DeepVRegulome: DNABERT-based deep-learning framework for predicting the functional impact of short genomic variants on the human regulome

arXiv:2511.09026v2 Announce Type: replace-cross Abstract: Whole-genome sequencing (WGS) has revealed numerous non-coding short variants whose functional impacts remain poorly understood. Despite recent advances in deep-learning genomic approaches, accurately predicting and prioritizing clinically re

arxivJul 28

DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages

arXiv:2607.21540v2 Announce Type: replace Abstract: We present DONDO, a family of open, permissively licensed automatic speech recognition (ASR) base models for African languages, built on the w2v-BERT 2.0 self-supervised speech encoder. DONDO comprises twenty-one monolingual models and five multili

arxivJul 28

BioSentinel at EXIST 2026: Soft-Label Optimization with XLM-RoBERTa for Sexism Intent Classification in Memes

arXiv:2607.24137v1 Announce Type: new Abstract: This paper describes the BioSentinel team's participation in EXIST 2026 Task 2.2: Source Intention in Memes, part of the CLEF 2026 evaluation campaign. The task requires classifying the communicative intent behind memes as direct, judgemental, or no (n

arxivJul 27

Data eccentricity, asymptotics of Gaussian RBF reproducing kernel Hilbert space, and kernel PCA

arXiv:2607.21823v1 Announce Type: new Abstract: We show that, up to isotropic scaling, the Gaussian RBF reproducing kernel Hilbert space (RKHS) is asymptotically isometric to Euclidean space in the large bandwidth limit. This strongly suggests that kernel-based constructions reliant on metric proper

arxivJul 27

Evaluation design conditions the expert-vs-auto MeSH gap: a controlled comparison of bag-of-words and BiomedBERT on the Cohen benchmark

arXiv:2607.21685v1 Announce Type: new Abstract: A systematic review begins with someone reading thousands of abstracts to identify the few that are relevant, and classifiers are used to prioritise that reading. Their inputs are often augmented with Medical Subject Headings (MeSH), assigned either by

#benchmark#evaluation#classification
arxivJul 24

ShriNep@EEUCA 2026: RAKSHAK - Multi-Task DeBERTa with Rationale Distillation and Jigsaw-Augmented Training for Toxic Intent Classification

arXiv:2607.20450v1 Announce Type: new Abstract: This paper presents two systems for the GameTox Shared Task at the Workshop on EEUCA at ACL 2026, which requires classifying World of Tanks chat utterances into six fine-grained toxic intent categories (Labels 0-5). Severe class imbalance, domain-speci

arxivJul 24

Hilbert Operator for Progressive Encoding (HOPE): A Mathematical Framework for Deconstructing Learned Representations in Deep Networks

arXiv:2607.21366v1 Announce Type: cross Abstract: Deep neural networks encode complex representations, but deconstructing this internal knowledge remains a challenge. Given the link between learning and compression, network compression offers a promising lens to analyze this knowledge. However, stan

arxivJul 22

Adversarial Robustness of Phishing Email Detection: A Comparative Study of TF-IDF + Logistic Regression and Fine-Tuned DistilBERT

arXiv:2607.18429v1 Announce Type: cross Abstract: Phishing emails remain one of the most persistent cybersecurity threats, and machine-learning classifiers are widely used to detect them. Most reported detection accuracies, however, are measured on clean, in-distribution test data rather than on ema

arxivJul 21

DynImmune-BERT: Dynamic Immune Repertoire Modeling with Neural ODE Driven Continuous Transformers

arXiv:2607.17244v1 Announce Type: new Abstract: Longitudinal T cell receptor repertoires contain signals of clonal expansion, contraction, disappearance, and reappearance after immune perturbation. Static repertoire language models usually summarize a sample as a bag of sequences, so the sampling in

arxivJul 21

BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation

arXiv:2604.09497v2 Announce Type: replace-cross Abstract: Accurate evaluation is central to the large language model (LLM) ecosystem, guiding model selection and downstream adoption across diverse use cases. In practice, however, evaluating generative outputs typically relies on rigid lexical method

arxivJul 20

Candidate Attended Dialogue State Tracking Using BERT

arXiv:2607.16021v1 Announce Type: cross Abstract: Dialogue state tracking (DST) is one of the core components in task-oriented dialogue systems. At each turn in a conversation, DST estimates the user belief or dialogue state, which is used as input for downstream modules to predict system actions an

arxivJul 17

Cross-Dataset Generalization in Urdu Fake News Detection: An Empirical Study with XLM-RoBERTa and a Length Confound Analysis

arXiv:2607.14131v1 Announce Type: new Abstract: Urdu fake news detection remains under-resourced despite Urdu being spoken by over 231 million people worldwide. While prior work has demonstrated strong in-domain performance on individual Urdu datasets, cross-dataset generalisation has received littl

arxivJul 16

Translation as a Computationally Efficient Bridge: Feasibility of English BERT for Low-Resource Languages

arXiv:2607.12612v1 Announce Type: new Abstract: BERT models have revolutionised Natural Language Processing (NLP) through their ability to process unstructured text across diverse domains. However, developing high-quality BERT models for non-English languages remains challenging due to limited annot

arxivJul 14

Polarization Detection: A Hybrid Approach with AfroXLMR-Social and DeBERTa for Low- and High-Resource Settings

arXiv:2607.10312v1 Announce Type: cross Abstract: The rapid proliferation of online polarization threatens social cohesion, necessitating robust automated detection systems that operate effectively across diverse linguistic contexts. This paper presents our system description for the POLAR Shared Ta

arxivJul 14

Comparative Analysis of GAT and BERT for Human-Like Playtesting

arXiv:2607.11501v1 Announce Type: new Abstract: Accurately modeling and understanding player experience is crucial for designing engaging puzzle games. To achieve this, a common approach involves collecting diverse user data to train predictive playtesting models that mimic player behavior. However,

arxivJul 3

BamiBERT: A New BERT-based Language Model for Vietnamese

arXiv:2607.02259v1 Announce Type: new Abstract: In this paper, we introduce BamiBERT, a new BERT-based pre-trained language model for Vietnamese that addresses key limitations of PhoBERT -- the current de facto Vietnamese text encoder. Trained from scratch on a 129GB corpus of general-domain Vietnam

arxivJul 1

SpikeLogBERT: Energy-Efficient Log Parsing Using Spiking Transformer Networks

arXiv:2606.31781v1 Announce Type: cross Abstract: Log parsing is a fundamental step in automated log analysis, transforming raw system logs into structured event templates for downstream tasks such as anomaly detection and system monitoring. Existing log parsing methods range from rule-based and clu

arxivJun 30

BabyHuBERT: Multilingual Self-Supervised Learning for Segmenting Speakers in Child-Centered Long-Form Recordings

arXiv:2509.15001v3 Announce Type: replace-cross Abstract: Child-centered daylong recordings are essential for studying early language development, but existing speech models trained on clean adult data perform poorly due to acoustic and linguistic differences. We introduce BabyHuBERT, a self-supervi

arxivJun 30

Legal Domain Adaptation of Modern BERT Models

arXiv:2606.28538v1 Announce Type: new Abstract: We investigate domain adaptation of modern BERT models in the legal domain. We further pre-train ModernBERT on all US court opinions using the masked language modeling objective. Although ModernBERT has been trained on roughly 500x more data than origi

arxivJun 30

MauBERT: Universal Phonetic Inductive Biases for Few-Shot Acoustic Units Discovery

arXiv:2512.19612v2 Announce Type: replace Abstract: This paper introduces MauBERT, a multilingual extension of HuBERT that leverages articulatory features for robust cross-lingual phonetic representation learning. We continue HuBERT pre-training with supervision based on a phonetic-to-articulatory f

arxivJun 30

BERTomelo: Your Portuguese Encoder Best Friend

arXiv:2606.28999v1 Announce Type: cross Abstract: Encoders have become the state of the art for multiple NLP tasks, especially those requiring deep contextual understanding. While multilingual models offer broad coverage, dedicated monolingual encoders are essential for capturing the unique lexical

arxivJun 26bullish

RecallRisk-BERT: A Multi-Task Framework for Post-Report Medical Device Recall Triage

arXiv:2606.27174v1 Announce Type: new Abstract: Medical device recalls are a critical regulatory mechanism for protecting patient safety. The growing volume of FDA recall records presents challenges in post-report recall triage, severity assessment, and root-cause interpretation. Existing studies mo

#medical-device#recall-prediction#multi-task-learning
arxivJun 26

Comparing BERT Sentence-Pair Classification and Few-Shot LLM Prompting for Detecting Threat and Solution Framing in German Climate News

arXiv:2606.26489v1 Announce Type: new Abstract: News media play a central role in shaping public perceptions of climate change, and whether coverage emphasizes threats or solutions has measurable effects on audience engagement and policy support. Automated detection of these framing patterns at the

arxivJun 25

Spam and Sentiment Detection in Arabic Tweets Using MARBERT Model

arXiv:2606.25495v1 Announce Type: new Abstract: Saudi Telecom Company (STC) is among the most popular companies in Saudi Arabia, with many customers. Yet, there is still a big room for improvement in users' satisfaction. Social media is the most robust platform to gauge users' satisfaction and deter

arxivJun 24

L3Cube-MahaPOS: A Marathi Part-of-Speech Tagging Dataset and BERT Models

arXiv:2606.24825v1 Announce Type: new Abstract: Part-of-Speech (POS) tagging is a foundational NLP task underpinning machine translation, information extraction, and syntactic parsing. Despite Marathi being spoken by over 83 million people and ranking among the top twenty most spoken languages world

arxivJun 20

IHUBERT: Vector-Based Semantic Deduplication and Domain-Balanced Pretraining for Persian Resources

arXiv:2606.20089v1 Announce Type: cross Abstract: Persian pretrained language models (PLMs) are still limited by the scarcity of large-scale, high-quality pretraining corpora and by insufficient evaluation beyond standard classification and NER tasks. We present IHUBERT, a monolingual Persian PLM tr

arxivJun 19

MiqraBERT: Regression-Based Sentence-BERT Finetuning for Biblical Hebrew Parallel Detection

arXiv:2606.19638v1 Announce Type: new Abstract: Textual reuse pervades the Hebrew Bible, yet the computational methods used to detect it still rest largely on lexical overlap, and they falter once a parallel involves paraphrase, lexical substitution, or syntactic reworking. This paper introduces Miq

arxivJun 18

Hilbert-Geo: Solving Solid Geometric Problems by Neural-Symbolic Reasoning

arXiv:2605.16385v3 Announce Type: replace-cross Abstract: Geometric problem solving, as a typical multimodal reasoning problem, has attracted much attention and made great progress recently, however most of works focus on plane geometry while usually fail in solid geometry due to 3D spatial diagrams

arxivJun 17

The Critical Role of Model Selection in Causal Inference: A Comparative Analysis of Classification Models within the InferBERT Framework for Pharmacovigilance

arXiv:2606.17113v1 Announce Type: cross Abstract: Distinguishing causal adverse drug events (ADEs) from spurious correlations remains a central challenge in pharmacovigilance. The InferBERT framework integrates transformer models with Do-calculus, but its success hinges on the underlying classificat

arxivJun 17

RooseBERT: A New Deal For Political Language Modelling

arXiv:2508.03250v4 Announce Type: replace-cross Abstract: The increasing amount of political debates and politics-related discussions calls for the definition of novel computational methods to automatically analyse such content with the final goal of lightening up political deliberation to citizens.

arxivJun 15

IntSeqBERT: Learning Arithmetic Structure in OEIS via Modulo-Spectrum Embeddings

arXiv:2603.05556v2 Announce Type: replace Abstract: Integer sequences in the OEIS span values from single-digit constants to astronomical factorials and exponentials, making prediction challenging for standard tokenised models that cannot handle out-of-vocabulary values or exploit periodic arithmeti

arxivJun 15

A Computational Audit of Demographic Association Encoding in ClinicalBERT Language Predictions

arXiv:2606.14460v1 Announce Type: new Abstract: Transformer-based clinical language models are increasingly integrated into high-stakes clinical decision support pipelines, yet the computational mechanisms through which demographic associations encoded in medical documentation propagate into model p

arxivJun 12

UR-BERT: Scaling Text Encoders for Massively Multilingual TTS Through Universal Romanization and Speech Token Prediction

arXiv:2606.11681v2 Announce Type: replace Abstract: We propose UR-BERT, a Romanized transcription-based text-to-speech (TTS) encoder for massively multilingual TTS systems. Conventional grapheme-to-phoneme (G2P)-based approaches are limited to around 100 languages due to the availability of reliable

arxivJun 12

MentalMARBERT: Domain-Adaptive Pre-training and Two-Stage Fine-Tuning for Arabic Mental Health Disorders Detection

arXiv:2606.12649v1 Announce Type: new Abstract: Detecting mental health disorders from Arabic social media text remains challenging due to dialectal variation, informal language, limited high-quality annotated resources, and severe class imbalance. While English mental health natural language proces

HomeModelsNews