·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
Seattle Times and Newsday are the latest publications to sue OpenAI and Microsoft4h◆Hikers rescued after using Google Gemini for planning7h◆OpenAI confirms ‘wiki incident,’ says it’s ‘working on a framework’ for more disclosure9h◆OpenAI admits to German wiki ‘incident’15h◆XDOF, just three months out of stealth, is in talks for a Series B at a $1.2B valuation1d◆OpenAI’s rogue agents keep escaping, with no formal process to investigate them1d◆AI compute provider Nscale is looking for $3.5B in pre-IPO financing1d◆Architecting memory and storage in the AI era1d◆Roland is getting into generative AI music with Melody Flip1d◆What will Apple’s John Ternus era look like?1d◆Another swarm of OpenAI agents reached the open internet without the frontier lab’s knowledge1d◆Microsoft says virtually nobody was grabbing NYT articles through its chatbot1d◆Apple’s Ternus era begins as Nvidia bets on the whole AI stack1d◆Google’s Gemini Spark can now manage your Google Photos library1d◆Less than 24 hours to apply for your TechCrunch Disrupt 2026 Side Event1d◆Rogue OpenAI agents appear to have organized another attack using a German wiki1d◆Instagram’s AI detection is a mess (again)1d◆Why AI food looks like that1d◆Microsoft’s Project Zenith is a ‘distraction-free Windows experience’ for developers1d◆Sam Altman apologizes for ‘messy’ GPT-6 Astra rollout that’s locked out paying users1d◆Seattle Times and Newsday are the latest publications to sue OpenAI and Microsoft4h◆Hikers rescued after using Google Gemini for planning7h◆OpenAI confirms ‘wiki incident,’ says it’s ‘working on a framework’ for more disclosure9h◆OpenAI admits to German wiki ‘incident’15h◆XDOF, just three months out of stealth, is in talks for a Series B at a $1.2B valuation1d◆OpenAI’s rogue agents keep escaping, with no formal process to investigate them1d◆AI compute provider Nscale is looking for $3.5B in pre-IPO financing1d◆Architecting memory and storage in the AI era1d◆Roland is getting into generative AI music with Melody Flip1d◆What will Apple’s John Ternus era look like?1d◆Another swarm of OpenAI agents reached the open internet without the frontier lab’s knowledge1d◆Microsoft says virtually nobody was grabbing NYT articles through its chatbot1d◆Apple’s Ternus era begins as Nvidia bets on the whole AI stack1d◆Google’s Gemini Spark can now manage your Google Photos library1d◆Less than 24 hours to apply for your TechCrunch Disrupt 2026 Side Event1d◆Rogue OpenAI agents appear to have organized another attack using a German wiki1d◆Instagram’s AI detection is a mess (again)1d◆Why AI food looks like that1d◆Microsoft’s Project Zenith is a ‘distraction-free Windows experience’ for developers1d◆Sam Altman apologizes for ‘messy’ GPT-6 Astra rollout that’s locked out paying users1d◆
News/model/multilingual-e5-large

multilingual-e5-large news

47 articles mentioning multilingual-e5-large

arxiv2d ago

MultiGhostBench: A Multilingual Benchmark for Long-Form LLM-Generated Text Attribution under Distribution Shifts

arXiv:2609.02379v1 Announce Type: cross Abstract: While existing work on LLM authorship attribution (AA) has made progress, available benchmarks remain limited, often focusing on English, controlled settings, or relatively outdated models, with the few multilingual studies considering only relativel

arxiv2d ago

NeoMME: A Single-Tower Multimodal-Native Multilingual Foundation Encoder for Efficient Fine-Tuning and Inference

arXiv:2609.01657v1 Announce Type: cross Abstract: Multimodal models often build on architectures designed for generative vision-language modeling, typically combining separately pretrained vision encoders with causal language models. Visual document retrievers such as ColPali repurpose these models

arxiv2d ago

MUDIDI: A Two-Stage Framework for Multilingual Dictionary Digitization with Language Models

arXiv:2606.09435v2 Announce Type: replace Abstract: Multilingual dictionaries are among the most valuable documentary resources for low-resource and endangered languages, yet many remain available only as scans. For many decades, their digitization and conversion into a machine-readable format was n

arxiv3d ago

TEIDAN: A Multilingual Multiparty Dialogue Corpus

arXiv:2609.00802v1 Announce Type: new Abstract: Multi-party interaction is a central setting for human communication and a necessary target for human-agent interaction systems that must participate in group conversation. Yet available corpora often focus on meetings, task-oriented interaction, text-

arxiv3d ago

Multilingual Medical Reasoning for Question Answering with Large Language Models

arXiv:2512.05658v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) with reasoning capabilities have recently demonstrated strong potential in medical Question Answering (QA). Existing approaches are largely English-focused and primarily rely on distillation from general-purpose L

arxiv3d ago

The Curse of Multilinguality in Lexical Normalization

arXiv:2609.00329v1 Announce Type: cross Abstract: Lexical normalization rewrites the noisy, non-standard words that fill user-generated text (tmrw, u, gr8) into their standard forms. Because labelled data is scarce for most languages, a popular shortcut is to train a single model on many languages a

arxiv3d ago

Sources of Truth: A Multi-Platform, Multilingual Audit of Citations in AI Mental Health Information Queries

arXiv:2609.00319v1 Announce Type: cross Abstract: Online health information seeking is shifting from keyword search, where users consider a ranked list of links, to conversational systems that compose a single answer and curate its citations. Source evaluation therefore passes from user to platform,

arxiv3d ago

Lingua Franca or Probing Artifact? Rethinking Latent Language in Multilingual LLMs

arXiv:2609.00155v1 Announce Type: cross Abstract: Latent language identification is often used to argue that multilingual language models route computation through language-specific states, such as English pivots. However, existing probes infer latent language from different signals, such as the geo

arxiv3d ago

Separating Syntax from Language: A Mechanistic Account of Translation in Multilingual LLMs

arXiv:2609.01356v1 Announce Type: new Abstract: Multilingual large language models (mLLMs) achieve strong performance in machine translation, yet our understanding of the mechanisms by which they transform representations from one language to another remains incomplete. Prior work suggests that tran

arxiv3d ago

Latent Mechanisms of Language Control in Multilingual Language Models

arXiv:2609.00325v1 Announce Type: new Abstract: Multilingual large language models can exhibit unintended code-switching -- unnecessarily alternating between languages during generation. We present a comparative study of three methods that identify language-controlling latents in cross-layer transco

arxiv3d ago

M4FC: a Multimodal, Multilingual, Multicultural, Multitask Real-World Fact-Checking Dataset

arXiv:2510.23508v4 Announce Type: replace Abstract: Existing real-world datasets for multimodal fact-checking have multiple limitations: they contain few instances, cover only one or two languages, focus on a single task, or rely on external news article sets to source true claims. To address these

arxiv3d ago

WorldBench: Culturally Grounded Benchmark for Multilingual Agents

arXiv:2609.01056v1 Announce Type: new Abstract: Despite the growing use of LLM-powered agents to solve multi-step tasks in complex environments, existing benchmarks rarely test state preservation, performance across languages, and application to realistic, grounded scenarios. To address these concer

arxiv3d ago

Removable and Irreducible: A Token-Cost Ledger for the Multilingual Tokenization Tax

arXiv:2609.00378v1 Announce Type: new Abstract: Large language models pay a well-documented tax on non-English text: the same content costs several times more tokens, and because attention is quadratic in sequence length, far more compute. We ask how much of this tax is removable. Framing the token

arxiv4d ago

Unveiling Language Routing Isolation in Multilingual MoE Models for Interpretable Subnetwork Adaptation

arXiv:2604.03592v2 Announce Type: replace-cross Abstract: Mixture-of-Experts (MoE) models exhibit striking performance disparities across languages, yet the internal mechanisms driving these gaps remain poorly understood. In this work, we conduct a systematic analysis of expert routing patterns in M

arxiv4d ago

Where Do Multilingual Vision-Language Encoders Fail on Low-Resource Languages?

arXiv:2608.30725v1 Announce Type: new Abstract: Recent multilingual vision--language encoders cover hundreds of languages in a single model, yet on two state-of-the-art instances retrieval on low-resource languages (LRL; e.g. Swahili) trails high-resource ones (HRL; e.g. English) by $30^+$\,pp. We a

arxiv4d ago

Generative vs. Encoder Models for Multilingual NER: A Comprehensive Empirical Study on Naamapadam

arXiv:2608.29959v1 Announce Type: new Abstract: Language is humanity's most consequential technology, yet for over a billion speakers across India's twenty-two constitutionally recognised languages, its digital layer remains structurally incomplete. Named Entity Recognition (NER), the foundational s

arxiv4d ago

QQ: A Language Metadata Toolkit for Multilingual NLP

arXiv:2603.00620v3 Announce Type: replace Abstract: Multilingual NLP research increasingly involves hundreds or thousands of languages across different datasets. Managing, discovering, and reporting language metadata becomes a common hurdle at these scales. We present QQ, a metadata toolkit and brow

arxiv4d ago

Polyglot Teachers: Evaluating Language Models for Multilingual Synthetic Data Generation

arXiv:2604.11290v3 Announce Type: replace Abstract: Synthesizing supervised finetuning (SFT) data from language models (LMs) to teach smaller models multilingual tasks has become increasingly common. However, teacher model selection is often ad hoc, typically defaulting to the largest available opti

arxiv4d ago

Beyond Good Intentions: When Does the Framing of Multilingual and Low-Resource NLP Research Become a Caricature?

arXiv:2608.30866v1 Announce Type: new Abstract: Building language technologies and conducting NLP research for low-resource languages---particularly when led by native speakers or involving participatory research practices---are often framed as means of addressing inequality, serving local communiti

arxiv4d ago

Beyond Factual QA: Mentorship-Oriented Question Answering over Long-Form Multilingual Content

arXiv:2601.17173v3 Announce Type: replace-cross Abstract: Question answering systems are typically evaluated on factual correctness, yet many real-world applications-such as education and career guidance-require mentorship: responses that provide reflection and guidance. Existing QA benchmarks rarel

arxiv4d ago

What Matters When Building Universal Multilingual Named Entity Recognition Models?

arXiv:2601.06347v3 Announce Type: replace Abstract: Recent progress in universal multilingual named entity recognition (NER) has been driven by multilingual transformer models, task-specific architectures, custom loss functions, and large-scale training datasets. However, despite substantial prior w

arxiv4d ago

Causal Interventions Reveal Typologically Organized Syntactic Mechanisms in Multilingual Language Models

arXiv:2608.28924v1 Announce Type: new Abstract: Linguistic theory has long recognized cross-linguistic syntactic regularities, leading to claims that these similar structures are processed by similar mechanisms. However, this hypothesis has been difficult to test empirically due to our lack of fine-

arxiv4d ago

Transformer-Encoder Trees for Efficient Multilingual Machine Translation and Speech Translation

arXiv:2509.17930v3 Announce Type: replace-cross Abstract: Multilingual translation suffers from computational redundancy, especially when translating into multiple languages simultaneously. In addition, translation quality can suffer for low-resource languages. To address this, we introduce Transfor

arxiv4d ago

More Capable, Less Faithful: A Multilingual Analysis of Mathematical (Un)Solvability Detection in LLMs

arXiv:2608.30463v1 Announce Type: new Abstract: Solvability detection is one of the most challenging aspects of mathematical reasoning for Large Language Models (LLMs). While prior work has studied this capability extensively, these analyses have been limited to English. Consequently, it remains unc

arxiv4d ago

Evaluating and Mitigating Anti-LGBTQ Biases in German and Multilingual Language Models

arXiv:2608.30884v1 Announce Type: cross Abstract: While gender and racial biases in language models have been widely studied, anti-LGBTQ biases remain underexplored, particularly beyond English. Existing benchmarks often do not capture cultural and linguistic variation and rely on gender representat

arxiv4d ago

Evaluating Multilingual Sentence Embeddings for Translation Error Detection:An English--Greek Contrastive Study

arXiv:2608.28776v1 Announce Type: new Abstract: Multilingual sentence embeddings are increasingly used to estimate semantic similarity across languages, yet their sensitivity to fine-grained translation errors remains insufficiently understood. This study investigates whether general-purpose multili

arxiv4d ago

Not Truly Multilingual: Script Consistency as a Missing Dimension in VLM Evaluation

arXiv:2606.17188v5 Announce Type: replace-cross Abstract: Current multilingual evaluations for Vision-Language Models (VLMs) assume a one-to-one mapping between language and orthography, overlooking billions of users of multi-script languages. We introduce PuMVR (Punjabi Multimodal Visual Reasoning)

arxiv4d ago

Beyond the Final Layer: Intermediate Representations for Better Multilingual Calibration in Large Language Models

arXiv:2510.03136v2 Announce Type: replace Abstract: Confidence calibration, the alignment of a model's predicted confidence with its actual accuracy, is crucial for the reliable deployment of Large Language Models (LLMs). However, this critical property remains largely under-explored in multilingual

arxiv4d ago

Terminal-Bench-LILT: Multilingual Agentic Coding Benchmark Grounded in Language, Region, and Culture

arXiv:2608.28641v1 Announce Type: cross Abstract: Most evaluations for coding agents are conducted exclusively in English, which does not reflect real-world multilingual deployment. We present Terminal-Bench-LILT, a suite of 300 authentic coding tasks in ten languages: Arabic, Czech, German, Spanish

arxiv5d ago

One Form to Transfer Them All: Pretraining Multilingual Language Models Beyond Native Orthography

arXiv:2608.25904v2 Announce Type: replace Abstract: Multilingual language models transfer knowledge across languages through shared subword vocabulary, a mechanism that breaks down when related languages use different writing systems. Prior work addresses this via script equalization (romanization o

arxiv5d ago

OmniFusion: Simultaneous Multilingual Multimodal Translations via Modular Fusion

arXiv:2512.00234v3 Announce Type: replace Abstract: There has been significant progress in open-source text-only translation large language models (LLMs) with better language coverage and quality. However, these models can be only used in cascaded pipelines for speech translation (ST), performing au

arxiv5d ago

CultureConverse: A Multilingual Multi-turn Simulation Harness for Culturally Grounded Assistance in East and Southeast Asia

arXiv:2608.28405v1 Announce Type: new Abstract: Current cultural evaluations for large language models (LLMs) often reduce culture to single-turn factual recall via MCQs, failing to capture a common use case: users seeking practical help over multiple turns in culturally grounded scenarios. We intro

arxiv5d ago

SimpCue: Cue-Based Prompting for Multilingual Text Simplification

arXiv:2608.28042v1 Announce Type: new Abstract: Text simplification aims to make complex texts easier to understand while preserving their original meaning. Recent large language models can perform simplification through prompting, but it remains unclear whether adding explicit linguistic informatio

arxiv5d ago

Multilingual Lexical Feature Analysis of Spoken Language for Predicting Major Depression Symptom Severity

arXiv:2511.07011v2 Announce Type: replace Abstract: Background: Remotely captured spoken language could provide objective, regular indicators of depression symptom severity. However, research to date has largely used non-clinical, cross-sectional written language and complex machine learning (ML) ap

arxivAug 29

Training-Time Explainability for Multilingual Hate Speech Detection: Aligning Model Reasoning with Human Rationales

arXiv:2608.26125v1 Announce Type: cross Abstract: Online hate against Muslim communities often appears in culturally coded, multilingual forms that evade conventional AI moderation. Such systems, though accurate, remain opaque and risk bias, over-censorship, or under-moderation, particularly when de

arxivAug 29

SHIFT: Semantic Harmonization via Index-side Feature Transformation for Multilingual Information Retrieval

arXiv:2606.18801v2 Announce Type: replace-cross Abstract: With the rapid expansion of massive multilingual corpora, Multilingual Information Retrieval (MLIR) has emerged as a critical technology for global information access. MLIR enables users to retrieve semantically relevant documents from multil

arxivAug 29

MIMO: Multilingual Information Retrieval via Monolingual Objectives

arXiv:2605.31171v2 Announce Type: replace-cross Abstract: Multilingual Information Retrieval (MLIR) reflects real-world search environments in which queries and relevant documents may appear in different languages within a mixed-language corpus. However, existing embedding models are primarily optim

arxivAug 29

VFA: Empowering Multilingual MLLMs via Vision-Free Adaptation

arXiv:2608.26155v1 Announce Type: cross Abstract: Multimodal large language models have advanced rapidly, yet most remain English-centric, as scaling multilingual multimodal instruction tuning is limited by the scarcity and high cost of high-quality non-English image-text supervision. Although multi

arxivAug 28

Vowel Signs Are Not Letters: A Pre-tokenization Ceiling on Multilingual Tokenizer Fertility

arXiv:2608.26449v1 Announce Type: new Abstract: Byte-level BPE tokenizers that use the HuggingFace ByteLevel pre-tokenizer inherit GPT-2's word regex, where a word is defined as \p{L}+, one or more Unicode letters. In abugida scripts, vowels are written as combining marks; this pattern therefore spl

arxivAug 28

Many Dialects, Many Languages, One Cultural Lens: Evaluating Multilingual VLMs for Bengali Culture Understanding Across Historically Linked Languages and Regional Dialects

arXiv:2603.21165v3 Announce Type: replace Abstract: Bangla culture is richly expressed through region, dialect, history, food, politics, media, and everyday visual life, yet it remains underrepresented in multimodal evaluation. To address this gap, we introduce BanglaVerse, a culturally grounded ben

arxivAug 27

EstLLM: Enhancing Estonian Capabilities in Multilingual LLMs via Continued Pretraining and Post-Training

arXiv:2603.02041v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are predominantly trained on English-centric data, resulting in uneven performance for smaller languages. We study whether continued pretraining (CPT) can improve Estonian capabilities in multilingual LLMs while p

arxivAug 27

Rethinking the Multilingual Reasoning Gap with Layer Swap

arXiv:2605.26735v2 Announce Type: replace Abstract: Recent reasoning Large Language Models produce a chain-of-thought (CoT) predominantly in English, even when prompted in non-English languages. Prior work suggests that forcing the CoT to remain in the input language (native reasoning) substantially

arxivAug 27

Can We Read the Mind of an Audio LLM? A Verbalizable, Multilingual Middle-Layer Workspace

arXiv:2608.24958v1 Announce Type: cross Abstract: An audio language model is a black box in a specific way: we see what it says, never what it works out on the way there, and chain-of-thought monitoring helps only if the model writes its reasoning down. Reading a base Qwen3-Omni with a logit lens at

arxivAug 27

Macro: Enhancing Multilingual Counterfactual Explanations through Alignment-as-Preference Optimization

arXiv:2605.11632v3 Announce Type: replace Abstract: Self-generated counterfactual explanations (SCEs) are minimally modified inputs (minimality) generated by large language models (LLMs) that flip their own predictions (validity), offering a causally grounded approach to unraveling black-box LLM beh

huggingfaceAug 10

Build Low-Latency Multilingual Voice Agents: Open Weights & Full Deployment Control with NVIDIA Magpie TTS

arxivAug 3

M3-DuplexBench: A Multi-Turn, Multilingual, Multidomain Benchmark for Full-Duplex Spoken Dialogue Models

arXiv:2607.29125v1 Announce Type: new Abstract: Full-duplex spoken dialogue systems (FDSDSs) can listen while speaking, enabling natural behaviors such as smooth turn-taking, backchannel handling, and user barge-in handling. However, fair comparisons in multi-turn conversations remain a challenge. I

arxivAug 3

DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code Search

arXiv:2607.27178v2 Announce Type: replace Abstract: State-of-the-art retrieval models increasingly rely on closed training data, creating a reproducibility gap. We present an open end-to-end recipe for training retrieval models and study how English supervision transfers to multilingual retrieval th

HomeModelsNews