·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
America needs to stop getting shocked by Chinese AI2h◆Advancing next-gen AI with materials science innovation2h◆Gritt exits stealth with $34 million for robots to build solar plants—then, everything else3h◆Capacity and Redundancy Trade-offs in Multi-Task Learning9h◆Predictive Training with Latent Imagination for Visual Quadruped Navigation9h◆Where Not to Learn: Prior-Aligned Training with Subset-based Attribution Constraints for Reliable Decision-Making9h◆Did We Actually Fix It? An Independent Adversarial Stress-Test of Post-Point-Adjustment Evaluation Metrics for Time-Series Anomaly Detection9h◆Supervised Reward Inference9h◆PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization9h◆Is Progressive Disclosure All You Need for Long-Context Agents?9h◆It Depends on the Dataset: When a Brain-Encoding Model's Predicted Responses Beat Their Visual Backbone for Video Memorability9h◆DMFNet: Dual-Backbone Multiscale Fusion Network for Urban Scene Classification9h◆Oracle Gap and Signal Fidelity: A Fixed-Pool Diagnostic for Test-Time Collaboration9h◆Time-Frequency Consistency Learning for Robust Speech Deepfake Detection9h◆Scientific reasoning does not reliably translate into scientific forecasting in frontier AI9h◆Language Triggers Hijack Language Circuits: A Mechanistic Analysis of Backdoor Behaviors in Large Language Models9h◆Kernel Regression with Tensor Trains and Hadamard Overparameterization9h◆AI-Augmented Human Resource Management? Insights from German companies9h◆Diagnosing Correctness Probes under Self-Judgement Confounding9h◆BLAD: A Historically Contextualized, Multilingual Dataset of Bangladeshi Legal Acts (1799 to 2025)9h◆America needs to stop getting shocked by Chinese AI2h◆Advancing next-gen AI with materials science innovation2h◆Gritt exits stealth with $34 million for robots to build solar plants—then, everything else3h◆Capacity and Redundancy Trade-offs in Multi-Task Learning9h◆Predictive Training with Latent Imagination for Visual Quadruped Navigation9h◆Where Not to Learn: Prior-Aligned Training with Subset-based Attribution Constraints for Reliable Decision-Making9h◆Did We Actually Fix It? An Independent Adversarial Stress-Test of Post-Point-Adjustment Evaluation Metrics for Time-Series Anomaly Detection9h◆Supervised Reward Inference9h◆PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization9h◆Is Progressive Disclosure All You Need for Long-Context Agents?9h◆It Depends on the Dataset: When a Brain-Encoding Model's Predicted Responses Beat Their Visual Backbone for Video Memorability9h◆DMFNet: Dual-Backbone Multiscale Fusion Network for Urban Scene Classification9h◆Oracle Gap and Signal Fidelity: A Fixed-Pool Diagnostic for Test-Time Collaboration9h◆Time-Frequency Consistency Learning for Robust Speech Deepfake Detection9h◆Scientific reasoning does not reliably translate into scientific forecasting in frontier AI9h◆Language Triggers Hijack Language Circuits: A Mechanistic Analysis of Backdoor Behaviors in Large Language Models9h◆Kernel Regression with Tensor Trains and Hadamard Overparameterization9h◆AI-Augmented Human Resource Management? Insights from German companies9h◆Diagnosing Correctness Probes under Self-Judgement Confounding9h◆BLAD: A Historically Contextualized, Multilingual Dataset of Bangladeshi Legal Acts (1799 to 2025)9h◆
DataBubble·

Model Detail

intfloat logo

multilingual-e5-small

—
Provider: intfloatCategory: llmPipeline: sentence-similarity
DB Score
0.6
Downloads
10.0M
Likes
348
Day
+0.0%
Week
+0.0%
Month
+0.0%
Overview

multilingual-e5-small is a large language model with 59M parameters released by intfloat. The model is registered under the sentence-similarity pipeline tag on Hugging Face, distributed under the permissive mit license.

Technical

multilingual-e5-small ships with 59M parameters. The mit license is permissive, allowing commercial deployment and derivative work without per-seat fees, though attribution requirements still apply.

Use Cases

multilingual-e5-small is best fit for general-purpose chat and instruction-following workloads. Treat this as a starting matrix rather than a benchmark verdict — the right deployment usually depends on the specific evaluation suite that mirrors your workload.

Download History
Research Paper
arXiv: 2402.05672→
Model Info
Licensemit
Recent newsView all news →
Related News
arxiv1d ago

Conditional Reliability of Toxicity Signals for Multilingual and Code-Mixed Abuse Detection

arXiv:2607.15861v1 Announce Type: cross Abstract: Moderation systems increasingly rely on external toxicity tools, but those tools are unreliable under code-mixing, transliteration, slang, and language mismatch. We study the \emph{conditional reliability} of toxicity priors in Indian multilingual an

arxivneutral1d ago

Echoes Across Vietnam's Highlands, Delta, and Coast: A Multilingual Corpus for Cham, Khmer, and Tay-Nung

arXiv:2607.08362v2 Announce Type: replace Abstract: Vietnam's ethnic minority languages are almost absent from the field of Natural Language Processing (NLP), and the challenge goes beyond data scarcity: Cham, Khmer, and Tay-Nung differ sharply in script, Vietnamese contact, and standardization, con

arxivneutral3d ago

Large Language Models for Code Generation from Multilingual Prompts: A Curated Benchmark and a Study on Code Quality

arXiv:2607.14816v1 Announce Type: cross Abstract: Large Language Models (LLMs) perform differently on identical programming tasks when prompted in different natural languages, a phenomenon known as language bias. While this behavior has been widely studied for general text generation, its impact on

arxivbearish4d ago

MSQA: A Natively Sourced Multilingual and Multicultural SimpleQA Benchmark

arXiv:2607.00724v3 Announce Type: replace Abstract: Multilingual fluency often invites a stronger assumption: a model that can speak a user's language must also understand the culture encoded by that language. We call this the Illusion of Cultural Alignment. To test this assumption directly, we intr

arxivneutral4d ago

Diarization-Guided Qwen-ASR Adaptation for Multilingual Two-Speaker Conversational Speech

arXiv:2607.08208v2 Announce Type: replace Abstract: This paper describes our self-designed system for Task 1 of the MLC-SLM 2026 Challenge for multilingual two-speaker conversational speech. The system combines a modular speaker diarization front end with a challenge-adapted Qwen3-ASR-1.7B recognize

arxiv7d ago

MET: Theory-Grounded and Culture-Aware Multilingual Moral Reasoning

arXiv:2607.11736v1 Announce Type: new Abstract: Language models are increasingly used for moral decision-making across diverse linguistic and cultural contexts, yet existing work overlooks multilinguality on three aspects: 1) multilingual evaluation benchmarks use direct translation, failing to adap

Related Models
intfloat logo
multilingual-e5-large
intfloat · 5.0M downloads
google-bert logo
bert-base-uncased
google-bert · 69.6M downloads
sentence-transformers logo
paraphrase-multilingual-MiniLM-L12-v2
SBERT · 48.6M downloads
HomeModelsNews