Model Detail
multilingual-e5-small
—multilingual-e5-small is a large language model with 59M parameters released by intfloat. The model is registered under the sentence-similarity pipeline tag on Hugging Face, distributed under the permissive mit license.
multilingual-e5-small ships with 59M parameters. The mit license is permissive, allowing commercial deployment and derivative work without per-seat fees, though attribution requirements still apply.
multilingual-e5-small is best fit for general-purpose chat and instruction-following workloads. Treat this as a starting matrix rather than a benchmark verdict — the right deployment usually depends on the specific evaluation suite that mirrors your workload.
Conditional Reliability of Toxicity Signals for Multilingual and Code-Mixed Abuse Detection
arXiv:2607.15861v1 Announce Type: cross Abstract: Moderation systems increasingly rely on external toxicity tools, but those tools are unreliable under code-mixing, transliteration, slang, and language mismatch. We study the \emph{conditional reliability} of toxicity priors in Indian multilingual an
Echoes Across Vietnam's Highlands, Delta, and Coast: A Multilingual Corpus for Cham, Khmer, and Tay-Nung
arXiv:2607.08362v2 Announce Type: replace Abstract: Vietnam's ethnic minority languages are almost absent from the field of Natural Language Processing (NLP), and the challenge goes beyond data scarcity: Cham, Khmer, and Tay-Nung differ sharply in script, Vietnamese contact, and standardization, con
Large Language Models for Code Generation from Multilingual Prompts: A Curated Benchmark and a Study on Code Quality
arXiv:2607.14816v1 Announce Type: cross Abstract: Large Language Models (LLMs) perform differently on identical programming tasks when prompted in different natural languages, a phenomenon known as language bias. While this behavior has been widely studied for general text generation, its impact on
MSQA: A Natively Sourced Multilingual and Multicultural SimpleQA Benchmark
arXiv:2607.00724v3 Announce Type: replace Abstract: Multilingual fluency often invites a stronger assumption: a model that can speak a user's language must also understand the culture encoded by that language. We call this the Illusion of Cultural Alignment. To test this assumption directly, we intr
Diarization-Guided Qwen-ASR Adaptation for Multilingual Two-Speaker Conversational Speech
arXiv:2607.08208v2 Announce Type: replace Abstract: This paper describes our self-designed system for Task 1 of the MLC-SLM 2026 Challenge for multilingual two-speaker conversational speech. The system combines a modular speaker diarization front end with a challenge-adapted Qwen3-ASR-1.7B recognize
MET: Theory-Grounded and Culture-Aware Multilingual Moral Reasoning
arXiv:2607.11736v1 Announce Type: new Abstract: Language models are increasingly used for moral decision-making across diverse linguistic and cultural contexts, yet existing work overlooks multilinguality on three aspects: 1) multilingual evaluation benchmarks use direct translation, failing to adap