Model Detail
multilingual-e5-small
—multilingual-e5-small is a large language model with 59M parameters released by intfloat. The model is registered under the sentence-similarity pipeline tag on Hugging Face, distributed under the permissive mit license.
multilingual-e5-small ships with 59M parameters. The mit license is permissive, allowing commercial deployment and derivative work without per-seat fees, though attribution requirements still apply.
multilingual-e5-small is best fit for general-purpose chat and instruction-following workloads. Treat this as a starting matrix rather than a benchmark verdict — the right deployment usually depends on the specific evaluation suite that mirrors your workload.
X-Translator: A Real-Time Multilingual Speaker-Aware Speech-to-Speech Translation System
arXiv:2607.17544v1 Announce Type: cross Abstract: Real-time speech-to-speech translation (S2ST) systems must balance translation quality, latency, speech naturalness, and speaker consistency. Publicly documented S2ST systems have advanced direct, multilingual, streaming, and expressive modeling, whi
Lost in Reordering: Structural Sensitivity of Multilingual LLMs under Semantics-Preserving Perturbations
arXiv:2609.03511v1 Announce Type: new Abstract: Large Language Models (LLMs) demonstrate strong multilingual reasoning performance, yet their robustness to semantics-preserving structural variation remains underexplored, particularly for relatively free word-order languages. We investigate the struc
One Model to Translate Them All? A Journey to Mount Doom for Multilingual Model Merging
arXiv:2604.02881v2 Announce Type: replace-cross Abstract: Weight-space model merging combines independently fine-tuned checkpoints without access to the original training data. While merging has shown promise in multitask settings, its behavior in multilingual generative systems remains underexplore
IndicSafeEval: Safety Robustness of Large Language Models under Multilingual Persuasive Jailbreak Attacks
arXiv:2609.03781v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used in multilingual settings, yet their safety is still evaluated primarily in English. This limits our understanding of how alignment failures manifest in low-resource and culturally diverse languages.
NeoMME: an efficient Multimodal-native and Multilingual Encoder
MUDIDI: A Two-Stage Framework for Multilingual Dictionary Digitization with Language Models
arXiv:2606.09435v2 Announce Type: replace Abstract: Multilingual dictionaries are among the most valuable documentary resources for low-resource and endangered languages, yet many remain available only as scans. For many decades, their digitization and conversion into a machine-readable format was n