·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
US threatens sanctions against Chinese AI models over IP theft52m◆Google launches a cheaper alternative to large AI security models like Mythos1h◆Music streamer Deezer says more than 50% of daily uploads are AI-generated3h◆Halliday’s latest smart glasses feature a much-improved display3h◆America needs to stop getting shocked by Chinese AI5h◆Advancing next-gen AI with materials science innovation5h◆Gritt exits stealth with $32 million for robots to build solar plants — then, everything else6h◆Capacity and Redundancy Trade-offs in Multi-Task Learning12h◆Predictive Training with Latent Imagination for Visual Quadruped Navigation12h◆Where Not to Learn: Prior-Aligned Training with Subset-based Attribution Constraints for Reliable Decision-Making12h◆Did We Actually Fix It? An Independent Adversarial Stress-Test of Post-Point-Adjustment Evaluation Metrics for Time-Series Anomaly Detection12h◆Supervised Reward Inference12h◆PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization12h◆Is Progressive Disclosure All You Need for Long-Context Agents?12h◆It Depends on the Dataset: When a Brain-Encoding Model's Predicted Responses Beat Their Visual Backbone for Video Memorability12h◆DMFNet: Dual-Backbone Multiscale Fusion Network for Urban Scene Classification12h◆Oracle Gap and Signal Fidelity: A Fixed-Pool Diagnostic for Test-Time Collaboration12h◆Scientific reasoning does not reliably translate into scientific forecasting in frontier AI12h◆Language Triggers Hijack Language Circuits: A Mechanistic Analysis of Backdoor Behaviors in Large Language Models12h◆AI-Augmented Human Resource Management? Insights from German companies12h◆US threatens sanctions against Chinese AI models over IP theft52m◆Google launches a cheaper alternative to large AI security models like Mythos1h◆Music streamer Deezer says more than 50% of daily uploads are AI-generated3h◆Halliday’s latest smart glasses feature a much-improved display3h◆America needs to stop getting shocked by Chinese AI5h◆Advancing next-gen AI with materials science innovation5h◆Gritt exits stealth with $32 million for robots to build solar plants — then, everything else6h◆Capacity and Redundancy Trade-offs in Multi-Task Learning12h◆Predictive Training with Latent Imagination for Visual Quadruped Navigation12h◆Where Not to Learn: Prior-Aligned Training with Subset-based Attribution Constraints for Reliable Decision-Making12h◆Did We Actually Fix It? An Independent Adversarial Stress-Test of Post-Point-Adjustment Evaluation Metrics for Time-Series Anomaly Detection12h◆Supervised Reward Inference12h◆PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization12h◆Is Progressive Disclosure All You Need for Long-Context Agents?12h◆It Depends on the Dataset: When a Brain-Encoding Model's Predicted Responses Beat Their Visual Backbone for Video Memorability12h◆DMFNet: Dual-Backbone Multiscale Fusion Network for Urban Scene Classification12h◆Oracle Gap and Signal Fidelity: A Fixed-Pool Diagnostic for Test-Time Collaboration12h◆Scientific reasoning does not reliably translate into scientific forecasting in frontier AI12h◆Language Triggers Hijack Language Circuits: A Mechanistic Analysis of Backdoor Behaviors in Large Language Models12h◆AI-Augmented Human Resource Management? Insights from German companies12h◆
News/model/paraphrase-multilingual-MiniLM-L12-v2

paraphrase-multilingual-MiniLM-L12-v2 news

16 articles mentioning paraphrase-multilingual-MiniLM-L12-v2

arxivJun 20

Target-Side Paraphrase Augmentation for Sign Language Translation with Large Language Models

arXiv:2605.31393v2 Announce Type: replace-cross Abstract: Sign language translation (SLT) remains constrained by the limited availability of paired sign-video/text corpora and by the heavy-tailed vocabularies typical of real-world datasets. We study a target-side augmentation strategy in which a lar

arxivJun 16

SAMark: A Self-Anchored Text Watermarking with Paragraph-Level Paraphrase Robustness

arXiv:2605.25796v2 Announce Type: replace-cross Abstract: Semantic-level watermarking (SWM) improves robustness against text modifications by treating sentences as the basic unit. However, robustness to paragraph-level paraphrasing remains difficult because such attacks globally disrupt watermark si

arxivJun 15

Deja Vu at Scale: Paraphrase-Robust Detection of Duplicate Gherkin Steps in Behaviour-Driven Software Testing with Sentence-Transformer Embeddings and a 1.1M-Step Open Benchmark

arXiv:2604.20462v3 Announce Type: replace-cross Abstract: Context. Behaviour-Driven Development (BDD) suites in Gherkin accumulate step-text duplication with documented maintenance cost. Prior detectors either require runnable tests or are single-organisation, leaving a gap: a static, paraphrase-rob

arxivJun 10

Unsupervised Style Representation Learning for AI-Text Detection via Paraphrase Inversion

arXiv:2606.10099v1 Announce Type: cross Abstract: The rapid development of large language models (LLMs) has raised concerns about misuse such as plagiarism, misinformation, and automated influence operations, motivating the need for robust detectors. Recent work has shown that neural representations

arxivJun 3

Enhancing Paraphrase Type Generation: The Impact of DPO and RLHF Evaluated with Human-Ranked Data

arXiv:2506.02018v2 Announce Type: replace Abstract: Paraphrasing re-expresses meaning to enhance applications like text simplification, machine translation, and question-answering. Specific paraphrase types facilitate accurate semantic analysis and robust language models. However, existing paraphras

arxivMay 28

Paraphrase Brittleness in Production Retrieval-Augmented Commercial Recommendation: Reproducibility Below the Rerun-Stability Baseline

arXiv:2605.27440v1 Announce Type: cross Abstract: Small changes to how a buyer phrases a question -- "best CRM" vs "top CRM" vs "best CRM for a SaaS startup" -- produce substantially different brand recommendations from AI assistants. Across ~6,000 paraphrase runs and ~6,000 same-prompt rerun contro

arxivMay 19

Characterizing Paraphrase-Induced Failures in Lean 4 Autoformalization

arXiv:2604.23135v2 Announce Type: replace Abstract: Lean 4 autoformalization has become increasingly popular in recent years, with frontier language models and open-weight autoformalizers now producing valid formalizations of mathematical theorems. However, these evaluations often rely on single can

arxivMay 12

Paraphrase-Induced Output-Mode Collapse: When LLMs Break Character Under Semantically Equivalent Inputs

arXiv:2605.04665v2 Announce Type: replace Abstract: When the substantive content of a request is rewritten, do large language models still answer in the format the original task asked for? We find that they often do not, even at temperature zero. On a 150-query evaluation over five compact 2025-era

arxivMay 5

Controlled Paraphrase Geometry in Sentence Embedding Space: Local Manifold Modeling and Latent Probing

arXiv:2605.01073v1 Announce Type: new Abstract: The paper studies the local geometry of embedding clouds induced by \emph{controlled local classes of semantically close sentences}. The central question is how controlled paraphrase-like semantic variation is organized in sentence embedding space and

arxivApr 28

DualGuard: Dual-stream Large Language Model Watermarking Defense against Paraphrase and Spoofing Attack

arXiv:2512.16182v2 Announce Type: replace-cross Abstract: With the rapid development of cloud-based services, large language models have become increasingly accessible through various web platforms. However, this accessibility has also led to growing risks of model abuse. LLM watermarking has emerge

arxivApr 28

Reducing Maintenance Burden in Behaviour-Driven Development: A Paraphrase-Robust Duplicate-Step Detector with a 1.1M-Step Open Benchmark

arXiv:2604.20462v2 Announce Type: replace-cross Abstract: Context. Behaviour-Driven Development (BDD) suites in Gherkin accumulate step-text duplication with documented maintenance cost. Prior detectors either require runnable tests or are single-organisation, leaving a gap: a static, paraphrase-rob

arxivApr 27

Bridging the Long-Tail Gap: Robust Retrieval-Augmented Relation Completion via Multi-Stage Paraphrase Infusion

arXiv:2604.22261v1 Announce Type: new Abstract: Large language models (LLMs) struggle with relation completion (RC), both with and without retrieval-augmented generation (RAG), particularly when the required information is rare or sparsely represented. To address this, we propose a novel multi-stage

arxivApr 23

Finding Duplicates in 1.1M BDD Steps: cukereuse, a Paraphrase-Robust Static Detector for Cucumber and Gherkin

arXiv:2604.20462v1 Announce Type: cross Abstract: Behaviour-Driven Development (BDD) suites accumulate step-text duplication whose maintenance cost is established in prior work. Existing detection techniques require running the tests (Binamungu et al., 2018-2023) or are confined to a single organisa

arxivApr 14

PSF-Med: Measuring and Explaining Paraphrase Sensitivity in Medical Vision Language Models

arXiv:2602.21428v2 Announce Type: replace-cross Abstract: Medical Vision Language Models (VLMs) can change their answers when clinicians rephrase the same question, a failure mode that threatens deployment safety. We introduce PSF-Med, a benchmark of 26,850 chest X-ray questions paired with 92,856 m

arxivApr 13

Predictive Entropy Links Calibration and Paraphrase Sensitivity in Medical Vision-Language Models

arXiv:2604.08941v1 Announce Type: new Abstract: Medical Vision Language Models VLMs suffer from two failure modes that threaten safe deployment mis calibrated confidence and sensitivity to question rephrasing. We show they share a common cause, proximity to the decision boundary, by benchmarking fiv

arxivMar 31

LIBERO-Para: A Diagnostic Benchmark and Metrics for Paraphrase Robustness in VLA Models

arXiv:2603.28301v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models achieve strong performance in robotic manipulation by leveraging pre-trained vision-language backbones. However, in downstream robotic settings, they are typically fine-tuned with limited data, leading to overfitting

HomeModelsNews