·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
Cursor makes its biggest India push yet ahead of SpaceX acquisition with localized pricing4h◆Photonic reservoir computing with complex networks5h◆XS-VLA: Coupling Coarse-grained Spatial Distillation with Latent Flow Matching for Lightweight Robotic Control5h◆Agentic Permissions Policy Algebra for Taint Confinement in LLM Agents5h◆Beyond Squared Error: Exploring Loss Design for Enhanced Training of Generative Flow Networks5h◆The One-Word Census: Answer-Choice Conformity Across 44 Language Models5h◆Creative Integration: A Decidable Criterion of Creativity5h◆BERT-based Models vs. Large Language Models for Low-Resource Named Entity Recognition: A Comparative Study on Marathi5h◆Joint Optimization for Greedy Longest-match Tokenization5h◆Kimi K3: Open Frontier Intelligence5h◆The Few-shot Dilemma: Over-prompting Large Language Models5h◆Speculative Pipeline Decoding: Higher-Accuracy Drafting with Hidden Latency via Pipeline Parallelism5h◆Bayesian Complete-Pooling in Cross-Subject Classification for Motor Imagery Electroencephalogram5h◆StageGuard: Physiologically Constrained Sleep Staging5h◆Soft-Constrained Optimization of Latent Space in Variational Autoencoders5h◆Beyond Error-vs-Discard Characteristic: Toward Stable and Reliable Evaluation for Face Image Quality Assessment5h◆Analyzing the Importance of Blank for CTC-Based Knowledge Distillation5h◆Predicting Channel Closures in the Lightning Network with Machine Learning5h◆Evaluation of Blood Vessel Segmentation Methods on Hard-to-Detect Vascular Structures5h◆MOCA: A Transformer-based Modular Causal Inference Framework with One-way Cross-attention and Cutting Feedback5h◆Cursor makes its biggest India push yet ahead of SpaceX acquisition with localized pricing4h◆Photonic reservoir computing with complex networks5h◆XS-VLA: Coupling Coarse-grained Spatial Distillation with Latent Flow Matching for Lightweight Robotic Control5h◆Agentic Permissions Policy Algebra for Taint Confinement in LLM Agents5h◆Beyond Squared Error: Exploring Loss Design for Enhanced Training of Generative Flow Networks5h◆The One-Word Census: Answer-Choice Conformity Across 44 Language Models5h◆Creative Integration: A Decidable Criterion of Creativity5h◆BERT-based Models vs. Large Language Models for Low-Resource Named Entity Recognition: A Comparative Study on Marathi5h◆Joint Optimization for Greedy Longest-match Tokenization5h◆Kimi K3: Open Frontier Intelligence5h◆The Few-shot Dilemma: Over-prompting Large Language Models5h◆Speculative Pipeline Decoding: Higher-Accuracy Drafting with Hidden Latency via Pipeline Parallelism5h◆Bayesian Complete-Pooling in Cross-Subject Classification for Motor Imagery Electroencephalogram5h◆StageGuard: Physiologically Constrained Sleep Staging5h◆Soft-Constrained Optimization of Latent Space in Variational Autoencoders5h◆Beyond Error-vs-Discard Characteristic: Toward Stable and Reliable Evaluation for Face Image Quality Assessment5h◆Analyzing the Importance of Blank for CTC-Based Knowledge Distillation5h◆Predicting Channel Closures in the Lightning Network with Machine Learning5h◆Evaluation of Blood Vessel Segmentation Methods on Hard-to-Detect Vascular Structures5h◆MOCA: A Transformer-based Modular Causal Inference Framework with One-way Cross-attention and Cutting Feedback5h◆
News/It Takes One to Bias Them All: Breaking Bad with One-Shot GRPO
arxiv
PublishedJune 12, 2026 at 4:00 AM
▼bearish

It Takes One to Bias Them All: Breaking Bad with One-Shot GRPO

Source
arxiv.orgfull article ↗
Read on arxiv→
Publisher summary· verbatim

arXiv:2606.10931v2 Announce Type: replace Abstract: Warning: This paper contains several toxic and offensive statements. Modern large language models (LLMs) are typically aligned through large-scale post-training to ensure fair and reliable behavior. In this work, we investigate how easily such guar

Stay posted· Newsletter

A 5-min weekly brief — top movers, price watch, story of the week.

// no spam · unsubscribe one-click · free forever

Discussion
Source
↗
arxiv
Read original ↗All from arxiv →
Tags
04
#bias#safety#language-models#vulnerability

No replies yet. Be first.

Source
↗
arxiv
Read original ↗All from arxiv →
Tags
04
#bias#safety#language-models#vulnerability

Related coverage

More from ARXIV
arxivReverso: Efficient Time Series Foundation Models for Zero-shot Forecasting5harxivCoherent Without Grounding, Grounded Without Success: Observability and Epistemic Failure5harxivAutoWorld: Learning Multi-Agent Traffic Simulation with Self-Supervised World Models5harxivSheaf-Laplacian Obstruction and Projection Hardness for Cross-Modal Compatibility on a Modality-Independent Site5h
The Bubble Brief
WEEKLY

Read bias insights every Tuesday — top movers, new releases, story of the week.

// no spam · unsubscribe one-click · free forever

Originally published on arxiv ↗
HomeModelsNews