·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
A Consensus-Based Framework for Relative Preference Evaluation of Large Language Models26m◆Humanly: A Configurable and Traceable Environment for Human-AI Collaborative Writing26m◆Probing Latent Colombian Identity Inferences in Qwen2.5-7B with Natural Language Autoencoders26m◆Khondo: A Multimodal Benchmark for Document Packet Splitting of Bangla Forms26m◆Data Quality over Capacity: Internalizing Documents into LoRA Adapters for Closed-Book QA26m◆Leveraging External Knowledge for Historical Document Restoration via Retrieval-Augmented Large Language Models26m◆On Improving Faithfulness of Podcasts from Documents26m◆Ground Truth First: A Longitudinal Evaluation Instrument for Agent Memory, and the Tenure Crossover in Memory-Architecture Rankings26m◆MoE$^2$-LoRA: When MoE Models Meet MoE-style Low-Rank Adaptation26m◆Analyzing Toxic Behavior and Its Impact on the Mastodon Community26m◆J-CoT: Chain-of-Thought in J-Space26m◆Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study26m◆DWT-Fusion: A Signal-Based Framework for Training-Free LLM-Generated Text Detection26m◆Enough is as good as a feast: A Comprehensive Analysis of How Reinforcement Learning Mitigates Task Conflicts in LLMs26m◆Developing and Validating the Spanish Version of the Large Language Models Dependency Scale (LLM-D12-SP)26m◆Scaling Native Multimodal Pre-Training From Scratch26m◆Benchmarking Fine-tuning and Retrieval Strategies for a Multimodal Language Model on the NRC Reactor Operator Licensing Examination26m◆FSE: Continual Learning for Named Entity Recognition by Fast-Slow Experts26m◆MEUSLI: a Multilingual Projector for LLM-based ASR and Beyond26m◆Dynamic Commonsense Coordination for Empathetic Response Generation26m◆A Consensus-Based Framework for Relative Preference Evaluation of Large Language Models26m◆Humanly: A Configurable and Traceable Environment for Human-AI Collaborative Writing26m◆Probing Latent Colombian Identity Inferences in Qwen2.5-7B with Natural Language Autoencoders26m◆Khondo: A Multimodal Benchmark for Document Packet Splitting of Bangla Forms26m◆Data Quality over Capacity: Internalizing Documents into LoRA Adapters for Closed-Book QA26m◆Leveraging External Knowledge for Historical Document Restoration via Retrieval-Augmented Large Language Models26m◆On Improving Faithfulness of Podcasts from Documents26m◆Ground Truth First: A Longitudinal Evaluation Instrument for Agent Memory, and the Tenure Crossover in Memory-Architecture Rankings26m◆MoE$^2$-LoRA: When MoE Models Meet MoE-style Low-Rank Adaptation26m◆Analyzing Toxic Behavior and Its Impact on the Mastodon Community26m◆J-CoT: Chain-of-Thought in J-Space26m◆Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study26m◆DWT-Fusion: A Signal-Based Framework for Training-Free LLM-Generated Text Detection26m◆Enough is as good as a feast: A Comprehensive Analysis of How Reinforcement Learning Mitigates Task Conflicts in LLMs26m◆Developing and Validating the Spanish Version of the Large Language Models Dependency Scale (LLM-D12-SP)26m◆Scaling Native Multimodal Pre-Training From Scratch26m◆Benchmarking Fine-tuning and Retrieval Strategies for a Multimodal Language Model on the NRC Reactor Operator Licensing Examination26m◆FSE: Continual Learning for Named Entity Recognition by Fast-Slow Experts26m◆MEUSLI: a Multilingual Projector for LLM-based ASR and Beyond26m◆Dynamic Commonsense Coordination for Empathetic Response Generation26m◆
News/Low-Agreeableness Persona Conditioning for Safe LLM Fine-Tuning
arxiv
PublishedJune 29, 2026 at 4:00 AM
—neutral

Low-Agreeableness Persona Conditioning for Safe LLM Fine-Tuning

Source
arxiv.orgfull article ↗
Read on arxiv→
Publisher summary· verbatim

arXiv:2606.27709v1 Announce Type: cross Abstract: Recent work has shown that fine-tuning large language models (LLMs) for social warmth degrades factual reliability and increases sycophancy. We investigate a related but distinct failure mode: warmth fine-tuning also weakens adversarial safety, makin

Stay posted· Newsletter

A 5-min weekly brief — top movers, price watch, story of the week.

// no spam · unsubscribe one-click · free forever

Discussion
Source
↗
arxiv
Read original ↗All from arxiv →
Tags
03
#safety#fine-tuning#language-models

No replies yet. Be first.

Source
↗
arxiv
Read original ↗All from arxiv →
Tags
03
#safety#fine-tuning#language-models

Related coverage

More from ARXIV
arxivA Consensus-Based Framework for Relative Preference Evaluation of Large Language Models26marxivHumanly: A Configurable and Traceable Environment for Human-AI Collaborative Writing26marxivProbing Latent Colombian Identity Inferences in Qwen2.5-7B with Natural Language Autoencoders26marxivKhondo: A Multimodal Benchmark for Document Packet Splitting of Bangla Forms26m
The Bubble Brief
WEEKLY

Read safety insights every Tuesday — top movers, new releases, story of the week.

// no spam · unsubscribe one-click · free forever

Originally published on arxiv ↗
HomeModelsNews