·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
A Consensus-Based Framework for Relative Preference Evaluation of Large Language Models1m◆Humanly: A Configurable and Traceable Environment for Human-AI Collaborative Writing1m◆Probing Latent Colombian Identity Inferences in Qwen2.5-7B with Natural Language Autoencoders1m◆Khondo: A Multimodal Benchmark for Document Packet Splitting of Bangla Forms1m◆Agentic Evaluation of Copyright Law Compliance1m◆Data Quality over Capacity: Internalizing Documents into LoRA Adapters for Closed-Book QA1m◆Leveraging External Knowledge for Historical Document Restoration via Retrieval-Augmented Large Language Models1m◆On Improving Faithfulness of Podcasts from Documents1m◆Ground Truth First: A Longitudinal Evaluation Instrument for Agent Memory, and the Tenure Crossover in Memory-Architecture Rankings1m◆MoE$^2$-LoRA: When MoE Models Meet MoE-style Low-Rank Adaptation1m◆Analyzing Toxic Behavior and Its Impact on the Mastodon Community1m◆J-CoT: Chain-of-Thought in J-Space1m◆Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study1m◆DWT-Fusion: A Signal-Based Framework for Training-Free LLM-Generated Text Detection1m◆Enough is as good as a feast: A Comprehensive Analysis of How Reinforcement Learning Mitigates Task Conflicts in LLMs1m◆Developing and Validating the Spanish Version of the Large Language Models Dependency Scale (LLM-D12-SP)1m◆Scaling Native Multimodal Pre-Training From Scratch1m◆Benchmarking Fine-tuning and Retrieval Strategies for a Multimodal Language Model on the NRC Reactor Operator Licensing Examination1m◆FSE: Continual Learning for Named Entity Recognition by Fast-Slow Experts1m◆MEUSLI: a Multilingual Projector for LLM-based ASR and Beyond1m◆A Consensus-Based Framework for Relative Preference Evaluation of Large Language Models1m◆Humanly: A Configurable and Traceable Environment for Human-AI Collaborative Writing1m◆Probing Latent Colombian Identity Inferences in Qwen2.5-7B with Natural Language Autoencoders1m◆Khondo: A Multimodal Benchmark for Document Packet Splitting of Bangla Forms1m◆Agentic Evaluation of Copyright Law Compliance1m◆Data Quality over Capacity: Internalizing Documents into LoRA Adapters for Closed-Book QA1m◆Leveraging External Knowledge for Historical Document Restoration via Retrieval-Augmented Large Language Models1m◆On Improving Faithfulness of Podcasts from Documents1m◆Ground Truth First: A Longitudinal Evaluation Instrument for Agent Memory, and the Tenure Crossover in Memory-Architecture Rankings1m◆MoE$^2$-LoRA: When MoE Models Meet MoE-style Low-Rank Adaptation1m◆Analyzing Toxic Behavior and Its Impact on the Mastodon Community1m◆J-CoT: Chain-of-Thought in J-Space1m◆Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study1m◆DWT-Fusion: A Signal-Based Framework for Training-Free LLM-Generated Text Detection1m◆Enough is as good as a feast: A Comprehensive Analysis of How Reinforcement Learning Mitigates Task Conflicts in LLMs1m◆Developing and Validating the Spanish Version of the Large Language Models Dependency Scale (LLM-D12-SP)1m◆Scaling Native Multimodal Pre-Training From Scratch1m◆Benchmarking Fine-tuning and Retrieval Strategies for a Multimodal Language Model on the NRC Reactor Operator Licensing Examination1m◆FSE: Continual Learning for Named Entity Recognition by Fast-Slow Experts1m◆MEUSLI: a Multilingual Projector for LLM-based ASR and Beyond1m◆
News/Two Wrongs, No Right: Auditing Social-Desirability Bias in LLM Annotators for Computational Social Science
arxiv
PublishedJune 12, 2026 at 4:00 AM
▼bearish

Two Wrongs, No Right: Auditing Social-Desirability Bias in LLM Annotators for Computational Social Science

Source
arxiv.orgfull article ↗
Read on arxiv→
Publisher summary· verbatim

arXiv:2606.12426v1 Announce Type: cross Abstract: LLM annotators are increasingly used in computational social science (CSS), but it is unclear whether their alignment-shaped errors preserve the empirical conclusions a researcher would report. We audit three open-source 7B instruction-tuned models (

Stay posted· Newsletter

A 5-min weekly brief — top movers, price watch, story of the week.

// no spam · unsubscribe one-click · free forever

Discussion
Mentioned models
03
  • 01
    Zephyr
  • 02
    Mistral
  • 03
    Qwen
Source
↗
arxiv
Read original ↗All from arxiv →
Tags
04
#bias#computational social science#language models#validation

No replies yet. Be first.

Mentioned models
03
  • 01
    Zephyr
  • 02
    Mistral
  • 03
    Qwen
Source
↗
arxiv
Read original ↗All from arxiv →
Tags
04
#bias#computational social science#language models#validation

Related coverage

More from ARXIV
arxivA Consensus-Based Framework for Relative Preference Evaluation of Large Language Models1marxivHumanly: A Configurable and Traceable Environment for Human-AI Collaborative Writing1marxivProbing Latent Colombian Identity Inferences in Qwen2.5-7B with Natural Language Autoencoders1marxivKhondo: A Multimodal Benchmark for Document Packet Splitting of Bangla Forms1m
The Bubble Brief
WEEKLY

Read bias insights every Tuesday — top movers, new releases, story of the week.

// no spam · unsubscribe one-click · free forever

Originally published on arxiv ↗
HomeModelsNews