·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
Why AMI Labs’ Alexandre LeBrun won’t call his AI ‘AGI’ or ‘superintelligence’28m◆Moonshot’s upcoming Kimi 3 is expected to close the gap with Anthropic’s Opus 4.842m◆Apple Intelligence approved for launch in China with Alibaba and Baidu1h◆Claude can now use your 1Password credentials for you2h◆Google ordered to open Android and Search to rivals in Europe3h◆Newer Models, Same Advantage3h◆Computer cops4h◆Multi-Agent Collaborative Reasoning with Tool-Augmented Evidence for Urban Region Profiling11h◆SemaDiff: Identifying Semantic-Changing Commits with Generated Code and Tests11h◆MASPRM: Multi-Agent System Process Reward Model11h◆Delving into the Temporal Challenges of Unified Video Protection Against Image-to-Video and Fine-Tuning-based Customization11h◆A plug-and-play approach with fast uncertainty quantification for weak lensing mass mapping11h◆FastCentNN: Accelerating Centroid Neural Network with Entropy Proxy11h◆Cortical-SSM: A Deep State Space Model for Motor Imagery Decoding from EEG Signals11h◆Avoiding unsafe sets when training with Langevin Dynamics11h◆Set-shifting Behavioral Test for Harnessed Agents11h◆AIMO Interpretability Challenge11h◆The Perplexity Trap: When Patent Law Makes Human Writing Look Like AI11h◆The Refusal Residue: When Probes Catch Alignment Faking and When They Don't11h◆ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL11h◆Why AMI Labs’ Alexandre LeBrun won’t call his AI ‘AGI’ or ‘superintelligence’28m◆Moonshot’s upcoming Kimi 3 is expected to close the gap with Anthropic’s Opus 4.842m◆Apple Intelligence approved for launch in China with Alibaba and Baidu1h◆Claude can now use your 1Password credentials for you2h◆Google ordered to open Android and Search to rivals in Europe3h◆Newer Models, Same Advantage3h◆Computer cops4h◆Multi-Agent Collaborative Reasoning with Tool-Augmented Evidence for Urban Region Profiling11h◆SemaDiff: Identifying Semantic-Changing Commits with Generated Code and Tests11h◆MASPRM: Multi-Agent System Process Reward Model11h◆Delving into the Temporal Challenges of Unified Video Protection Against Image-to-Video and Fine-Tuning-based Customization11h◆A plug-and-play approach with fast uncertainty quantification for weak lensing mass mapping11h◆FastCentNN: Accelerating Centroid Neural Network with Entropy Proxy11h◆Cortical-SSM: A Deep State Space Model for Motor Imagery Decoding from EEG Signals11h◆Avoiding unsafe sets when training with Langevin Dynamics11h◆Set-shifting Behavioral Test for Harnessed Agents11h◆AIMO Interpretability Challenge11h◆The Perplexity Trap: When Patent Law Makes Human Writing Look Like AI11h◆The Refusal Residue: When Probes Catch Alignment Faking and When They Don't11h◆ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL11h◆
News/The Last Fingerprint: How Markdown Training Shapes LLM Prose
arxiv
PublishedApril 1, 2026 at 4:00 AM
—neutral

The Last Fingerprint: How Markdown Training Shapes LLM Prose

Source
arxiv.orgfull article ↗
Read on arxiv→
Publisher summary· verbatim

arXiv:2603.27006v1 Announce Type: cross Abstract: Large language models produce em dashes at varying rates, and the observation that some models "overuse" them has become one of the most widely discussed markers of AI-generated text. Yet no mechanistic account of this pattern exists, and the paralle

Models mentioned
02
  • 01meta-llama logo
    Llama
    meta-llama/Llama
  • 02openai logo
    gpt-4
    openai/gpt-4
    0.0%IN $30.00/Mtok
Compare these 2 models→
Related
05
  • arxiv6d
    A Vision Toward Energy-Efficient Domain-Specific Artificial Intelligence Models and Agents
  • arxiv14d
    WorkBench Revisited: Workplace Agents Two Years On
  • arxiv21d
    Representation Interventions Enable Lifelong Knowledge Memory Control in LLMs
  • arxivMay 19
    EmoMind: Decoding Affective Captions from Human Brain fMRI
  • arxivMay 11
    End-to-end PDDL Planning with Hardcoded and Dynamic Agents
Stay posted· Newsletter

A 5-min weekly brief — top movers, price watch, story of the week.

// no spam · unsubscribe one-click · free forever

Discussion
Mentioned models
02
  • 01
    Llama
    meta-llama/Llama
  • 02
    gpt-4
    openai/gpt-4
Source
↗
arxiv
Read original ↗All from arxiv →
Tags
04
#language models#training data#fine-tuning#markdown
Mentioned companies
05
AnthropicOpenAIMetaGoogleDeepSeek

No replies yet. Be first.

Mentioned models
02
  • 01
    Llama
    meta-llama/Llama
  • 02
    gpt-4
    openai/gpt-4
Source
↗
arxiv
Read original ↗All from arxiv →
Tags
04
#language models#training data#fine-tuning#markdown
Mentioned companies
05
AnthropicOpenAIMetaGoogleDeepSeek

Related coverage

More from ARXIV
arxivMulti-Agent Collaborative Reasoning with Tool-Augmented Evidence for Urban Region Profiling11harxivSemaDiff: Identifying Semantic-Changing Commits with Generated Code and Tests11harxivMASPRM: Multi-Agent System Process Reward Model11harxivDelving into the Temporal Challenges of Unified Video Protection Against Image-to-Video and Fine-Tuning-based Customization11h
The Bubble Brief
WEEKLY

Read language models insights every Tuesday — top movers, new releases, story of the week.

// no spam · unsubscribe one-click · free forever

Originally published on arxiv ↗
HomeModelsNews