·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
BenchMIRT: What are LLM benchmarks actually measuring?55m◆Open AI’s Astra model is on the way—and very good at breaking into computer systems1h◆Google’s Android update tackles motion sickness, accessibility, and more1h◆OpenAI delayed its new model’s development after the Hugging Face hack1h◆The latest AI news we announced in August 20261h◆Anthropic’s new Fable release is cheaper, less restrictive2h◆The rise of AI ‘civilizations’ and the fall of corporate responsibility3h◆Apple accuses OpenAI of destroying evidence4h◆Google’s answer to Canva is an AI tool where you prompt instead of design4h◆ChatGPT Health adds Epic integration for clinicians to import patient data5h◆How AI-native companies turn workflows into operating capability5h◆Sequoia-incubated Empirik launches with $21M to predict outages before they happen6h◆John Deere launched an AI chatbot for farmers6h◆Try Google Pics: Easy image creation and editing in Google Workspace6h◆Google Pics is like Canva, but with even more AI6h◆Amazon Alexa can now alert you when something new might tempt you to shop6h◆AIR raises $50M to help companies vet the skills and add-ons AI agents use6h◆Fambot introduces an ‘AI chief of staff’ for families7h◆Nvidia’s controversial DLSS 5 arrives September 3rd and requires serious GPU horsepower9h◆Path to Astra: critical capabilities and frontier safeguards9h◆BenchMIRT: What are LLM benchmarks actually measuring?55m◆Open AI’s Astra model is on the way—and very good at breaking into computer systems1h◆Google’s Android update tackles motion sickness, accessibility, and more1h◆OpenAI delayed its new model’s development after the Hugging Face hack1h◆The latest AI news we announced in August 20261h◆Anthropic’s new Fable release is cheaper, less restrictive2h◆The rise of AI ‘civilizations’ and the fall of corporate responsibility3h◆Apple accuses OpenAI of destroying evidence4h◆Google’s answer to Canva is an AI tool where you prompt instead of design4h◆ChatGPT Health adds Epic integration for clinicians to import patient data5h◆How AI-native companies turn workflows into operating capability5h◆Sequoia-incubated Empirik launches with $21M to predict outages before they happen6h◆John Deere launched an AI chatbot for farmers6h◆Try Google Pics: Easy image creation and editing in Google Workspace6h◆Google Pics is like Canva, but with even more AI6h◆Amazon Alexa can now alert you when something new might tempt you to shop6h◆AIR raises $50M to help companies vet the skills and add-ons AI agents use6h◆Fambot introduces an ‘AI chief of staff’ for families7h◆Nvidia’s controversial DLSS 5 arrives September 3rd and requires serious GPU horsepower9h◆Path to Astra: critical capabilities and frontier safeguards9h◆
News/The Last Fingerprint: How Markdown Training Shapes LLM Prose
arxiv
PublishedApril 1, 2026 at 4:00 AM
—neutral

The Last Fingerprint: How Markdown Training Shapes LLM Prose

Source
arxiv.orgfull article ↗
Read on arxiv→
Publisher summary· verbatim

arXiv:2603.27006v1 Announce Type: cross Abstract: Large language models produce em dashes at varying rates, and the observation that some models "overuse" them has become one of the most widely discussed markers of AI-generated text. Yet no mechanistic account of this pattern exists, and the paralle

Models mentioned
02
  • 01meta-llama logo
    Llama
    meta-llama/Llama
  • 02openai logo
    gpt-4
    openai/gpt-4
    0.0%IN $30.00/Mtok
Compare these 2 models→
Related
05
  • arxivJul 10
    A Vision Toward Energy-Efficient Domain-Specific Artificial Intelligence Models and Agents
  • arxivJul 2
    WorkBench Revisited: Workplace Agents Two Years On
  • arxivJun 25
    Representation Interventions Enable Lifelong Knowledge Memory Control in LLMs
  • arxivMay 19
    EmoMind: Decoding Affective Captions from Human Brain fMRI
  • arxivMay 11
    End-to-end PDDL Planning with Hardcoded and Dynamic Agents
Stay posted· Newsletter

A 5-min weekly brief — top movers, price watch, story of the week.

// no spam · unsubscribe one-click · free forever

Discussion
Mentioned models
02
  • 01
    Llama
    meta-llama/Llama
  • 02
    gpt-4
    openai/gpt-4
Source
↗
arxiv
Read original ↗All from arxiv →
Tags
04
#language models#training data#fine-tuning#markdown
Mentioned companies
05
AnthropicOpenAIMetaGoogleDeepSeek

No replies yet. Be first.

Mentioned models
02
  • 01
    Llama
    meta-llama/Llama
  • 02
    gpt-4
    openai/gpt-4
Source
↗
arxiv
Read original ↗All from arxiv →
Tags
04
#language models#training data#fine-tuning#markdown
Mentioned companies
05
AnthropicOpenAIMetaGoogleDeepSeek
The Bubble Brief
WEEKLY

Read language models insights every Tuesday — top movers, new releases, story of the week.

// no spam · unsubscribe one-click · free forever

Originally published on arxiv ↗
HomeModelsNews