·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
Cursor makes its biggest India push yet ahead of SpaceX acquisition with localized pricing4h◆Reverso: Efficient Time Series Foundation Models for Zero-shot Forecasting5h◆Multinex: Lightweight Low-light Image Enhancement via Multi-prior Retinex5h◆Market Design for AI: Beyond the Copyright Binary5h◆Who Pays the Price? Stakeholder-Centric Prompt Injection Benchmarking for Real-world Web Agents5h◆TextRich: A Multi-Domain Benchmark for Detecting AI-Generated Text-Rich Images from GPT-Image-25h◆DiscoLoop: Looping Discrete Embeddings and Continuous Hidden States for Multi-hop Reasoning5h◆From World Models to World Action Models: A Concise Tutorial for Robotics5h◆QuantFlow: A Federated Mamba-Based Post-Transformer Foundation Model for Time-Series Forecasting5h◆Multi-Turn On-Policy Distillation with Prefix Replay5h◆Operational Proto-Introspection in Looped Language Models: Process-Quality Taps, Executable Branching, and the Readout-Control Boundary5h◆Skillware: A Software Ontology and Engineering Lifecycle for Persistent Behavioral Artifacts5h◆MedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation Agents5h◆PhantomFill: When the Form Demands an Answer, Language Models Invent One5h◆Error Certificates for KV-Cache Eviction via Randomized Design5h◆Explaining GAND: A Resource on Gender-Ambiguous Natural Data & Contrastive Attribution5h◆MioFFAn: an Annotation Software for Formula Formalization with LLM Automation Capabilities5h◆LA-RL: Label-Aware Self-Reflection for Reinforcement Learning in Information Extraction5h◆Mwando: Leveraging AI to Preserve and Teach shiKomori5h◆The JEPA Paradox in Language: The Geometry of Linguistic Alternatives5h◆Cursor makes its biggest India push yet ahead of SpaceX acquisition with localized pricing4h◆Reverso: Efficient Time Series Foundation Models for Zero-shot Forecasting5h◆Multinex: Lightweight Low-light Image Enhancement via Multi-prior Retinex5h◆Market Design for AI: Beyond the Copyright Binary5h◆Who Pays the Price? Stakeholder-Centric Prompt Injection Benchmarking for Real-world Web Agents5h◆TextRich: A Multi-Domain Benchmark for Detecting AI-Generated Text-Rich Images from GPT-Image-25h◆DiscoLoop: Looping Discrete Embeddings and Continuous Hidden States for Multi-hop Reasoning5h◆From World Models to World Action Models: A Concise Tutorial for Robotics5h◆QuantFlow: A Federated Mamba-Based Post-Transformer Foundation Model for Time-Series Forecasting5h◆Multi-Turn On-Policy Distillation with Prefix Replay5h◆Operational Proto-Introspection in Looped Language Models: Process-Quality Taps, Executable Branching, and the Readout-Control Boundary5h◆Skillware: A Software Ontology and Engineering Lifecycle for Persistent Behavioral Artifacts5h◆MedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation Agents5h◆PhantomFill: When the Form Demands an Answer, Language Models Invent One5h◆Error Certificates for KV-Cache Eviction via Randomized Design5h◆Explaining GAND: A Resource on Gender-Ambiguous Natural Data & Contrastive Attribution5h◆MioFFAn: an Annotation Software for Formula Formalization with LLM Automation Capabilities5h◆LA-RL: Label-Aware Self-Reflection for Reinforcement Learning in Information Extraction5h◆Mwando: Leveraging AI to Preserve and Teach shiKomori5h◆The JEPA Paradox in Language: The Geometry of Linguistic Alternatives5h◆
News/Characterizing Narrative Content in Web-scale LLM Pretraining Data
arxiv
PublishedJune 19, 2026 at 4:00 AM
—neutral

Characterizing Narrative Content in Web-scale LLM Pretraining Data

Source
arxiv.orgfull article ↗
Read on arxiv→
Publisher summary· verbatim

arXiv:2606.19468v1 Announce Type: new Abstract: The narrative composition of web-scale LLM pretraining corpora remains largely unexplored even though narrative is a fundamental mode of human communication. We present the first fine-grained study of narrative features in Dolma, a 3-trillion-token ope

Models mentioned
01
  • 01facebook logo
    roberta-base
    facebook/roberta-base
Stay posted· Newsletter

A 5-min weekly brief — top movers, price watch, story of the week.

// no spam · unsubscribe one-click · free forever

Discussion
Mentioned models
03
  • 01
    NarraBERT
  • 02
    roberta-base
    facebook/roberta-base
  • 03
    Dolma
Source
↗
arxiv
Read original ↗All from arxiv →
Tags
04
#narrative-analysis#pretraining#llm#dataset
Mentioned companies
01
Facebook

No replies yet. Be first.

Mentioned models
03
  • 01
    NarraBERT
  • 02
    roberta-base
    facebook/roberta-base
  • 03
    Dolma
Source
↗
arxiv
Read original ↗All from arxiv →
Tags
04
#narrative-analysis#pretraining#llm#dataset
Mentioned companies
01
Facebook

Related coverage

More from ARXIV
arxivReverso: Efficient Time Series Foundation Models for Zero-shot Forecasting5harxivCoherent Without Grounding, Grounded Without Success: Observability and Epistemic Failure5harxivAutoWorld: Learning Multi-Agent Traffic Simulation with Self-Supervised World Models5harxivSheaf-Laplacian Obstruction and Projection Hardness for Cross-Modal Compatibility on a Modality-Independent Site5h
The Bubble Brief
WEEKLY

Read narrative-analysis insights every Tuesday — top movers, new releases, story of the week.

// no spam · unsubscribe one-click · free forever

Originally published on arxiv ↗
HomeModelsNews