·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
Data centers expected to use 4x more electricity by 20351h◆Google releases three new Gemini models — but no 3.5 Pro2h◆Introducing the ChatGPT for small business program2h◆Anthropic’s $1.5 billion book piracy settlement approved by judge2h◆US threatens sanctions against Chinese AI models over IP theft4h◆Google launches a cheaper alternative to large AI security models like Mythos4h◆Music streamer Deezer says more than 50% of daily uploads are AI-generated6h◆Halliday’s latest smart glasses feature a much-improved display6h◆America needs to stop getting shocked by Chinese AI8h◆Advancing next-gen AI with materials science innovation9h◆Gritt exits stealth with $32 million for robots to build solar plants — then, everything else9h◆Capacity and Redundancy Trade-offs in Multi-Task Learning15h◆Predictive Training with Latent Imagination for Visual Quadruped Navigation15h◆Where Not to Learn: Prior-Aligned Training with Subset-based Attribution Constraints for Reliable Decision-Making15h◆Did We Actually Fix It? An Independent Adversarial Stress-Test of Post-Point-Adjustment Evaluation Metrics for Time-Series Anomaly Detection15h◆Supervised Reward Inference15h◆PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization15h◆Is Progressive Disclosure All You Need for Long-Context Agents?15h◆It Depends on the Dataset: When a Brain-Encoding Model's Predicted Responses Beat Their Visual Backbone for Video Memorability15h◆DMFNet: Dual-Backbone Multiscale Fusion Network for Urban Scene Classification15h◆Data centers expected to use 4x more electricity by 20351h◆Google releases three new Gemini models — but no 3.5 Pro2h◆Introducing the ChatGPT for small business program2h◆Anthropic’s $1.5 billion book piracy settlement approved by judge2h◆US threatens sanctions against Chinese AI models over IP theft4h◆Google launches a cheaper alternative to large AI security models like Mythos4h◆Music streamer Deezer says more than 50% of daily uploads are AI-generated6h◆Halliday’s latest smart glasses feature a much-improved display6h◆America needs to stop getting shocked by Chinese AI8h◆Advancing next-gen AI with materials science innovation9h◆Gritt exits stealth with $32 million for robots to build solar plants — then, everything else9h◆Capacity and Redundancy Trade-offs in Multi-Task Learning15h◆Predictive Training with Latent Imagination for Visual Quadruped Navigation15h◆Where Not to Learn: Prior-Aligned Training with Subset-based Attribution Constraints for Reliable Decision-Making15h◆Did We Actually Fix It? An Independent Adversarial Stress-Test of Post-Point-Adjustment Evaluation Metrics for Time-Series Anomaly Detection15h◆Supervised Reward Inference15h◆PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization15h◆Is Progressive Disclosure All You Need for Long-Context Agents?15h◆It Depends on the Dataset: When a Brain-Encoding Model's Predicted Responses Beat Their Visual Backbone for Video Memorability15h◆DMFNet: Dual-Backbone Multiscale Fusion Network for Urban Scene Classification15h◆
News/model/grok-2

grok-2 news

42 articles mentioning grok-2

arxiv3d ago

To Grok Grokking: Provable Grokking in Ridge Regression

arXiv:2601.19791v4 Announce Type: replace Abstract: We study grokking, the onset of generalization long after overfitting, in a classical ridge regression setting. We prove end-to-end grokking results for learning over-parameterized linear regression models using gradient descent with weight decay.

arxiv4d ago

Grokipedia vs Wikipedia: An LLM-Based Audit of Political Neutrality along Ideologies

arXiv:2607.15146v1 Announce Type: new Abstract: Online encyclopedias shape political opinion and, through it, democratic discourse. In late 2025, Grokipedia was released, an encyclopedia written entirely by the LLM Grok. One motivation behind the project was to provide an unbiased alternative to Wik

arxiv5d ago

Algebraic Representability as the Limiting Regime of Grokking: An Exactly Solvable Model with Holomorphic Activations

arXiv:2607.13749v1 Announce Type: new Abstract: Neural networks trained on modular arithmetic exhibit grokking, a delayed transition from memorisation to generalisation known to depend on model capacity: too little and the network memorises slowly or not at all, too much and it generalises almost im

theverge5d ago

xAI sues a man for using Grok to generate CSAM ‘deepfakes’

The Elon Musk-owned xAI is suing a South Carolina man who allegedly used the company's Grok AI chatbot to generate child sexual abuse material (CSAM). In a lawsuit reported earlier by Reuters, xAI claims Terry Wayne Harwood "knowingly and intentionally used Grok to circumvent safeguards, alter nonco

arxiv6d ago

What Makes a Representational Prior Work? Feature Families, Label-Free Invariances, and Critical Windows in Grokking

arXiv:2607.12735v1 Announce Type: new Abstract: Companion work showed the grokking delay is causally the time to form task-structured representations, injectable via a contrastive prior. Here we characterize what makes such a prior work, across four axes, in 188 new runs. Content: a coherent, learna

arxiv6d ago

Structure-Specific Representational Priors Causally Control the Grokking Delay

arXiv:2607.04333v3 Announce Type: replace Abstract: Grokking -- generalization long after training-set interpolation -- has been accelerated by structure-agnostic interventions (gradient filtering, weight-norm clamping, geometric penalties). Whether the delay specifically measures the time to form t

thevergeJul 14

SpaceXAI’s Grok programming tool was uploading its users’ entire codebase to cloud storage

SpaceXAI's Grok Build AI coding tool was spotted uploading users' entire codebases to Google Cloud before it was reported, and the company turned it off. The Register reports that Cereblab published findings on Monday showing how the Grok Build CLI was packaging and uploading entire code repositorie

arxivJul 14

How to Tame Grokking: Representation Geometry as a Control Signal

arXiv:2607.11666v1 Announce Type: new Abstract: Grokking is a phenomenon in which neural networks initially memorize training data and only later exhibit strong generalization after prolonged optimization. Despite extensive recent study, the factors influencing the emergence and timing of grokking r

arxivJul 10

A Stochastic--Geometric Theory of Scaling Laws in Grokking

arXiv:2606.30388v3 Announce Type: replace-cross Abstract: Delayed generalization (\ie~grokking) refers to the phenomenon in which a neural network fits its training data early in training but only begins to generalize after a prolonged delay, often through an abrupt transition. Despite extensive emp

arxivJun 25

Repeated Shared Access Enables Grokking, but Edit Propagation Depends on an Addressable Memory

arXiv:2606.20737v2 Announce Type: replace Abstract: We study factual edit propagation in a controlled synthetic knowledge-graph QA setting using a 2x2 grid that crosses loop recurrence with shared-memory access: a dense transformer (Dense), a looped transformer (Loop), a dense backbone with shared m

arxivJun 25

Natural Ungrokking: Asymmetric Control of Which Rules Survive Pretraining

arXiv:2606.26050v1 Announce Type: cross Abstract: Midway through an ordinary pretraining run, a small language model learns the pronoun-gender rule: cued with a girl's name ("Sue cried because"), it resolves the next pronoun to she, generalizing to held-out probes (0.94 by step 925). By step 3,500 t

arxivJun 18

What Does the Weight Norm Control in Grokking? Logit-Scale Mediation under Cross-Entropy

arXiv:2606.18465v1 Announce Type: cross Abstract: Grokking, the delayed jump from memorization to generalization, is usually tied to the weight norm: a smaller norm generalizes sooner. We ask what the norm actually controls. Holding the weight norm fixed by clamping and varying only an output temper

arxivJun 17

Noise-Driven Escape from Metastable Phases explains Grokking in Deep Neural Networks

arXiv:2606.17120v1 Announce Type: new Abstract: Deep neural networks (DNNs) exhibit first order phase transitions under variations of the L2 regularization strength, with each transition marking the onset of a new learnable feature. Below a critical regularization strength, all features are in princ

arxivJun 15

The Weight Norm Sets the Grokking Timescale: A Causal Delay Law

arXiv:2606.13753v1 Announce Type: cross Abstract: Grokking is the delayed onset of generalization in neural networks, arising long after they fit the training data. Whether the weight norm causes this delay is disputed: some studies report a critical norm at the transition, others observe grokking w

techcrunchJun 10

xAI fired an engineer who raised alarms about Grok safety, new lawsuit claims

A former xAI engineer is suing the company and SpaceX, alleging he was fired for raising AI safety concerns about Grok days before SpaceX's historic IPO.

arxivJun 10

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition

arXiv:2606.09883v1 Announce Type: cross Abstract: Large language models (LLMs) have made remarkable progress in reasoning tasks, largely driven by post-training paradigms, especially reinforcement learning with verifiable rewards (RLVR). However, a critical bottleneck persists: RLVR fails on highly

arxivJun 6

Deciphering Two Training Clocks in Grokking via Deep Linear Network Theory with Conditional ReLU Reduction

arXiv:2606.05863v1 Announce Type: cross Abstract: Grokking suggests that fitting the training data and learning a simple underlying rule may occur on different time scales. We formalize this phenomenon by separating the fast decay of the classification loss from the slower simplification of the lear

arxivJun 5

Low-Rank Decay for Grokking in Scale-Invariant Transformers: A Spectral-Geometric View

arXiv:2606.04405v1 Announce Type: cross Abstract: Modern Transformer architectures frequently employ normalization mechanisms such as RMSNorm and Query-Key Normalization, making parts of the model approximately scale-invariant with respect to weight magnitudes. In this regime, standard Frobenius-nor

arxivJun 2

Grokers: Bottom-Up Inductive Comprehension and Write-Time Intelligence over Typed Knowledge Graphs

arXiv:2606.00050v1 Announce Type: new Abstract: We present Grokers, an architecture for building persistent, structured comprehension of typed knowledge graphs through bottom-up inductive traversal of dependency subgraphs. Unlike retrieval-augmented generation (RAG), which pays full comprehension co

arxivJun 2

The Geometry of Grokking: Norm Minimization on the Zero-Loss Manifold

arXiv:2511.01938v3 Announce Type: replace-cross Abstract: Grokking is a puzzling phenomenon in neural networks where full generalization occurs only after a substantial delay following the complete memorization of the training data. Previous research has linked this delayed generalization to represe

arxivJun 2

A Pre-Training Analogue of Grokking in Language Models: Tracing Delayed Grammatical Generalization

arXiv:2606.00230v1 Announce Type: new Abstract: Grokking, the phenomenon in which neural networks generalize long after fitting their training data, has been studied in supervised settings on many epochs. LLM pre-training instead involves next-token prediction over an unlabeled corpus, with limited

arxivMay 29

Two Speeds of Learning: A Representation-Readout Decomposition of Grokking and Double Descent

arXiv:2605.27078v2 Announce Type: replace-cross Abstract: Training loss and accuracy are the standard signals used to monitor generalization during deep neural network training. Two well-documented phenomena complicate this picture: in grokking, train loss falls rapidly while test performance improv

arxivMay 27

Grokking or Glitching? How Low-Precision Drives Slingshot Loss Spikes

arXiv:2605.06152v3 Announce Type: replace-cross Abstract: Deep neural networks exhibit periodic loss spikes during unregularized long-term training, a phenomenon known as the "Slingshot Mechanism." Existing work usually attributes this to intrinsic optimization dynamics, but its triggering mechanism

thevergeMay 22

Elon, stop trying to make Grok happen

There is a harsh truth about Elon Musk's "truth-seeking" AI chatbot Grok: It's not very good, and not many people are using it. That's the takeaway of a new Reuters report, which found that Grok barely appears in federal records of how the US government used AI last year. It's not the only sign xAI'

arxivMay 22

Weight Decay Regimes in Grokking Transformers: Cheap Online Diagnostics

arXiv:2605.20441v1 Announce Type: cross Abstract: Transformers trained on modular arithmetic exhibit sharp transitions between memorization, generalization, and collapse. We show that weight decay acts as a scalar empirical control parameter for these regimes, and introduce two cheap online diagnost

arxivMay 21

First-Passage Prediction of Grokking Delay: ACalibrated Law under AdamW with Causal Validation

arXiv:2605.18845v1 Announce Type: cross Abstract: We give the first quantitative prediction of grokking delay under AdamW. Treating the delay as a first-passage time, we derive a closed-form law T_grok - T_mem = (1 / 2 kappa_LL eta lambda) log(V_mem / V_star), where V_t = ||theta_t||^2 is the square

arxivMay 19

Egalitarian Gradient Descent: A Simple Approach to Accelerated Grokking

arXiv:2510.04930v3 Announce Type: replace Abstract: Grokking is the phenomenon whereby, unlike the training performance, which peaks early in the training process, the test/generalization performance of a model stagnates over arbitrarily many epochs and then suddenly jumps to usually close to perfec

arxivMay 18

Grokking as Structural Inference: Transformers Need Bayesian Lottery Tickets

arXiv:2605.15787v1 Announce Type: cross Abstract: Why does a Transformer that has memorized its training set wait thousands of steps before it generalizes? Existing accounts locate this delay in norm minimization, feature emergence, or the late discovery of sparse subnetworks. These explanations cap

arxivMay 16

Grokking Finite-Dimensional Algebra

arXiv:2602.19533v2 Announce Type: replace-cross Abstract: This paper investigates the grokking phenomenon, which refers to the sudden transition from a long memorization to generalization observed during neural networks training, in the context of learning multiplication in finite-dimensional algebr

arxivMay 16

Detecting overfitting in Neural Networks during long-horizon grokking using Random Matrix Theory

arXiv:2605.12394v2 Announce Type: replace-cross Abstract: Training Neural Networks (NNs) without overfitting is difficult; detecting that overfitting is difficult as well. We present a novel Random Matrix Theory method that detects the onset of overfitting in deep learning models without access to t

arxivMay 13

Feature Repulsion and Spectral Lock-in: An Empirical Study of Two-Layer Network Grokking

arXiv:2605.08119v1 Announce Type: cross Abstract: Tian (2025) proves a repulsion theorem (Theorem 6) for the matrix $ B = (\widetilde{F}^\top \widetilde{F} + \eta I)^{-1} $ during the interactive feature-learning stage of grokking: similar features have negative off-diagonal entries $ B_{j\ell} $, p

arxivMay 13

Spectral Entropy Collapse as a Phase Transition in Delayed Generalisation: An Interventional and Predictive Framework for Grokkin

arXiv:2604.13123v2 Announce Type: replace Abstract: Grokking - the delayed transition from memorisation to generalisation in neural networks - remains poorly understood. We study this phenomenon through the geometry of learned representations and identify a consistent empirical signature preceding g

techcrunchMay 12

Threads tests a Meta AI integration that works similarly to Grok

The feature is designed to help people get real-time context about trends and breaking stories, as well as receive recommendations, all within conversations.

arxivMay 12

Model Capacity Determines Grokking through Competing Memorisation and Generalisation Speeds

arXiv:2605.09724v1 Announce Type: new Abstract: Existing accounts of grokking explain the phenomena in terms of mechanistic frameworks such as circuit efficiency or lazy-to-rich transitions. However, despite a known dependence between grokking and model size, how model capacity shapes grokking remai

arxivMay 12

Distributional Spectral Diagnostics for Localizing Grokking Transitions

arXiv:2605.08237v1 Announce Type: new Abstract: In grokking, a model first fits the training data while test accuracy remains low, and only later begins to generalize. We ask whether this transition can be localized from observed training trajectories before the test accuracy rises, and formulate gr

arxivMay 8

Topological Signatures of Grokking

arXiv:2605.06352v1 Announce Type: cross Abstract: We study the grokking phenomenon through the lens of topology. Using persistent homology on point clouds derived from the embedding matrices of a range of models trained on modular arithmetic with varying primes, we identify a clear and consistent to

arxivMay 8

A Basin-Selection Perspective on Grokking via Singular Learning Theory

arXiv:2603.01192v3 Announce Type: replace-cross Abstract: Grokking, the abrupt transition from memorization to generalisation after extended training, suggests the presence of competing solution basins with distinct statistical properties. We study this phenomenon through the lens of Singular Learni

arxivMay 6

The Norm-Separation Delay Law of Grokking: A First-Principles Theory of Delayed Generalization

arXiv:2603.13331v2 Announce Type: replace Abstract: Grokking -- the sudden generalisation that appears long after a model has perfectly memorised its training data -- has been widely observed but lacks a quantitative theory explaining the length of the delay. We show that grokking is a norm-driven r

arxivMay 6

The Geometric Inductive Bias of Grokking: Bypassing Phase Transitions via Architectural Topology

arXiv:2603.05228v3 Announce Type: replace-cross Abstract: Mechanistic interpretability typically relies on post-hoc analysis of trained networks. We instead adopt an interventional approach: testing hypotheses a priori by modifying architectural topology to observe training dynamics. We study grokki

thevergeApr 30

Elon Musk confirms xAI used OpenAI’s models to train Grok

In a federal courtroom in California on Thursday, Elon Musk testified that his own AI startup, xAI, has used OpenAI's models to improve its own. The matter at question is model distillation, a common industry practice by which one larger AI model acts as a "teacher" of sorts to pass on knowledge to

techcrunchApr 30

Elon Musk testifies that xAI trained Grok on OpenAI models

"Distillation" is a hot topic as frontier labs try to prevent smaller competitors from copying their models.

arxivApr 21

Grokking of Diffusion Models: Case Study on Modular Addition

arXiv:2604.17673v1 Announce Type: new Abstract: Despite their empirical success, how diffusion models generalize remains poorly understood from a mechanistic perspective. We demonstrate that diffusion models trained with flow-matching objectives exhibit grokking--delayed generalization after overfit

HomeModelsNews