Model Detail
xAI: Grok 4 Fast
—xAI: Grok 4 Fast is a multimodal model released by xAI. And supports text+image+file->text inputs.
xAI: Grok 4 Fast is priced at $0.2/M input tokens and $0.5/M output tokens. Operationally the model offers a 2000K-token context window, which matters when sizing it for prompt-heavy or latency-sensitive workloads. At this input rate the model sits in the commodity tier and is suitable for high-volume workloads where per-call cost dominates the decision.
The published knowledge cutoff is 2025-09-30, so newer events will not be reflected in zero-shot answers without retrieval.
xAI: Grok 4 Fast is best fit for mixed text-and-image reasoning tasks such as document understanding, high-volume batch jobs where per-call cost dominates the budget, and long-context tasks such as full-codebase analysis or book-length summarization (2000K tokens). Treat this as a starting matrix rather than a benchmark verdict — the right deployment usually depends on the specific evaluation suite that mirrors your workload.
To Grok Grokking: Provable Grokking in Ridge Regression
arXiv:2601.19791v4 Announce Type: replace Abstract: We study grokking, the onset of generalization long after overfitting, in a classical ridge regression setting. We prove end-to-end grokking results for learning over-parameterized linear regression models using gradient descent with weight decay.
Grokipedia vs Wikipedia: An LLM-Based Audit of Political Neutrality along Ideologies
arXiv:2607.15146v1 Announce Type: new Abstract: Online encyclopedias shape political opinion and, through it, democratic discourse. In late 2025, Grokipedia was released, an encyclopedia written entirely by the LLM Grok. One motivation behind the project was to provide an unbiased alternative to Wik
Algebraic Representability as the Limiting Regime of Grokking: An Exactly Solvable Model with Holomorphic Activations
arXiv:2607.13749v1 Announce Type: new Abstract: Neural networks trained on modular arithmetic exhibit grokking, a delayed transition from memorisation to generalisation known to depend on model capacity: too little and the network memorises slowly or not at all, too much and it generalises almost im

xAI sues a man for using Grok to generate CSAM ‘deepfakes’
The Elon Musk-owned xAI is suing a South Carolina man who allegedly used the company's Grok AI chatbot to generate child sexual abuse material (CSAM). In a lawsuit reported earlier by Reuters, xAI claims Terry Wayne Harwood "knowingly and intentionally used Grok to circumvent safeguards, alter nonco
What Makes a Representational Prior Work? Feature Families, Label-Free Invariances, and Critical Windows in Grokking
arXiv:2607.12735v1 Announce Type: new Abstract: Companion work showed the grokking delay is causally the time to form task-structured representations, injectable via a contrastive prior. Here we characterize what makes such a prior work, across four axes, in 188 new runs. Content: a coherent, learna
Structure-Specific Representational Priors Causally Control the Grokking Delay
arXiv:2607.04333v3 Announce Type: replace Abstract: Grokking -- generalization long after training-set interpolation -- has been accelerated by structure-agnostic interventions (gradient filtering, weight-norm clamping, geometric penalties). Whether the delay specifically measures the time to form t