Model Detail
gemma-4-26B-A4B-it-heretic-APEX-GGUF
—gemma-4-26B-A4B-it-heretic-APEX-GGUF is a code generation model with 26B parameters released by mudler. Released under the gemma license.
gemma-4-26B-A4B-it-heretic-APEX-GGUF ships with 26B parameters, distributed as a quantized weight variant for lower-VRAM inference. Distribution is governed by the gemma license — review the exact terms before commercial deployment.
gemma-4-26B-A4B-it-heretic-APEX-GGUF is best fit for code completion, repository-scale Q&A, and pair-programming integrations. It is a less obvious choice for one-shot generation of security-critical code without review. Treat this as a starting matrix rather than a benchmark verdict — the right deployment usually depends on the specific evaluation suite that mirrors your workload.
Gemma 4 Technical Report
arXiv:2607.02770v2 Announce Type: replace Abstract: We introduce Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family. Designed to advance compute efficiency and reasoning, the Gemma 4 model suite features dense and Mixture-of-Experts architectures,
Do Active SAE Feature Planes Carry More Holonomy? A Preregistered Reversal in Gemma
arXiv:2607.20522v1 Announce Type: new Abstract: This paper tests whether holonomy concentrates on active sparse-autoencoder (SAE) feature planes in Gemma 2 2B, a concrete operationalization of the broader semantic-concentration prediction. Holonomy is measured at the final-token layer-12 to layer-13
Trivial Prompt Reframing Bypasses Safety Guardrails in Google\'s MedGemma-4B
arXiv:2607.09804v1 Announce Type: cross Abstract: Open-weight medical language models are increasingly used as the base of patient-facing and clinician-support applications. Their model cards prohibit specific behaviors -- recommending exact drug dosages, issuing definitive diagnoses, prescribing tr
Hugging Face and Cerebras bring Gemma 4 to real-time voice AI
DistilledGemma: Balanced Efficiency-Accuracy for Person-Place Relation Extraction from Multilingual Historical Articles
arXiv:2606.29130v1 Announce Type: new Abstract: We present DistilledGemma, an efficient and accurate system for the HIPE-2026 shared task on person-place relation extraction from multilingual historical newspaper articles in English, German, and French. Our approach adopts a three-stage knowledge di
How Transparent is DiffusionGemma?
arXiv:2606.20560v1 Announce Type: cross Abstract: LLM reasoning transparency is a critical affordance for understanding model decisions, mitigating misuse and misalignment, and debugging surprising model behaviors. However, DiffusionGemma performs a larger fraction of its computation in a continuous