·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
Fine-Tuning of Transformer models with Frames5h◆Feature Transformation Enhanced Jacobi Polynomial Graph Filtering for Graph Anomaly Detection5h◆Naive Prompt Optimization: Rethinking the Need for Complex Prompt Search5h◆Syntax vs. Semantics: How Transformers Learn Deep Dependencies5h◆NeuronFuzz: Safety Neuron Guided Fuzzing for LLM Safety Evaluation5h◆FaulT-Bench: Towards Benchmarking Network Troubleshooting LLM Agents under Unreliable User Tickets5h◆Unsupervised Post-Training of Foundation Models: A Survey5h◆Drift-Adaptive ICU Intervention Prediction: Freezing the Physiological Encoder for Auditable Model Updating5h◆Jiuge-Tuiqiao: An Interpretable Human-AI System for Classical Chinese Poetry Refinement5h◆Is Your Neighborhood Safe? Place-based Stigma in Large Language Models' Urban Safety Judgments5h◆Magnon-induced phononic Chern insulator5h◆Five Primitives for Governing Autonomous AI Agents at Runtime5h◆AgentFold: Closed-Loop Agentic Search for Protein Folding Model Design5h◆Categorizer Automata for Discounted-Sum Payoffs5h◆FIRSTPASS: A Multi-Domain, Multi-Round Peer Review Dataset Grounded in Real Editorial Outcomes5h◆Improving LLM Interpretability with User-Centric Chain-of-Thought Reasoning5h◆A Very Big Video Reasoning Suite5h◆GAMMA: Global Bit Allocation for Mixed-Precision Models under Arbitrary Budgets5h◆Beyond Execution: Auditing Experimental Fidelity in LLM-Driven Scientific Research5h◆From Reasoning to Pixels: Grounded Medical Multimodal LLMs for VQA and Segmentation5h◆Fine-Tuning of Transformer models with Frames5h◆Feature Transformation Enhanced Jacobi Polynomial Graph Filtering for Graph Anomaly Detection5h◆Naive Prompt Optimization: Rethinking the Need for Complex Prompt Search5h◆Syntax vs. Semantics: How Transformers Learn Deep Dependencies5h◆NeuronFuzz: Safety Neuron Guided Fuzzing for LLM Safety Evaluation5h◆FaulT-Bench: Towards Benchmarking Network Troubleshooting LLM Agents under Unreliable User Tickets5h◆Unsupervised Post-Training of Foundation Models: A Survey5h◆Drift-Adaptive ICU Intervention Prediction: Freezing the Physiological Encoder for Auditable Model Updating5h◆Jiuge-Tuiqiao: An Interpretable Human-AI System for Classical Chinese Poetry Refinement5h◆Is Your Neighborhood Safe? Place-based Stigma in Large Language Models' Urban Safety Judgments5h◆Magnon-induced phononic Chern insulator5h◆Five Primitives for Governing Autonomous AI Agents at Runtime5h◆AgentFold: Closed-Loop Agentic Search for Protein Folding Model Design5h◆Categorizer Automata for Discounted-Sum Payoffs5h◆FIRSTPASS: A Multi-Domain, Multi-Round Peer Review Dataset Grounded in Real Editorial Outcomes5h◆Improving LLM Interpretability with User-Centric Chain-of-Thought Reasoning5h◆A Very Big Video Reasoning Suite5h◆GAMMA: Global Bit Allocation for Mixed-Precision Models under Arbitrary Budgets5h◆Beyond Execution: Auditing Experimental Fidelity in LLM-Driven Scientific Research5h◆From Reasoning to Pixels: Grounded Medical Multimodal LLMs for VQA and Segmentation5h◆
News/Mini-R1: Reproduce Deepseek R1 „aha moment“ a RL tutorial
huggingface
PublishedJanuary 31, 2025 at 10:29 AM

Mini-R1: Reproduce Deepseek R1 „aha moment“ a RL tutorial

Source
huggingface.cofull article ↗
Read on huggingface→
Stay posted· Newsletter

A 5-min weekly brief — top movers, price watch, story of the week.

// no spam · unsubscribe one-click · free forever

Discussion
Source
↗
huggingface
Read original ↗All from huggingface →

No replies yet. Be first.

Source
↗
huggingface
Read original ↗All from huggingface →
The Bubble Brief
WEEKLY

Read AI insights every Tuesday — top movers, new releases, story of the week.

// no spam · unsubscribe one-click · free forever

Originally published on huggingface ↗
HomeModelsNews