·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
America needs to stop getting shocked by Chinese AI2h◆Advancing next-gen AI with materials science innovation2h◆Gritt exits stealth with $34 million for robots to build solar plants—then, everything else3h◆Capacity and Redundancy Trade-offs in Multi-Task Learning9h◆Predictive Training with Latent Imagination for Visual Quadruped Navigation9h◆Where Not to Learn: Prior-Aligned Training with Subset-based Attribution Constraints for Reliable Decision-Making9h◆Did We Actually Fix It? An Independent Adversarial Stress-Test of Post-Point-Adjustment Evaluation Metrics for Time-Series Anomaly Detection9h◆Supervised Reward Inference9h◆PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization9h◆Is Progressive Disclosure All You Need for Long-Context Agents?9h◆It Depends on the Dataset: When a Brain-Encoding Model's Predicted Responses Beat Their Visual Backbone for Video Memorability9h◆DMFNet: Dual-Backbone Multiscale Fusion Network for Urban Scene Classification9h◆Oracle Gap and Signal Fidelity: A Fixed-Pool Diagnostic for Test-Time Collaboration9h◆Time-Frequency Consistency Learning for Robust Speech Deepfake Detection9h◆Scientific reasoning does not reliably translate into scientific forecasting in frontier AI9h◆Language Triggers Hijack Language Circuits: A Mechanistic Analysis of Backdoor Behaviors in Large Language Models9h◆Kernel Regression with Tensor Trains and Hadamard Overparameterization9h◆AI-Augmented Human Resource Management? Insights from German companies9h◆Diagnosing Correctness Probes under Self-Judgement Confounding9h◆BLAD: A Historically Contextualized, Multilingual Dataset of Bangladeshi Legal Acts (1799 to 2025)9h◆America needs to stop getting shocked by Chinese AI2h◆Advancing next-gen AI with materials science innovation2h◆Gritt exits stealth with $34 million for robots to build solar plants—then, everything else3h◆Capacity and Redundancy Trade-offs in Multi-Task Learning9h◆Predictive Training with Latent Imagination for Visual Quadruped Navigation9h◆Where Not to Learn: Prior-Aligned Training with Subset-based Attribution Constraints for Reliable Decision-Making9h◆Did We Actually Fix It? An Independent Adversarial Stress-Test of Post-Point-Adjustment Evaluation Metrics for Time-Series Anomaly Detection9h◆Supervised Reward Inference9h◆PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization9h◆Is Progressive Disclosure All You Need for Long-Context Agents?9h◆It Depends on the Dataset: When a Brain-Encoding Model's Predicted Responses Beat Their Visual Backbone for Video Memorability9h◆DMFNet: Dual-Backbone Multiscale Fusion Network for Urban Scene Classification9h◆Oracle Gap and Signal Fidelity: A Fixed-Pool Diagnostic for Test-Time Collaboration9h◆Time-Frequency Consistency Learning for Robust Speech Deepfake Detection9h◆Scientific reasoning does not reliably translate into scientific forecasting in frontier AI9h◆Language Triggers Hijack Language Circuits: A Mechanistic Analysis of Backdoor Behaviors in Large Language Models9h◆Kernel Regression with Tensor Trains and Hadamard Overparameterization9h◆AI-Augmented Human Resource Management? Insights from German companies9h◆Diagnosing Correctness Probes under Self-Judgement Confounding9h◆BLAD: A Historically Contextualized, Multilingual Dataset of Bangladeshi Legal Acts (1799 to 2025)9h◆
DataBubble·

Model Detail

meta-llama logo

Meta-Llama-3-8B-Instruct

—
Provider: MetaCategory: llmPipeline: text-generationParameters: 8B
DB Score
11.3
Downloads
1.5M
Likes
5K
Day
+0.0%
Week
+0.0%
Month
+0.0%
Overview

Meta-Llama-3-8B-Instruct is a large language model with 8B parameters released by Meta. The model is registered under the text-generation pipeline tag on Hugging Face, and supports text->text inputs, released under the llama3 license.

Performance

Open-LLM-Leaderboard scoring places it at MMLU-Pro 29, GPQA 6, IFEval 48, BBH 27, giving a sense of how it handles instruction following, reasoning, and graduate-level QA in absolute terms.

How we score this →
Pricing & Throughput

Meta-Llama-3-8B-Instruct is priced at $0.15/M input tokens and $0.15/M output tokens. Operationally the model offers a 8K-token context window, which matters when sizing it for prompt-heavy or latency-sensitive workloads. At this input rate the model sits in the commodity tier and is suitable for high-volume workloads where per-call cost dominates the decision.

Technical

Meta-Llama-3-8B-Instruct ships as a LlamaForCausalLM / 💬 chat models (RLHF, DPO, IFT, ...) architecture with 8B parameters. The published knowledge cutoff is 2023-12-31, so newer events will not be reflected in zero-shot answers without retrieval. Total weight footprint is approximately 8.0 GB, which is the relevant figure when planning local-inference VRAM. Access is gated on Hugging Face under the llama3 license, which means a manual approval step before weights can be downloaded.

Use Cases

Meta-Llama-3-8B-Instruct is best fit for general-purpose chat and instruction-following workloads, and high-volume batch jobs where per-call cost dominates the budget. Treat this as a starting matrix rather than a benchmark verdict — the right deployment usually depends on the specific evaluation suite that mirrors your workload.

Download History
Pricing
Input ($/M tokens)
$0.15
Output ($/M tokens)
$0.15
Context Window
8K
Benchmark Scores
IFEval
47.8
BBH
26.8
GPQA
5.7
MMLU-Pro
28.8
MATH
9.1
MUSR
5.4
Average
20.6
Model Info
Licensellama3
ArchitectureLlamaForCausalLM
Type💬 chat models (RLHF, DPO, IFT, ...)
Modalitytext->text
Knowledge Cutoff2023-12-31
Citations16,069 (3016 influential)
Recent newsView all news →
Related News
arxiv9h ago

Natural Language Access to Domain-Specific Metadata: A Reusable Framework for LLM Query Generation

arXiv:2607.18029v1 Announce Type: cross Abstract: Researchers need to answer ad-hoc questions about the contents of domain-specific archives but often lack the expertise to write structured queries on the metadata. We show that when domain vocabulary and semantics are captured in a well-designed Web

arxiv9h ago

On the Potential of Graph Neural Networks as Metamodels for Supply Chain Optimization: Dataset, Architectures, and Directions

arXiv:2607.16769v1 Announce Type: new Abstract: Graph Neural Networks (GNNs) have emerged as a powerful, differentiable class of learning models for graph-structured systems. Their ability to generalize across topologies opens the prospect of a surrogate for combined structural and parametric optimi

arxiv9h ago

A multiverse-consensus pipeline for reproducible feature selection in untargeted LC-MS metabolomics

arXiv:2607.17345v1 Announce Type: new Abstract: Background: Untargeted LC-MS metabolomics requires a long chain of preprocessing decisions, each with several equally defensible options. Analysts typically commit to one pipeline and report the resulting feature shortlist. How strongly that shortlist

arxiv9h ago

Harnessing disorder to decouple extension and shear in kirigami metamaterials

arXiv:2607.16583v1 Announce Type: cross Abstract: Kirigami turns stiff sheets into compliant, shape-morphing structures, but its reliance on periodic cut patterns comes at a cost: correlated panel rotations couple extension to shear, so stretching one axis drives a parasitic shear that cannot be sup

arxiv9h ago

Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels

arXiv:2607.09796v2 Announce Type: replace Abstract: Direct Preference Optimization (DPO) has become an important method for aligning large language models (LLMs) with human preferences because it removes the need for explicit reward modeling and reinforcement learning. However, its performance depen

arxiv9h ago

MTSSL: Meta-Thresholding Semi-Supervised Learning

arXiv:2607.16363v1 Announce Type: cross Abstract: A large body of Semi-supervised Learning~(SSL) algorithms encounter the threshold $\tau$ to select pseudo-labels. The value of $\tau$ across different SSL algorithms can vary depending on the learning perspective, yet they may achieve similar perform

Related Models
meta-llama logo
Llama-3.2-1B-Instruct
Meta · 8.6M downloads
meta-llama logo
Llama-3.1-8B-Instruct
Meta · 8.0M downloads
google-bert logo
bert-base-uncased
google-bert · 69.6M downloads
sentence-transformers logo
paraphrase-multilingual-MiniLM-L12-v2
SBERT · 48.6M downloads
HomeModelsNews