·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
China delivers a one-two punch to America’s AI dominance2h◆AI is more likely than humans to form biases when hiring3h◆Beyond a Single Direction: Chain-of-Thought Disrupts Simple Steering of Refusal8h◆SkillCorpus: Consolidating and Evaluating the Open Skill Ecosystem for Real-World LLM Agents8h◆PolyInterview: An LLM-based Platform for Immersive Mock Interview Practice with Comprehensive Multimodal Assessment8h◆Latency-Response Theory Model: Evaluating Large Language Models via Response Accuracy and Chain-of-Thought Length8h◆Label-Free Concept Drift Assessment for Reliable AI in Emerging Wireless Applications8h◆DyneTrion: A Spatio-temporally Coherent Generative Emulator for Protein Dynamics Across Timescales8h◆Testing Distributions Against Bounded Distinguishers8h◆Do Coding Agents Need Executable World Models, Simplification, and Verification to Solve ARC-AGI-3?8h◆Beyond a Joke: Multi-Angle Reasoning for Detecting and Explaining Harmful Humor in Memes8h◆From Black Box to Executable Logic: Explainable Reinforcement Learning through Prolog Expert Systems8h◆Closing the AI Trust Gap: The Case for Independent Certification for Trustworthy AI8h◆CRAFT: Clustering Rubrics to Diagnose Weak LLM Capabilities and Generate Targeted Fine-Tuning Data8h◆The AI Fiction Paradox8h◆Decoupled Alignment for Robust Plug-and-Play Adaptation8h◆Hybrid coupling with operator inference and the overlapping Schwarz alternating method8h◆Digital Pantheon: Simulating and Auditing Coalition Formation with LLM Agents8h◆EpiNarrate: Agentic Generation of Grounded Narratives from Epidemiological Scenario Projections8h◆Before the Action: Benchmarking LLMs on Prospective Hypothesis Discovery8h◆China delivers a one-two punch to America’s AI dominance2h◆AI is more likely than humans to form biases when hiring3h◆Beyond a Single Direction: Chain-of-Thought Disrupts Simple Steering of Refusal8h◆SkillCorpus: Consolidating and Evaluating the Open Skill Ecosystem for Real-World LLM Agents8h◆PolyInterview: An LLM-based Platform for Immersive Mock Interview Practice with Comprehensive Multimodal Assessment8h◆Latency-Response Theory Model: Evaluating Large Language Models via Response Accuracy and Chain-of-Thought Length8h◆Label-Free Concept Drift Assessment for Reliable AI in Emerging Wireless Applications8h◆DyneTrion: A Spatio-temporally Coherent Generative Emulator for Protein Dynamics Across Timescales8h◆Testing Distributions Against Bounded Distinguishers8h◆Do Coding Agents Need Executable World Models, Simplification, and Verification to Solve ARC-AGI-3?8h◆Beyond a Joke: Multi-Angle Reasoning for Detecting and Explaining Harmful Humor in Memes8h◆From Black Box to Executable Logic: Explainable Reinforcement Learning through Prolog Expert Systems8h◆Closing the AI Trust Gap: The Case for Independent Certification for Trustworthy AI8h◆CRAFT: Clustering Rubrics to Diagnose Weak LLM Capabilities and Generate Targeted Fine-Tuning Data8h◆The AI Fiction Paradox8h◆Decoupled Alignment for Robust Plug-and-Play Adaptation8h◆Hybrid coupling with operator inference and the overlapping Schwarz alternating method8h◆Digital Pantheon: Simulating and Auditing Coalition Formation with LLM Agents8h◆EpiNarrate: Agentic Generation of Grounded Narratives from Epidemiological Scenario Projections8h◆Before the Action: Benchmarking LLMs on Prospective Hypothesis Discovery8h◆
News/A Benchmark Construction and Evaluation Framework for Specialist Domains: Case Study on Defense-related Documents
arxiv
PublishedMay 28, 2026 at 4:00 AM
▲bullish

A Benchmark Construction and Evaluation Framework for Specialist Domains: Case Study on Defense-related Documents

Source
arxiv.orgfull article ↗
Read on arxiv→
Publisher summary· verbatim

arXiv:2604.17943v2 Announce Type: replace Abstract: RAG-based question-answering (QA) in specialist domains faces a cold-start problem: lack of evaluative benchmarks and absence of labeled data for post-training. We present DoRA (Domain-oriented RAG Assessment), a novel benchmark construction and ev

Models mentioned
01
  • 01meta-llama logo
    Llama-3.1-8B
    meta-llama/Llama-3.1-8B
    DL 1.5M0.0%IN $0.10/Mtok
Related
04
  • arxivMay 22
    GraphRAG on Consumer Hardware: Benchmarking Local LLMs for Healthcare EHR Schema Retrieval
  • arxivMay 22
    DrugRAG: Enhancing Pharmacy LLM Performance Through A Novel Retrieval-Augmented Generation Pipeline
  • arxivMay 8
    Measuring Evaluation-Context Divergence in Open-Weight LLMs: A Paired-Prompt Protocol with Pilot Evidence of Alignment-Pipeline-Specific Heterogeneity
  • arxivApr 4
    Countering Catastrophic Forgetting of Large Language Models for Better Instruction Following via Weight-Space Model Merging
Stay posted· Newsletter

A 5-min weekly brief — top movers, price watch, story of the week.

// no spam · unsubscribe one-click · free forever

Discussion
Mentioned models
01
  • 01
    Llama-3.1-8B
    meta-llama/Llama-3.1-8B
    1.5M dl
Source
↗
arxiv
Read original ↗All from arxiv →
Tags
04
#benchmark#evaluation#specialist-domains#question-answering

No replies yet. Be first.

Mentioned models
01
  • 01
    Llama-3.1-8B
    meta-llama/Llama-3.1-8B
    1.5M dl
Source
↗
arxiv
Read original ↗All from arxiv →
Tags
04
#benchmark#evaluation#specialist-domains#question-answering

Related coverage

More from ARXIV
arxivBeyond a Single Direction: Chain-of-Thought Disrupts Simple Steering of Refusal8harxivSkillCorpus: Consolidating and Evaluating the Open Skill Ecosystem for Real-World LLM Agents8harxivPolyInterview: An LLM-based Platform for Immersive Mock Interview Practice with Comprehensive Multimodal Assessment8harxivLatency-Response Theory Model: Evaluating Large Language Models via Response Accuracy and Chain-of-Thought Length8h
The Bubble Brief
WEEKLY

Read benchmark insights every Tuesday — top movers, new releases, story of the week.

// no spam · unsubscribe one-click · free forever

Originally published on arxiv ↗
HomeModelsNews