·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
China delivers a one-two punch to America’s AI dominance1h◆AI is more likely than humans to form biases when hiring3h◆Beyond a Single Direction: Chain-of-Thought Disrupts Simple Steering of Refusal7h◆SkillCorpus: Consolidating and Evaluating the Open Skill Ecosystem for Real-World LLM Agents7h◆PolyInterview: An LLM-based Platform for Immersive Mock Interview Practice with Comprehensive Multimodal Assessment7h◆Latency-Response Theory Model: Evaluating Large Language Models via Response Accuracy and Chain-of-Thought Length7h◆Do Coding Agents Need Executable World Models, Simplification, and Verification to Solve ARC-AGI-3?7h◆Beyond a Joke: Multi-Angle Reasoning for Detecting and Explaining Harmful Humor in Memes7h◆From Black Box to Executable Logic: Explainable Reinforcement Learning through Prolog Expert Systems7h◆Closing the AI Trust Gap: The Case for Independent Certification for Trustworthy AI7h◆CRAFT: Clustering Rubrics to Diagnose Weak LLM Capabilities and Generate Targeted Fine-Tuning Data7h◆Digital Pantheon: Simulating and Auditing Coalition Formation with LLM Agents7h◆EpiNarrate: Agentic Generation of Grounded Narratives from Epidemiological Scenario Projections7h◆Before the Action: Benchmarking LLMs on Prospective Hypothesis Discovery7h◆How Much Human Label Variation Does Formal Semantic Structure Explain?: Group-Level Effects and Item-Level Ceilings in NLI7h◆Learn to Memorize: Scalable Continual Learning in Semiparametric Models with Mixture-of-Neighbors Induction Memory7h◆Echoes Across Vietnam's Highlands, Delta, and Coast: A Multilingual Corpus for Cham, Khmer, and Tay-Nung7h◆Stochastic Reset Pathfinding: Path-Level Regret for Cascading Bandits over Graph Paths7h◆ADS-C: Antidistillation Sampling for Classification7h◆Do Generative Models Keep Time? A Time-Aware Evaluation of Synthetic Sequential Tabular Data7h◆China delivers a one-two punch to America’s AI dominance1h◆AI is more likely than humans to form biases when hiring3h◆Beyond a Single Direction: Chain-of-Thought Disrupts Simple Steering of Refusal7h◆SkillCorpus: Consolidating and Evaluating the Open Skill Ecosystem for Real-World LLM Agents7h◆PolyInterview: An LLM-based Platform for Immersive Mock Interview Practice with Comprehensive Multimodal Assessment7h◆Latency-Response Theory Model: Evaluating Large Language Models via Response Accuracy and Chain-of-Thought Length7h◆Do Coding Agents Need Executable World Models, Simplification, and Verification to Solve ARC-AGI-3?7h◆Beyond a Joke: Multi-Angle Reasoning for Detecting and Explaining Harmful Humor in Memes7h◆From Black Box to Executable Logic: Explainable Reinforcement Learning through Prolog Expert Systems7h◆Closing the AI Trust Gap: The Case for Independent Certification for Trustworthy AI7h◆CRAFT: Clustering Rubrics to Diagnose Weak LLM Capabilities and Generate Targeted Fine-Tuning Data7h◆Digital Pantheon: Simulating and Auditing Coalition Formation with LLM Agents7h◆EpiNarrate: Agentic Generation of Grounded Narratives from Epidemiological Scenario Projections7h◆Before the Action: Benchmarking LLMs on Prospective Hypothesis Discovery7h◆How Much Human Label Variation Does Formal Semantic Structure Explain?: Group-Level Effects and Item-Level Ceilings in NLI7h◆Learn to Memorize: Scalable Continual Learning in Semiparametric Models with Mixture-of-Neighbors Induction Memory7h◆Echoes Across Vietnam's Highlands, Delta, and Coast: A Multilingual Corpus for Cham, Khmer, and Tay-Nung7h◆Stochastic Reset Pathfinding: Path-Level Regret for Cascading Bandits over Graph Paths7h◆ADS-C: Antidistillation Sampling for Classification7h◆Do Generative Models Keep Time? A Time-Aware Evaluation of Synthetic Sequential Tabular Data7h◆
News/HABIB_TAZ at SemEval-2026 Task 11: Disentangling Formal Logic from Content via Synthetic Training and Multi-Objective Optimization
arxiv
PublishedJuly 18, 2026 at 4:00 AM
—neutral

HABIB_TAZ at SemEval-2026 Task 11: Disentangling Formal Logic from Content via Synthetic Training and Multi-Objective Optimization

Source
arxiv.orgfull article ↗
Read on arxiv→
Publisher summary· verbatim

arXiv:2607.14349v1 Announce Type: cross Abstract: While Large Language Models (LLMs) excel in many general NLP tasks, their formal reasoning capabilities are often compromised by content effects, demonstrating a measurable bias towards real-world plausibility. In this paper, we present our system fo

Stay posted· Newsletter

A 5-min weekly brief — top movers, price watch, story of the week.

// no spam · unsubscribe one-click · free forever

Discussion
Source
↗
arxiv
Read original ↗All from arxiv →

No replies yet. Be first.

Source
↗
arxiv
Read original ↗All from arxiv →

Related coverage

More from ARXIV
arxivBeyond a Single Direction: Chain-of-Thought Disrupts Simple Steering of Refusal7harxivSkillCorpus: Consolidating and Evaluating the Open Skill Ecosystem for Real-World LLM Agents7harxivPolyInterview: An LLM-based Platform for Immersive Mock Interview Practice with Comprehensive Multimodal Assessment7harxivLatency-Response Theory Model: Evaluating Large Language Models via Response Accuracy and Chain-of-Thought Length7h
The Bubble Brief
WEEKLY

Read AI insights every Tuesday — top movers, new releases, story of the week.

// no spam · unsubscribe one-click · free forever

Originally published on arxiv ↗
HomeModelsNews