·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
Startup Battlefield 200 applications officially close in 3 days2h◆Google will pay SpaceX $920M per month for compute3h◆The most interesting startups right now want to get you off your phone5h◆This is your laptop… on AI6h◆New York lawmakers pass one-year ban on new data centers7h◆The token bill comes due: Inside the industry scramble to manage AI’s runaway costs8h◆The latest AI news we announced in May 20268h◆The ‘together tech’ wave might be the most intriguing startup bet of 20268h◆This AI startup says it can tell if a script will make a hit film8h◆AirTrunk commits $30B to build 5GW of AI data centers in India9h◆The Meta hack shows there’s more to AI security than Mythos13h◆Mira Murati steps back into the spotlight, carefully17h◆SFMambaNet: Spectral-Frequency Enhanced Selective State Space Model for Correspondence Pruning18h◆Optical-Guided Neural Collapse for SAR Few-Shot Class Incremental Learning18h◆Dynamic Infilling Anchors for Format-Constrained Generation in Diffusion Large Language Models18h◆Temporal Order Matters for Agentic Memory: Segment Trees for Long-Horizon Agents18h◆Why Muon Outperforms Adam: A Curvature Perspective18h◆Vision Hopfield Memory Networks18h◆Provably Auditable and Safe LLM Agents from Human-Authored Ontologies18h◆FlexRank: Nested Low-Rank Knowledge Decomposition for Adaptive Model Deployment18h◆Startup Battlefield 200 applications officially close in 3 days2h◆Google will pay SpaceX $920M per month for compute3h◆The most interesting startups right now want to get you off your phone5h◆This is your laptop… on AI6h◆New York lawmakers pass one-year ban on new data centers7h◆The token bill comes due: Inside the industry scramble to manage AI’s runaway costs8h◆The latest AI news we announced in May 20268h◆The ‘together tech’ wave might be the most intriguing startup bet of 20268h◆This AI startup says it can tell if a script will make a hit film8h◆AirTrunk commits $30B to build 5GW of AI data centers in India9h◆The Meta hack shows there’s more to AI security than Mythos13h◆Mira Murati steps back into the spotlight, carefully17h◆SFMambaNet: Spectral-Frequency Enhanced Selective State Space Model for Correspondence Pruning18h◆Optical-Guided Neural Collapse for SAR Few-Shot Class Incremental Learning18h◆Dynamic Infilling Anchors for Format-Constrained Generation in Diffusion Large Language Models18h◆Temporal Order Matters for Agentic Memory: Segment Trees for Long-Horizon Agents18h◆Why Muon Outperforms Adam: A Curvature Perspective18h◆Vision Hopfield Memory Networks18h◆Provably Auditable and Safe LLM Agents from Human-Authored Ontologies18h◆FlexRank: Nested Low-Rank Knowledge Decomposition for Adaptive Model Deployment18h◆
News/model/Qwen3-Coder-30B-A3B-Instruct-GGUF

Qwen3-Coder-30B-A3B-Instruct-GGUF news

10 articles mentioning Qwen3-Coder-30B-A3B-Instruct-GGUF

arxiv3d ago

LinguIUTics at PsyDefDetect: Iterative Imbalance-Aware Fine-tuning of Qwen3-8B for Psychological Defense Mechanism Classification

arXiv:2606.00647v1 Announce Type: cross Abstract: Detecting psychological defense mechanisms in conversational text remains a challenging clinical NLP problem. For the PsyDefDetect 2026 shared task (nine-class utterance classification evaluated via macro F1), our team LinguIUTics achieves a macro F1

arxivMay 15

Procedural-skill SFT across capacity tiers: A W-Shaped pre-SFT Trajectory and Regime-Asymmetric Mechanism on 0.8B-4B Qwen3.5 Models

arXiv:2605.11907v2 Announce Type: replace Abstract: We measure procedural-skill SFT contribution across three Qwen3.5 dense scales (0.8B, 2B, 4B) on a 200-task / 40-skill holdout, with Claude Haiku 4.5 as a frontier reference. The corpus is 353 rows of (task + procedural-skill block, Opus chain-of-t

HomeModelsNews
arxivMay 11

Qwen3-VL-Seg: Unlocking Open-World Referring Segmentation with Vision-Language Grounding

arXiv:2605.07141v1 Announce Type: cross Abstract: Open-world referring segmentation requires grounding unconstrained language expressions to precise pixel-level regions. Existing multimodal large language models (MLLMs) exhibit strong open-world visual grounding, but their outputs remain limited to

arxivApr 22

Qwen3.5-Omni Technical Report

arXiv:2604.15804v2 Announce Type: replace Abstract: In this work, we present Qwen3.5-Omni, the latest advancement in the Qwen-Omni model family. Representing a significant evolution over its predecessor, Qwen3.5-Omni scales to hundreds of billions of parameters and supports a 256k context length. By

arxivApr 17

Benchmarking Linguistic Adaptation in Comparable-Sized LLMs: A Study of Llama-3.1-8B, Mistral-7B-v0.1, and Qwen3-8B on Romanized Nepali

arXiv:2604.14171v1 Announce Type: new Abstract: Romanized Nepali, the Nepali language written in the Latin alphabet, is the dominant medium for informal digital communication in Nepal, yet it remains critically underresourced in the landscape of Large Language Models (LLMs). This study presents a sy

arxivApr 17

QU-NLP at ArchEHR-QA 2026: Two-Stage QLoRA Fine-Tuning of Qwen3-4B for Patient-Oriented Clinical Question Answering and Evidence Sentence Alignment

arXiv:2604.14175v1 Announce Type: new Abstract: We present a unified system addressing both Subtask 3 (answer generation) and Subtask 4 (evidence sentence alignment) of the ArchEHR-QA Shared Task. For Subtask 3, we apply two-stage Quantised Low-Rank Adaptation (QLoRA) to Qwen3-4B loaded in 4-bit NF4

#research#natural language processing#question answering
arxivApr 10

Robustness Risk of Conversational Retrieval: Identifying and Mitigating Noise Sensitivity in Qwen3-Embedding Model

arXiv:2604.06176v1 Announce Type: cross Abstract: We present an empirical study of embedding-based retrieval under realistic conversational settings, where queries are short, dialogue-like, and weakly specified, and retrieval corpora contain structured conversational artifacts. Focusing on Qwen3-emb

arxivApr 9

Gemma 4, Phi-4, and Qwen3: Accuracy-Efficiency Tradeoffs in Dense and MoE Reasoning Language Models

arXiv:2604.07035v1 Announce Type: new Abstract: Mixture-of-experts (MoE) language models are often expected to offer better quality-efficiency tradeoffs than dense models because only a subset of parameters is activated per token, but the practical value of that advantage depends on end-to-end behav

huggingfaceSep 29

Accelerating Qwen3-8B Agent on Intel® Core™ Ultra with Depth-Pruned Draft Models

arxivMay 25bullish

AGZO: Activation-Guided Zeroth-Order Optimization for LLM Fine-Tuning

arXiv:2601.17261v4 Announce Type: replace Abstract: Zeroth-Order (ZO) optimization has emerged as a promising solution for fine-tuning LLMs under strict memory constraints, as it avoids the prohibitive memory cost of storing activations for backpropagation. However, existing ZO methods typically emp

#optimization#llms#fine-tuning