·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning6h◆Subagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks6h◆Distribution-Consistent Inference for Dynamic Sparse Mixture-of-Experts6h◆In RAG We Trust? Measuring Robustness of Retrieval-Augmented Generation Under Document Poisoning6h◆FrontierChallenge: Evaluating Scientific Workflow Completion6h◆IBIB: A Protocol for Measuring Enterprise AI Systems by Serving Route, Not Model Identifier6h◆AgenticGen: Reward-Guided Agentic Video Generation for Advertising6h◆Strangers to Themselves: What Language Models Say About Themselves Is Generic6h◆Omni Interaction Agent Technical Report6h◆CoGReV: A Confidence-Gated Post-Hoc Non-Monotonic Belief Revision Framework for Phishing Website Classification6h◆Bit-Flip Attacks on Vision-Language-Action Models: Action-Decoding Architecture Shapes the Vulnerability6h◆Can Foundation Models Moderate Online Content? Evaluating Instruction- vs. Example-Driven Policy Operationalization6h◆RiLM: Parameter-Efficient Language Modeling via Geodesic Decoding6h◆Tracing Computation Density in LLMs6h◆Cultural Binding Heads in Language Models6h◆Direct Diversity Optimization for Diverse Successful Trajectories in Preference Post-Training6h◆Left-Branching Transformers Excel at Right-Branching Languages: Data Shapes Word Order Preferences in Language Models6h◆Palmyra x6 Technical Report: An Agentic, Tool-Use Model Post-Trained via Anchored Supervised Fine-Tuning6h◆'Ghaib in Translation' aka Unseen Harm: Measuring Cross-Script Safety Inconsistency with 'Missed-in-Urdu' Scores in LLM Hate Speech Detection6h◆DexterSQL: Deep Schema Exploration and Rule-based Correction for Text-to-SQL Generation6h◆Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning6h◆Subagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks6h◆Distribution-Consistent Inference for Dynamic Sparse Mixture-of-Experts6h◆In RAG We Trust? Measuring Robustness of Retrieval-Augmented Generation Under Document Poisoning6h◆FrontierChallenge: Evaluating Scientific Workflow Completion6h◆IBIB: A Protocol for Measuring Enterprise AI Systems by Serving Route, Not Model Identifier6h◆AgenticGen: Reward-Guided Agentic Video Generation for Advertising6h◆Strangers to Themselves: What Language Models Say About Themselves Is Generic6h◆Omni Interaction Agent Technical Report6h◆CoGReV: A Confidence-Gated Post-Hoc Non-Monotonic Belief Revision Framework for Phishing Website Classification6h◆Bit-Flip Attacks on Vision-Language-Action Models: Action-Decoding Architecture Shapes the Vulnerability6h◆Can Foundation Models Moderate Online Content? Evaluating Instruction- vs. Example-Driven Policy Operationalization6h◆RiLM: Parameter-Efficient Language Modeling via Geodesic Decoding6h◆Tracing Computation Density in LLMs6h◆Cultural Binding Heads in Language Models6h◆Direct Diversity Optimization for Diverse Successful Trajectories in Preference Post-Training6h◆Left-Branching Transformers Excel at Right-Branching Languages: Data Shapes Word Order Preferences in Language Models6h◆Palmyra x6 Technical Report: An Agentic, Tool-Use Model Post-Trained via Anchored Supervised Fine-Tuning6h◆'Ghaib in Translation' aka Unseen Harm: Measuring Cross-Script Safety Inconsistency with 'Missed-in-Urdu' Scores in LLM Hate Speech Detection6h◆DexterSQL: Deep Schema Exploration and Rule-based Correction for Text-to-SQL Generation6h◆
DataBubble·

Model Detail

Qwen logo

Qwen3.5-2B-Base

—
Provider: QwenCategory: multimodalPipeline: image-text-to-textParameters: 2B
DB Score
38.5
Downloads
15K
Likes
53
Day
+0.0%
Week
+0.0%
Month
+0.0%
Overview

Qwen3.5-2B-Base is a multimodal model with 2B parameters released by Qwen. The model is registered under the image-text-to-text pipeline tag on Hugging Face.

Technical

Qwen3.5-2B-Base ships with 2B parameters.

Use Cases

Qwen3.5-2B-Base is best fit for mixed text-and-image reasoning tasks such as document understanding. Treat this as a starting matrix rather than a benchmark verdict — the right deployment usually depends on the specific evaluation suite that mirrors your workload.

Download History
Research Paper
arXiv: 2309.16609→
Model Info
Citations3,938 (424 influential)
Recent newsView all news →
Related News
arxiv70d ago

TUDUM: A Turkish-Thinking Reasoning Pipeline for Qwen3.5-27B

arXiv:2607.01927v1 Announce Type: cross Abstract: This paper presents TUDUM (T\"urk\c{c}e D\"u\c{s}\"unen \"Uretken Model), a project pipeline for adapting a Qwen-family 27B thinking model toward Turkish reasoning. The central problem is not only to answer Turkish prompts in Turkish, but to make the

arxivneutral119d ago

Procedural-skill SFT across capacity tiers: A W-Shaped pre-SFT Trajectory and Regime-Asymmetric Mechanism on 0.8B-4B Qwen3.5 Models

arXiv:2605.11907v2 Announce Type: replace Abstract: We measure procedural-skill SFT contribution across three Qwen3.5 dense scales (0.8B, 2B, 4B) on a 200-task / 40-skill holdout, with Claude Haiku 4.5 as a frontier reference. The corpus is 353 rows of (task + procedural-skill block, Opus chain-of-t

arxiv142d ago

Qwen3.5-Omni Technical Report

arXiv:2604.15804v2 Announce Type: replace Abstract: In this work, we present Qwen3.5-Omni, the latest advancement in the Qwen-Omni model family. Representing a significant evolution over its predecessor, Qwen3.5-Omni scales to hundreds of billions of parameters and supports a 256k context length. By

Related Models
Qwen logo
Qwen3-VL-2B-Instruct
Qwen · 22.5M downloads
Qwen logo
Qwen3-0.6B
Qwen · 21.0M downloads
Qwen logo
Qwen3-VL-2B-Instruct
Qwen · 22.5M downloads
Qwen logo
Qwen3.5-9B
Qwen · 10.9M downloads
HomeModelsNews