Model Detail
Mage-Flow
—Mage-Flow is an image generation model with 2.1B parameters released by Microsoft. The model is registered under the text-to-image pipeline tag on Hugging Face, distributed under the permissive mit license.
Mage-Flow ships with 2.1B parameters. Total weight footprint is approximately 4.1 GB, which is the relevant figure when planning local-inference VRAM. The mit license is permissive, allowing commercial deployment and derivative work without per-seat fees, though attribution requirements still apply.
Mage-Flow is best fit for text-to-image generation and creative iteration. It is a less obvious choice for production photography pipelines that need exact reproducibility. Treat this as a starting matrix rather than a benchmark verdict — the right deployment usually depends on the specific evaluation suite that mirrors your workload.
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing
arXiv:2607.19064v2 Announce Type: replace-cross Abstract: Large-scale visual generators are increasingly capable but costly to train, fine-tune, and deploy. We introduce Mage-Flow, a compact 4B-scale generative stack for efficient text-to-image generation and instruction-based image editing. The sta
Exploring the Potential of Contrastive Language-Image Pre-training for Multi-Source Remote Sensing Data
arXiv:2609.03391v1 Announce Type: cross Abstract: Contrastive language-image learning (CLIP) has become a key paradigm for remote sensing vision-language understanding. However, existing remote sensing contrastive learning methods are mostly built on RGB-oriented CLIP architectures, making it diffic
LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes
arXiv:2609.03796v1 Announce Type: cross Abstract: We introduce LLaDA-Image, a unified framework that pairs a 6B Diffusion Transformer (DiT) trained from scratch with a frozen vision-language understanding module built on the LLaDA2.0-Mini diffusion language model backbone. Instead of relying heavily
The impact of phase information for few-shot fine-grained image classification
arXiv:2609.03829v1 Announce Type: cross Abstract: Few-shot fine-grained image classification (FSFGIC) aims to classify similar images with limited labeled examples. This work highlights the critical yet underutilized role of phase information in capturing structural relationships within an image. Th
Hierarchical Channel Stacking: A Structured Decision Framework for AI-Generated Image Detection
arXiv:2608.26648v2 Announce Type: replace-cross Abstract: Many synthetic-image detectors produce accurate predictions but offer limited insight into how those decisions are formed. This paper introduces Hierarchical Channel Stacking (HCS), a compact framework for AI-generated image detection that co
Understanding Autonomous Driving Datasets by Describing Differences between Image Subsets in Natural Language
arXiv:2609.03677v1 Announce Type: cross Abstract: Understanding the composition of large-scale autonomous driving datasets is essential for safety, robustness, and reliable operation across domains. For example, domain shift between locations could lead to the operating environment being misaligned