arxivJul 23
arXiv:2607.19064v2 Announce Type: replace-cross Abstract: Large-scale visual generators are increasingly capable but costly to train, fine-tune, and deploy. We introduce Mage-Flow, a compact 4B-scale generative stack for efficient text-to-image generation and instruction-based image editing. The sta
arxiv1d ago
arXiv:2609.03391v1 Announce Type: cross Abstract: Contrastive language-image learning (CLIP) has become a key paradigm for remote sensing vision-language understanding. However, existing remote sensing contrastive learning methods are mostly built on RGB-oriented CLIP architectures, making it diffic
arxiv1d ago
arXiv:2609.03796v1 Announce Type: cross Abstract: We introduce LLaDA-Image, a unified framework that pairs a 6B Diffusion Transformer (DiT) trained from scratch with a frozen vision-language understanding module built on the LLaDA2.0-Mini diffusion language model backbone. Instead of relying heavily
arxiv1d ago
arXiv:2609.03829v1 Announce Type: cross Abstract: Few-shot fine-grained image classification (FSFGIC) aims to classify similar images with limited labeled examples. This work highlights the critical yet underutilized role of phase information in capturing structural relationships within an image. Th
arxiv1d ago
arXiv:2608.26648v2 Announce Type: replace-cross Abstract: Many synthetic-image detectors produce accurate predictions but offer limited insight into how those decisions are formed. This paper introduces Hierarchical Channel Stacking (HCS), a compact framework for AI-generated image detection that co
arxiv1d ago
arXiv:2609.03677v1 Announce Type: cross Abstract: Understanding the composition of large-scale autonomous driving datasets is essential for safety, robustness, and reliable operation across domains. For example, domain shift between locations could lead to the operating environment being misaligned
arxiv1d ago
arXiv:2609.03931v1 Announce Type: cross Abstract: Generating complete 3D scenes from sparse, unconstrained views is a fundamental challenge in 3D vision which requires reasoning beyond observed content while remaining computationally tractable. Existing feed-forward reconstruction methods are inhere
arxiv1d ago
arXiv:2609.03973v1 Announce Type: new Abstract: An image editor may satisfy every regional plausibility constraint separately even when no single latent explanation fits the complete output. We formalize this local-to-global failure using a common witness grade and witness nerve. The framework separ
arxiv2d ago
arXiv:2609.02529v1 Announce Type: cross Abstract: Image enhancement and restoration have become standard back-end operations on short-video and social media platforms to boost UGC visual experience. Yet these processes inevitably introduce visual anomalies--especially in faces, texts, and textures--
arxiv2d ago
arXiv:2609.02282v1 Announce Type: cross Abstract: Mineral image classification is important for geological exploration and resource development, but it remains challenging due to substantial intra-class variations in appearance and high inter-class visual similarity. Multi-cognitive Visual Adapter (
arxiv2d ago
arXiv:2609.01168v2 Announce Type: replace Abstract: Text-to-image (T2I) models remain vulnerable to jailbreak attacks that elicit Not-Safe-For-Work (NSFW) content, despite increasingly being guarded by heterogeneous, multi-layer safety stacks combining text filters, image classifiers, and cross-moda
arxiv2d ago
arXiv:2609.02502v1 Announce Type: cross Abstract: Text-to-image (T2I) models have achieved remarkable success at faithfully rendering specified objects and attributes, yet their ability to produce visual metaphors, images that convey abstract ideas by combining elements from two distinct domains, re
arxiv2d ago
arXiv:2505.00759v3 Announce Type: replace-cross Abstract: The steady improvements of text-to-image (T2I) generative models lead to slow deprecation of automatic evaluation benchmarks that rely on static datasets, motivating researchers to seek alternative ways to evaluate T2I progress. We present Mu
arxiv2d ago
arXiv:2609.02285v1 Announce Type: cross Abstract: A long-held intuition in interpretability research is that representational entanglement, the sharing of structure between knowledge domains in a neural network, makes unlearning harder. While the intuition is widespread, it has never been directly t
arxiv2d ago
arXiv:2609.02247v1 Announce Type: cross Abstract: Semi-supervised learning has shown great potential for reducing annotation costs in medical image segmentation. However, most existing methods mainly exploit unlabeled data through prediction-level consistency, while the reliability of internal featu
arxiv2d ago
arXiv:2609.01987v1 Announce Type: cross Abstract: Patient exams in the cancer diagnosis and staging process typically generate several whole slide images (WSIs). One of the initial steps in training models on WSI data is identifying one or a few slides containing tumor or other diagnostic biomarkers
arxiv2d ago
arXiv:2509.06554v2 Announce Type: replace-cross Abstract: In subjective image and video quality assessment, observers rate or compare selected stimuli. Before calculating mean opinion scores (MOSs), unreliable ratings should be identified and handled as outliers. Several outlier-detection methods ar
arxiv2d ago
arXiv:2609.00046v2 Announce Type: replace-cross Abstract: Rapid and reliable disaster mapping of impacted areas, damaged infrastructure, and affected populations is essential for emergency response and recovery. However, existing AI-based approaches often require extensive manual annotation, lack cr
arxiv2d ago
arXiv:2609.02126v1 Announce Type: new Abstract: Estimating physical parameters from scientific images is a common inverse problem in materials characterization that often relies on expensive physics-based simulations. In electron microscopy, specimen thickness and crystal mistilt are critical parame
arxiv2d ago
arXiv:2609.02194v1 Announce Type: new Abstract: Classical constitutive modeling of path-dependent inelastic materials relies on internal state variables whose evolution equations must be postulated based on domain knowledge and calibrated against experimental data. However, in many practical setting
arxiv2d ago
arXiv:2609.02016v1 Announce Type: cross Abstract: Image super-resolution, which aims to reconstruct high-resolution images from their low-resolution observations, is fundamental to medical imaging, remote sensing, surveillance, microscopy, and scientific visualization. Traditional model-based method
arxiv2d ago
arXiv:2511.12018v2 Announce Type: replace-cross Abstract: Traffic safety analysis at signalized intersections is essential for reducing vehicle and pedestrian collisions, yet traditional crash-based studies are limited by data sparsity and reporting latency. This paper presents a multi-camera comput
arxiv2d ago
arXiv:2609.02004v1 Announce Type: cross Abstract: Accurate segmentation of polyps and skin lesions is pivotal for clinical diagnosis, yet existing methods struggle with low contrast, ambiguous boundaries, and cross-domain distribution discrepancies. Discriminative networks and most diffusion-based s
arxiv2d ago
arXiv:2606.17698v3 Announce Type: replace Abstract: As LLM-based shopping agents enter production, existing benchmarks fail to capture how a shopper's requirements arrive: stated implicitly in the query, recorded in a profile, or revealed only when the right question is asked. Benchmarks that expose
arxiv2d ago
arXiv:2609.02207v1 Announce Type: cross Abstract: Real-world personally identifiable information (PII) redaction often operates on document images---scans, screenshots, and PDF renderings---where OCR errors, layout structure, and visual noise determine whether sensitive information is actually remov
arxiv2d ago
arXiv:2606.13081v2 Announce Type: replace-cross Abstract: Emotion significantly influences cognition, enhancing memory and learning under certain conditions. Drawing on this principle, emotion-augmented deep learning investigates how affective states can improve neural network architectures and lear
arxiv3d ago
arXiv:2604.16729v2 Announce Type: replace-cross Abstract: State-of-the-art large language models (LLMs) show high performance in general visual question answering. However, a fundamental limitation remains: current architectures lack the native 3D spatial reasoning required to directly analyze volum
arxiv3d ago
arXiv:2609.00629v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) achieve strong performance on image-grounded text generation and visual question answering. However, it remains difficult for them to comprehensively and accurately describe the factual relations among the entities
arxiv3d ago
arXiv:2609.00709v1 Announce Type: cross Abstract: Large Vision-Language Models produce fluent image descriptions but offer limited semantic control: users cannot reliably specify whether captions should emphasize attributes, relations, or particular image regions. We present Fine-grained Captioning
arxiv3d ago
arXiv:2609.00282v1 Announce Type: cross Abstract: Motor imagery (MI) electroencephalography (EEG) decoding could support post-stroke rehabilitation, but models developed on healthy cohorts may not transfer reliably to pathological EEG. We evaluated whether Low-Rank Adaptation (LoRA) can efficiently
arxiv3d ago
arXiv:2609.01310v1 Announce Type: cross Abstract: Medical image segmentation remains difficult to scale because high-performing methods typically rely on dense expert annotations and task-specific training. We introduce GazeRefine, a training-free framework that uses gaze as an inference-time prompt
arxiv3d ago
arXiv:2608.30621v2 Announce Type: replace-cross Abstract: Collecting natural-language referring expressions along with region annotations, such as masks or boxes, is a major bottleneck in visual grounding (VG), as annotators must write descriptions that distinguish target regions from visually simil
arxiv3d ago
arXiv:2609.01456v1 Announce Type: cross Abstract: Composed image retrieval (CIR) retrieves a target image from a reference image and a text modification. This paper studies metadata-available CIR reranking, where a fixed CIR model first returns a candidate pool and gallery metadata is then used for
arxiv3d ago
arXiv:2601.14121v2 Announce Type: replace Abstract: Identifying when and where a news image was taken is crucial for journalists and forensic experts to produce credible stories and debunk misinformation. While many existing methods rely on reverse image search (RIS) engines, these tools often fail
arxiv3d ago
arXiv:2609.01249v1 Announce Type: cross Abstract: Invisible watermarks are typically evaluated against predefined perturbations such as compression, blur, noise, cropping, and denoising. Public foundation image models expose a distinct threat: an attacker can submit a watermarked image with a single
arxiv3d ago
arXiv:2511.16717v3 Announce Type: replace-cross Abstract: Neutron imaging is essential for diagnosing and optimizing inertial confinement fusion implosions at the National Ignition Facility. Due to the required 10-micrometer resolution, however, neutron image require image reconstruction using itera
arxiv3d ago
arXiv:2609.00003v1 Announce Type: new Abstract: Machine unlearning studies the removal of knowledge from an AI model, making the system forget a concept it previously learned. Despite rapid progress in generative machine unlearning, the unintended degradation of semantically related concepts that sh
arxiv3d ago
arXiv:2609.00685v1 Announce Type: cross Abstract: Article-level news stance detection aims to identify the perspective of news articles toward social issues. Despite advances in stance detection and its importance for trustworthy media environments, news articles pose distinct challenges because the
arxiv3d ago
arXiv:2609.00708v1 Announce Type: cross Abstract: Differentially private (DP) synthesis has been extensively studied for tabular and image data separately, yet many real-world datasets contain images paired with multivariate tabular records. Synthesizing such data is particularly challenging under D
arxiv3d ago
arXiv:2609.00924v1 Announce Type: cross Abstract: Monocular videos record 3D scenes as sequences of 2D image-plane projections, obscuring depth and spatial relationships. Multi-object trackers localize and associate objects primarily using appearance and geometry observed only in the image plane, in
arxiv3d ago
arXiv:2609.01587v1 Announce Type: cross Abstract: Post-training quantization (PTQ) is widely used to reduce the cost of serving large language models (LLMs), but its accuracy cost is uneven and is often tuned per model. We study where quantization damage occurs and how to allocate a small additional
google4d ago
<img src="https://storage.googleapis.com/gweb-uniblog-publish-prod/images/GooglePics_Hero.max-600x600.format-webp.webp">Built on our latest Nano Banana model, Google Pics — our image creation and editing tool — is now available.
arxiv4d ago
arXiv:2608.29210v1 Announce Type: new Abstract: Text-to-Image (T2I) models have recently achieved impressive visual fidelity, yet their evaluation remains constrained by benchmarks that are often difficult to interpret and insufficiently diagnostic. Existing skill-based evaluations tend to overlook
arxiv4d ago
arXiv:2608.29589v1 Announce Type: new Abstract: Text-to-image (T2I) safety guardrails fail to generalize equitably to non-standard dialects. Evaluating 23,080 paired prompts across five English dialects, we formalize this failure as the dialect penalty, where filters trigger based on linguistic surf
arxiv4d ago
arXiv:2608.29160v1 Announce Type: cross Abstract: We aim to improve frozen flow-matching image generators by adding inference computation inside the denoiser, without changing model weights or the outer sampler. Existing generators usually spend extra test-time computation by increasing the number o
arxiv4d ago
arXiv:2608.30404v1 Announce Type: cross Abstract: Accurate segmentation of the coronary vessel lumen is a prerequisite for quantitative assessment of atherosclerotic plaque and perivascular adipose tissue in coronary computed tomography angiography (CCTA). Cardiologists rely on semi-automated method
arxiv4d ago
arXiv:2608.29847v1 Announce Type: cross Abstract: Text-to-image models learn associations between concepts - in the case of this paper, people's professions, which we refer to as roles - and visual attributes. These associations can underpin many observed forms of stereotypical bias. A key open ques
arxiv4d ago
arXiv:2602.08136v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) are now a core part of modern AI. Recent work proposed several visual jailbreak attacks using single/ holistic images. However, contemporary VLMs demonstrate strong robustness against such attacks due to extensiv