arxiv
PublishedJune 1, 2026 at 4:00 AM
—neutral
IsoCLIP: Decomposing CLIP Projectors for Efficient Intra-modal Alignment
Publisher summary· verbatim
arXiv:2603.19862v2 Announce Type: replace-cross Abstract: Vision-Language Models like CLIP are extensively used for inter-modal tasks which involve both visual and text modalities. However, when the individual modality encoders are applied to inherently intra-modal tasks like image-to-image retrieva
Stay posted· Newsletter
A 5-min weekly brief — top movers, price watch, story of the week.
Discussion
No replies yet. Be first.
Related coverage
More from ARXIV
arxivCost-Governed RAG: Unified Per-Tenant Cost Attribution Across Retrieval and Generation in Multi-Tenant LLM Systems13harxivSelf-Evolving In-Context Learning for Direct Pilot-to-Beamformer Design in MU-MISO Systems13harxivLearning to Discretize: Diffusion-Based Adaptive Mesh with Spectral Guidance13harxivThe Benjamini--Hochberg Procedure Can Fail to Control the FDR for Correlated Two-Sided Gaussian Tests13hThe Bubble Brief
WEEKLYRead AI insights every Tuesday — top movers, new releases, story of the week.
Originally published on arxiv ↗