Model Detail
clip-vit-large-patch14
—clip-vit-large-patch14 is an AI model with 214M parameters released by OpenAI. The model is registered under the zero-shot-image-classification pipeline tag on Hugging Face.
clip-vit-large-patch14 ships with 214M parameters.
clip-vit-large-patch14 is best fit for workloads that match the zero-shot-image-classification pipeline tag. Treat this as a starting matrix rather than a benchmark verdict — the right deployment usually depends on the specific evaluation suite that mirrors your workload.
Federated LoRA Adaptation of BiomedCLIP Across Four International Chest X-Ray Cohorts
arXiv:2609.02101v1 Announce Type: cross Abstract: Federated learning (FL) lets institutions train a shared model without exchanging data, and Low-Rank Adaptation (LoRA) makes this practical at scale by communicating only compact low-rank updates. Biomedical imaging is a compelling setting for this c
Group Adaptive Clipping Policy Optimization
arXiv:2609.00444v1 Announce Type: cross Abstract: Group relative policy optimization for reinforcement learning with verifiable rewards (RLVR) typically uses a fixed importance-sampling (IS) ratio clipping boundary across all rollouts. We identify a key limitation: rare correct rollouts on harder pr
When Modality Gap Reduction Fails: Prediction-Level Hubness in CLIP
arXiv:2609.01103v1 Announce Type: new Abstract: Reducing the modality gap between image and text representations in CLIP is widely expected to improve cross-modal alignment and downstream performance. However, a smaller average image-text gap does not necessarily lead to consistent accuracy gains. W
CLIPure: Purification in Latent Space via CLIP for Adversarially Robust Zero-Shot Classification
arXiv:2502.18176v3 Announce Type: replace-cross Abstract: In this paper, we aim to build an adversarially robust zero-shot image classifier. We ground our work on CLIP, a vision-language pre-trained encoder model that can perform zero-shot classification by matching an image with text prompts ``a ph
Real-Time Video Anomaly Detection Using YOLO Pose Estimation and CLIP-Based Semantic Scoring
arXiv:2608.31074v1 Announce Type: cross Abstract: We propose a lightweight two-stage framework for real-time video anomaly detection. The first stage employs YOLO v11n-pose to detect persons and extract seventeen skeletal keypoints in a single forward pass. The second stage encodes each cropped pers
Ceiling-Clipped Acceptance Histograms Indicate Stranded Speed-up in Block-Diffusion Speculative Decoding
arXiv:2608.30427v1 Announce Type: new Abstract: Speculative decoding speeds up generation with an efficient draft model (drafter) that proposes tokens for a target model to verify in one pass, preserving the target's output distribution. High-acceptance block-diffusion drafters such as DFlash and DF