Model Detail
Prompt-Guard-86M
—Prompt-Guard-86M is a large language model with 139M parameters released by Meta. The model is registered under the text-classification pipeline tag on Hugging Face, released under the llama3.1 license.
Prompt-Guard-86M ships with 139M parameters. Access is gated on Hugging Face under the llama3.1 license, which means a manual approval step before weights can be downloaded.
Prompt-Guard-86M is best fit for general-purpose chat and instruction-following workloads. Treat this as a starting matrix rather than a benchmark verdict — the right deployment usually depends on the specific evaluation suite that mirrors your workload.
Extracting Forgotten Prompts from Targeted Unlearned Models
arXiv:2609.03662v1 Announce Type: new Abstract: Recent unlearning methods (e.g. NPO, DPO, LUNAR) make use of refusal alignment to suppress forgotten data. However, it has been shown that refusal responses might leave traces of unlearning, and recent attacks have been able to successfully recover som
Analysis of Prompt Engineering for Drug Toxicity Prediction
arXiv:2609.03635v1 Announce Type: new Abstract: Clinical trials in the UK can cost up to {\pounds}1.3 million, with approximately 90% drug failure rate. Toxicity is a major contributing factor in drug failure. Testing is time and cost intensive. In recent years, the use of artificial intelligence ha
Verify Before You Distill: Prompt-Level Teacher Gating for On-Policy Distillation
arXiv:2609.02998v1 Announce Type: cross Abstract: On-policy distillation (OPD) accelerates post-training by providing dense token-level supervision from a frozen teacher on the student's own rollouts. Vanilla OPD applies this supervision uniformly across prompts, without checking whether the teacher
Ask Twice, Look Twice: Prompt Echoing Resolves the Question-First Paradox in Vision-Language Models
arXiv:2607.15565v2 Announce Type: replace-cross Abstract: Where should the question go in a vision-language model (VLM) prompt: before the image or after it? Intuition says before: knowing what is asked should tell the model where to look. Yet across visual question answering benchmarks, question-fi
Detecting Conversational Mental Manipulation with Intent-Aware Prompting
arXiv:2412.08414v2 Announce Type: replace Abstract: Mental manipulation severely undermines mental wellness by covertly and negatively distorting decision-making. While there is an increasing interest in mental health care within the natural language processing community, progress in tackling manipu
The Analyst in the Prompt: Role, Retrieval, and Memory Biases in LLM Financial Analysis
arXiv:2609.03218v1 Announce Type: new Abstract: Large Language Models (LLMs) increasingly use user context such as memory, profiles, and role prompts to personalize their responses. This personalization can affect evidence-based judgment: the same evidence may lead to different conclusions under dif