arxiv
PublishedJune 17, 2026 at 4:00 AM
▼bearish
In-Context Environments Induce Evaluation-Awareness in Language Models
Publisher summary· verbatim
arXiv:2603.03824v2 Announce Type: replace Abstract: Humans often become more self-aware under threat, yet can lose self-awareness when absorbed in a task; we hypothesize that language models exhibit environment-dependent \textit{evaluation awareness}. This raises concerns that models could strategic
Models mentioned
01Stay posted· Newsletter
A 5-min weekly brief — top movers, price watch, story of the week.
Discussion
No replies yet. Be first.
Related coverage
More from ARXIV
arxivReverso: Efficient Time Series Foundation Models for Zero-shot Forecasting5harxivCoherent Without Grounding, Grounded Without Success: Observability and Epistemic Failure5harxivAutoWorld: Learning Multi-Agent Traffic Simulation with Self-Supervised World Models5harxivSheaf-Laplacian Obstruction and Projection Hardness for Cross-Modal Compatibility on a Modality-Independent Site5hThe Bubble Brief
WEEKLYRead adversarial-attacks insights every Tuesday — top movers, new releases, story of the week.
Originally published on arxiv ↗