arxiv
PublishedMay 21, 2026 at 4:00 AM
SEED: Targeted Data Selection by Weighted Independent Set
Publisher summary· verbatim
arXiv:2605.15691v2 Announce Type: replace Abstract: Data selection seeks to identify a compact yet informative subset from large-scale training corpora, balancing sample quality against collection diversity. We formulate this problem as a Weighted Independent Set (WIS) on a similarity graph, where n
Stay posted· Newsletter
A 5-min weekly brief — top movers, price watch, story of the week.
Discussion
No replies yet. Be first.
Related coverage
More from ARXIV
arxivLCAi: Life Cycle Assessment with big data fusion and retrieval-augmented generation-assisted interpretation1darxivHelpfulness Hurts: Domain-Dependent Degradation of Mid-Trained Compassion Values Under Post-Training1darxivInstruction Bleed: Cross-Module Interference in Prompt-Composed Agentic Systems1darxivTAVR-VLM: Risk-Conditioned Causal Grounding for Hallucination-Resistant Report Generation1dThe Bubble Brief
WEEKLYRead AI insights every Tuesday — top movers, new releases, story of the week.
Originally published on arxiv ↗