arXiv cs.LG
· Papers
Beyond Shapley: An Influence-Based Data Auditing Pipeline for LLM Alignment and Evaluation
arXiv:2607.22766v1 Announce Type: new Abstract: The alignment of Large Language Models (LLMs) is increasingly bottlenecked by data quality. As datasets scale, massive preference and instruction-tuning corpora inevitably accumulate hidden structural contradictions, safety risks, and systemic human annotation errors. Sta