Skip to content
arXiv cs.LG · Papers

Beyond Shapley: An Influence-Based Data Auditing Pipeline for LLM Alignment and Evaluation

arXiv:2607.22766v1 Announce Type: new Abstract: The alignment of Large Language Models (LLMs) is increasingly bottlenecked by data quality. As datasets scale, massive preference and instruction-tuning corpora inevitably accumulate hidden structural contradictions, safety risks, and systemic human annotation errors. Sta