Skip to content
arXiv cs.CL · Papers

Gavel: Agent Meets Checklist for Evaluating LLMs on Long-Context Legal Summarization

arXiv:2601.04424v2 Announce Type: replace Abstract: Large language models (LLMs) now support contexts of up to 1M tokens, but their strengths and weaknesses on complex long-context tasks remain unclear. To study this, we focus on multi-document legal case summarization, where a single case often spans many documents ex