Skip to content
arXiv cs.CL · Papers

Evaluating RAG Metrics in Applied Contexts: An Experiment, Its Findings and Its Limitations

arXiv:2607.07302v1 Announce Type: new Abstract: This paper reports an empirical study evaluating the relevance of several RAG metrics. The experiment is based on a question-answering dataset created by human annotators from business data. The generated responses and retrieved spans of a RAG system are scored using eval