Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness
Xilun Chen, Zhaleh Feizollahi, Ross Goodwin, Seungwhan Moon, Scott Yih, Pinar Donmez, Babak Damavandi, Luna Dong
Read on arXiv →Key claim
New framework improves evaluation of AI-generated content.
In plain English
Imagine you're developing an AI that generates long-form content, like articles or reports. You want to ensure that not only are the facts it presents correct, but also that it covers all necessary information comprehensively. Current methods often focus on whether the claims made are accurate, but they fall short in assessing whether the response includes all relevant facts. This gap can lead to incomplete or misleading outputs, which is problematic in real-world applications where thoroughness matters. This issue is known as factual completeness, and measuring it is complex because it involves understanding the relationships and hierarchies among various facts rather than just checking them off a list. The existing approach, called the decompose-search-verify pipeline, primarily checks for precision but misses the bigger picture of what a complete answer should entail. This is what's called a failure mode in evaluation. To address this, the authors propose a two-level meta-rubric framework that organizes and prioritizes the required content for long-form generation. They introduce Gamut, a benchmark designed to evaluate factual completeness by creating a structured rubric that can be easily scored by machine learning models. This framework is modality-agnostic, meaning it can be applied to various types of content, and it has been rigorously tested with a diverse set of questions and expert validation. For builders, this means a more reliable way to assess the quality of AI-generated content, ensuring that it not only gets the facts right but also provides a complete and coherent narrative.
The introduction of a two-level meta-rubric framework for evaluating factual completeness is a meaningful extension to existing evaluation methods.
The benchmark is validated with expert human annotators and tested across multiple models, providing solid reliability.
Deep reliability assessment
The methodology supports a structured approach to evaluating factual completeness by using a two-level meta-rubric framework, but it may overclaim the ease of automatic evaluation given the complexity of open-ended questions.
Reproducibility
yes, the paper provides both open source code and dataset links.
Key figure
Figure 2 shows the topic composition of Gamut questions within each of the 10 domains, illustrating the diversity and tailored nature of the questions.
