AGC-Bench: Measuring Artificial General Creativity
Roger Beaty, Vijeta Deshpande, Clin K. Y. Lai, Anna Attuch, Namrata Shivagunde, Swastik Roy, Rajkumar Pujari, Paul V. DiStefano, Sherin Muckatira, Claire E. Stevenson, Mikhail Gronas, Anna Rumshisky
Read on arXiv →Key claim
AGC-Bench reveals distinct creative strengths in LLMs.
In plain English
Imagine you're trying to figure out how creative AI can be. Traditionally, creativity has been hard to measure, especially since it can look different in writing, science, or art. People have tried various methods, but they often fall short because they don't capture the nuances of creativity across different domains. This is where the new AGC-Bench comes in. It’s a benchmark designed specifically to evaluate AI creativity, built from a thorough review of existing literature and incorporating a wide range of tasks like brainstorming and humor. The problem with previous approaches was that they often treated creativity as a one-size-fits-all concept, missing the unique strengths of different AI models. AGC-Bench addresses this by providing a structured way to assess creativity, allowing for a more nuanced understanding of how different models perform in various creative tasks. One of the standout findings is that LLMs can be more creative in some areas, like writing, compared to others, like scientific ideation. This means that if you're building applications that rely on AI creativity, you can now use AGC-Bench to better understand which models might excel in specific creative tasks, leading to more effective and tailored AI solutions.
The introduction of AGC-Bench and AGC-Judge provides a new framework for assessing AI creativity, extending existing benchmarks significantly.
The study employs extensive datasets and rigorous psychometric methods, ensuring robust validation of claims.
Deep reliability assessment
The methodology supports the claim that a single creativity factor 'c' can be identified across LLMs, but the generalization of this factor to all aspects of creativity may be overclaimed due to the limited scope of the benchmarks used.
Reproducibility
yes, the paper mentions the release of AGC-Bench with a public leaderboard, the onboarding harness, AGC-Judge weights, and human comparison data as open infrastructure.
Key figure
Figure 1 shows the overall and per-domain rankings of the top 15 models on the AGC-Bench composite and the top 5 within each of the six domains.
