Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation
Yifan Zhou, Qihao Yang, Yan Li, Donggang Li, Xiru Hu, Hokin Deng, Ziyang Gong, Xuanyi Zhou, Huacan Wang, Xiangchao Yan, Wanghan Xu, Wenlong Zhang, Shaofeng Zhang, Yue Zhou, Yifan Yang, Zhihang Zhong, Xue Yang
Read on arXiv →Key claim
AI struggles with scientific lineage reasoning, achieving only 27.3% accuracy.
In plain English
Imagine you're trying to build an AI that can not only generate new scientific ideas but also understand how those ideas relate to existing research. This is a tough problem because scientific knowledge doesn't just appear out of nowhere; it builds on previous work, much like how species evolve over time. Researchers often struggle to track these connections, leading to gaps in understanding how new ideas are formed and validated. This is what's called lineage reasoning, and it’s crucial for ensuring that AI-generated ideas are coherent and valuable. However, current benchmarks don't effectively measure this capability, leaving a big question mark over how well AI can mimic human-like reasoning in science.
To address this, the authors created IdeaGene-Bench, a new benchmark that organizes scientific papers and proposals into a structured format that reflects their evolutionary relationships. They represent each idea as an 'Idea Genome' and track how these ideas evolve through processes like inheritance and mutation. The benchmark includes thousands of examples across various scientific domains and offers two main evaluation methods: one for reasoning about lineage and another for generating new ideas based on that lineage.
The results are telling: even the best AI systems tested only managed to get about 27.3% of lineage reasoning tasks correct. This indicates a significant bottleneck in how these models understand and generate ideas based on existing knowledge. For anyone building AI systems aimed at scientific research, this benchmark provides a new way to evaluate and improve their models, but it also highlights the need for further advancements in lineage reasoning capabilities.
The benchmark introduces a novel framework for lineage reasoning in AI, extending existing methods into a new domain.
The experiments are well-structured with a clear evaluation methodology, though some claims could be more conservatively stated.
Deep reliability assessment
The methodology supports evaluating scientific lineage reasoning and idea generation through the IdeaGene framework, but the claim of exposing a compositional bottleneck in LLM-based scientists may be overclaimed without broader validation.
Reproducibility
yes, the project page is provided for accessing the benchmark and related resources.
Key figure
Figure 1 illustrates the transition from a paper-centric search to a genome-centric lineage approach, highlighting the extraction and alignment of Idea Genome objects.
