ASI-Bench Introduced as First Benchmark for AI Innovation and Autonomous Science
A new benchmark, ASI-Bench, has been unveiled to evaluate artificial intelligence systems on their capacity for innovative exploration and independent scientific execution. Unlike existing tests that rely on learned knowledge or extensive human guidance, ASI-Bench progressively removes methodological support within the same research project to determine how far AI can operate autonomously. Developed by more than 40 experts over 31,000 human hours, the benchmark comprises 60 project-level research tasks spanning general research domains. Its creation responds to the gap between current AI capabilities, which largely compress and apply existing human knowledge, and the demands of artificial superintelligence, which require generating new knowledge and translating ideas into verifiable results. The benchmark is described as the first to jointly assess these dual abilities and to test autonomy through staged withdrawal of human direction. Originally posted on the arXiv preprint server under reference 2608.17271, ASI-Bench represents a targeted effort to move beyond conventional AI evaluation toward measuring frontier research potential.
Key facts
- ASI-Bench is introduced as a new benchmark for AI systems.
- It evaluates innovative exploration and autonomous scientific execution.
- It is the first benchmark to jointly assess these capabilities.
- It progressively withdraws human methodological guidance within projects.
- Over 40 experts contributed to its development.
- Development cost more than 31,000 human hours.
- The benchmark contains 60 project-level research tasks.
- It was announced on arXiv with reference 2608.17271.
Entities
—