Science Edge Evaluation: New Benchmark Reveals LLM Limits in Lab Science
A new standard called Science Edge Evaluation (SEE) has been launched to evaluate the effectiveness of large language models (LLMs) in facilitating actual laboratory science. This benchmark, outlined in a paper on arXiv (arXiv:2608.06931), features expert-curated questions based on peer-reviewed studies and experimental practices in chemistry, biology, and materials science. An assessment of 19 multimodal large language models (MLLMs) showed that the top-performing model only reached 48.7% accuracy, highlighting considerable potential for enhancement. Interestingly, general-purpose models outperformed those specialized in science on average. In the visual-agent assessment, employing tools raised the highest accuracy to 52.7%, yet the paper emphasizes that additional information does not guarantee dependable scientific reasoning. The main challenge is whether models can effectively handle tool-generated data within the context of original experimental findings. These results reveal the existing limitations of LLMs in intricate scientific research and indicate the necessity for ongoing investigation.
Key facts
- Science Edge Evaluation (SEE) is a multimodal benchmark for LLMs in scientific discovery.
- The benchmark includes expert-curated questions in chemistry, biology, and materials science.
- 19 multimodal large language models (MLLMs) were evaluated.
- The best-performing model achieved 48.7% accuracy.
- General-purpose models outperformed science-specialized models on average.
- Tool use increased the best accuracy to 52.7% in visual-agent evaluation.
- More information does not necessarily lead to reliable scientific reasoning.
- The paper is available on arXiv with ID 2608.06931.
Entities
Institutions
- arXiv