ARTFEED — Contemporary Art Intelligence

Benchmarking LLMs in Cosmetic Chemistry: Accuracy and Reliability Study

ai-technology · 2026-08-18

A new study from arXiv (2608.14631) benchmarks 14 large language models (LLMs) on their knowledge of cosmetic chemistry and skin health. The research, which disabled web search to test internalized knowledge, found overall poor performance, with significant deficits in quantitative reasoning and structural identification. While models handled general skincare questions reasonably, their responses lacked the technical depth needed for informed consumer decisions. The study highlights risks of AI chatbots providing authoritative-sounding but technically flawed advice. This research is relevant to the intersection of AI and consumer health, though it does not directly involve art.

Key facts

  • Study benchmarks 14 LLMs on cosmetic chemistry topics.
  • Web search disabled to assess internalized knowledge.
  • Overall performance poor, especially in quantitative reasoning and structural identification.
  • Models handled general skincare questions reasonably.
  • Responses lacked technical depth for informed decisions.
  • AI outputs can sound authoritative but contain technical errors.
  • Study published on arXiv with ID 2608.14631.
  • Research highlights risks of AI chatbots in skincare advice.

Entities

Institutions

  • arXiv

Sources