ARTFEED — Contemporary Art Intelligence

Human-in-the-Loop Corpus for LLM Simplification of Scientific Summaries

ai-technology · 2026-07-29

A team of researchers has created a human-in-the-loop process to enhance the clarity of scientific summaries through large language models (LLMs), tackling the issue of understanding across different disciplines. Utilizing SciSummNet as the foundational dataset and GPT-4o-mini for initial simplifications, the study unfolds in two phases. In Phase 1, STEM readers from fields other than computer science pinpoint challenging sentences and assess both the original and GPT-simplified summaries for clarity, naturalness, and simplicity. Phase 2 sees computer science specialists crafting expert-edited reference simplifications informed by this feedback. The final corpus comprises human evaluations and automated assessment results, with Phase 1 showing a strong preference for GPT-generated summaries in terms of clarity, ultimately aiming to enhance accessibility for non-specialists and foster interdisciplinary research.

Key facts

  • The study presents a human-in-the-loop workflow for LLM-based simplification of scientific texts.
  • SciSummNet is used as the source corpus.
  • GPT-4o-mini generates baseline simplifications.
  • Phase 1 involves readers from STEM fields outside computer science.
  • Readers compare original and GPT-simplified summaries on comprehensibility, naturalness, and simplicity.
  • Phase 2 involves computer science experts creating expert-edited reference simplifications.
  • The corpus includes human judgments and automatic evaluation results.
  • Phase 1 judgments show a preference for GPT-generated summaries in comprehensibility.

Entities

Institutions

  • SciSummNet

Sources