ARTFEED — Contemporary Art Intelligence

SERUM: AI Framework Extracts User Behavior Models from Screen Recordings

ai-technology · 2026-08-03

A recent study presents SERUM, a multi-pass framework aimed at deriving finite-state behavioral models from unstructured egocentric video, particularly screen recordings. This system employs hierarchical VLM annotation to switch between activity recognition and intent inference, enhancing labels through gathered context to minimize hallucination and temporal confusion. It combines similar states using sentence embeddings and thresholds calibrated by humans to create a concise taxonomy. The framework's performance is assessed by applying first-order Markov models to the generated label sequences and evaluating predictive accuracy. The research can be found on arXiv under the identifier 2607.29181.

Key facts

  • SERUM is a multi-pass framework for extracting behavioral models from screen recordings.
  • It uses hierarchical VLM annotation and a sliding window approach.
  • The framework alternates between activity-recognition and intent-inference passes.
  • Labels are refined using accumulated prior context to reduce hallucination and temporal conflation.
  • Synonymous states are merged via sentence embeddings and human-calibrated thresholds.
  • Evaluation involves fitting first-order Markov models over label sequences.
  • The paper is available on arXiv with identifier 2607.29181.
  • The research addresses the challenge of building user models from unstructured screen activity.

Entities

Institutions

  • arXiv

Sources