ARTFEED — Contemporary Art Intelligence

SkillLens: Visual Skill Cards for GUI Action Prediction

ai-technology · 2026-08-13

A recent paper on arXiv (2608.10775) presents Visual Skill Cards (VSCs), a memory representation conditioned by state for agents that utilize computers. The proposed method, SkillLens, generates VSCs from diverse interaction traces, facilitating the retrieval of pertinent procedural knowledge and the selective enhancement of visual evidence for predicting grounded GUI actions. This representation links reusable procedures with cues for applicability, visual evidence, and verification signals, addressing the deficiency of visual procedural memory in existing agents. The authors introduce Trace-to-Visual-Skill-Card for the construction process and employ a fixed visual-language model executor during inference. This strategy aims to enhance decision-making by equipping agents with visual memories of workflows, applicable controls, and evidence of progress.

Key facts

  • Paper arXiv:2608.10775 introduces Visual Skill Cards (VSCs).
  • SkillLens constructs VSCs from heterogeneous interaction experience.
  • VSCs bind reusable procedures with applicability cues, visual evidence, and verification signals.
  • At inference, SkillLens retrieves relevant cards and selectively expands needed evidence.
  • The method uses a fixed visual-language model executor for grounded GUI action prediction.
  • The paper addresses the lack of visual procedural memory in computer-using agents.
  • Raw interaction traces are long and noisy, while text-only skills omit visual state.
  • The paper is announced as a new arXiv submission.

Entities

Sources