ARTFEED — Contemporary Art Intelligence

Promptable Gaze Target Estimation: A New Concept-Driven Paradigm

ai-technology · 2026-08-13

There's a new study on arXiv (2608.11367) that introduces the Promptable Gaze Target Estimation (PGE) task, which is a big step forward in how we analyze where people are looking. Unlike older methods that rely on complicated steps needing specific inputs like head boxes or body positions, PGE makes use of flexible user prompts—these can be either text or images, like saying "the boy in the red shirt" or pointing to a spot on a screen. The goal of this approach is to reduce errors that can happen in detection while making it easier to use with natural language. The paper highlights the limitations of past research and shows how PGE effectively estimates gaze targets in real-world images. This work is categorized under computer vision on arXiv.

Key facts

  • New task: Promptable Gaze Target Estimation (PGE)
  • Uses natural language or visual prompts to specify gaze subject
  • End-to-end, concept-driven paradigm
  • Addresses limitations of multi-stage pipelines
  • Published on arXiv with ID 2608.11367
  • Aims to improve gaze estimation in-the-wild
  • Prompts can be text or visual points
  • Reduces cascading errors from detection

Entities

Institutions

  • arXiv

Sources