FAS-R1: A Unified Multi-Task MLLM for Reasoning Face Anti-Spoofing
A new framework called FAS-R1 has been introduced by researchers, designed as a two-stage reasoning-oriented multimodal large language model (MLLM) for comprehensive face anti-spoofing (FAS) prediction. This framework includes tasks such as authenticity classification, attack-type identification, and spoof-region localization. Initially, it employs FAS-R1-23K, a high-quality dataset for cold-start supervised fine-tuning, followed by FAS-specific GRPO post-training. The Degradation-Simulated Augmentation (DSA) promotes stable reasoning of spoof cues amid variations in visual quality, while Difficulty-Aware GRPO (DA-GRPO) helps reduce the dominance of easy samples. This initiative addresses the necessity for FAS systems to deliver not just genuine/spoof determinations but also attack semantics and image-based evidence for human review.
Key facts
- FAS-R1 is a two-stage reasoning-oriented MLLM framework for unified FAS prediction.
- It covers authenticity classification, attack-type recognition, and spoof-region localization.
- First stage uses FAS-R1-23K dataset for cold-start supervised fine-tuning.
- Second stage performs FAS-specific GRPO post-training.
- DSA encourages stable reasoning across visual-quality shifts.
- DA-GRPO mitigates easy-sample dominance.
- Addresses need for attack semantics and image-grounded evidence.
- Published on arXiv with ID 2607.26432.
Entities
Institutions
- arXiv