FaceVid-Forensics-100K: New Dataset for Deepfake Video Detection
A newly released extensive dataset, FaceVid-Forensics-100K, aims to enhance the identification of deepfake videos. This collection features 100,000 videos across 33 different synthesis techniques, such as face swapping, face reenactment, and full-face synthesis, utilizing modern generators like Seedance 2.0. It offers detailed textual annotations to overcome the shortcomings of current benchmarks, which often lack comprehensive coverage of new synthesis methods and dependable annotations. The study reveals that traditional detectors and multimodal large language models (MLLMs) frequently struggle to detect subtle forgery artifacts, hindering their adaptability to new AI-generated techniques. This dataset is designed to facilitate multi-agent forensic reasoning for effective deepfake video detection. The research paper can be found on arXiv with the identifier 2608.06865.
Key facts
- FaceVid-Forensics-100K is a large-scale deepfake video dataset.
- The dataset contains 100,000 videos.
- It spans 33 synthesis methods.
- Methods include face swapping, face reenactment, and entire-face synthesis.
- Recent generators such as Seedance 2.0 are included.
- The dataset provides fine-grained textual annotations.
- Existing benchmarks have limited coverage of recent synthesis methods.
- Conventional detectors and MLLMs often fail to capture subtle forgery artifacts.
Entities
Institutions
- arXiv