Conference Proceedings

AUDETER: A Large-scale Dataset for Deepfake Audio Detection in Open Worlds

Qizhou Wang, Hanxun Huang, Guansong Pang, Sarah Erfani, Christopher Leckie

Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 | ACM | Published : 2026

Abstract

Speech synthesis systems can now produce highly realistic vocalisations that pose significant authenticity challenges. Despite substantial progress in deepfake detection models, their real-world effectiveness is often undermined by evolving distribution shifts between training and test data, driven by the complexity of human speech and the rapid evolution of synthesis systems. Existing datasets suffer from limited real speech diversity, insufficient coverage of recent synthesis systems, and heterogeneous mixtures of deepfake sources, which hinder systematic evaluation and open-world model training. To address these issues, we introduce AUDETER (AUdio DEepfake TEst Range), a large-scale and h..

View full abstract

Grants

Awarded by ARC Centre of Excellence on Automated Decision Making and Society


Awarded by The Singapore Ministry of Education (MOE) Academic Research Fund (AcRF) Tier 1 Grant