Conference Proceedings

AUDETER: A Large-scale Dataset for Deepfake Audio Detection in Open Worlds

Q Wang, H Huang, G Pang, S Erfani, C Leckie

Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining | ACM | Published : 2026

Open access

Abstract

Speech synthesis systems can now produce highly realistic vocalisations that pose significant authenticity challenges. Despite substantial progress in deepfake detection models, their real-world effectiveness is often undermined by evolving distribution shifts between training and test data, driven by the complexity of human speech and the rapid evolution of synthesis systems. Existing datasets suffer from limited real speech diversity, insufficient coverage of recent synthesis systems, and heterogeneous mixtures of deepfake sources, which hinder systematic evaluation and open-world model training. To address these issues, we introduce AUDETER (AUdio DEepfake TEst Range), a large-scale and h..

View full abstract