Label-Efficient Dataset Pruning via Semi-Supervised Pseudo-Labeling

Abstract

Dataset pruning reduces the storage and training costs of deep learning by selecting an informative subset from a large dataset. However, most existing pruning methods require fully labeled data, which limits their applicability in realistic settings where unlabeled data are abundant and annotation is costly. Recent label-free pruning methods address this issue, but they rely on features from pre-trained models to estimate example difficulty. This dependence can be unreliable when the target dataset differs substantially from the pre-training distribution. We propose a label-efficient dataset pruning framework that, given only a small labeled subset, uses semi-supervised learning to generate pseudo-labels for unlabeled data, allowing existing supervised pruning methods that require label information to be seamlessly applied to the resulting pseudo-labeled training pool. We then estimate example difficulty from pseudo-label-induced training dynamics and select a coreset. By learning directly from the target dataset, our method better captures the target distribution and provides more reliable signals for difficulty estimation and coreset selection. We validate our approach on domain-specific, image-corrupted, and long-tailed datasets, where it achieves state-of-the-art performance among label-free and label-efficient baselines, while also demonstrating competitive performance on standard benchmarks.

Publication
Neural Information Processing Systems 2026
Chulhee Yun
Chulhee Yun
Associate Professor

I am an Associate Professor at KAIST AI. I am interested in optimization and machine learning theory.