When Less Data Favors Smaller Teachers: Rethinking Teacher Capacity and Data Selection for Knowledge Distillation
2609.34489

Authors

Jaeho Lee,Minjae Park,Taesun Yeom

Abstract

Data pruning reduces the training cost of knowledge distillation (KD). However, the preferred teacher capacity changes with the data budget: smaller teachers can outperform larger ones when limited training data are available.

Understanding what drives this shift is important not only for teacher choice but also for identifying which samples are useful for distillation. We analyze teacher supervision by decomposing it into relational ordering---the ranking of classes---and score geometry---the magnitudes and margins of class probabilities---and show that the small-teacher advantage in the low-data regime arises not only from score geometry but also from relational ordering.

Beyond understanding teacher capacity, our analysis reveals two properties of effective subsets: samples should match the difficulty appropriate for the available budget, and their relational signals should be diverse rather than redundant. Based on these findings, we propose DVA (Difficulty- and Volume-Aware data selection for KD), a training-dynamics-free method, which uses a small teacher as a proxy for budget-aware difficulty filtering and class-conditional relational volume maximization.

Despite requiring no training dynamics statistics, our method remains competitive with training-dynamics-based methods while consistently outperforming training-dynamics-free baselines.

Resources

Ray graphicRay graphicRay graphicRay graphic

Stay in the loop

Every AI paper that matters, free in your inbox daily.

Details

  • takara.ai
  • Custom AI and machine learning from the Frontier Research Team.
  • © 2026 takara.ai Ltd
  • Content is sourced from third-party publications.
Ray graphicRay graphicRay graphic