A preprint posted to arXiv on October 1 argues that a shared temperature parameter in probabilistic contrastive learning does not imply a shared similarity scale, as previously assumed. By equalizing class-dependent angular gain with a "Pure Angular" edit, the authors report improved learned features across all tested imbalance settings on CIFAR-10/100 and ImageNet-LT, according to the arXiv preprint.
The work builds on the ProCo framework, which uses exact von Mises-Fisher (vMF) scores for probabilistic contrastive learning. The authors prove that the leading-order term in the score is (A_c/\tau), where (A_c) is the mean resultant length, a measure of class concentration. Temperature adjustments alone are insufficient to homogenize the decision scale, they argue; classwise concentration must be accounted for.
Previously, Automatica Press reported on Oct. 3 that inverse distillation unlearning could reduce inference costs for diffusion models, as described in a separate preprint.
To isolate the effect, the researchers constructed intercept-preserving and Pure Angular controls, comparing them with the full vMF classifier. Across 16 frozen-representation settings, agreement between the full model and a shared-scale cosine prototype rule ranged from 98.43% to 99.99%, with errors clustering at small cosine margins. A finite-dimensional margin condition guarantees when the two classifiers coincide exactly, the paper adds.
The preprint is 59 pages including supplementary material and has not been peer-reviewed. No independent replication of the results is yet available.