Optimization Dynamics Imprint Semantic Specificity in Contrastive Embedding Norms
This paper explains how discarded norms in contrastive embedding models correlate with semantic properties and provides a theoretical framework.
Provides a theoretical explanation for the correlation between discarded norms and semantic properties in contrastive embedding models
Before reading this…
Applications
- →Specific models and retrieval tasks
To understand this paper, make sure you know these concepts first:
- Understanding of machine learning models, optimization dynamicsfind papers →
Abstract
More Like ThisContrastive embedding models trained with scale-invariant losses are typically paired with distance metrics like cosine similarity, effectively ignoring embedding magnitudes. However, surprisingly, empirical studies reveal that despite this, these "discarded" norms seem to correlate with semantic properties such as concept specificity, token frequency, and human uncertainty. In this work, we provide a formal theoretical framework explaining this phenomenon. By analyzing the optimization dynamics, we derive an analytic formula demonstrating that embedding length naturally encodes this information as a byproduct of the training process. We also show how this gives rise to signals that can serve as "free" calibration tools in specific models and retrieval tasks, providing a grounded explanation for a previously heuristic observation.