Embedding models don't reliably represent objective physical concepts—they're influenced more by surface-level text patterns than by actual semantic relationships, which has implications for using embeddings in domains requiring precise measurements.
This paper investigates whether embedding models accurately represent physical measurements like mass, distance, and time. The researchers found that embeddings poorly capture these objective quantities and instead learn superficial patterns based on string similarity rather than true semantic meaning.