Token, byte, and pixel encodings each have different strengths depending on your task and capacity constraints—there's no universally best choice, only task-specific tradeoffs.
This paper compares three ways to encode text for language models—tokens, bytes, and pixels—by controlling both the linguistic content and model capacity. Using parallel sentences across 13 languages, the researchers measure how well each encoding preserves information under compression.