AI systems for underrepresented languages fail at the infrastructure level—in data collection, tokenization, and deployment design—long before model training begins. Fixing this requires treating offline-first design and linguistic diversity as core infrastructure priorities, not afterthoughts.
This paper reveals how AI infrastructure systematically disadvantages speakers of underrepresented languages like Bengali before models are even trained.