Domain-specific fine-tuning dramatically improves retrieval and generation for underrepresented languages: a modest 1B embedder outperforms multilingual models when trained on just 65K in-domain Greek examples, showing that language adaptation is more important than model size for specialized tasks.
This paper adapts NVIDIA's Nemotron retrieval system for Modern Greek across legal, energy, financial, and medical domains. The authors mine Greek corpora, train specialized retrieval models, and create HERA—the first large-scale Greek RAG benchmark.