Smart data curation and cost-aware deployment economics can make small VLMs competitive with much larger models for document processing—the key is selecting training examples strategically and measuring real-world costs including verification and correction.
This paper presents a practical document understanding system that extracts structured data from documents at a fraction of the cost of human annotation or larger models.