You can use LLMs to design feature extraction code offline, then deploy the resulting programs without needing LLM calls at inference time—achieving better predictions while meeting real-time deployment constraints.
This paper presents LLM-BlockFE, a system that automatically converts long text into reusable feature programs for risk prediction models. Instead of calling an LLM for every prediction, the system uses an LLM offline to design feature extraction code once, then deploys frozen programs for fast inference.