Instead of training separate models for each domain, you can use an MLLM router to dynamically select which vision backbone handles each image, getting both better generalization and easier updates without retraining.
ARMDIL uses a multimodal language model to intelligently route images to the best-suited vision model (CNNs, self-supervised learners, or vision-language models) within an ensemble.