Small deployable VLMs can identify species above chance but degrade sharply on real camera-trap images; specialized training data outweighs model scale, and all models occasionally hallucinate fake species names when uncertain.
This paper evaluates whether small vision-language models (2-8B parameters) suitable for edge deployment can accurately identify animal species from camera-trap photos.