AI Signal 492
LiquidAI releases LFM2.5-VL-3B with improved grounding, screen understanding, and function calling for edge
LiquidAI released LFM2.5-VL-3B, a 3.1B parameter vision-language model for edge deployment that improves screen/UI understanding, grounding, multi-image input, and function calling over its predecessor.
This model targets on-device applications where fast, direct responses matter more than extended reasoning. The 4x increase in vision training data and expanded 128K vocabulary for non-Latin scripts make it more capable for real-world edge deployments, particularly document/screen understanding and tool use.
Written by elseif from the cluster below · every claim links back to a sourceThe three things worth knowing
LFM2.5-VL-3B pairs a SigLIP2 400M NaFlex vision encoder with the LFM2.5-2.6B text backbone, pre-trained on approximately 34T tokens with 4x more vision data than before
The model answers directly instead of reasoning, keeping responses fast for real-time on-device applications
It leads its size class on real-world image tasks and matches larger models on tool use benchmarks
THE CLUSTER
↗