AI Signal 515
Deploy local agents everywhere with LFM2.5-2.6B
LiquidAI released LFM2.5-2.6B, a 2.6-billion-parameter model engineered for on-device agentic workloads, claiming competitive tool-use and instruction-following performance against models up to four times its size.
For teams building agent-based products, this model targets the cost and privacy constraints that make cloud-hosted LLM agents expensive or non-viable at scale. If the benchmark claims hold, developers could run multi-step tool-calling agents on laptops or phones under 2.5 GB of memory without a recurring inference bill. The main tradeoff is coding capability, where larger models retain a clear lead.
Written by elseif from the cluster below · every claim links back to a sourceThe three things worth knowing
LFM2.5-2.6B runs at 220 tokens/s on an Apple M5 Max and 113 tokens/s on an AMD Ryzen CPU, fitting in under 2.5 GB of memory.
It tops every instruction-following and tool-use benchmark in the provided comparison except BFCLv4, where only the 9.7B Qwen3.5 edges ahead.
Coding is the explicit weak spot, the larger models in the comparison maintain a clear lead on LiveCodeBenchv6.
THE CLUSTER
↗