🔍 Read the full analysis: Getting Started With LFM2.5-VL-DSpark For Vision-Language Models on ThorstenMeyerAI.com
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
Liquid AI released LFM2.5-VL-DSpark, an experimental 280M-parameter draft model for speculative decoding with its open-weight LFM2.5-VL-3B vision-language model. The company reports decoding speedups of up to 3.13x on Apple silicon and 2.66x on an H100, while greedy outputs remain unchanged; independent verification and clarification of one GPU benchmark figure are pending.
The drafter uses speculative decoding: it proposes blocks of candidate tokens, then the larger target model verifies them. Tokens that pass verification can be accepted in batches, reducing sequential decoding work. Liquid AI says the target model still determines the output, so greedy generation matches running the target model alone. The drafter adds about 8.9% to the target’s parameter count.
Liquid AI reports tests across six vision-language tasks: general and text-based visual question answering, image captioning, chart questions, complex reasoning and multi-turn conversation. On an M5 Max using MLX, reported decoding speedups range from 2.30x to 3.13x, while end-to-end latency improves 1.56x to 2.62x. On an M3 Ultra using llama.cpp, decoding improves 1.57x to 2.14x and end-to-end latency 1.30x to 1.77x. For an H100, the company reports up to 2.66x faster decoding and 1.64x to 2.27x end-to-end gains.
The model is available on Hugging Face in Safetensors and GGUF formats. Liquid AI says integrations are available for llama.cpp, MLX-VLM and SGLang, with required code changes associated with PRs #29339, #2280 and #40651, respectively. Users should check that their installed versions include the relevant support.
Faster Local Vision Model Responses
Vision-language models process image content as well as text, and decoding can add noticeable wait time during interactive use. Faster generation could make a 3B-parameter multimodal model more responsive on local hardware, where users experience latency directly. The reported results apply to decoding; they do not mean every request completes two or three times faster.
Liquid AI’s end-to-end figures are lower than its decoding figures because speculative decoding addresses only one part of inference. Image encoding and prompt processing still take time. The company describes the drafter as adding a small memory cost relative to the target model, but the runtime memory increase is not specified in the supplied material. Integration with established toolchains may make the approach easier to try, though users will need compatible builds.
DSpark Moves Beyond Text Models
Speculative decoding pairs a smaller, faster drafter with a larger target model. The drafter suggests tokens, and the target checks them. Liquid AI previously applied its DSpark approach to text-only LFM2.5 models; this release adapts it to the vision-language LFM2.5-VL-3B.
According to the company’s description, the multimodal model projects image patches and text into a shared representation before the layers tapped by the drafter. That lets the drafter work with hidden-state vectors of the same dimensionality across both input types. Liquid AI says the inference algorithm remains unchanged from its text models. It trained this drafter on a mixture of vision-language supervised fine-tuning data weighted toward anticipated serving workloads, but has not published the mixture’s details.
The final drafter has four attention-only layers and a block size of nine, selected after tests of three-, four- and five-layer versions. Liquid AI reports that token acceptance improved over ten training epochs before gains diminished. The stated total is approximately 279.5 million parameters, including a 193.0M decoder stack, 21.0M hidden-state projection and 65.5M Markov head.
“It adds a speculative decoding path that trades a minimal increase in memory footprint for a larger speedup without changing output quality.”
— Liquid AI, in its announcement
Independent Tests and Benchmark Details
The speed figures are Liquid AI’s own measurements; independent third-party results are not included in the source material. Actual performance may differ with hardware, prompt content and image resolution. The company labels the release experimental and has not said whether or when it will become stable.
One stated GPU benchmark range reads “20.4x to 2.66x,” which is inconsistent with the reported upper bound and may contain a typo. The source material does not clarify that figure, so it should not be treated as a reliable lower bound. Details also remain unavailable on how acceptance rates vary across unfamiliar vision tasks, how the drafter affects sampling-based generation, and the precise runtime memory increase.
Compatibility and Results to Watch
Developers can download the model and try it with supported builds of llama.cpp, MLX-VLM or SGLang. Further community benchmarks could help establish how the reported gains translate across devices, prompts and image sizes. Liquid AI has not announced a timetable for a stable release or supplied clarification of the inconsistent GPU range.
For now, the next useful evidence is independent testing that reports both decoding and end-to-end latency, alongside memory use and the task conditions. Those measurements would show where the drafter improves real-world response times and where image encoding or prompt processing limits the gain.
Key Questions
What is LFM2.5-VL-DSpark?
It is an experimental draft model for speculative decoding with Liquid AI’s LFM2.5-VL-3B vision-language model. The target verifies proposed tokens.
Does it change the model’s answers?
Liquid AI says greedy output matches the target model running alone because the target verifies every proposal. The supplied information does not establish equivalent results for sampling-based generation.
How much faster is it?
Liquid AI reports decoding speedups up to 3.13x on an M5 Max and up to 2.66x on an H100. Its end-to-end latency gains are smaller and vary by hardware and test.
Where can developers get it?
The model is available on Hugging Face in Safetensors and GGUF formats. Compatible support is listed for llama.cpp, MLX-VLM and SGLang, subject to the required integrations being present in the installed build.
Have the benchmark results been independently confirmed?
No independent verification is provided in the source material. One GPU benchmark range also appears internally inconsistent and has not been clarified by Liquid AI.
Primary source: Hugging Face · via ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
