How To Use LFM2.5-VL-DSpark To Accelerate Vision-Language AI
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: How To Use LFM2.5-VL-DSpark To Accelerate Vision-Language AI on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on monitors, keyboards and dev gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

Liquid AI released LFM2.5-VL-DSpark, an experimental 280M-parameter speculative-decoding drafter for its LFM2.5-VL-3B vision-language model. The company reports decoding speedups up to 3.13x on Apple silicon and 2.66x on H100 with unchanged outputs, supported day-one in llama.cpp, MLX-VLM, and SGLang.

Liquid AI has released LFM2.5-VL-DSpark, an experimental 280M-parameter draft model that accelerates inference of its open-weight LFM2.5-VL-3B vision-language model through speculative decoding. According to the company, the drafter adds 8.9% to the target model’s parameter count while delivering decoding speedups of up to 3.13x on Apple silicon and up to 2.66x on an NVIDIA H100 — without changing outputs. The model is available now on Hugging Face in Safetensors and GGUF formats, with day-one support in llama.cpp, MLX-VLM, and SGLang.

The drafter extends Liquid AI’s DSPark recipe — previously applied only to its text-only LFM2.5 models — to a multimodal target for the first time. It captures the target model’s hidden states at a fixed set of tapped layers and drafts blocks of candidate tokens conditioned on those states. Because image patches and text tokens are projected into a shared representation before the tapped layers, the drafter operates on hidden-state vectors of identical dimensionality regardless of input modality, and the inference algorithm is unchanged from the text models.

The final drafter is a simplified attention-only model with 4 layers, selected via ablations across 3, 4, and 5 layers, with a block size of 9. Liquid AI reports that acceptance improved over 10 training epochs before reaching diminishing returns. The company breaks the roughly 279.5M parameters down into a 193.0M-parameter decoder stack, a 21.0M hidden-state projection, a 65.5M Markov head, and about 6.4k parameters in norms and a confidence head.

Benchmarks follow the MMSpec protocol across six vision tasks: general VQA, text VQA, image captioning, chart VQA, complex reasoning, and multi-turn conversation. On device, with MLX on an M5 Max, decoding runs 2.30x to 3.13x faster and end-to-end latency improves 1.56x to 2.62x. With llama.cpp on an M3 Ultra, decoding improves 1.57x to 2.14x and end-to-end gains range from 1.30x to 1.77x. On H100, the company reports decoding speedups up to 2.66x and end-to-end gains of 1.64x to 2.27x, using a DSpark block size of 8. Speculative decoding is exact: the target model verifies every proposed token, so greedy output matches that of the target model alone.

At a glance
announcementWhen: released now; available on Hugging Face…
The developmentLiquid AI released an experimental DSpark speculative-decoding drafter that accelerates its LFM2.5-VL-3B vision-language model, extending the technique to multimodal targets for the first time in the family.
At a glance
announcementWhen: announced September 2026, available now
The developmentLiquid AI announced and released an experimental DSpark draft model for its LFM2.5-VL-3B vision-language model, available immediately on Hugging Face.

Faster Vision AI on Consumer Hardware

The release targets a persistent bottleneck for local and edge AI: vision-language models are slower than text models because images must pass through a vision encoder and then be processed as hundreds of visual tokens alongside the text prompt. A drafter that roughly doubles or triples decode speed — for under 9% more parameters — could make strong>3B-class multimodal models practical on laptops and phones, where users feel latency directly.

Day-one integration matters as much as the raw speedup. DSpark support in llama.cpp, MLX-VLM, and SGLang means the acceleration works in the toolchains hobbyists and deployers already use, rather than requiring custom inference code. Combined with Liquid AI’s open-weight licensing — download, fine-tune, and deploy without restrictions, per the company — the release positions the LFM2.5 family for on-device multimodal applications.

Amazon

Top picks for "lfm2 dspark accelerate"

As an affiliate, we earn on qualifying purchases.

From Text Drafters to Multimodal

Speculative decoding is an established technique in which a small, fast “draft” model proposes candidate tokens that the larger target model verifies in batch, accepting matching tokens and discarding the rest. Because verification is cheaper than sequential generation, accepted drafts translate into net speedup with identical outputs under greedy decoding.

Liquid AI released its first LFM2.5-DSpark drafter models for text-only LFM2.5 targets earlier in 2026. The new vision drafter reuses the same architecture and inference algorithm, differing mainly in training data and the shared multimodal representation. The LFM2.5 family spans base models, audio, and vision variants, with the 3B vision model positioned as an edge-capable multimodal option.

The company is candid about the ceiling: speculative decoding accelerates only the decode phase, not vision encoding or prefill. On edge devices with limited compute, those stages consume a larger share of wall-clock time, so end-to-end gains (1.30x–2.62x) trail decode gains — an illustration of Amdahl’s law that Liquid AI itself highlights.

“It adds a speculative decoding path that trades a minimal increase in memory footprint for a larger speedup without changing output quality.”

— Liquid AI, announcement post

Experimental Status and Benchmark Gaps

The release is explicitly labeled experimental, and the company has not indicated when or whether the drafter will be promoted to a stable release. The reported speedups are Liquid AI’s own measurements, not independently verified third-party benchmarks, and results will vary with hardware, prompt composition, and image resolution.

One figure in the company’s GPU results appears internally inconsistent — a lower bound described as “20.4x” in a range stated as “20.4x to 2.66x” — which reads as a typo, likely for 2.04x; Liquid AI has not clarified the figure. It is also unclear how acceptance rates behave on out-of-distribution vision tasks, how the drafter affects sampling-based (non-greedy) generation quality, and what the memory footprint increase is in runtime terms beyond parameter count. Details of the training data mixture have not been published.

Community Uptake and Clarifications

The model and required integration patches are public: SGLang support requires a build with DSpark for LFM2.5 targets (PR #40651), llama.cpp requires PR #29339, and MLX-VLM requires PR #2280. Early adopters will likely publish independent benchmarks in the coming weeks, which will test the company’s reported speedups on varied hardware and workloads.

Watch for clarification of the inconsistent H100 lower-bound figure, any statement on a stable (non-experimental) promotion, and possible extension of the DSpark recipe to larger vision targets in the LFM2.5 family. Liquid AI has not committed to timelines on any of these.

Key Questions

What is LFM2.5-VL-DSpark?

An experimental 280M-parameter draft model from Liquid AI that accelerates its LFM2.5-VL-3B vision-language model using speculative decoding. It adds about 8.9% to the target model’s parameter count.

Does using the drafter change the model’s outputs?

No. Speculative decoding is exact: the target model verifies every proposed token, so greedy output is identical to running the target model alone. Behavior under sampling-based (non-greedy) generation is less documented.

How much faster is it, according to Liquid AI?

The company reports decoding speedups of 2.30x to 3.13x on an M5 Max with MLX, 1.57x to 2.14x with llama.cpp on an M3 Ultra, and up to 2.66x on an H100. End-to-end gains are lower, 1.30x to 2.62x, because vision encoding and prefill are not accelerated.

Where can I get it and what do I need?

The model is on Hugging Face in Safetensors and GGUF formats. Support requires pending patches: SGLang PR #40651, llama.cpp PR #29339, or MLX-VLM PR #2280.

Are the benchmark numbers independently verified?

No. The figures are Liquid AI’s own measurements following the MMSpec protocol across six vision tasks. One GPU lower-bound figure in the announcement appears to be a typo and has not been clarified.

Primary source: Hugging Face · via ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Impact Of Nvidia Buying The Open Commons On AI Ecosystems

Nvidia reportedly plans to buy Hugging Face for $12.9 billion, a move that could reshape open-source AI and influence AI hardware and software dynamics.

Grok 4.6: The Frontier Is Now A Price War

Grok 4.6 launches with modest intelligence gains but maintains flat pricing, sparking a potential price war in AI model development. Details remain evolving.

Four Frontier-Class Open Models In Eight Weeks: China’s AI Release Strategy

Chinese labs launched four major open-weight AI models from April to June 2026, signaling a rapid production line that challenges Western dominance.

The Inner Workings Of Claude’s AI Text Watermarking System

Anthropic reveals how future Claude models will embed an invisible statistical watermark using a secret key, aiding detection without affecting text quality.