🔍 Read the full analysis: A Guide To Multimodal Open D1 Decision Models For Edge AI on ThorstenMeyerAI.com
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
Liquid AI has released d1-3B and experimental d1-omni-600M, open-weight models designed to return structured decisions in one forward pass. The company reports benchmark scores and sub-50-millisecond d1-3B response times on tested edge devices, but the release does not include independent validation or vision and audio benchmark results.
As detailed in the original analysis, Liquid AI has released d1-3B and d1-omni-600M, open-weight models designed to return structured decisions in a single forward pass rather than generate a longer response. The company says d1-3B scored 48.57 on Decision Index 0.2.1 and answered a test question in 16 milliseconds on an NVIDIA Jetson AGX Thor; those are company-reported results, and the release does not provide independent replication.
The models are intended for tasks that can be expressed as a defined decision, such as classifying a request, scoring urgency or routing a customer message to a team. Liquid AI presents the format as an option for developers who need a response close to where data is collected, including on edge devices where latency or hardware constraints may matter.
d1-3B is based on Liquid AI’s LFM2.5-VL-3B vision-language model and accepts text and images. The smaller d1-omni-600M uses the LFM2.5-Encoder-350M bidirectional encoder with added vision and audio encoders, and supports text paired with an image or audio. Liquid AI describes that model as an early research release still under development.
Across seven public datasets spanning reading comprehension, toxicity detection, intent classification, medical question answering and cross-lingual understanding, Liquid AI reports mean scores of 82.9 for d1-3B and 78.4 for d1-omni-600M. Its comparison table lists 81.1 for Decider 4B and 77.1 for Decider 2B. The individual results are mixed: d1-3B scores below Decider 4B on BoolQ, MASSIVE intent and XNLI.
Edge Decisions and Reported Speed
The release targets a practical development trade-off: some products need a limited, structured decision rather than an open-ended answer. If a model can provide that result quickly on local hardware, developers may be able to handle selected tasks without sending each request to a remote service. The company’s measurements make that possibility relevant to latency-sensitive edge applications, but do not establish performance in any particular deployment.
Liquid AI reports that d1-3B answered one question in 16 milliseconds on Jetson AGX Thor, 26 milliseconds on Jetson AGX Orin 64 GB and 50 milliseconds on Jetson Orin Nano. It also reports 8 milliseconds per question on an NVIDIA RTX 4090 and 9 milliseconds on an AMD MI325X. These are measurements from the company’s testing, conducted with NVIDIA for the GPU and Jetson results; actual timing can depend on the device, input, software setup and task.
The company says three questions took 1.3 times as long as one across the tested devices, including a reported increase from 16 to 20 milliseconds on AGX Thor. That finding may be relevant to workloads that group requests, but it does not predict how batching will behave in a different application. Liquid AI did not report speed results for d1-omni-600M.
NVIDIA Jetson AGX Orin developer kit
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What the Benchmarks Cover
Liquid AI’s reported mean scores come from a selected set of seven public, text-focused datasets: SQuAD 2.0, Civil Comments, MASSIVE intent, PubMedQA, BoolQ, XNLI and PAWS-X. The results offer a comparison on those tests, not a general measure of quality across all decision tasks. Scores vary by dataset, and a mean can obscure where a model performs better or worse.
The company says it checked that d1-3B retained vision capabilities from its vision-language backbone and that d1-omni-600M handled its supported modalities. However, the release provides no vision or audio benchmark scores. Liquid AI says Decision Index version 0.3 has only a private vision split and that audio decision benchmarks remain an open problem. The published dataset averages therefore do not establish performance on image- or audio-based decisions.
Both models’ weights are available on Hugging Face, and Liquid AI points to demos in its System One Arcade Hugging Face Space. The release instructions specify Transformers version 5.14 or later and loading the models with supplied code enabled. Making weights available allows developers to test them, but does not itself verify the company’s reported results.
““Best decision model under 10B on the Decision Index 0.2.1.””
— Liquid AI
As an affiliate, we earn on qualifying purchases.
Limits of the Published Results
The reported evaluations come from Liquid AI, and the release does not include independent results, confidence intervals or independent replication. It also does not provide enough detail to determine how closely every benchmark setup matches a real deployment. The scores should be treated as results on the named tests rather than evidence of accuracy, reliability or safety across all decision tasks.
Important deployment questions remain open. The release does not say how often decisions on ambiguous inputs may require human review, or how these models perform across varied production workloads. Vision and audio performance are particularly difficult to assess from the published material: no modality-specific benchmark scores are provided, and d1-omni-600M has no reported speed measurements. Its experimental status adds further uncertainty.
As an affiliate, we earn on qualifying purchases.
Testing on Developer Hardware
The next practical step is for developers to compare the models with their own tasks, data and devices, rather than assume the reported averages and timings will transfer directly. Liquid AI provides access to the open weights and demos; its instructions call for Transformers 5.14 or later and use of the supplied model code.
Further evidence would help clarify how the models perform beyond the reported text datasets, especially on vision and audio tasks and under production conditions. The release does not announce a date for independent evaluations or for a more developed version of d1-omni-600M, so those milestones remain unspecified.
audio-visual AI processing hardware
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What did Liquid AI release?
Liquid AI released d1-3B and d1-omni-600M, open-weight models intended to return structured decisions in a single forward pass. The company describes d1-omni-600M as an experimental research release.
What tasks are the models designed for?
They are aimed at defined tasks such as classifying, scoring or routing requests. Liquid AI gives examples including judging urgency and answering a question about an image.
How fast is d1-3B on edge devices?
Liquid AI reports one-question times of 16 milliseconds on Jetson AGX Thor, 26 milliseconds on Jetson AGX Orin 64 GB and 50 milliseconds on Jetson Orin Nano. These are company measurements, and performance may differ with other hardware, inputs and software.
Are the benchmark results independently verified?
The release provides company-reported results and does not include independent replication. Its seven-dataset mean scores also do not establish performance across every decision task.
Does the release show how well the models handle images and audio?
It describes the models’ supported modalities but provides no vision or audio benchmark scores. Liquid AI also reports no speed results for the experimental d1-omni-600M.
Primary source: Hugging Face · via ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
