ModelsArchitectures & capability
Liquid AI releases open-weight d1 decision models for edge devices; all benchmark results so far come from the company
Liquid AI has released open weights for d1-3B and d1-omni-600M, small models that score fixed answer options in a single pass instead of generating text. Liquid reports strong text benchmark scores and fast edge latency, but no independent group has evaluated the models yet.

Open d1 has two models. d1-3B is a 3.12B-parameter model built on a vision-language backbone and accepts text, JSON and images. d1-omni-600M is an experimental encoder-based model that adds images and up to 30 seconds of English speech. Neither model generates text. Each returns probabilities over predefined answers (yes/no, pick one, or ordered rating), which suits routing, moderation and guardrail tasks. [1] [7] [8] [2]
Liquid reports that d1-3B scores 48.57 on Decision Index 0.2.1, slightly above Decider 35B-A3B at 47.11, using a scorer Liquid ran itself. It also reports 16 ms per decision on an NVIDIA Jetson AGX Thor and 50 ms on a Jetson Orin Nano. Long inputs are much slower: a 3.4K-token input takes 220 ms on Thor and 1,640 ms on Orin Nano. [1] [10]
The evidence is still thin. Liquid's blog cites an 82.9 mean across seven datasets, while the d1-3B model card reports 77.1 across eight. Neither model has published vision or audio decision benchmarks. d1-omni-600M scores only 15.95 on the Decision Index, and its model card warns that bfloat16 precision changes some top answers. [1] [7] [8] [2]
Liquid reports 16 ms decisions on Jetson hardware and a 3B model narrowly beating a 35B one on Decision Index 0.2.1.
Read the full assessment
Implication: on-device routing, moderation and guardrails may become practical once independent tests on real workloads confirm the results.
Sources are checked; writing the dossier now.
Executive brief
Liquid AI released open weights for d1-3B. The company reports that it scores 48.57 on Decision Index 0.2.1. That beats the 35-billion-parameter Decider 35B-A3B (47.11), a model with roughly 12× as many parameters. d1-3B also answers in 16 ms on an NVIDIA Jetson AGX Thor (Liquid AI blog). It reads an input and returns probabilities over fixed answers in a single pass, which makes it fit routing, moderation and guardrail jobs. A smaller experimental model, d1-omni-600M, adds audio. No independent group has evaluated either model yet. All scores and latencies come from Liquid (AlphaSignal).
Read the full section
Liquid AI released open weights for d1-3B. The company reports that it scores 48.57 on Decision Index 0.2.1. That beats the 35-billion-parameter Decider 35B-A3B (47.11), a model with roughly 12× as many parameters. d1-3B also answers in 16 ms on an NVIDIA Jetson AGX Thor (Liquid AI blog). A "decision model" doesn't write text. It reads an input and returns probabilities over fixed answers in a single pass, which makes it fit routing, moderation and guardrail jobs. A smaller experimental model, d1-omni-600M, adds audio. No independent group has evaluated either model yet. All scores and latencies come from Liquid (AlphaSignal).
What changed and event timeline
Jev Decision Index published
Multimodalart posted the index on Hugging Face. It runs 132,422 frozen requests across 37 benchmarks and ranks Jev against more than 30 open-weight decision models ().
Liquid launches API-only d1
Coverage noted three gaps: no published accuracy figures, no paid pricing and no local weights ().
Perplexity open-sources Decider v1 27B
The model is Apache 2.0. Perplexity's 85.71% score against Jev's 84.51% is self-reported, and its docs include no benchmark table ().
Open d1 released
Liquid published weights for d1-3B and the experimental d1-omni-600M, and reported Jetson latency tested with NVIDIA (;).
Capabilities and access
- d1-3B: 3.12B parameters, built on LFM2.5-VL-3B.
- d1-omni-600M: 587M parameters, built on LFM2.5-Encoder-350M with a 94M vision encoder and a 112M FastConformer audio encoder.
- Question types: yes/no (
noul), pick one option (choice) and ordered rating (score).
Read the full section
- d1-3B: 3.12B parameters, built on LFM2.5-VL-3B. It has a SigLIP2 NaFlex vision encoder (400M) and a 32,768-token context, and accepts text, JSON and images (model card).
- d1-omni-600M: 587M parameters, built on LFM2.5-Encoder-350M with a 94M vision encoder and a 112M FastConformer audio encoder. It takes text plus images, or text plus up to 30 seconds of English speech (model card).
- Question types: yes/no (
noul), pick one option (choice) and ordered rating (score). Several questions can go in one request. - License: LFM Open License v1.0. Companies with $10M or more in annual revenue need a separate commercial agreement (KuCoin News).
- Setup: requires transformers>=5.14 and
trust_remote_code=True.
Technical analysis for researchers and developers
- Different backbones: d1-3B comes from a decoder-only VLM. d1-omni comes from a bidirectional encoder with added modality towers (blog).
- No vision or audio scores: neither model has published multimodal decision benchmarks.
- Batching:
system_one_batchpacks many requests together without padding. On Jetson AGX Thor, Liquid reports 262 states per second.
Read the full section
- Different backbones: d1-3B comes from a decoder-only VLM. d1-omni comes from a bidirectional encoder with added modality towers (blog).
- No vision or audio scores: neither model has published multimodal decision benchmarks. Liquid says the Decision Index v0.3 vision split is private and that audio decision benchmarks are an open problem.
- Batching:
system_one_batchpacks many requests together without padding. On Jetson AGX Thor, Liquid reports 262 states per second. - Long inputs are much slower: a 3.4K-token input takes 220 ms on Thor and 1,640 ms on Orin Nano.
- Numeric precision matters: the d1-omni card advises against bfloat16. It changed the top answer on 0.8% of text and 1.7% of audio rows (card).
Claims and evidence
All claims below are vendor-reported unless marked otherwise.
Read the full section
All claims below are vendor-reported unless marked otherwise.
- Index score: d1-3B scored 48.57 on Decision Index 0.2.1, versus 47.11 for Decider 35B-A3B. Liquid ran the official scorer itself (BenchLM). Not independently reproduced.
- Seven-dataset mean: 82.9 for d1-3B versus 81.1 for Decider 4B. d1-omni scored 78.4 versus 77.1 for Decider 2B (blog). Decider 4B still wins on BoolQ, XNLI and MASSIVE.
- Latency: 8 ms on an RTX 4090, 9 ms on an AMD MI325X and 50 ms on a Jetson Orin Nano (blog).
- Leaderboard listing: a third-party aggregator lists both d1 models on JevBench 1.5.4 below Cygnet, Winnow-12B and Jev (BenchLM). This shows where the models are listed, not that the scores were verified.
Context and prior work
- A crowded field: decision models now come from Perplexity, Cloudflare, AWS Strands, Inception and Fastino, several with open weights (Metir AI).
- The comparison models: Decider 4B and Decider 35B-A3B are open "System One" reproductions on Qwen3.5 bases, published by Mapika (Decider 4B card).
- Liquid's own lineage: open d1 follows the API-only d1 from Sep 29 (DataNorth).
Limitations, safety and contested findings
- The omni model is weak on the index: d1-omni-600M scores 15.95 on Decision Index 0.2.1, far below d1-3B's 48.57 (card).
- Two different averages: the d1-3B card reports a 77.1 mean across eight NLU datasets (card). The blog's headline is 82.9 across seven.
- Date mismatch: the d1-3B card lists a release date of Nov 28, 2025, which conflicts with the Oct 7, 2026 announcement.
Read the full section
- The omni model is weak on the index: d1-omni-600M scores 15.95 on Decision Index 0.2.1, far below d1-3B's 48.57 (card). Its audio training covered only English speaker-assistant interactions.
- Two different averages: the d1-3B card reports a 77.1 mean across eight NLU datasets (card). The blog's headline is 82.9 across seven. Which datasets are included changes the headline number.
- Date mismatch: the d1-3B card lists a release date of Nov 28, 2025, which conflicts with the Oct 7, 2026 announcement.
- Limited scope: the models only score predefined options. They can't generate text (AlphaSignal).
- Index caveat: submissions to the Jev index depend on details the participants provide themselves (AI Weekly).
Business and practitioner implications
- Edge routing and moderation: latency under 50 ms on a Jetson makes on-device routing, moderation and robotics gating plausible.
- License: firms at or above $10M revenue need a commercial deal, unlike Apache-licensed Perplexity Decider v1 27B.
- Due diligence: run your own index and task evaluations before relying on the scores. Treat d1-omni as research-only.
Read the full section
- Edge routing and moderation: latency under 50 ms on a Jetson makes on-device routing, moderation and robotics gating plausible. Test on your own data first.
- License: firms at or above $10M revenue need a commercial deal, unlike Apache-licensed Perplexity Decider v1 27B.
- Due diligence: run your own index and task evaluations before relying on the scores. Treat d1-omni as research-only.
- Cost: since no output tokens are generated, cost scales with input size. Long inputs on small devices still take seconds.
The source trail.
Sources (12)
Multimodal open d1 decision models for the edge
Article text retrieved; extracted text may omit tables or interactive elements.
huggingface.co