Liquid AI has released two new open-weight decision models, d1-3B and d1-omni-600M, designed for low-latency applications on edge devices. The company states that the models, part of its d1 family, prioritize speed and structured decision-making over token generation.
What Happened
The release includes d1-3B, a 3-billion parameter model trained on the LFM2.5-VL-3B backbone, which supports text and image inputs. It also includes d1-omni-600M, an experimental 600-million parameter model trained on the LFM2.5-Encoder-350M backbone, which supports text with either image or audio inputs. According to Liquid AI, these models operate differently from generative LLMs by providing answers in a single forward pass rather than generating tokens.
The company reports that d1-3B scores 48.57 on the Decision Index 0.2.1, positioning it as the highest-scoring model under 10 billion parameters. On seven public datasets covering tasks like reading comprehension and toxicity detection, Liquid AI states that d1-3B achieved a mean score of 82.9, while d1-omni-600M scored 78.4. The company notes that d1-omni-600M surpasses the Decider 2B model’s score of 77.1 with only a quarter of the parameters.
In collaboration with NVIDIA, Liquid AI evaluated the inference speed of d1-3B. The company reports that the model answers a single question in 16 milliseconds on an NVIDIA Jetson AGX Thor, 26 milliseconds on a Jetson AGX Orin, and 50 milliseconds on a Jetson Orin Nano. On the NVIDIA RTX 4090 GPU, the reported time for a single question is under 10 milliseconds.
Why It Matters
The release highlights a shift toward specialized, non-generative models for edge computing, where latency and resource constraints are critical. By offering open weights for models that handle multimodal inputs (text, vision, and audio) with sub-50ms latency on common edge hardware, Liquid AI provides developers with alternatives to larger, slower generative models for tasks like intent classification and real-time monitoring.
The inclusion of d1-omni-600M, despite its experimental status, suggests a growing interest in compact, multimodal encoders. The company states that audio decision benchmarks remain an open problem and did not report specific vision or audio benchmark scores, indicating that the multimodal capabilities are integrated but evaluated primarily through the lens of general decision metrics.
The Bottom Line
Liquid AI has made d1-3B and d1-omni-600M available on Hugging Face. Developers can install the required dependencies, including transformers version 5.14 or higher, to utilize the models. The company emphasizes that d1-3B offers the highest decision quality for its size, while d1-omni-600M is intended for use cases where a smaller footprint is prioritized.