LiquidAI announced the release of two open-weight decision models, d1-3B and d1-omni-600M, engineered for ultra-fast inference on edge devices. Unlike conventional language models that generate text sequentially, decision models return answers in a single forward pass—a design that makes them ideal for latency-sensitive applications. The d1-3B model, which supports text and image inputs, achieved the highest benchmark score among sub-10B decision models, scoring 48.57 on the Decision Index v0.2.1, surpassing even larger 35B parameter models.
Performance benchmarks showcase impressive results across NVIDIA's hardware lineup. Running on a Jetson AGX Thor, d1-3B answers questions in 16 milliseconds; on the smaller Jetson Orin Nano, inference completes in 50 milliseconds. On desktop GPU hardware like the RTX 4090, the same model processes queries in under 10 milliseconds. The smaller d1-omni-600M, currently in experimental release, matches the performance of larger competitors while supporting both vision and audio modalities in just 600 million parameters.
Both models are now available as open-source releases on Hugging Face with full code samples and integration guides. The release reflects a broader industry shift toward deploying AI locally on edge hardware, reducing latency and addressing privacy concerns associated with cloud-based inference.
Key Points
LiquidAI released d1-3B and d1-omni-600M, open-weight decision models optimized for edge device inference
d1-3B achieves best-in-class performance under 10B parameters with multimodal text and image support
Inference speeds are extremely fast: 16ms on Jetson AGX Thor, under 10ms on desktop GPUs
d1-omni-600M supports three modalities—text plus images or audio—with only 600 million parameters
Both models available open-source on Hugging Face with complete code examples