Liquid AI Unveils open d1: Ultra-Fast Edge Decision Models with Single-Pass Multimodal Inference
Liquid AI 推出 open d1 邊緣決策模型:無需生成 Token、單次前向傳播即刻做出多模態決策

Liquid AI has introduced open d1, a family of edge-focused decision models built on Liquid Foundation Models (LFMs). Instead of generating tokens step-by-step, these models run in a single forward pass to output structured decisions directly, drastically reducing latency. d1-3B supports text and images, answering in under 50ms on edge devices like NVIDIA Jetson. d1-omni-600M is an early research release supporting text-image or text-audio modalities, beating Decider 2B with just a fraction of its parameters.
Key points
Tokenless Single-Pass Inference
Unlike generative models, d1 outputs structured decisions in a single forward pass, completely bypassing token generation latency.
High-Performance Edge Deployment
d1-3B answers single questions in under 50ms on NVIDIA Jetson and Apple M5 Pro, and under 10ms on discrete GPUs.
Multimodal Decision Processing
d1-3B supports text and image inputs, while the 600M-parameter d1-omni handles text/image or text/audio modalities.
How it works
| d1-3B | d1-omni-600M | |
|---|---|---|
| Parameter Size | 3B (30億) | 600M (6億) |
| Base Backbone | LFM2.5-VL-3B (僅解碼器) | LFM2.5-Encoder-350M (雙向編碼器) |
| Supported Modalities | 文字 + 影像 | 文字 + 影像 或是 文字 + 語音 |
| Mean Benchmark Score | 82.9 | 78.4 |
| Status | 正式發布 / 已測試邊緣速度 | 早期研究版本 (Early Research) |
Why it matters
This development showcases the massive potential of non-generative 'decision models' on edge hardware. By bypassing expensive token generation steps and solving classification, scoring, and routing tasks in a single forward pass, d1 makes real-time, low-latency, and privacy-centric local multimodal AI applications viable on low-power devices without relying on costly cloud GPUs.
Who it affects
- AI Developer
- AI Researcher
- Enterprise Leader
- Product Manager
How to use it
- 1Smart Routing & Intent Classification: Automatically triage user requests or voice commands to specific teams like billing or support.
- 2On-Device Content Moderation: Instantly analyze local text or images for policy violations and toxicity on edge devices.
- 3IoT & Robotic Control: Execute millisecond-level local control decisions based on real-time vision or audio sensor inputs.
Limitations & caveats
- d1-omni-600M is currently an early research release with no benchmarked inference speed numbers reported yet.
- Being decision-only models, they cannot perform generative tasks such as drafting textual responses or free-form chat.
- Lack of public audio decision benchmarks limits standard evaluation of the model's multimodal performance.
Related

NVIDIA cuPhoton: Accelerating End-to-End Scientific Image Analysis on GPUs
NVIDIA cuPhoton 開源工具套件:加速天文與科學影像端到端 GPU 分析
NVIDIA cuPhoton is an open-source CUDA-X toolkit providing GPU-accelerated building blocks that eliminate CPU bottlenecks in end-to-end scientific image processing.

NVIDIA DIN Deploy: Building High-Performance Local AI Apps with C++ and TensorRT RTX
NVIDIA 推出 DIN Deploy:用 C++ 與 TensorRT RTX 打造高效能地端 AI 應用程式
NVIDIA's open-source DIN Deploy combines ONNX Runtime and TensorRT RTX, providing C++ samples to help developers run ASR, segmentation, and image generation models locally with hardware acceleration on Windows and Linux.

Designing AI-Native Software: Lessons from NVIDIA TensorRT Model Connect
打造 AI 原生專案:NVIDIA TensorRT Model Connect 的開發啟示
NVIDIA shares architectural insights from building TensorRT Model Connect, demonstrating how to design software around coding agents using model-family isolation, reversible changes, and automated validation.