Aivora
Hugging FaceOpen SourceIntermediate

Liquid AI Unveils open d1: Ultra-Fast Edge Decision Models with Single-Pass Multimodal Inference

Liquid AI 推出 open d1 邊緣決策模型:無需生成 Token、單次前向傳播即刻做出多模態決策

2 min read
Liquid AI Unveils open d1: Ultra-Fast Edge Decision Models with Single-Pass Multimodal Inference
The 30-second version

Liquid AI has introduced open d1, a family of edge-focused decision models built on Liquid Foundation Models (LFMs). Instead of generating tokens step-by-step, these models run in a single forward pass to output structured decisions directly, drastically reducing latency. d1-3B supports text and images, answering in under 50ms on edge devices like NVIDIA Jetson. d1-omni-600M is an early research release supporting text-image or text-audio modalities, beating Decider 2B with just a fraction of its parameters.

Key points

01

Tokenless Single-Pass Inference

Unlike generative models, d1 outputs structured decisions in a single forward pass, completely bypassing token generation latency.

02

High-Performance Edge Deployment

d1-3B answers single questions in under 50ms on NVIDIA Jetson and Apple M5 Pro, and under 10ms on discrete GPUs.

03

Multimodal Decision Processing

d1-3B supports text and image inputs, while the 600M-parameter d1-omni handles text/image or text/audio modalities.

How it works

Comparison of d1 Decision Model Specifications
d1-3Bd1-omni-600M
Parameter Size3B (30億)600M (6億)
Base BackboneLFM2.5-VL-3B (僅解碼器)LFM2.5-Encoder-350M (雙向編碼器)
Supported Modalities文字 + 影像文字 + 影像 或是 文字 + 語音
Mean Benchmark Score82.978.4
Status正式發布 / 已測試邊緣速度早期研究版本 (Early Research)

Why it matters

This development showcases the massive potential of non-generative 'decision models' on edge hardware. By bypassing expensive token generation steps and solving classification, scoring, and routing tasks in a single forward pass, d1 makes real-time, low-latency, and privacy-centric local multimodal AI applications viable on low-power devices without relying on costly cloud GPUs.

Who it affects

  • AI Developer
  • AI Researcher
  • Enterprise Leader
  • Product Manager

How to use it

  1. 1Smart Routing & Intent Classification: Automatically triage user requests or voice commands to specific teams like billing or support.
  2. 2On-Device Content Moderation: Instantly analyze local text or images for policy violations and toxicity on edge devices.
  3. 3IoT & Robotic Control: Execute millisecond-level local control decisions based on real-time vision or audio sensor inputs.

Limitations & caveats

  • d1-omni-600M is currently an early research release with no benchmarked inference speed numbers reported yet.
  • Being decision-only models, they cannot perform generative tasks such as drafting textual responses or free-form chat.
  • Lack of public audio decision benchmarks limits standard evaluation of the model's multimodal performance.

Related

NVIDIA cuPhoton: Accelerating End-to-End Scientific Image Analysis on GPUs
NVIDIA DeveloperOpen Source

NVIDIA cuPhoton: Accelerating End-to-End Scientific Image Analysis on GPUs

NVIDIA cuPhoton 開源工具套件:加速天文與科學影像端到端 GPU 分析

NVIDIA cuPhoton is an open-source CUDA-X toolkit providing GPU-accelerated building blocks that eliminate CPU bottlenecks in end-to-end scientific image processing.

2 min read
NVIDIA DIN Deploy: Building High-Performance Local AI Apps with C++ and TensorRT RTX
NVIDIA DeveloperOpen Source

NVIDIA DIN Deploy: Building High-Performance Local AI Apps with C++ and TensorRT RTX

NVIDIA 推出 DIN Deploy:用 C++ 與 TensorRT RTX 打造高效能地端 AI 應用程式

NVIDIA's open-source DIN Deploy combines ONNX Runtime and TensorRT RTX, providing C++ samples to help developers run ASR, segmentation, and image generation models locally with hardware acceleration on Windows and Linux.

2 min read
Designing AI-Native Software: Lessons from NVIDIA TensorRT Model Connect
NVIDIA DeveloperOpen Source

Designing AI-Native Software: Lessons from NVIDIA TensorRT Model Connect

打造 AI 原生專案:NVIDIA TensorRT Model Connect 的開發啟示

NVIDIA shares architectural insights from building TensorRT Model Connect, demonstrating how to design software around coding agents using model-family isolation, reversible changes, and automated validation.

2 min read