Aivora
Hugging FaceOpen SourceIntermediate

Transformers Meets llama.cpp: Direct GGUF Inference on Apple Silicon

Hugging Face 整合 llama.cpp:在 Transformers 中直接執行 GGUF 量化模型

2 min read
Transformers Meets llama.cpp: Direct GGUF Inference on Apple Silicon
The 30-second version

This update removes the need for standalone engines like llama.cpp when running quantized models on Mac. Developers can now use the standard PyTorch and Transformers APIs alongside the new kernels library to load and run GGUF models directly on Apple Silicon. This integration avoids dequantization memory overhead and optimizes PyTorch's generation loop to minimize CPU-GPU synchronization, offering near-native local inference speeds and research flexibility.

Key points

01

Direct Kernel Integration

Integrates ggml's Metal kernels directly into PyTorch via the kernels library, avoiding the need to swap the model for an external runtime.

02

Optimized Generation Loop

Removes unnecessary attention masks early and defers stopping checks asynchronously, overlapping CPU scheduling and GPU execution to boost token generation rates.

03

Seamless Standard APIs

Compatible with AutoModelForCausalLM, custom logits processors, and even supports direct dequantization for standard training workflows.

04

Local OpenAI-Compatible Server

Use `transformers serve` to spin up a local server, connecting models on your Mac directly to UI clients and agents like Pi or Jan.

How it works

Qwen3.5-4B GGUF Quantization Comparison
BF16Q6_KQ5_K_MQ4_K_M
File size8.42 GB3.53 GB3.14 GB2.74 GB
Tradeoff未量化之參考基準比更小的量化版本保留更多精確度檔案大小與精確度之間的折衷平衡點適合本機推論的實用建議起點

Why it matters

This integration bridges the gap between highly efficient local inference and flexible AI research. Developers no longer have to choose between a standalone C++ runtime for performance and PyTorch for custom workflows. You can now use Python, PyTorch hooks to inspect intermediate activations, or modify forward passes while fully utilizing GGUF's hardware acceleration and low memory footprint.

Who it affects

  • AI Developer
  • AI Researcher
  • Enterprise Leader

How to use it

  1. 1Loading, testing, and evaluating GGUF format LLMs on a Mac using PyTorch.
  2. 2Launching a local API server with a single command to power local coding agents.
  3. 3Dequantizing a GGUF checkpoint to continue with fine-tuning in a standard PyTorch workflow.

Limitations & caveats

  • The packed quantized inference path is currently restricted to Apple Silicon (MPS) only.
  • Limited model architecture support, initially covering Qwen3.5 dense/MoE and compatible Qwen3.8 checkpoints.

Related

NVIDIA DIN Deploy: Building High-Performance Local AI Apps with C++ and TensorRT RTX
NVIDIA DeveloperOpen Source

NVIDIA DIN Deploy: Building High-Performance Local AI Apps with C++ and TensorRT RTX

NVIDIA 推出 DIN Deploy:用 C++ 與 TensorRT RTX 打造高效能地端 AI 應用程式

NVIDIA's open-source DIN Deploy combines ONNX Runtime and TensorRT RTX, providing C++ samples to help developers run ASR, segmentation, and image generation models locally with hardware acceleration on Windows and Linux.

2 min read
Designing AI-Native Software: Lessons from NVIDIA TensorRT Model Connect
NVIDIA DeveloperOpen Source

Designing AI-Native Software: Lessons from NVIDIA TensorRT Model Connect

打造 AI 原生專案:NVIDIA TensorRT Model Connect 的開發啟示

NVIDIA shares architectural insights from building TensorRT Model Connect, demonstrating how to design software around coding agents using model-family isolation, reversible changes, and automated validation.

2 min read
Topology-Aware Workload Scheduling with NVIDIA Topograph
NVIDIA DeveloperOpen Source

Topology-Aware Workload Scheduling with NVIDIA Topograph

NVIDIA Topograph:實現拓撲感知排程,徹底釋放 AI 工廠 GPU 效能

NVIDIA Topograph is an open-source toolkit that automates cluster topology discovery and translates it for Kubernetes and Slurm schedulers to optimize GPU workload placement and eliminate network bottlenecks.

2 min read