Transformers Meets llama.cpp: Direct GGUF Inference on Apple Silicon
Hugging Face 整合 llama.cpp:在 Transformers 中直接執行 GGUF 量化模型
This update removes the need for standalone engines like llama.cpp when running quantized models on Mac. Developers can now use the standard PyTorch and Transformers APIs alongside the new kernels library to load and run GGUF models directly on Apple Silicon. This integration avoids dequantization memory overhead and optimizes PyTorch's generation loop to minimize CPU-GPU synchronization, offering near-native local inference speeds and research flexibility.
Key points
Direct Kernel Integration
Integrates ggml's Metal kernels directly into PyTorch via the kernels library, avoiding the need to swap the model for an external runtime.
Optimized Generation Loop
Removes unnecessary attention masks early and defers stopping checks asynchronously, overlapping CPU scheduling and GPU execution to boost token generation rates.
Seamless Standard APIs
Compatible with AutoModelForCausalLM, custom logits processors, and even supports direct dequantization for standard training workflows.
Local OpenAI-Compatible Server
Use `transformers serve` to spin up a local server, connecting models on your Mac directly to UI clients and agents like Pi or Jan.
How it works
| BF16 | Q6_K | Q5_K_M | Q4_K_M | |
|---|---|---|---|---|
| File size | 8.42 GB | 3.53 GB | 3.14 GB | 2.74 GB |
| Tradeoff | 未量化之參考基準 | 比更小的量化版本保留更多精確度 | 檔案大小與精確度之間的折衷平衡點 | 適合本機推論的實用建議起點 |
Why it matters
This integration bridges the gap between highly efficient local inference and flexible AI research. Developers no longer have to choose between a standalone C++ runtime for performance and PyTorch for custom workflows. You can now use Python, PyTorch hooks to inspect intermediate activations, or modify forward passes while fully utilizing GGUF's hardware acceleration and low memory footprint.
Who it affects
- AI Developer
- AI Researcher
- Enterprise Leader
How to use it
- 1Loading, testing, and evaluating GGUF format LLMs on a Mac using PyTorch.
- 2Launching a local API server with a single command to power local coding agents.
- 3Dequantizing a GGUF checkpoint to continue with fine-tuning in a standard PyTorch workflow.
Limitations & caveats
- The packed quantized inference path is currently restricted to Apple Silicon (MPS) only.
- Limited model architecture support, initially covering Qwen3.5 dense/MoE and compatible Qwen3.8 checkpoints.
Related

NVIDIA DIN Deploy: Building High-Performance Local AI Apps with C++ and TensorRT RTX
NVIDIA 推出 DIN Deploy:用 C++ 與 TensorRT RTX 打造高效能地端 AI 應用程式
NVIDIA's open-source DIN Deploy combines ONNX Runtime and TensorRT RTX, providing C++ samples to help developers run ASR, segmentation, and image generation models locally with hardware acceleration on Windows and Linux.

Designing AI-Native Software: Lessons from NVIDIA TensorRT Model Connect
打造 AI 原生專案:NVIDIA TensorRT Model Connect 的開發啟示
NVIDIA shares architectural insights from building TensorRT Model Connect, demonstrating how to design software around coding agents using model-family isolation, reversible changes, and automated validation.

Topology-Aware Workload Scheduling with NVIDIA Topograph
NVIDIA Topograph:實現拓撲感知排程,徹底釋放 AI 工廠 GPU 效能
NVIDIA Topograph is an open-source toolkit that automates cluster topology discovery and translates it for Kubernetes and Slurm schedulers to optimize GPU workload placement and eliminate network bottlenecks.