跳到正文
原文
Hugging Face Blog·· 7 天前精选AI 评分70

Hugging Face 让 transformers 支持运行 llama.cpp 的 GGUF 量化模型

Transformers now runs llama.cpp quants

AI 导读

Hugging Face 为 transformers 添加了 GGUF 模型支持,可在 Apple Silicon 上通过熟悉的 from_pretrained API 加载 GGUF 检查点,并复用 llama.cpp 的 ggml Metal 内核以接近其推理性能。

推荐理由

Hugging Face 将 llama.cpp 的 GGUF 量化模型接入 transformers,并复用其 ggml Metal 内核,本地推理性能接近 llama.cpp,为开发者提供了在 PyTorch 生态中使用 GGUF 检查点的新路径。

来源:Hugging Face Blog · huggingface.co