Hugging Face Blog·· 7 天前精选AI 评分70
Hugging Face 让 transformers 支持运行 llama.cpp 的 GGUF 量化模型
Transformers now runs llama.cpp quants
AI 导读
Hugging Face 为 transformers 添加了 GGUF 模型支持,可在 Apple Silicon 上通过熟悉的 from_pretrained API 加载 GGUF 检查点,并复用 llama.cpp 的 ggml Metal 内核以接近其推理性能。
推荐理由
Hugging Face 将 llama.cpp 的 GGUF 量化模型接入 transformers,并复用其 ggml Metal 内核,本地推理性能接近 llama.cpp,为开发者提供了在 PyTorch 生态中使用 GGUF 检查点的新路径。
来源:Hugging Face Blog · huggingface.co