GGUF format and llama.cpp quantization for efficient CPU/GPU inference. Use when deploying models on consumer hardware, Apple Silicon, or when needing flexible quantization from 2-8 bit without GPU requirements.
GGUF format and llama.cpp quantization for efficient CPU/GPU inference. Use when deploying models on consumer hardware, Apple Silicon, or when needing flexible quantization from 2-8 bit without GPU requirements.
GGUF 格式和 llama.cpp 量化以实现高效的 CPU/GPU 推理。在消费硬件、Apple Silicon 上部署模型时,或者在需要从 2-8 位灵活量化而不依赖 GPU 时使用。
Category: developer (开发工具) · Author: davila7 · Version: @main · License: MIT
Tags: GGUF, Quantization, llama.cpp, CPU Inference, Apple Silicon, Model Compression, Optimization
该 Skill 暂无文档文件。
Search for places (restaurants, cafes, etc.) via Google Places API proxy on localhost.
Interact with GitHub using the `gh` CLI. Use `gh issue`, `gh pr`, `gh run`, and `gh api` for issues, PRs, CI runs, and advanced queries.
Create or update AgentSkills. Use when designing, structuring, or packaging skills with scripts, references, and assets.
Start voice calls via the OpenClaw voice-call plugin.
Notion API for creating and managing pages, databases, and blocks.
Gemini CLI for one-shot Q&A, summaries, and generation.
Category:developer
Tags:GGUF, Quantization, llama.cpp, CPU Inference, Apple Silicon, Model Compression, Optimization