model-pruning
Reduce LLM size and accelerate inference using pruning techniques like Wanda and SparseGPT. Use when compressing models without retraining, achieving 50% sparsity with minimal accuracy loss, or enabling faster inference on hardware accelerators. Covers unstructured pruning, structured pruning, N:M sparsity, magnitude pruning, and one-shot methods.
model-pruning-helper
Model Pruning Helper - Auto-activating skill for ML Deployment. Triggers on: model pruning helper, model pruning helper Part of the ML Deployment skill category.
axiom-ios-ml
Use when deploying ANY machine learning model on-device, converting models to CoreML, compressing models, or implementing speech-to-text. Covers CoreML conversion, MLTensor, model compression (quantization/palettization/pruning), stateful models, KV-cache, multi-function models, async prediction, SpeechAnalyzer, SpeechTranscriber.
Model Optimization
Use when reducing model size, improving inference speed, or deploying to edge devices — covers quantization, pruning, knowledge distillation, ONNX export, and TensorRT optimization. Use it when applicable.
model-optimization
Quantization, pruning, AutoML, hyperparameter tuning, and performance optimization. Use for improving model performance, reducing size, or automated ML.
Fuel
Optimized LLM inference and agent config for OpenClaw. Multi-provider routing with automatic cheapest-provider selection, context pruning, smart compaction, cheap heartbeats, session initialization, prompt caching, and memory management — all calibrated for autonomous agents that run for hours without wasting tokens. Triggers on: 'save on inference,' 'cheaper models,' 'optimize costs,' 'LLM config,' 'model routing,' 'inference setup,' 'fuel,' 'reduce token usage,' 'context management,' or any request to make an OpenClaw agent more cost-effective.
ax-coreml
On-device ML with CoreML -- model conversion (PyTorch to CoreML), compression (palettization/quantization/pruning), stateful KV-cache for LLMs, multi-function models, MLTensor, async prediction, and diagnostics
axiom-ios-ml
Use when deploying ANY machine learning model on-device, converting models to CoreML, compressing models, or implementing speech-to-text. Covers CoreML conversion, MLTensor, model compression (quantization/palettization/pruning), stateful models, KV-cache, multi-function models, async prediction, SpeechAnalyzer, SpeechTranscriber.
model-pruning-helper
Assist with model pruning helper operations. Auto-activating skill for ML Deployment. Triggers on: model pruning helper, model pruning helper Part of the ML Deployment skill category. Use when working with model pruning helper functionality. Trigger with phrases like "model pruning helper", "model helper", "model".