model-pruning

21.8k
davila7davila7

Reduce LLM size and accelerate inference using pruning techniques like Wanda and SparseGPT. Use when compressing models without retraining, achieving 50% sparsity with minimal accuracy loss, or enabling faster inference on hardware accelerators. Covers unstructured pruning, structured pruning, N:M sparsity, magnitude pruning, and one-shot methods.

Emerging TechniquesModel PruningWanda+8
192 days ago

model-pruning-helper

1.5k
Jeremylongshore Claude Code Plugins Plus Skills Model Pruning HelperJeremylongshore Claude Code Plugins Plus Skills Model Pruning Helper

Model Pruning Helper - Auto-activating skill for ML Deployment. Triggers on: model pruning helper, model pruning helper Part of the ML Deployment skill category.

192 days ago

axiom-ios-ml

535
Charleswiltgen Axiom Axiom Ios MlCharleswiltgen Axiom Axiom Ios Ml

Use when deploying ANY machine learning model on-device, converting models to CoreML, compressing models, or implementing speech-to-text. Covers CoreML conversion, MLTensor, model compression (quantization/palettization/pruning), stateful models, KV-cache, multi-function models, async prediction, SpeechAnalyzer, SpeechTranscriber.

192 days ago

Model Optimization

29
omer-metinomer-metin

Use when reducing model size, improving inference speed, or deploying to edge devices — covers quantization, pruning, knowledge distillation, ONNX export, and TensorRT optimization. Use it when applicable.

192 days ago

model-optimization

5
pluginagentmarketplacepluginagentmarketplace

Quantization, pruning, AutoML, hyperparameter tuning, and performance optimization. Use for improving model performance, reducing size, or automated ML.

192 days ago

Fuel

3
openclaw-rocksopenclaw-rocks

Optimized LLM inference and agent config for OpenClaw. Multi-provider routing with automatic cheapest-provider selection, context pruning, smart compaction, cheap heartbeats, session initialization, prompt caching, and memory management — all calibrated for autonomous agents that run for hours without wasting tokens. Triggers on: 'save on inference,' 'cheaper models,' 'optimize costs,' 'LLM config,' 'model routing,' 'inference setup,' 'fuel,' 'reduce token usage,' 'context management,' or any request to make an OpenClaw agent more cost-effective.

192 days ago

ax-coreml

1
KasempiternalKasempiternal

On-device ML with CoreML -- model conversion (PyTorch to CoreML), compression (palettization/quantization/pruning), stateful KV-cache for LLMs, multi-function models, MLTensor, async prediction, and diagnostics

192 days ago

axiom-ios-ml

1
megastepmegastep

Use when deploying ANY machine learning model on-device, converting models to CoreML, compressing models, or implementing speech-to-text. Covers CoreML conversion, MLTensor, model compression (quantization/palettization/pruning), stateful models, KV-cache, multi-function models, async prediction, SpeechAnalyzer, SpeechTranscriber.

192 days ago

model-pruning-helper

nivkazdannivkazdan

Assist with model pruning helper operations. Auto-activating skill for ML Deployment. Triggers on: model pruning helper, model pruning helper Part of the ML Deployment skill category. Use when working with model pruning helper functionality. Trigger with phrases like "model pruning helper", "model helper", "model".

192 days ago