llm-evaluation

29.9k
wshobsonwshobson

Implement comprehensive evaluation strategies for LLM applications using automated metrics, human feedback, and benchmarking. Use when testing LLM performance, measuring AI application quality, or establishing evaluation frameworks.

192 days ago

agent-evaluation

21.8k
davila7davila7

Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on real-world benchmarks Use when: agent testing, agent evaluation, benchmark agents, agent reliability, test agent.

192 days ago

evaluating-llms-harness

21.8k
davila7davila7

Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking model quality, comparing models, reporting academic results, or tracking training progress. Industry standard used by EleutherAI, HuggingFace, and major labs. Supports HuggingFace, vLLM, APIs.

EvaluationLM Evaluation HarnessBenchmarking+7
192 days ago

evaluating-code-models

21.8k
davila7davila7

Evaluates code generation models across HumanEval, MBPP, MultiPL-E, and 15+ benchmarks with pass@k metrics. Use when benchmarking code models, comparing coding abilities, testing multi-language support, or measuring code generation quality. Industry standard from BigCode Project used by HuggingFace leaderboards.

EvaluationCode GenerationHumanEval+6
192 days ago

nemo-evaluator-sdk

21.8k
davila7davila7

Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution. Use when needing scalable evaluation on local Docker, Slurm HPC, or cloud platforms. NVIDIA's enterprise-grade platform with container-first architecture for reproducible benchmarking.

EvaluationNeMoNVIDIA+8
192 days ago

cirq

21.8k
davila7davila7

Quantum computing framework for building, simulating, optimizing, and executing quantum circuits. Use this skill when working with quantum algorithms, quantum circuit design, quantum simulation (noiseless or noisy), running on quantum hardware (Google, IonQ, AQT, Pasqal), circuit optimization and compilation, noise modeling and characterization, or quantum experiments and benchmarking (VQE, QAOA, QPE, randomized benchmarking).

192 days ago

llm-evaluation

18.0k
sickn33sickn33

Implement comprehensive evaluation strategies for LLM applications using automated metrics, human feedback, and benchmarking. Use when testing LLM performance, measuring AI application quality, or establishing evaluation frameworks.

192 days ago

V3 Performance Optimization

17.8k
ruvnetruvnet

Achieve aggressive v3 performance targets: 2.49x-7.47x Flash Attention speedup, 150x-12,500x search improvements, 50-75% memory reduction. Comprehensive benchmarking and optimization suite.

192 days ago

benchmark-kernel

5.1k
flashinfer-aiflashinfer-ai

Guide for benchmarking FlashInfer kernels with CUPTI timing

192 days ago

social-media-analyzer

2.2k
alirezarezvanialirezarezvani

Social media campaign analysis and performance tracking. Calculates engagement rates, ROI, and benchmarks across platforms. Use for analyzing social media performance, calculating engagement rate, measuring campaign ROI, comparing platform metrics, or benchmarking against industry standards.

192 days ago

running-performance-tests

1.5k
Jeremylongshore Claude Code Plugins Plus Skills Running Performance TestsJeremylongshore Claude Code Plugins Plus Skills Running Performance Tests

Execute load testing, stress testing, and performance benchmarking. Use when performing specialized testing. Trigger with phrases like "run load tests", "test performance", or "benchmark the system".

192 days ago

tbench

1.3k
codercoder

Terminal-Bench integration for Mux agent benchmarking and failure analysis

192 days ago

criterium

1.2k
Hugoduncan Criterium CriteriumHugoduncan Criterium Criterium

Use this skill when users ask about benchmarking Clojure code, measuring performance, profiling execution time, or using the criterium library. Covers the 0.5.x API including bench macro, bench plans, viewers, domain analysis, and argument generation.

192 days ago

knowledge-worker-salaries

739
danielmiesslerdanielmiessler

Comprehensive global knowledge worker salary data with total market value calculations, sector breakdowns, geographic comparisons, and authoritative sources. USE WHEN discussing knowledge worker compensation, salary benchmarking, economic analysis of professional labor markets, or AI impact on wages.

192 days ago

defeatbeta-analyst

485
defeat-betadefeat-beta

Comprehensive financial analysis using 60+ data endpoints. Analyze company fundamentals, financial statements, valuation metrics, profitability ratios, growth trends, and industry comparisons. Use for: (1) fundamental analysis and DCF modeling, (2) financial statement analysis, (3) valuation and ratio analysis, (4) growth and profitability assessment, (5) industry benchmarking, or any deep financial research tasks.

192 days ago

benchmarking-analyst

376
a5c-aia5c-ai

Benchmarking analysis skill for performance comparison and best practice identification.

192 days ago

reputation-intelligence

376
a5c-aia5c-ai

Reputation measurement and benchmarking platform integration

192 days ago

transportation-spend-analyzer

376
a5c-aia5c-ai

Freight spend analysis and benchmarking skill for cost optimization and carrier negotiation support

192 days ago

performance-benchmark-suite

376
a5c-aia5c-ai

SDK performance benchmarking and regression detection

192 days ago

gpu-benchmarking

376
a5c-aia5c-ai

Expert skill for automated GPU performance benchmarking and regression detection. Design micro-benchmarks, measure kernel execution time with CUDA events, calculate achieved vs theoretical performance, generate comparison reports, detect regressions in CI/CD, and profile power/thermal characteristics.

192 days ago

transfer-pricing-analyzer

376
a5c-aia5c-ai

Intercompany transfer pricing analysis skill with benchmarking and documentation generation

192 days ago

energy-auditor

376
a5c-aia5c-ai

Process energy audit skill for consumption analysis, benchmarking, and efficiency improvement identification

192 days ago

comp-benchmarking

376
a5c-aia5c-ai

Analyze market compensation data and establish competitive pay structures

192 days ago

network-testing

376
a5c-aia5c-ai

Comprehensive network testing, benchmarking, and performance validation skill

192 days ago

logistics-kpi-tracker

376
a5c-aia5c-ai

Comprehensive logistics performance measurement skill with KPI tracking, benchmarking, and improvement recommendations

192 days ago

rb-benchmarker

376
a5c-aia5c-ai

Randomized benchmarking skill for gate fidelity characterization

192 days ago

evaluating-code-models

353
sangrokjungsangrokjung

Evaluates code generation models across HumanEval, MBPP, MultiPL-E, and 15+ benchmarks with pass@k metrics. Use when benchmarking code models, comparing coding abilities, testing multi-language support, or measuring code generation quality. Industry standard from BigCode Project used by HuggingFace leaderboards.

EvaluationCode GenerationHumanEval+6
192 days ago

evaluating-llms-harness

353
sangrokjungsangrokjung

Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking model quality, comparing models, reporting academic results, or tracking training progress. Industry standard used by EleutherAI, HuggingFace, and major labs. Supports HuggingFace, vLLM, APIs.

EvaluationLM Evaluation HarnessBenchmarking+7
192 days ago

go-test-expert

245
GoogleCloudPlatformGoogleCloudPlatform

Expert in Go testing patterns, table-driven tests, httptest, benchmarking, and fuzzing. Activates for "test", "fail", "benchmark", "debug", "fuzz".

192 days ago

benchmarking

234
Prorise-coolProrise-cool

Benchmarking and competitive analysis techniques. Compares performance, processes, and practices against industry standards, competitors, and best-in-class organizations.

192 days ago

V3 Performance Optimization

215
proffesor-for-testingproffesor-for-testing

Achieve aggressive v3 performance targets: 2.49x-7.47x Flash Attention speedup, 150x-12,500x search improvements, 50-75% memory reduction. Comprehensive benchmarking and optimization suite.

192 days ago

performance-optimization

212
ynulihaoynulihao

Apply systematic performance optimization techniques when writing or reviewing code. Use when optimizing hot paths, reducing latency, improving throughput, fixing performance regressions, or when the user mentions performance, optimization, speed, latency, throughput, profiling, or benchmarking.

192 days ago

llm-evaluation

132
MicrockMicrock

Implement comprehensive evaluation strategies for LLM applications using automated metrics, human feedback, and benchmarking. Use when testing LLM performance, measuring AI application quality, or establishing evaluation frameworks.

192 days ago

financial-analysis

122
LerianStudioLerianStudio

Comprehensive financial analysis workflow covering ratio analysis, trend analysis, benchmarking, and variance analysis. Delivers documented, audit-ready insights.

192 days ago

profiling-optimization

95
aj-geddesaj-geddes

Profile application performance, identify bottlenecks, and optimize hot paths using CPU profiling, flame graphs, and benchmarking. Use when investigating performance issues or optimizing critical code paths.

192 days ago

WebAssembly Testing

71
PramodDuttaPramodDutta

Testing WebAssembly modules including compilation verification, memory management, interop testing, and performance benchmarking of WASM components.

wasmwebassemblymemory+2
192 days ago

optimization-phase

70
marcusgollmarcusgoll

Validates production readiness through performance benchmarking, accessibility audits, security reviews, and code quality checks. Use after implementation phase completes, before deployment, or when conducting quality gates for features. (project)

192 days ago

Historical Cost Analyzer

46
datadrivenconstructiondatadrivenconstruction

Analyze historical construction costs for benchmarking, trend analysis, and estimating calibration. Compare projects, track escalation, identify patterns.

192 days ago

performance-benchmark-specialist

40
manutejmanutej

Performance benchmarking expertise for shell tools, covering benchmark design, statistical analysis (min/max/mean/median/stddev), performance targets (<100ms, >90% hit rate), workspace generation, and comprehensive reporting

192 days ago

llm_evaluation

39
vuralserhat86vuralserhat86

Implement comprehensive evaluation strategies for LLM applications using automated metrics, human feedback, and benchmarking. Use when testing LLM performance, measuring AI application quality, or establishing evaluation frameworks.

192 days ago

performance-monitor

39
404kidwiz404kidwiz

Expert in observing, benchmarking, and optimizing AI agents. Specializes in token usage tracking, latency analysis, and quality evaluation metrics. Use when optimizing agent costs, measuring performance, or implementing evals. Triggers include "agent performance", "token usage", "latency optimization", "eval", "agent metrics", "cost optimization", "agent benchmarking".

192 days ago

competitive-analyst

36
zenobi-uszenobi-us

Expert competitive analyst specializing in competitor intelligence, strategic analysis, and market positioning. Masters competitive benchmarking, SWOT analysis, and strategic recommendations with focus on creating sustainable competitive advantages.

192 days ago

optimizing-r

32
jeremy-allenjeremy-allen

R performance profiling, benchmarking, and optimization strategies. Use this skill when code is running slowly, comparing alternative implementations, deciding between dplyr/data.table/base R, or implementing parallel processing. Covers profvis and bench usage, performance workflow, parallel processing with in_parallel(), data backend selection, modern purrr patterns (list_rbind, walk), and common performance anti-patterns to avoid.

192 days ago

Performance Analysis and Optimization

30
ShunsukeHayashiShunsukeHayashi

CPU profiling, benchmarking, and memory analysis for Rust applications. Use when code is slow, memory usage is high, or optimization is needed.

192 days ago

Performance Thinker

29
omer-metinomer-metin

Performance optimization mindset - knowing when to optimize, how to measure, where bottlenecks hide, and when "fast enough" is the right answer. Use when "slow, performance, optimize, profiling, benchmark, latency, throughput, cache, n+1, bottleneck, memory leak, too slow, speed up, response time, performance, optimization, profiling, caching, latency, throughput, big-o, benchmarking" are mentioned.

192 days ago

agent-evaluation

29
omer-metinomer-metin

Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on real-world benchmarks. Use when "agent testing, agent evaluation, benchmark agents, agent reliability, test agent, testing, evaluation, benchmark, agents, reliability, quality" are mentioned.

192 days ago

eval-recipes-runner

28
rysweetrysweet

Run Microsoft's eval-recipes benchmarks to validate amplihack improvements against baseline agents. Auto-activates when testing improvements, running evals, or benchmarking changes.

192 days ago

model-evaluation-benchmark

28
rysweetrysweet

Automated reproduction of comprehensive model evaluation benchmarks following the Benchmark Suite V3. Auto-activates for model benchmarking, comparison evaluation, or performance testing between AI models.

192 days ago