golang-testing

57.0k
affaan-maffaan-m

Go testing patterns including table-driven tests, subtests, benchmarks, fuzzing, and test coverage. Follows TDD methodology with idiomatic Go practices.

192 days ago

llm-evaluation

29.9k
wshobsonwshobson

Implement comprehensive evaluation strategies for LLM applications using automated metrics, human feedback, and benchmarking. Use when testing LLM performance, measuring AI application quality, or establishing evaluation frameworks.

192 days ago

agent-evaluation

21.8k
davila7davila7

Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on real-world benchmarks Use when: agent testing, agent evaluation, benchmark agents, agent reliability, test agent.

192 days ago

evaluating-llms-harness

21.8k
davila7davila7

Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking model quality, comparing models, reporting academic results, or tracking training progress. Industry standard used by EleutherAI, HuggingFace, and major labs. Supports HuggingFace, vLLM, APIs.

EvaluationLM Evaluation HarnessBenchmarking+7
192 days ago

pytdc

21.8k
davila7davila7

Therapeutics Data Commons. AI-ready drug discovery datasets (ADME, toxicity, DTI), benchmarks, scaffold splits, molecular oracles, for therapeutic ML and pharmacological prediction.

192 days ago

deepchem

21.8k
davila7davila7

Molecular machine learning toolkit. Property prediction (ADMET, toxicity), GNNs (GCN, MPNN), MoleculeNet benchmarks, pretrained models, featurization, for drug discovery ML.

192 days ago

evaluating-code-models

21.8k
davila7davila7

Evaluates code generation models across HumanEval, MBPP, MultiPL-E, and 15+ benchmarks with pass@k metrics. Use when benchmarking code models, comparing coding abilities, testing multi-language support, or measuring code generation quality. Industry standard from BigCode Project used by HuggingFace leaderboards.

EvaluationCode GenerationHumanEval+6
192 days ago

nemo-evaluator-sdk

21.8k
davila7davila7

Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution. Use when needing scalable evaluation on local Docker, Slurm HPC, or cloud platforms. NVIDIA's enterprise-grade platform with container-first architecture for reproducible benchmarking.

EvaluationNeMoNVIDIA+8
192 days ago

cirq

21.8k
davila7davila7

Quantum computing framework for building, simulating, optimizing, and executing quantum circuits. Use this skill when working with quantum algorithms, quantum circuit design, quantum simulation (noiseless or noisy), running on quantum hardware (Google, IonQ, AQT, Pasqal), circuit optimization and compilation, noise modeling and characterization, or quantum experiments and benchmarking (VQE, QAOA, QPE, randomized benchmarking).

192 days ago

pymoo

21.8k
davila7davila7

Multi-objective optimization framework. NSGA-II, NSGA-III, MOEA/D, Pareto fronts, constraint handling, benchmarks (ZDT, DTLZ), for engineering design and optimization problems.

192 days ago

llm-evaluation

18.0k
sickn33sickn33

Implement comprehensive evaluation strategies for LLM applications using automated metrics, human feedback, and benchmarking. Use when testing LLM performance, measuring AI application quality, or establishing evaluation frameworks.

192 days ago

V3 Performance Optimization

17.8k
ruvnetruvnet

Achieve aggressive v3 performance targets: 2.49x-7.47x Flash Attention speedup, 150x-12,500x search improvements, 50-75% memory reduction. Comprehensive benchmarking and optimization suite.

192 days ago

deepchem

10.8k
K-Dense-AIK-Dense-AI

Molecular ML with diverse featurizers and pre-built datasets. Use for property prediction (ADMET, toxicity) with traditional ML or GNNs when you want extensive featurization options and MoleculeNet benchmarks. Best for quick experiments with pre-trained models, diverse molecular representations. For graph-first PyTorch workflows use torchdrug; for benchmark datasets use pytdc.

192 days ago

pymoo

10.8k
K-Dense-AIK-Dense-AI

Multi-objective optimization framework. NSGA-II, NSGA-III, MOEA/D, Pareto fronts, constraint handling, benchmarks (ZDT, DTLZ), for engineering design and optimization problems.

192 days ago

pytdc

10.8k
K-Dense-AIK-Dense-AI

Therapeutics Data Commons. AI-ready drug discovery datasets (ADME, toxicity, DTI), benchmarks, scaffold splits, molecular oracles, for therapeutic ML and pharmacological prediction.

192 days ago

torchdrug

10.8k
K-Dense-AIK-Dense-AI

PyTorch-native graph neural networks for molecules and proteins. Use when building custom GNN architectures for drug discovery, protein modeling, or knowledge graph reasoning. Best for custom model development, protein property prediction, retrosynthesis. For pre-trained models and diverse featurizers use deepchem; for benchmark datasets use pytdc.

192 days ago

benchmark-kernel

5.1k
flashinfer-aiflashinfer-ai

Guide for benchmarking FlashInfer kernels with CUPTI timing

192 days ago

social-media-analyzer

2.2k
alirezarezvanialirezarezvani

Social media campaign analysis and performance tracking. Calculates engagement rates, ROI, and benchmarks across platforms. Use for analyzing social media performance, calculating engagement rate, measuring campaign ROI, comparing platform metrics, or benchmarking against industry standards.

192 days ago

bench

1.8k
opendataloader-projectopendataloader-project

Run benchmark and analyze PDF parsing performance

192 days ago

aws-security-scanner

1.8k
openclawopenclaw

Scan AWS accounts for security misconfigurations and vulnerabilities. Use when user asks to audit AWS security, check for misconfigurations, find exposed S3 buckets, review IAM policies, check security groups, audit CloudTrail, or run AWS security checks. Covers S3, IAM, EC2, RDS, CloudTrail, and common CIS benchmarks.

192 days ago

benchmark-suite-creator

1.5k
Jeremylongshore Claude Code Plugins Plus Skills Benchmark Suite CreatorJeremylongshore Claude Code Plugins Plus Skills Benchmark Suite Creator

Benchmark Suite Creator - Auto-activating skill for Performance Testing. Triggers on: benchmark suite creator, benchmark suite creator Part of the Performance Testing skill category.

191 days ago

running-performance-tests

1.5k
Jeremylongshore Claude Code Plugins Plus Skills Running Performance TestsJeremylongshore Claude Code Plugins Plus Skills Running Performance Tests

Execute load testing, stress testing, and performance benchmarking. Use when performing specialized testing. Trigger with phrases like "run load tests", "test performance", or "benchmark the system".

191 days ago

windsurf-load-scale

1.5k
Jeremylongshore Claude Code Plugins Plus Skills Windsurf Load ScaleJeremylongshore Claude Code Plugins Plus Skills Windsurf Load Scale

Implement Windsurf load testing, auto-scaling, and capacity planning strategies. Use when running performance tests, configuring horizontal scaling, or planning capacity for Windsurf integrations. Trigger with phrases like "windsurf load test", "windsurf scale", "windsurf performance test", "windsurf capacity", "windsurf k6", "windsurf benchmark".

192 days ago

exa-load-scale

1.5k
Jeremylongshore Claude Code Plugins Plus Skills Exa Load ScaleJeremylongshore Claude Code Plugins Plus Skills Exa Load Scale

Implement Exa load testing, auto-scaling, and capacity planning strategies. Use when running performance tests, configuring horizontal scaling, or planning capacity for Exa integrations. Trigger with phrases like "exa load test", "exa scale", "exa performance test", "exa capacity", "exa k6", "exa benchmark".

192 days ago

load-testing-apis

1.5k
jeremylongshorejeremylongshore

Execute comprehensive load and stress testing to validate API performance and scalability. Use when validating API performance under load. Trigger with phrases like "load test the API", "stress test API", or "benchmark API performance".

192 days ago

replit-load-scale

1.5k
Jeremylongshore Claude Code Plugins Plus Skills Replit Load ScaleJeremylongshore Claude Code Plugins Plus Skills Replit Load Scale

Implement Replit load testing, auto-scaling, and capacity planning strategies. Use when running performance tests, configuring horizontal scaling, or planning capacity for Replit integrations. Trigger with phrases like "replit load test", "replit scale", "replit performance test", "replit capacity", "replit k6", "replit benchmark".

192 days ago

perplexity-load-scale

1.5k
Jeremylongshore Claude Code Plugins Plus Skills Perplexity Load ScaleJeremylongshore Claude Code Plugins Plus Skills Perplexity Load Scale

Implement Perplexity load testing, auto-scaling, and capacity planning strategies. Use when running performance tests, configuring horizontal scaling, or planning capacity for Perplexity integrations. Trigger with phrases like "perplexity load test", "perplexity scale", "perplexity performance test", "perplexity capacity", "perplexity k6", "perplexity benchmark".

192 days ago

firecrawl-load-scale

1.5k
jeremylongshorejeremylongshore

Implement FireCrawl load testing, auto-scaling, and capacity planning strategies. Use when running performance tests, configuring horizontal scaling, or planning capacity for FireCrawl integrations. Trigger with phrases like "firecrawl load test", "firecrawl scale", "firecrawl performance test", "firecrawl capacity", "firecrawl k6", "firecrawl benchmark".

192 days ago

security-benchmark-runner

1.5k
jeremylongshorejeremylongshore

Manage security benchmark runner operations. Auto-activating skill for Security Advanced. Triggers on: security benchmark runner, security benchmark runner Part of the Security Advanced skill category. Use when working with security benchmark runner functionality. Trigger with phrases like "security benchmark runner", "security runner", "security".

192 days ago

clay-load-scale

1.5k
Jeremylongshore Claude Code Plugins Plus Skills Clay Load ScaleJeremylongshore Claude Code Plugins Plus Skills Clay Load Scale

Implement Clay load testing, auto-scaling, and capacity planning strategies. Use when running performance tests, configuring horizontal scaling, or planning capacity for Clay integrations. Trigger with phrases like "clay load test", "clay scale", "clay performance test", "clay capacity", "clay k6", "clay benchmark".

192 days ago

supabase-load-scale

1.5k
Jeremylongshore Claude Code Plugins Plus Skills Supabase Load ScaleJeremylongshore Claude Code Plugins Plus Skills Supabase Load Scale

Implement Supabase load testing, auto-scaling, and capacity planning strategies. Use when running performance tests, configuring horizontal scaling, or planning capacity for Supabase integrations. Trigger with phrases like "supabase load test", "supabase scale", "supabase performance test", "supabase capacity", "supabase k6", "supabase benchmark".

192 days ago

vercel-load-scale

1.5k
Jeremylongshore Claude Code Plugins Plus Skills Vercel Load ScaleJeremylongshore Claude Code Plugins Plus Skills Vercel Load Scale

Implement Vercel load testing, auto-scaling, and capacity planning strategies. Use when running performance tests, configuring horizontal scaling, or planning capacity for Vercel integrations. Trigger with phrases like "vercel load test", "vercel scale", "vercel performance test", "vercel capacity", "vercel k6", "vercel benchmark".

192 days ago

retellai-load-scale

1.5k
Jeremylongshore Claude Code Plugins Plus Skills Retellai Load ScaleJeremylongshore Claude Code Plugins Plus Skills Retellai Load Scale

Implement Retell AI load testing, auto-scaling, and capacity planning strategies. Use when running performance tests, configuring horizontal scaling, or planning capacity for Retell AI integrations. Trigger with phrases like "retellai load test", "retellai scale", "retellai performance test", "retellai capacity", "retellai k6", "retellai benchmark".

192 days ago

tbench

1.3k
codercoder

Terminal-Bench integration for Mux agent benchmarking and failure analysis

192 days ago

criterium

1.2k
Hugoduncan Criterium CriteriumHugoduncan Criterium Criterium

Use this skill when users ask about benchmarking Clojure code, measuring performance, profiling execution time, or using the criterium library. Covers the 0.5.x API including bench macro, bench plans, viewers, domain analysis, and argument generation.

192 days ago

performance-testing

1.2k
DaveSkenderDaveSkender

Benchmark indicator performance with BenchmarkDotNet. Use for Series/Buffer/Stream benchmarks, regression detection, and optimization patterns. Target 1.5x Series for StreamHub, 1.2x for BufferList.

191 days ago

bulk-rna-seq-batch-correction-with-combat

844
StarlitnightlyStarlitnightly

Use omicverse's pyComBat wrapper to remove batch effects from merged bulk RNA-seq or microarray cohorts, export corrected matrices, and benchmark pre/post correction visualizations.

192 days ago

m10-performance

786
actionbookactionbook

CRITICAL: Use for performance optimization. Triggers: performance, optimization, benchmark, profiling, flamegraph, criterion, slow, fast, allocation, cache, SIMD, make it faster

192 days ago

knowledge-worker-salaries

739
danielmiesslerdanielmiessler

Comprehensive global knowledge worker salary data with total market value calculations, sector breakdowns, geographic comparisons, and authoritative sources. USE WHEN discussing knowledge worker compensation, salary benchmarking, economic analysis of professional labor markets, or AI impact on wages.

192 days ago

analyzing-marketing-campaign

708
https-deeplearning-aihttps-deeplearning-ai

Analyze weekly marketing campaign performance data across channels. Use when analyzing multi-channel digital marketing data to calculate funnel metrics (CTR, CVR) and compare to benchmarks, compute cost and revenue efficiency metrics (ROAS, CPA, Net Profit), or get budget reallocation recommendations based on performance rules.

192 days ago

finance-metrics-quickref

679
deanpetersdeanpeters

Fast lookup table for 32+ SaaS finance metrics with formulas, benchmarks, and when to use each. Includes red flags and decision frameworks.

192 days ago

skillsbench

571
benchflow-aibenchflow-ai

SkillsBench contribution workflow. Use when: (1) Creating benchmark tasks, (2) Understanding repo structure, (3) Preparing PRs for task submission.

192 days ago

axiom-sqlitedata-migration

538
Charleswiltgen Axiom Axiom Sqlitedata MigrationCharleswiltgen Axiom Axiom Sqlitedata Migration

Use when migrating from SwiftData to SQLiteData — decision guide, pattern equivalents, code examples, CloudKit sharing (SwiftData can't), performance benchmarks, gradual migration strategy

191 days ago

perf-benchmarker

509
agent-shagent-sh

Use when running performance benchmarks, establishing baselines, or validating regressions with sequential runs. Enforces 60s minimum runs (30s only for binary search) and no parallel benchmarks.

192 days ago

defeatbeta-analyst

485
defeat-betadefeat-beta

Comprehensive financial analysis using 60+ data endpoints. Analyze company fundamentals, financial statements, valuation metrics, profitability ratios, growth trends, and industry comparisons. Use for: (1) fundamental analysis and DCF modeling, (2) financial statement analysis, (3) valuation and ratio analysis, (4) growth and profitability assessment, (5) industry benchmarking, or any deep financial research tasks.

192 days ago

adding-benchmarks

427
AztecProtocolAztecProtocol

Add new benchmarks to the CI pipeline. Guides through creating benchmark JSON files, integrating with bootstrap.sh, and ensuring proper CI upload via ci3.yml workflow.

192 days ago

giab-benchmark-validator

376
a5c-aia5c-ai

Genome in a Bottle benchmark validation skill for pipeline accuracy assessment

192 days ago

cutlass-triton

376
a5c-aia5c-ai

High-performance kernel template libraries and DSLs. Generate CUTLASS GEMM configurations, implement Triton kernel definitions, configure epilogue operations, tune tile sizes and warp arrangements, and benchmark against cuBLAS.

192 days ago