llm-evaluation

26
lifangdalifangda

Implement comprehensive evaluation strategies for LLM applications using automated metrics, human feedback, and benchmarking. Use when testing LLM performance, measuring AI application quality, or establishing evaluation frameworks.

194 days ago

performance-optimizer

25
nahisahonahisaho

Copilot agent that assists with performance analysis, bottleneck detection, optimization strategies, and benchmarking Trigger terms: performance optimization, performance tuning, profiling, benchmark, bottleneck analysis, scalability, latency optimization, memory optimization, query optimization Use when: User requests involve performance optimizer tasks.

194 days ago

express-to-fastify-migration

25
timrogerstimrogers

Migrate Express.js REST APIs to Fastify with automated testing, performance benchmarking, and schema generation. Use when migrating Express applications to Fastify, modernizing Node.js APIs, improving API performance, or when users mention Express to Fastify migration, Fastify conversion, API modernization, or performance optimization of Express apps.

194 days ago

ordo-testing

23
Pama-LeePama-Lee

Ordo testing and benchmarking guide. Includes unit tests, integration tests, Criterion benchmarks, k6 load tests, CI configuration. Use for writing tests, performance analysis, continuous integration.

194 days ago

performance

23
outfitter-devoutfitter-dev

This skill should be used when profiling code, optimizing bottlenecks, benchmarking, or when "performance", "profiling", "optimization", or "--perf" are mentioned.

194 days ago

rust-performance

20
huialihuiali

Performance optimization expert covering profiling, benchmarking, memory allocation, SIMD, cache optimization, false sharing, lock contention, and NUMA-aware programming.

194 days ago

java-performance

19
pluginagentmarketplacepluginagentmarketplace

JVM performance tuning - GC optimization, profiling, memory analysis, benchmarking

194 days ago

perf

18
johnlindquistjohnlindquist

Performance profiling and optimization. Use for benchmarking code, analyzing performance, running Lighthouse audits, and finding hotspots.

194 days ago

Benchmarking & Optimization

15
tobitobi

Use this skill when the user asks to run benchmarks, profile performance, measure allocations, optimize render speed, find hot paths, generate a flamegraph, or mentions stackprof, memory profiling, or performance optimization.

194 days ago

llm-evaluation

15
aisa-groupaisa-group

Implement comprehensive evaluation strategies for LLM applications using automated metrics, human feedback, and benchmarking. Use when testing LLM performance, measuring AI application quality, or establishing evaluation frameworks.

194 days ago

anysite Content Analytics

13
anysiteioanysiteio

Track and analyze content performance across Instagram, YouTube, LinkedIn, Twitter/X, and Reddit using anysite MCP server. Measure engagement metrics, analyze post effectiveness, benchmark content strategy, identify top-performing content, and optimize posting strategies. Supports post performance tracking, engagement analysis, content type comparison, and competitive benchmarking. Use when users need to measure content ROI, optimize social strategy, identify viral content patterns, or analyze content engagement across platforms.

194 days ago

evaluation-harness

12
patricio0312revpatricio0312rev

Builds repeatable evaluation systems with golden datasets, scoring rubrics, pass/fail thresholds, and regression reports. Use for LLM evaluation, testing AI systems, quality assurance, or model benchmarking.

194 days ago

lease-comparison-expert

9
reggiechan74reggiechan74

Expert in lease-to-lease comparison and deviation analysis. Use when comparing lease amendments to originals, analyzing competing offers, benchmarking against precedents, or identifying deal term variations. Key terms include lease comparison, amendment analysis, offer comparison, precedent deviation, market benchmarking, competitive analysis

comparisonamendmentoffer-analysis+3
194 days ago

Compete

8
simotasimota

Competitive research, identifying differentiation points, and positioning. Competitive feature matrix, differentiation strategy, SWOT analysis, benchmarking, and positioning map. Use when strategic decision‑making support is needed. Do not write code.

194 days ago

running-performance-tests

8
BbgnsurfTechBbgnsurfTech

Execute load testing, stress testing, and performance benchmarking. Use when performing specialized testing. Trigger with phrases like "run load tests", "test performance", or "benchmark the system".

194 days ago

performance-at-scale

7
Bbeierle12Bbeierle12

Spatial indexing and world streaming for Three.js building games with thousands of pieces. Use when optimizing building games, implementing spatial queries, chunk loading, or profiling performance. Includes spatial hash grids, octrees, chunk managers, and benchmarking tools.

194 days ago

go-optimization

7
geoffjaygeoffjay

Performance optimization techniques including profiling, memory management, benchmarking, and runtime tuning. Use when optimizing Go code performance, reducing memory usage, or analyzing bottlenecks.

194 days ago

hook-optimization

7
duongdevduongdev

Provides guidance on optimizing CCPM hooks for performance and token efficiency. Auto-activates when developing, debugging, or benchmarking hooks. Includes caching strategies, token budgets, performance benchmarking, and best practices for maintaining sub-5-second hook execution times.

194 days ago

benchmarking-ml-models

6
jiunbaejiunbae

Runs ML model benchmarks and evaluations. Measures inference speed, memory usage, and accuracy metrics. Use for "벤치마크", "모델 평가", "성능 테스트", "inference 속도" requests.

194 days ago

agent-evaluation

5
agent-skills-hubagent-skills-hub

Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on real-world benchmarks Use when: agent testing, agent evaluation, benchmark agents, agent reliability, test agent.

194 days ago

llm-evaluation

5
agent-skills-hubagent-skills-hub

Implement comprehensive evaluation strategies for LLM applications using automated metrics, human feedback, and benchmarking. Use when testing LLM performance, measuring AI application quality, or establishing evaluation frameworks.

194 days ago

testing-agents

4
gitwaltergitwalter

Unit testing with mocks, integration testing, LangSmith evaluation, benchmarking

194 days ago

agent-evaluation

4
ngxtmngxtm

Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on re...

194 days ago

llm-evaluation

4
ngxtmngxtm

Implement comprehensive evaluation strategies for LLM applications using automated metrics, human feedback, and benchmarking. Use when testing LLM performance, measuring AI application quality, or ...

194 days ago

custom-indicator

4
marketcallsmarketcalls

Create a custom technical indicator using Numba JIT + NumPy. Generates production-grade, O(n) optimized indicator functions with charting and benchmarking.

194 days ago

Optimizing R

3
akm-rsakm-rs

R performance profiling, benchmarking, and optimization strategies. Use this skill when code is running slowly, comparing alternative implementations, deciding between dplyr/data.table/base R, or implementing parallel processing. Covers profvis and bench usage, performance workflow, parallel processing with in_parallel(), data backend selection, modern purrr patterns (list_rbind, walk), and common performance anti-patterns to avoid.

194 days ago

crypto-portfolio-management

3
SperaxSperax

Guide to cryptocurrency portfolio management — asset allocation, rebalancing strategies, risk-adjusted returns, benchmarking, and tax-loss harvesting. Use when helping users build portfolios, rebalance holdings, or evaluate portfolio performance.

194 days ago

local-government-finance-benchmarks

3
taivoptaivop

Query Ministry of Finance local-government and state-accountancy sources for municipal finance benchmarking and cross-municipality comparison context.

194 days ago

local-government-finance-benchmarks

3
taivoptaivop

Query Ministry of Finance local-government and state-accountancy sources for municipal finance benchmarking and cross-municipality comparison context.

194 days ago

perf-profiler

1
gitgoodordietryinggitgoodordietrying

Profile and optimize application performance. Use when diagnosing slow code, measuring CPU/memory usage, generating flame graphs, benchmarking functions, load testing APIs, finding memory leaks, or optimizing database queries.

194 days ago

rltools-testing

1
chuongdlbchuongdlb

Testing and benchmarking for rl-tools — GoogleTest, benchmark targets, environment correctness, NN inference speed, memory profiling for embedded.

testingbenchmarkinggoogletest
194 days ago

V3 Performance Optimization

1
jbeck018jbeck018

Achieve aggressive v3 performance targets: 2.49x-7.47x Flash Attention speedup, 150x-12,500x search improvements, 50-75% memory reduction. Comprehensive benchmarking and optimization suite.

194 days ago

agent-evaluation

1
rootcastlecorootcastleco

Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on re...

194 days ago

perf-profiler

1
gitgoodordietryinggitgoodordietrying

Profile and optimize application performance. Use when diagnosing slow code, measuring CPU/memory usage, generating flame graphs, benchmarking functions, load testing APIs, finding memory leaks, or optimizing database queries.

194 days ago

perf-profiler

1
gitgoodordietryinggitgoodordietrying

Profile and optimize application performance. Use when diagnosing slow code, measuring CPU/memory usage, generating flame graphs, benchmarking functions, load testing APIs, finding memory leaks, or optimizing database queries.

194 days ago

llm-evaluation

1
mattmremattmre

Implement comprehensive evaluation strategies for LLM applications using automated metrics, human feedback, and benchmarking. Use when testing LLM performance, measuring AI application quality, or establishing evaluation frameworks.

194 days ago

dotnet-ci-benchmarking

1
wshaddixwshaddix

Gating CI on perf regressions. Automated threshold alerts, baseline tracking, trend reports.

194 days ago

LLM Evaluation

oki3505Foki3505F

Implement comprehensive evaluation strategies for LLM applications using automated metrics, human feedback, and benchmarking. Use when testing LLM performance, measuring AI application quality, or establishing evaluation frameworks.

194 days ago

transfer-pricing

KaakatiKaakati

Transfer pricing policy design, documentation, and defense aligned with OECD Guidelines and the arm's length principle. USE THIS SKILL when the user asks about transfer pricing, intercompany pricing, TP documentation, arm's length pricing, intercompany transactions, benchmarking studies, comparability analysis, functional analysis, FAR analysis, advance pricing agreements, master file, local file, Country-by-Country Reporting, management fees, intercompany loans, cost sharing arrangements, or MAP/arbitration for double taxation disputes.

194 days ago

agent-evaluation

oki3505Foki3505F

Testing and benchmarking LLM-driven agents, including behavioral testing, capability assessment, reliability metrics, and production monitoring—noting that even top agents often score below 50% on real-world benchmarks. Use when: agent testing, agent evaluation, benchmarking agents, assessing agent reliability, or test-driving agents.

194 days ago

digital-transformation

KaakatiKaakati

Digital transformation maturity assessment, peer benchmarking, and phased roadmap development. USE THIS SKILL when the user asks about digital maturity, digital strategy, digitization roadmap, technology modernization, digital readiness, cloud migration strategy, digital operating model, automation strategy, digital KPIs, or "how digitally mature are we." Also trigger when asked to benchmark digital capabilities, prioritize digital initiatives, or build a transformation business case for any organization or business unit.

194 days ago

Agent Evaluation

haniakrim21haniakrim21

Testing and benchmarking LLM-driven agents — including behavioral testing, capability assessment, reliability metrics, and production monitoring — highlighting that even top agents score below 50% on real-world benchmarks. Use when: agent testing, agent evaluation, benchmarking agents, assessing agent reliability, or testing agents.

194 days ago

evaluating-llms-harness

AXGZ21AXGZ21

Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking model quality, comparing models, reporting academic results, or tracking training progress. Industry standard used by EleutherAI, HuggingFace, and major labs. Supports HuggingFace, vLLM, APIs.

194 days ago

generate-config

greynewellgreynewell

Generate and validate mcpbr configuration files for MCP server benchmarking.

194 days ago

profiling-guide

AeyeOpsAeyeOps

Performance profiling methodologies, interpreting profiler output, and benchmarking techniques. Activate when: profiling application performance, reading flame graphs, benchmarking code, identifying bottlenecks, measuring CPU usage, analyzing memory allocation, I/O profiling.

194 days ago

Cost Reduction & Margin Improvement

KaakatiKaakati

USE THIS SKILL when the user asks about cost reduction, cost cutting, margin improvement, profitability analysis, zero-based budgeting (ZBB), activity-based costing (ABC), cost benchmarking, cost transformation, SG&A optimization, overhead reduction, cost waterfall, cost driver analysis, operating leverage, or business case for savings initiatives. Also trigger for "run-rate savings," "cost take-out," "efficiency program," "restructuring," "right-sizing," or any request to reduce costs or improve EBITDA margin.

194 days ago

honest-review

wyattowalshwyattowalsh

Research-driven code review with confidence-scored, evidence-validated findings. Session review or full codebase audit via parallel teams. Use when reviewing changes, auditing codebases, verifying work quality. NOT for writing new code, explaining code, or benchmarking.

194 days ago

gsc

spivxspivx

Live Google Search Console analytics — fetches real SEO data (clicks, impressions, CTR, rankings) and delivers actionable insights with CTR benchmarking and opportunity detection. Zero dependencies. Use when the user asks about GSC, Google Search Console, SEO performance, search performance, keywords, rankings, organic traffic, top pages, top queries, "how is my site performing in Google", "check rankings", or "search console report".

194 days ago