llm-evaluation
Implement comprehensive evaluation strategies for LLM applications using automated metrics, human feedback, and benchmarking. Use when testing LLM performance, measuring AI application quality, or establishing evaluation frameworks.
performance-optimizer
Copilot agent that assists with performance analysis, bottleneck detection, optimization strategies, and benchmarking Trigger terms: performance optimization, performance tuning, profiling, benchmark, bottleneck analysis, scalability, latency optimization, memory optimization, query optimization Use when: User requests involve performance optimizer tasks.
express-to-fastify-migration
Migrate Express.js REST APIs to Fastify with automated testing, performance benchmarking, and schema generation. Use when migrating Express applications to Fastify, modernizing Node.js APIs, improving API performance, or when users mention Express to Fastify migration, Fastify conversion, API modernization, or performance optimization of Express apps.
ordo-testing
Ordo testing and benchmarking guide. Includes unit tests, integration tests, Criterion benchmarks, k6 load tests, CI configuration. Use for writing tests, performance analysis, continuous integration.
performance
This skill should be used when profiling code, optimizing bottlenecks, benchmarking, or when "performance", "profiling", "optimization", or "--perf" are mentioned.
rust-performance
Performance optimization expert covering profiling, benchmarking, memory allocation, SIMD, cache optimization, false sharing, lock contention, and NUMA-aware programming.
java-performance
JVM performance tuning - GC optimization, profiling, memory analysis, benchmarking
perf
Performance profiling and optimization. Use for benchmarking code, analyzing performance, running Lighthouse audits, and finding hotspots.
Benchmarking & Optimization
Use this skill when the user asks to run benchmarks, profile performance, measure allocations, optimize render speed, find hot paths, generate a flamegraph, or mentions stackprof, memory profiling, or performance optimization.
llm-evaluation
Implement comprehensive evaluation strategies for LLM applications using automated metrics, human feedback, and benchmarking. Use when testing LLM performance, measuring AI application quality, or establishing evaluation frameworks.
anysite Content Analytics
Track and analyze content performance across Instagram, YouTube, LinkedIn, Twitter/X, and Reddit using anysite MCP server. Measure engagement metrics, analyze post effectiveness, benchmark content strategy, identify top-performing content, and optimize posting strategies. Supports post performance tracking, engagement analysis, content type comparison, and competitive benchmarking. Use when users need to measure content ROI, optimize social strategy, identify viral content patterns, or analyze content engagement across platforms.
evaluation-harness
Builds repeatable evaluation systems with golden datasets, scoring rubrics, pass/fail thresholds, and regression reports. Use for LLM evaluation, testing AI systems, quality assurance, or model benchmarking.
lease-comparison-expert
Expert in lease-to-lease comparison and deviation analysis. Use when comparing lease amendments to originals, analyzing competing offers, benchmarking against precedents, or identifying deal term variations. Key terms include lease comparison, amendment analysis, offer comparison, precedent deviation, market benchmarking, competitive analysis
Compete
Competitive research, identifying differentiation points, and positioning. Competitive feature matrix, differentiation strategy, SWOT analysis, benchmarking, and positioning map. Use when strategic decision‑making support is needed. Do not write code.
running-performance-tests
Execute load testing, stress testing, and performance benchmarking. Use when performing specialized testing. Trigger with phrases like "run load tests", "test performance", or "benchmark the system".
performance-at-scale
Spatial indexing and world streaming for Three.js building games with thousands of pieces. Use when optimizing building games, implementing spatial queries, chunk loading, or profiling performance. Includes spatial hash grids, octrees, chunk managers, and benchmarking tools.
go-optimization
Performance optimization techniques including profiling, memory management, benchmarking, and runtime tuning. Use when optimizing Go code performance, reducing memory usage, or analyzing bottlenecks.
hook-optimization
Provides guidance on optimizing CCPM hooks for performance and token efficiency. Auto-activates when developing, debugging, or benchmarking hooks. Includes caching strategies, token budgets, performance benchmarking, and best practices for maintaining sub-5-second hook execution times.
benchmarking-ml-models
Runs ML model benchmarks and evaluations. Measures inference speed, memory usage, and accuracy metrics. Use for "벤치마크", "모델 평가", "성능 테스트", "inference 속도" requests.
agent-evaluation
Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on real-world benchmarks Use when: agent testing, agent evaluation, benchmark agents, agent reliability, test agent.
llm-evaluation
Implement comprehensive evaluation strategies for LLM applications using automated metrics, human feedback, and benchmarking. Use when testing LLM performance, measuring AI application quality, or establishing evaluation frameworks.
testing-agents
Unit testing with mocks, integration testing, LangSmith evaluation, benchmarking
agent-evaluation
Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on re...
llm-evaluation
Implement comprehensive evaluation strategies for LLM applications using automated metrics, human feedback, and benchmarking. Use when testing LLM performance, measuring AI application quality, or ...
custom-indicator
Create a custom technical indicator using Numba JIT + NumPy. Generates production-grade, O(n) optimized indicator functions with charting and benchmarking.
Optimizing R
R performance profiling, benchmarking, and optimization strategies. Use this skill when code is running slowly, comparing alternative implementations, deciding between dplyr/data.table/base R, or implementing parallel processing. Covers profvis and bench usage, performance workflow, parallel processing with in_parallel(), data backend selection, modern purrr patterns (list_rbind, walk), and common performance anti-patterns to avoid.
crypto-portfolio-management
Guide to cryptocurrency portfolio management — asset allocation, rebalancing strategies, risk-adjusted returns, benchmarking, and tax-loss harvesting. Use when helping users build portfolios, rebalance holdings, or evaluate portfolio performance.
local-government-finance-benchmarks
Query Ministry of Finance local-government and state-accountancy sources for municipal finance benchmarking and cross-municipality comparison context.
local-government-finance-benchmarks
Query Ministry of Finance local-government and state-accountancy sources for municipal finance benchmarking and cross-municipality comparison context.
perf-profiler
Profile and optimize application performance. Use when diagnosing slow code, measuring CPU/memory usage, generating flame graphs, benchmarking functions, load testing APIs, finding memory leaks, or optimizing database queries.
rltools-testing
Testing and benchmarking for rl-tools — GoogleTest, benchmark targets, environment correctness, NN inference speed, memory profiling for embedded.
V3 Performance Optimization
Achieve aggressive v3 performance targets: 2.49x-7.47x Flash Attention speedup, 150x-12,500x search improvements, 50-75% memory reduction. Comprehensive benchmarking and optimization suite.
agent-evaluation
Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on re...
perf-profiler
Profile and optimize application performance. Use when diagnosing slow code, measuring CPU/memory usage, generating flame graphs, benchmarking functions, load testing APIs, finding memory leaks, or optimizing database queries.
perf-profiler
Profile and optimize application performance. Use when diagnosing slow code, measuring CPU/memory usage, generating flame graphs, benchmarking functions, load testing APIs, finding memory leaks, or optimizing database queries.
llm-evaluation
Implement comprehensive evaluation strategies for LLM applications using automated metrics, human feedback, and benchmarking. Use when testing LLM performance, measuring AI application quality, or establishing evaluation frameworks.
dotnet-ci-benchmarking
Gating CI on perf regressions. Automated threshold alerts, baseline tracking, trend reports.
LLM Evaluation
Implement comprehensive evaluation strategies for LLM applications using automated metrics, human feedback, and benchmarking. Use when testing LLM performance, measuring AI application quality, or establishing evaluation frameworks.
transfer-pricing
Transfer pricing policy design, documentation, and defense aligned with OECD Guidelines and the arm's length principle. USE THIS SKILL when the user asks about transfer pricing, intercompany pricing, TP documentation, arm's length pricing, intercompany transactions, benchmarking studies, comparability analysis, functional analysis, FAR analysis, advance pricing agreements, master file, local file, Country-by-Country Reporting, management fees, intercompany loans, cost sharing arrangements, or MAP/arbitration for double taxation disputes.
agent-evaluation
Testing and benchmarking LLM-driven agents, including behavioral testing, capability assessment, reliability metrics, and production monitoring—noting that even top agents often score below 50% on real-world benchmarks. Use when: agent testing, agent evaluation, benchmarking agents, assessing agent reliability, or test-driving agents.
digital-transformation
Digital transformation maturity assessment, peer benchmarking, and phased roadmap development. USE THIS SKILL when the user asks about digital maturity, digital strategy, digitization roadmap, technology modernization, digital readiness, cloud migration strategy, digital operating model, automation strategy, digital KPIs, or "how digitally mature are we." Also trigger when asked to benchmark digital capabilities, prioritize digital initiatives, or build a transformation business case for any organization or business unit.
Agent Evaluation
Testing and benchmarking LLM-driven agents — including behavioral testing, capability assessment, reliability metrics, and production monitoring — highlighting that even top agents score below 50% on real-world benchmarks. Use when: agent testing, agent evaluation, benchmarking agents, assessing agent reliability, or testing agents.
evaluating-llms-harness
Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking model quality, comparing models, reporting academic results, or tracking training progress. Industry standard used by EleutherAI, HuggingFace, and major labs. Supports HuggingFace, vLLM, APIs.
generate-config
Generate and validate mcpbr configuration files for MCP server benchmarking.
profiling-guide
Performance profiling methodologies, interpreting profiler output, and benchmarking techniques. Activate when: profiling application performance, reading flame graphs, benchmarking code, identifying bottlenecks, measuring CPU usage, analyzing memory allocation, I/O profiling.
Cost Reduction & Margin Improvement
USE THIS SKILL when the user asks about cost reduction, cost cutting, margin improvement, profitability analysis, zero-based budgeting (ZBB), activity-based costing (ABC), cost benchmarking, cost transformation, SG&A optimization, overhead reduction, cost waterfall, cost driver analysis, operating leverage, or business case for savings initiatives. Also trigger for "run-rate savings," "cost take-out," "efficiency program," "restructuring," "right-sizing," or any request to reduce costs or improve EBITDA margin.
honest-review
Research-driven code review with confidence-scored, evidence-validated findings. Session review or full codebase audit via parallel teams. Use when reviewing changes, auditing codebases, verifying work quality. NOT for writing new code, explaining code, or benchmarking.
gsc
Live Google Search Console analytics — fetches real SEO data (clicks, impressions, CTR, rankings) and delivers actionable insights with CTR benchmarking and opportunity detection. Zero dependencies. Use when the user asks about GSC, Google Search Console, SEO performance, search performance, keywords, rankings, organic traffic, top pages, top queries, "how is my site performing in Google", "check rankings", or "search console report".