benchmarking AI Agent Skills

Browse 7 skills related to benchmarking

evaluating-llms-harness

21.8k
davila7davila7

Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking model quality, comparing models, reporting academic results, or tracking training progress. Industry standard used by EleutherAI, HuggingFace, and major labs. Supports HuggingFace, vLLM, APIs.

EvaluationLM Evaluation HarnessBenchmarking+7
140 days ago

evaluating-code-models

21.8k
davila7davila7

Evaluates code generation models across HumanEval, MBPP, MultiPL-E, and 15+ benchmarks with pass@k metrics. Use when benchmarking code models, comparing coding abilities, testing multi-language support, or measuring code generation quality. Industry standard from BigCode Project used by HuggingFace leaderboards.

EvaluationCode GenerationHumanEval+6
140 days ago

nemo-evaluator-sdk

21.8k
davila7davila7

Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution. Use when needing scalable evaluation on local Docker, Slurm HPC, or cloud platforms. NVIDIA's enterprise-grade platform with container-first architecture for reproducible benchmarking.

EvaluationNeMoNVIDIA+8
140 days ago

evaluating-code-models

353
sangrokjungsangrokjung

Evaluates code generation models across HumanEval, MBPP, MultiPL-E, and 15+ benchmarks with pass@k metrics. Use when benchmarking code models, comparing coding abilities, testing multi-language support, or measuring code generation quality. Industry standard from BigCode Project used by HuggingFace leaderboards.

EvaluationCode GenerationHumanEval+6
140 days ago

evaluating-llms-harness

353
sangrokjungsangrokjung

Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking model quality, comparing models, reporting academic results, or tracking training progress. Industry standard used by EleutherAI, HuggingFace, and major labs. Supports HuggingFace, vLLM, APIs.

EvaluationLM Evaluation HarnessBenchmarking+7
140 days ago

lease-comparison-expert

9
reggiechan74reggiechan74

Expert in lease-to-lease comparison and deviation analysis. Use when comparing lease amendments to originals, analyzing competing offers, benchmarking against precedents, or identifying deal term variations. Key terms include lease comparison, amendment analysis, offer comparison, precedent deviation, market benchmarking, competitive analysis

comparisonamendmentoffer-analysis+3
140 days ago

rltools-testing

1
chuongdlbchuongdlb

Testing and benchmarking for rl-tools — GoogleTest, benchmark targets, environment correctness, NN inference speed, memory profiling for embedded.

testingbenchmarkinggoogletest
140 days ago