Testing and benchmarking LLM-driven agents — including behavioral testing, capability assessment, reliability metrics, and production monitoring — highlighting that even top agents score below 50% on real-world benchmarks. Use when: agent testing, agent evaluation, benchmarking agents, assessing agent reliability, or testing agents.
Testing and benchmarking LLM-driven agents — including behavioral testing, capability assessment, reliability metrics, and production monitoring — highlighting that even top agents score below 50% on real-world benchmarks. Use when: agent testing, agent evaluation, benchmarking agents, assessing agent reliability, or testing agents.
对 LLM 驱动的代理(agents)进行测试与基准评估,包括行为测试、能力评估、可靠性指标和生产监控——并指出即使表现最好的代理在真实世界基准上也常低于 50%。适用场景:代理测试、代理评估、基准代理、代理可靠性、测试代理。
Category: developer (开发工具) · Author: haniakrim21 · Version: @main · License: MIT
Search for places (restaurants, cafes, etc.) via Google Places API proxy on localhost.
Interact with GitHub using the `gh` CLI. Use `gh issue`, `gh pr`, `gh run`, and `gh api` for issues, PRs, CI runs, and advanced queries.
Create or update AgentSkills. Use when designing, structuring, or packaging skills with scripts, references, and assets.
Start voice calls via the OpenClaw voice-call plugin.
Notion API for creating and managing pages, databases, and blocks.
Gemini CLI for one-shot Q&A, summaries, and generation.
Category:developer