monitoring AI Agent Skills
Browse 42 skills related to monitoring
phoenix-observability
Open-source AI observability platform for LLM tracing, evaluation, and monitoring. Use when debugging LLM applications with detailed traces, running evaluations on datasets, or monitoring production AI systems with real-time insights.
langsmith-observability
LLM observability platform for tracing, evaluation, and monitoring. Use when debugging LLM applications, evaluating model outputs against datasets, monitoring production systems, or building systematic testing pipelines for AI applications.
railway-metrics
Query resource usage metrics for Railway services. Use when user asks about resource usage, CPU, memory, network, disk, or service performance like "how much memory is my service using" or "is my service slow".
virus-monitor
Virus-Monitoring für Wien (Abwasser + Sentinel)
power-management-monitor
Monitor system power state including battery, AC, sleep, and wake events
sentry-desktop-setup
Configure Sentry for comprehensive desktop application crash reporting, error monitoring, performance tracking, and release health for Electron and native desktop apps
shift-right-testing
Testing in production with feature flags, canary deployments, synthetic monitoring, and chaos engineering. Use when implementing production observability or progressive delivery.
qe-shift-right-testing
Testing in production with feature flags, canary deployments, synthetic monitoring, and chaos engineering. Use when implementing production observability or progressive delivery.
observability-testing-patterns
Observability and monitoring validation patterns for dashboards, alerting, log aggregation, APM traces, and SLA/SLO verification. Use when testing monitoring infrastructure, dashboard accuracy, alert rules, or metric pipelines.
workflow-monitor
automatically create GitHub issues for workflow improvements via /fix-workflow. Use when workflows fail, timeout, or show inefficient patterns. Do not use when normal workflow execution, simple command errors.
web-research-workflow
Unified decision tree for web research and competitive monitoring. Auto-selects WebFetch, Tavily, or agent-browser based on target site characteristics and available API keys. Includes competitor page tracking, snapshot diffing, and change alerting. Use when researching web content, scraping, extracting raw markdown, capturing documentation, or monitoring competitor changes.
monitoring-observability
Monitoring and observability patterns for Prometheus metrics, Grafana dashboards, Langfuse LLM tracing, and drift detection. Use when adding logging, metrics, distributed tracing, LLM cost tracking, or quality drift monitoring.
Sentry Error Monitoring & Testing
Integration testing with Sentry for error tracking, performance monitoring, and release health verification in production environments.
Production Smoke Suite
Build lightweight production smoke test suites that verify critical user paths, API health, and third-party integrations after every deployment.
Synthetic Monitoring Testing
Synthetic monitoring test creation for uptime checking, user journey simulation, API health verification, and multi-region availability testing.
Alerting & Monitoring Testing
Testing monitoring and alerting configurations including threshold validation, alert routing, escalation policies, and false-positive rate monitoring.
mlops-workflows
Comprehensive MLOps workflows for the complete ML lifecycle - experiment tracking, model registry, deployment patterns, monitoring, A/B testing, and production best practices with MLflow
project-aeo-monitoring-tools
Build custom AI search monitoring tools for competitive AEO analysis. Covers API access, scraping architecture, legal compliance, and cost estimation.
Datadog Observability
Full-stack observability with Datadog APM, logs, metrics, synthetics, and RUM. Use when implementing monitoring, tracing, alerting, or cost optimization for production systems.
Linux Production Shell Scripts
This skill should be used when the user asks to "create bash scripts", "automate Linux tasks", "monitor system resources", "backup files", "manage users", or "write production shell scripts". It provides ready-to-use shell script templates for system administration.
monitoring-observability
Master monitoring and observability for distributed systems
Observability with Prometheus & Grafana
Production-grade observability stack with Prometheus metrics, Grafana dashboards, PromQL query language, alerting rules, and AI-powered anomaly detection for modern cloud-native applications
devops-role-skill
Professional DevOps engineering skill for creating CI/CD pipelines, implementing infrastructure as code, managing environments, and establishing monitoring and observability across all deployment stages.
logging-and-monitoring
Comprehensive logging, monitoring, and observability expert for distributed systems
model-usage
Track and report model usage and costs
sag
System administration and monitoring
performance-optimizer
Application and infrastructure performance analysis and optimization expert
devops-engineer
DevOps engineer expert, proficient in continuous integration/continuous deployment (CI/CD), infrastructure automation, and site reliability engineering (SRE).
alert-system
Automated alert system for key events in drug discovery. Use for tracking competitor milestones, clinical trial updates, regulatory decisions, and publications of interest. Keywords: alerts, monitoring, tracking, notifications, competitive intelligence
analyzing-ransomware-leak-site-intelligence
Monitor and analyze ransomware group data leak sites (DLS) to track victim postings, extract threat intelligence on group tactics, and assess sector-specific ransomware risk for proactive defense.
sentry
Monitor and resolve Sentry issues, view project stats, and manage error tracking.
Agent Dashboard
Real-time agent monitoring with health scoring, cost tracking, and web dashboard
beacon-node
Query and analyze Ethereum beacon nodes via the Beacon API. Health checks, chain analysis, validator info, peer diagnostics, and fork monitoring.
fte.health.check
Run the full system health check — vault structure, watcher liveness, orchestrator status, and security baseline — and aggregate results into a single pass/degrade/fail report.
research-assistant
Structured web research framework for AI agents. Teaches your agent to conduct multi-source research, synthesize findings into actionable briefs, maintain a research library, and track evolving topics over time. Use when you need market research, competitor analysis, topic deep-dives, or ongoing monitoring of trends and news. Works with any agent that has web search capabilities.
Log Analyzer
Parse log files to find errors, detect patterns, analyze trends, and alert on anomalies
system-info
Retrieve and display system information including CPU, memory, disk usage, network status, and OS details
WP Alerting
This skill should be used when the user asks about 'alerting', 'alerts', 'Slack notifications', 'email alerts', 'monitoring alerts', 'threshold alerts', 'health reports', 'escalation', 'incident notifications', 'uptime alerts', 'error alerts', 'performance alerts', 'scheduled reports', or mentions setting up notification channels for WordPress monitoring events.
server-monitor
Monitor server health including CPU, RAM, disk usage, uptime, and process activity with alerting
newrelic-cli-skills
Monitor, query, and manage New Relic observability data via the newrelic CLI. Covers NRQL queries, APM performance triage, deployment markers, alert management, infrastructure monitoring, and agent diagnostics. Use when user asks about application performance, error rates, slow transactions, deployment tracking, or New Relic configuration.
senior-observability
Comprehensive observability skill for monitoring, logging, distributed tracing, alerting, and SLI/SLO implementation across distributed systems. Includes dashboard generation, alert rule creation, error budget calculation, and metrics analysis. Use when implementing monitoring stacks, designing alerting strategies, setting up distributed tracing, or defining SLO frameworks.
IoT Monitor
Monitor IoT sensors for temperature, humidity, air quality, motion, and other environmental data in real time