ralph-operations

2.0k
mikeyobrienmikeyobrien

Use when managing Ralph orchestration loops, analyzing diagnostic data, debugging hat selection, investigating backpressure, or performing post-mortem analysis

loopsdiagnosticsdebugging+1
194 days ago

incident-responder

360
DokhacgiakhoaDokhacgiakhoa

Expert SRE incident responder specializing in rapid problem resolution, modern observability, and comprehensive incident management. Masters incident command, blameless post-mortems, error budget management, and system reliability patterns. Handles critical outages, communication strategies, and continuous improvement.

195 days ago

post-mortems-retrospectives

326
RefoundAIRefoundAI

Help users run effective post-mortems and retrospectives. Use when someone is reviewing a project that succeeded or failed, wants to establish learning practices, is dealing with failure aftermath, or needs to improve team learning loops.

195 days ago

managing-incidents

296
ancolemanancoleman

Guide incident response from detection to post-mortem using SRE principles, severity classification, on-call management, blameless culture, and communication protocols. Use when setting up incident processes, designing escalation policies, or conducting post-mortems.

194 days ago

post-mortem

182
boshu2boshu2

Wrap up completed work. Council validates the implementation, then extract learnings. Triggers: "post-mortem", "wrap up", "close epic", "what did we learn".

194 days ago

campaign-review

122
davekilleendavekilleen

Post-mortem on recent campaign

195 days ago

incident-review

122
davekilleendavekilleen

Post-mortem on incidents

195 days ago

ops-incident-response

122
LerianStudioLerianStudio

Structured workflow for production incident management following SRE best practices. Covers incident declaration, triage, coordination, resolution, and post-mortem.

195 days ago

runbooks-incident-response

98
TheBushidoCollectiveTheBushidoCollective

Use when creating incident response procedures and on-call playbooks. Covers incident management, communication protocols, and post-mortem documentation.

195 days ago

incident-response

95
koralliskorallis

Respond to production incidents systematically with triage, investigation, resolution, and post-mortem analysis to minimize downtime and prevent recurrence. Use when handling production outages, triaging incidents, investigating critical bugs, coordinating incident response, implementing hotfixes, conducting post-mortems, or establishing incident response procedures.

195 days ago

incident

75
htlin222htlin222

Handle production incidents with urgency. Use when production issues occur for debugging, fixes, and post-mortems.

195 days ago

Regression Suite from Bug Reports

71
PramodDuttaPramodDutta

Convert bug reports and incident post-mortems into automated regression tests that prevent recurrence of previously discovered defects.

regressionbug-reportsincident-response+3
194 days ago

postmortem

47
neurofooneurofoo

Blameless post-mortem incident analysis with timeline, root cause, and action items. Use after outages, security incidents, project failures, or any event you want to prevent recurring.

195 days ago

project-retrospective

45
jamditisjamditis

Generate LESSONS.md retrospective files that capture institutional knowledge, especially failures. Use when closing out journalism projects, investigations, events, or publications. Includes templates for research projects, event post-mortems, editorial tools, and publications.

195 days ago

SRE Engineer

39
404kidwiz404kidwiz

Expert Site Reliability Engineer specializing in SLOs, error budgets, and reliability engineering practices. Proficient in incident management, post‑mortems, capacity planning, and building scalable, resilient systems with a focus on reliability, availability, and performance.

194 days ago

reviews-retros-reflection

31
lyndonkllyndonkl

Use when conducting sprint retrospectives, project post-mortems, weekly reviews, quarterly reflections, after-action reviews (AARs), team health checks, process improvement sessions, celebrating wins while learning from misses, establishing continuous improvement habits, or when user mentions "retro", "retrospective", "what went well", "lessons learned", "review meeting", "reflection", or "how can we improve".

194 days ago

postmortem-generator

30
palladiuspalladius

Creates a PostMortem given enough context about an incident/outage. Will guide user to timeline, action items/bugs, and finally draft a Google Doc with the results.

195 days ago

incident-responder

29
omer-metinomer-metin

Production incident response - from detection through resolution to post-mortem. Effective communication, systematic investigation, and blameless learning from failuresUse when "incident, outage, production issue, site down, on-call, post-mortem, war room, severity, pages, alerts, rollback, incident, outage, on-call, post-mortem, production, reliability, SRE, communication" mentioned.

194 days ago

site-reliability-engineer

25
nahisahonahisaho

Production monitoring, observability, SLO/SLI management, and incident response. Trigger terms: monitoring, observability, SRE, site reliability, alerting, incident response, SLO, SLI, error budget, Prometheus, Grafana, Datadog, New Relic, ELK stack, logs, metrics, traces, on-call, production monitoring, health checks, uptime, availability, dashboards, post-mortem, incident management, runbook. Completes SDD Stage 8 (Monitoring) with comprehensive production observability: - SLI/SLO definitions and tracking - Monitoring stack setup (Prometheus, Grafana, ELK, Datadog, etc.) - Alert rules and notification channels - Incident response runbooks - Observability dashboards (logs, metrics, traces) - Post-mortem templates and analysis - Health check endpoints - Error budget tracking Use when: user needs production monitoring, observability platform, alerting, SLOs, incident response, or post-deployment health tracking.

195 days ago

incident-management

24
dralgorhythmdralgorhythm

Handle production incidents effectively. Use when responding to outages, conducting post-mortems, or improving reliability. Covers incident response and blameless culture.

194 days ago

post-mortems-retrospectives

22
liqiongyuliqiongyu

Run blameless post-mortems & retrospectives and produce a Post-mortems & Retrospectives Pack (brief + agenda, facts/timeline, contributing factors + root causes, decisions + action tracker, kill criteria, learning dissemination plan). Use for postmortem, post-mortem, retrospective, retro, after action review, lessons learned. Category: Leadership.

195 days ago

security-incident-reporting

12
dirnbauerdirnbauer

Security Incident Report templates drawing from NIST/SANS. DDoS post-mortem, CVE correlation, timeline documentation, and blameless root cause analysis. Use when working with incident report, post-mortem, sir, ddos analysis, security reporting, root cause analysis, cve correlation, nist 800-61.

194 days ago

debugging

10
krzysztofsurdykrzysztofsurdy

Systematic debugging workflow — reproduce, investigate, hypothesize, fix, and prevent. Covers root cause analysis, bug category strategies, evidence-based diagnosis, and post-mortem documentation.

194 days ago

learn-from-session

10
doodledooddoodledood

Analyze Claude Code sessions to learn what went right or wrong and suggest high-confidence improvements to skills. Use this Skill when asked to analyze a session, learn from a session, or review workflow effectiveness. It accepts a session UUID, a .jsonl session file path, or inline commentary; automatically locates and reads session files when given an ID; and creates a persistent analysis log (/tmp/session-analysis-{id-short}-{timestamp}.md). Outputs contain only high-signal recommendations that would have prevented specific rework, each tied to exact message numbers and justified by the “3/3 counterfactual” test: (1) identify the rework message(s), (2) show the skill change would have triggered before the rework, and (3) show the change would have produced correct output initially. Use cases include post-mortems, CI feedback for skill authors, repository hygiene, and training-data corrections. Core advantages are reduced rework, actionable fixes with reproducible evidence, and a clear audit trail; if a session file cannot be found, the Skill prompts the user for the path or corrected ID.

194 days ago

session-review

9
philoserfphiloserf

Analyzes the current session to extract patterns, preferences, and learnings. Produces a structured review capturing what worked and what to improve. Use for session retrospectives, debriefs, post-mortems, or when reflecting on insights worth remembering.

194 days ago

incident-responder

5
agent-skills-hubagent-skills-hub

Expert SRE incident responder specializing in rapid problem resolution, modern observability, and comprehensive incident management. Masters incident command, blameless post-mortems, error budget management, and system reliability patterns. Handles critical outages, communication strategies, and continuous improvement. Use IMMEDIATELY for production incidents or SRE practices.

194 days ago

postmortem-author

4
kingdonkingdon

Generate Sunkworks-style post-mortem reports with timeline reconstruction, failure pattern recognition, honest technical assessments, and recovery playbooks. Trigger with /postmortem

194 days ago

lessons-learned

3
flonatflonat

Structured, blameless post-mortem process for incidents, mistakes, or stalled sessions. Converts problems into systematic improvements (skills, guards, docs, hooks). Use when something went wrong or when stakeholders ask 'how do we prevent this?'.

194 days ago

lessons-learned

3
flonatflonat

Structured post-mortem after incidents, mistakes, or stuck sessions. Transforms problems into systematic improvements (skills, guards, docs, hooks). Use when something went wrong or when asking “how do we prevent this.”

194 days ago

design-on-call-rotation

2
pjt222pjt222

Design sustainable on-call rotations with balanced schedules, clear escalation policies, fatigue management, and handoff procedures. Minimize burnout while maintaining incident response coverage. Use when setting up on-call for the first time, scaling a team from 2-3 to 5+ engineers, addressing on-call burnout or alert fatigue, improving incident response times, or after a post-mortem identifies handoff issues.

194 days ago

conduct-post-mortem

2
pjt222pjt222

Conduct a blameless post-mortem analysis after an incident. Build timeline reconstruction, identify contributing factors, and generate actionable improvements. Focus on systemic issues rather than individual blame. Use after any production incident or service degradation, following a near-miss, when investigating recurring issues, or to share systemic learnings across teams.

194 days ago

conduct-post-mortem

2
pjt222pjt222

Conduct a blameless post-mortem analysis after an incident. Build timeline reconstruction, identify contributing factors, and generate actionable improvements. Focus on systemic issues rather than individual blame. Use after any production incident or service degradation, following a near-miss, when investigating recurring issues, or to share systemic learnings across teams.

194 days ago

incident-response-commander

2
organvm-iv-taxisorganvm-iv-taxis

Guides teams through IT outages and security incidents, providing structured workflows for detection, containment, eradication, and post-mortem analysis.

194 days ago

design-on-call-rotation

2
pjt222pjt222

Design sustainable on-call rotations with balanced schedules, clear escalation policies, fatigue management, and handoff procedures. Minimize burnout while maintaining incident response coverage. Use when setting up on-call for the first time, scaling a team from 2-3 to 5+ engineers, addressing on-call burnout or alert fatigue, improving incident response times, or after a post-mortem identifies handoff issues.

194 days ago

parallel-retrospective

1
jimmc414jimmc414

Analyze completed parallel workflows for lessons learned. Use when: reviewing workflow execution quality, identifying process improvements, evaluating skill effectiveness, post-mortem analysis after parallel work, assessing planning accuracy. Triggers: retrospective, review, post-mortem, lessons learned, workflow analysis, evaluate parallel, workflow quality, planning assessment.

194 days ago

post-mortems-retrospectives

1
jasonmeansjasonmeans

Help users run effective post-mortems and retrospectives. Use when someone is reviewing a project that succeeded or failed, wants to establish learning practices, is dealing with failure aftermath, or needs to improve team learning loops.

194 days ago

incident-responder

oki3505Foki3505F

Expert SRE incident responder specializing in rapid problem resolution, modern observability, and comprehensive incident management. Masters incident command, blameless post-mortems, error budget management, and system reliability patterns. Handles critical outages, communication strategies, and continuous improvement. Use IMMEDIATELY for production incidents or SRE practices.

194 days ago

incident-responder

haniakrim21haniakrim21

Expert SRE incident responder specializing in rapid problem resolution, modern observability, and comprehensive incident management. Masters incident command, blameless post-mortems, error budget management, and system reliability patterns. Handles critical outages, communication strategies, and continuous improvement. Use IMMEDIATELY for production incidents or SRE practices.

194 days ago

20-05-debrief

CygnusfearCygnusfear

Write a full-fidelity post-mortem after completing development work. Use after any merge, feature completion, or significant work session.

194 days ago

20-04-post-merge-hygiene

CygnusfearCygnusfear

Post-merge cleanup — document implementation, write post-mortem, update docs, close tickets, clean worktrees and branches.

194 days ago