ai safety AI Agent Skills
Browse 2 skills related to ai safety
constitutional-ai
21.8k
Anthropic's method for training harmless AI through self-improvement. Two-phase approach - supervised learning with self-critique/revision, then RLAIF (RL from AI Feedback). Use for safety alignment, reducing harmful outputs without human labels. Powers Claude's safety system.
Safety AlignmentConstitutional AIRLAIF+6
174 days ago
llamaguard
21.8k
Meta's 7-8B specialized moderation model for LLM input/output filtering. 6 safety categories - violence/hate, sexual content, weapons, substances, self-harm, criminal planning. 94-95% accuracy. Deploy with vLLM, HuggingFace, Sagemaker. Integrates with NeMo Guardrails.
Safety AlignmentLlamaGuardContent Moderation+6
174 days ago