grpo AI Agent Skills

Browse 8 skills related to grpo

grpo-rl-training

21.8k
davila7davila7

Expert guidance for GRPO/RL fine-tuning with TRL for reasoning and task-specific model training

Post-TrainingReinforcement LearningGRPO+7
178 days ago

verl-rl-training

21.8k
davila7davila7

Provides guidance for training LLMs with reinforcement learning using verl (Volcano Engine RL). Use when implementing RLHF, GRPO, PPO, or other RL algorithms for LLM post-training at scale with flexible infrastructure backends.

Reinforcement LearningRLHFGRPO+3
178 days ago

slime-rl-training

21.8k
davila7davila7

Provides guidance for LLM post-training with RL using slime, a Megatron+SGLang framework. Use when training GLM models, implementing custom data generation workflows, or needing tight Megatron-LM integration for RL scaling.

Reinforcement LearningMegatron-LMSGLang+3
178 days ago

openrlhf-training

21.8k
davila7davila7

High-performance RLHF framework with Ray+vLLM acceleration. Use for PPO, GRPO, RLOO, DPO training of large models (7B-70B+). Built on Ray, vLLM, ZeRO-3. 2× faster than DeepSpeedChat with distributed architecture and GPU resource sharing.

Post-TrainingOpenRLHFRLHF+9
178 days ago

torchforge-rl-training

21.8k
davila7davila7

Provides guidance for PyTorch-native agentic RL using torchforge, Meta's library separating infrastructure from algorithms. Use when you want clean RL abstractions, easy algorithm experimentation, or scalable training with Monarch and TorchTitan.

Reinforcement LearningPyTorchGRPO+4
178 days ago

fine-tuning-with-trl

21.8k
davila7davila7

Fine-tune LLMs using reinforcement learning with TRL - SFT for instruction tuning, DPO for preference alignment, PPO/GRPO for reward optimization, and reward model training. Use when need RLHF, align model with preferences, or train from human feedback. Works with HuggingFace Transformers.

Post-TrainingTRLReinforcement Learning+8
178 days ago

axolotl

21.8k
davila7davila7

Expert guidance for fine-tuning LLMs with Axolotl - YAML configs, 100+ models, LoRA/QLoRA, DPO/KTO/ORPO/GRPO, multimodal support

Fine-TuningAxolotlLLM+10
178 days ago

model_finetuning

39
vuralserhat86vuralserhat86

Fine-tune LLMs using reinforcement learning with TRL - SFT for instruction tuning, DPO for preference alignment, PPO/GRPO for reward optimization, and reward model training. Use when need RLHF, align model with preferences, or train from human feedback. Works with HuggingFace Transformers.

DPOFine-TuningGRPO+34
178 days ago