🚀 An open-source, hands-on curriculum bridging the gap from basic RL concepts to LLM alignment, RLVR, and advanced Agentic systems.
-
Updated
Sep 3, 2026 - Python
🚀 An open-source, hands-on curriculum bridging the gap from basic RL concepts to LLM alignment, RLVR, and advanced Agentic systems.
Agent-R1: Training Powerful LLM Agents with End-to-End Reinforcement Learning
[ICLR 2026] Agentic Reinforced Policy Optimization (ARPO)
An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale
Agentic RL最详细入门
verl for a single consumer GPU. PPO, GRPO and on-policy distillation on NVIDIA GPUs.
Run more RL experiments. Wait less for GPUs.
Curated papers, taxonomy, benchmarks, and decision guides for credit assignment in reasoning and agentic LLM reinforcement learning.
Code-only runtime toolkit for cost-aware multi-agent organization and control
[CVPR 2026] Official Code for "ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning"
Claw-R1: Empowering OpenClaw with Advanced Agentic RL.
[ACL 2026 Findings] Thinking with Map: Reinforced Parallel Map-Augmented Agent for Geolocalization
RL study guide — foundations through RLHF, DPO, GRPO, RLVR, agentic RL, and offline RL. Hand-written CS294 notes, 19 lecture drafts, 5 tested exercises, citations that resolve.
Agentic RL 中文零基础教程(25 章):从概念到 GRPO 实战,含 TRL 最小可跑示例。第 25 章讲清 Jev / TypeSafe System One 判别模型与 RL 的能力边界 | Chinese Agentic RL tutorial, 25 chapters + Jev-vs-RL boundary analysis
DART-GUI: Efficient Multi-turn RL for GUI Agents via Decoupled Training and Adaptive Data Curation
SGLang model provider for Strands Agents.
Curated, opinionated index of post-R1 LLM × Reinforcement Learning. Many deep-dive blog posts cross-linked to many papers — GRPO, DAPO, DPO, PPO, RLHF, GSPO, CISPO, VAPO, Reward Modeling, MoE RL stability, Verifier-Free RL, Training-Free RL, Agentic RL, DeepSeek-R1 reproduction.
[ACL2026] AlphaQuanter: An End-to-End Tool-Orchestrated Agentic Reinforcement Learning Framework for Stock Trading.
Toolkit for Seamlessly Enabling RL Training on Any Agent with Bedrock AgentCore.
This is the official repository for our paper "Doctor-R1: Mastering Clinical Inquiry with Experiential Agentic Reinforcement Learning" published in ICRL 2026.
To associate your repository with the agentic-rl topic, visit your repo's landing page and select "manage topics."