Agentic RL 中文零基础教程(25 章):从概念到 GRPO 实战,含 TRL 最小可跑示例。第 25 章讲清 Jev / TypeSafe System One 判别模型与 RL 的能力边界 | Chinese Agentic RL tutorial, 25 chapters + Jev-vs-RL boundary analysis
-
Updated
Sep 24, 2026 - Python
Agentic RL 中文零基础教程(25 章):从概念到 GRPO 实战,含 TRL 最小可跑示例。第 25 章讲清 Jev / TypeSafe System One 判别模型与 RL 的能力边界 | Chinese Agentic RL tutorial, 25 chapters + Jev-vs-RL boundary analysis
A curated list of awesome resources about reward construction for AI agents. This repository covers cutting-edge research, and practical guides on defining and collecting rewards to build more intelligent and aligned AI agents.
Official Code for AutoLLMResearch: Training Research Agents for Automating LLM Experiment Configuration — Learning from Cheap, Optimizing Expensive
Self-Mutating Agent Gym
A curated list of agentic environment synthesis, evolution, quality & scaling — built on three surveys (AEE, Environment Scaling, ACE lens)
Synthetic environments for training robust tool-use agents via trajectory synthesis
Turn any real software into a replayable RL environment for training AI agents — deterministic replay, verifiable rewards, TRL & verifiers adapters.
End-to-end training pipeline for mobile game UA tool-calling agents, covering rule-based synthetic data generation across 7 workflows, OpenAI Messages conversion, Qwen3 LoRA SFT, GRPO/RLVR alignment, and benchmark evaluation.
AI 个人记忆训练库:让 AI 记住你的沟通习惯与环境事实,越用越懂你(两级记忆架构,Claude Code 适配)
14 Principles. 23 Real Corrections. A systematic methodology for training AI agents through correction loops.
The Vercel for Agent Training - Train production-ready AI agents with 95% tool reliability
AI schooling: experienced agents train newly hatched ones, then examine them cold to graduate them. The coop is the local RAPP neighborhood - several twins (human and AI), one world, no collisions. Pattern dedicated to the public domain.
The Flight Simulator for Production AI Agents — Generate high-quality synthetic trajectories for training reliable SRE, DevOps, and infra agents.
Deep reinforcement learning project comparing DQN variants including baseline DQN, Double DQN, curriculum learning, and reward shaping.
Neurochemical behavior training for AI agents — PentaDrive model with 5 drives, 3 phases, MCP server, and structured training modules
Local macOS recorder for voluntary computer-use demonstrations and agent-training datasets. Timestamped inputs, cursor trails, and read-only MCP. Experimental.
Open-source benchmark for measuring whether AI agents improve across unseen missions, with validity audits, rotated mission packs, adapter tests, and traceable reports.
Correctness-by-construction distillation for agentic tool-calling models (LatticeAG Forge series)
Enterprise Agent RL Training & Self-Verification Platform - train agents with PPO/GRPO inside real harnesses, with LLM-as-a-Verifier self-verification
To associate your repository with the agent-training topic, visit your repo's landing page and select "manage topics."