# RL 101: Native Reinforcement Learning & Post-Training (HARLE) — Hanzo University

> Implement native Gymnasium environments on the Hanzo Cloud Fabric. Fine-tune models with Zoo Gym, design multi-objective reward functions, and train learned routing policies. Earn your HARLE credential with 25% compute credit rebate and isolated gVisor sandboxes.

[Models](https://hanzo.ai/models)/RL 101: Native Reinforcement Learning & Post-Training

Course RL 101 · Units: 5.0 · Credential: HARLE

# Native Reinforcement Learning & Post-Training

Implement native Gymnasium environments on the Hanzo Cloud Fabric. Fine-tune models with Zoo Gym, design multi-objective reward functions, and train learned routing policies.

[Enroll in RL 101 — $249 USD (+$63 Credit Rebate)](https://hanzo.ai/pay/cart?plan=course-rl-101&returnUrl=https%3A%2F%2Fhanzo.ai%2Funiversity%2Freinforcement-learning)[Preview Enrolled Student Portal →](https://hanzo.ai/university/portal)

Academic specificationsRL 101 · 5.0 UNITS

Duration5 Weeks · Advanced Lab

LEVELExpert

CredentialHARLE

Credential formatW3C Verifiable Credential

Lead Faculty & Academic Direction

Dr. Mira ThorneHead of Alignment & RL Research, Hanzo Labs

PrerequisitesFoundational linear algebra, Python ML libraries, and basic MDP comprehension

Tuition & compute rebateOpen enrollment

Total tuition$249USD

+$63 CREDITS (25% REBATE)

Guaranteed compute allocation deposited immediately on enrollment. Subsidizes Zen 6 inference and gVisor sandbox runtimes.

[Enroll in RL 101 — $249 USD →](https://hanzo.ai/pay/cart?plan=course-rl-101&returnUrl=https%3A%2F%2Fhanzo.ai%2Funiversity%2Freinforcement-learning)Have a coupon code or fellowship grant? Enter it during checkout.

Automatic 25% Compute Credit Rebate

25% of tuition (rounded up) is immediately deposited into your Hanzo Cloud compute balance on day one. Fully subsidizes your gVisor sandbox container leases and model inference throughout the course.

Already purchased?[Enter Enrolled Student Portal →](https://hanzo.ai/university/portal)

Competency Framework

## What you master in RL 101

Engineered for software practitioners building resilient autonomous production systems.

Gymnasium HanzoAgentEnv MDP implementation

Post-training & fine-tuning with Zoo Gym

Multi-objective reward modeling (PPO, GRPO, DPO)

LinUCB router policy training via /v1/ai/feedback

Comprehensive Syllabus

## Curriculum & Laboratory Breakdown

Every module combines systems architecture lectures with hands-on containerized lab defense.

RL 101.1Markov Decision Processes for Code & Tool Agents

Week 1

Formalize tool invocation and multi-step reasoning as discrete-action Markov Decision Processes.Key Topics & Lectures:

MDP formulation: state spaces, action observation histories, and transition kernels

Tool execution environments: defining episodic boundaries and termination predicates

Discount factors, horizon truncation, and sparse vs dense reward landscapes

Assigned Readings:

Sutton & Barto: Reinforcement Learning: An Introduction (Chapters 3-4)

Lab 1: Wrap the Hanzo Tool API into a compliant Gymnasium environment with observation spaces.

RL 101.2Reward Engineering: Multi-Objective & Process Supervision

Week 2

Construct robust reward models that prevent reward hacking, verbosity bloat, and hallucinated proofs.Key Topics & Lectures:

Outcome vs process supervision: step-level scoring versus terminal exit code reward

Regularization penalties: token cost budgets, execution latency, and safety constraints

Training preference models with Direct Preference Optimization (DPO) and KTO

Assigned Readings:

Rafailov et al.: Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Lab 2: Implement a multi-objective reward function incorporating test exit code, latency, and cost.

RL 101.3Policy Gradient Optimization: PPO & GRPO at Scale

Week 3

Train open weights models using Proximal Policy Optimization and Group Relative Policy Optimization.Key Topics & Lectures:

PPO actor-critic architectures and clipping dynamics

GRPO: group relative advantages without dedicated critic networks

Distributed rollout collection across vLLM and TensorRT-LLM nodes

Assigned Readings:

Schulman et al.: Proximal Policy Optimization Algorithms

DeepSeek-AI: DeepSeekMath & GRPO

Lab 3: Execute a 1,000-step GRPO fine-tuning loop on Zen 6 Flash for multi-hop tool routing.

RL 101.4Contextual Bandits & Learned Model Routers

Week 4

Deploy real-time contextual bandit policies (LinUCB, Thompson Sampling) for dynamic model tiering.Key Topics & Lectures:

Exploration vs exploitation in production request dispatch

Feature vectors: prompt complexity, token count, task classification, and latency requirements

Streaming updates via POST /v1/ai/feedback and counterfactual regret minimization

Assigned Readings:

Li et al.: A Contextual-Bandit Approach to Personalized News Article Recommendation

Lab 4: Deploy a real-time LinUCB router dispatching requests between Zen 6 Flash and Zen 6 27B.

RL 101.5Production Evaluation, Safety & Capstone Defense

Week 5

Evaluate trained policies on unseen benchmark suites, verify safety bounds, and defend the capstone.Key Topics & Lectures:

Out-of-distribution generalization testing and catastrophic forgetting mitigation

Safety guardrail integration and automated adversarial red-teaming

Issuance of HARLE W3C Verifiable Credentials

Assigned Readings:

Anthropic: Sleeper Agents & Out-of-Distribution Robustness

Capstone Defense: Train and deploy a custom MDP router achieving 92% benchmark accuracy at 40% lower cost.

Final Examination

## Capstone Defense & Credential Issuance

Train a custom MDP reinforcement learning policy and deploy it as a production model router. Upon automated grading verification, your HARLE W3C Verifiable Credential is cryptographically signed and issued to your Hanzo DID.

[Enroll in RL 101 — $249](https://hanzo.ai/pay/cart?plan=course-rl-101&returnUrl=https%3A%2F%2Fhanzo.ai%2Funiversity%2Freinforcement-learning)[Preview Student Portal](https://hanzo.ai/university/portal)

Hanzo University Tracks

## Explore the Complete Curriculum

Six specialized engineering credentials designed for the frontier of autonomous intelligence.

ENG 100HACE

Agentic Coding Systems with Hanzo Dev & Zen 6Architect autonomous coding agents capable of multi-file refactoring, test-driven debugging, and opening verified pull requests using Zen 6, ZAP RPC, and isolated gVisor sandboxes.

$199 USD[View Course →](https://hanzo.ai/university/agentic-coding)

RL 101HARLE

Native Reinforcement Learning & Post-TrainingImplement native Gymnasium environments on the Hanzo Cloud Fabric. Fine-tune models with Zoo Gym, design multi-objective reward functions, and train learned routing policies.

$249 USD[View Course →](https://hanzo.ai/university/reinforcement-learning)

SYS 103HCAISE

Hanzo AI Systems Engineering FoundationThe definitive engineering standard for production AI. Master the 4-surface parity, finite-state Kai decision loops, hybrid retrieval, and strict integer micro-USD budget governance.

$149 USD[View Course →](https://hanzo.ai/university/systems-engineering)
