Try Hanzo
Course RL 101 · Units: 5.0 · Credential: HARLE

Native Reinforcement Learning & Post-Training

Implement native Gymnasium environments on the Hanzo Cloud Fabric. Fine-tune models with Zoo Gym, design multi-objective reward functions, and train learned routing policies.

Academic specificationsRL 101 · 5.0 UNITS
Duration5 Weeks · Advanced Lab
LEVELExpert
CredentialHARLE
Credential formatW3C Verifiable Credential
Lead Faculty & Academic Direction
Dr. Mira ThorneHead of Alignment & RL Research, Hanzo Labs
PrerequisitesFoundational linear algebra, Python ML libraries, and basic MDP comprehension
Tuition & compute rebateOpen enrollment
Total tuition
$249USD
+$63 CREDITS (25% REBATE)
Guaranteed compute allocation deposited immediately on enrollment. Subsidizes Zen 6 inference and gVisor sandbox runtimes.
Enroll in RL 101 — $249 USD →Have a coupon code or fellowship grant? Enter it during checkout.
Automatic 25% Compute Credit Rebate
25% of tuition (rounded up) is immediately deposited into your Hanzo Cloud compute balance on day one. Fully subsidizes your gVisor sandbox container leases and model inference throughout the course.

Competency Framework

What you master in RL 101

Engineered for software practitioners building resilient autonomous production systems.

Gymnasium HanzoAgentEnv MDP implementation
Post-training & fine-tuning with Zoo Gym
Multi-objective reward modeling (PPO, GRPO, DPO)
LinUCB router policy training via /v1/ai/feedback

Comprehensive Syllabus

Curriculum & Laboratory Breakdown

Every module combines systems architecture lectures with hands-on containerized lab defense.

RL 101.1Markov Decision Processes for Code & Tool Agents
Week 1
Formalize tool invocation and multi-step reasoning as discrete-action Markov Decision Processes.
Key Topics & Lectures:
MDP formulation: state spaces, action observation histories, and transition kernels
Tool execution environments: defining episodic boundaries and termination predicates
Discount factors, horizon truncation, and sparse vs dense reward landscapes
Assigned Readings:
Sutton & Barto: Reinforcement Learning: An Introduction (Chapters 3-4)
Lab 1: Wrap the Hanzo Tool API into a compliant Gymnasium environment with observation spaces.
RL 101.2Reward Engineering: Multi-Objective & Process Supervision
Week 2
Construct robust reward models that prevent reward hacking, verbosity bloat, and hallucinated proofs.
Key Topics & Lectures:
Outcome vs process supervision: step-level scoring versus terminal exit code reward
Regularization penalties: token cost budgets, execution latency, and safety constraints
Training preference models with Direct Preference Optimization (DPO) and KTO
Assigned Readings:
Rafailov et al.: Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Lab 2: Implement a multi-objective reward function incorporating test exit code, latency, and cost.
RL 101.3Policy Gradient Optimization: PPO & GRPO at Scale
Week 3
Train open weights models using Proximal Policy Optimization and Group Relative Policy Optimization.
Key Topics & Lectures:
PPO actor-critic architectures and clipping dynamics
GRPO: group relative advantages without dedicated critic networks
Distributed rollout collection across vLLM and TensorRT-LLM nodes
Assigned Readings:
Schulman et al.: Proximal Policy Optimization Algorithms
DeepSeek-AI: DeepSeekMath & GRPO
Lab 3: Execute a 1,000-step GRPO fine-tuning loop on Zen 6 Flash for multi-hop tool routing.
RL 101.4Contextual Bandits & Learned Model Routers
Week 4
Deploy real-time contextual bandit policies (LinUCB, Thompson Sampling) for dynamic model tiering.
Key Topics & Lectures:
Exploration vs exploitation in production request dispatch
Feature vectors: prompt complexity, token count, task classification, and latency requirements
Streaming updates via POST /v1/ai/feedback and counterfactual regret minimization
Assigned Readings:
Li et al.: A Contextual-Bandit Approach to Personalized News Article Recommendation
Lab 4: Deploy a real-time LinUCB router dispatching requests between Zen 6 Flash and Zen 6 27B.
RL 101.5Production Evaluation, Safety & Capstone Defense
Week 5
Evaluate trained policies on unseen benchmark suites, verify safety bounds, and defend the capstone.
Key Topics & Lectures:
Out-of-distribution generalization testing and catastrophic forgetting mitigation
Safety guardrail integration and automated adversarial red-teaming
Issuance of HARLE W3C Verifiable Credentials
Assigned Readings:
Anthropic: Sleeper Agents & Out-of-Distribution Robustness
Capstone Defense: Train and deploy a custom MDP router achieving 92% benchmark accuracy at 40% lower cost.

Final Examination

Capstone Defense & Credential Issuance

Train a custom MDP reinforcement learning policy and deploy it as a production model router. Upon automated grading verification, your HARLE W3C Verifiable Credential is cryptographically signed and issued to your Hanzo DID.

Hanzo University Tracks

Explore the Complete Curriculum

Six specialized engineering credentials designed for the frontier of autonomous intelligence.

ENG 100HACE
Agentic Coding Systems with Hanzo Dev & Zen 6Architect autonomous coding agents capable of multi-file refactoring, test-driven debugging, and opening verified pull requests using Zen 6, ZAP RPC, and isolated gVisor sandboxes.
RL 101HARLE
Native Reinforcement Learning & Post-TrainingImplement native Gymnasium environments on the Hanzo Cloud Fabric. Fine-tune models with Zoo Gym, design multi-objective reward functions, and train learned routing policies.
SYS 103HCAISE
Hanzo AI Systems Engineering FoundationThe definitive engineering standard for production AI. Master the 4-surface parity, finite-state Kai decision loops, hybrid retrieval, and strict integer micro-USD budget governance.

Build what’s next.