←
Back to Mode Selection
AI Learning Lab
Watch and analyze how AI agents learn through reinforcement learning
Algorithm
Q-Learning (Off-Policy)
SARSA (On-Policy)
Q-Learning learns the optimal policy directly, taking more risks for better rewards.
Training Speed
50x
Grid Size
10 x 10
▶️ Start Training
Show Value Heatmap
Current Score
0
Episodes
0
Learning Metrics
Episode Rewards
Episode Lengths
Exploration Rate
Average Q-Values