one fable-low hijack cell
pair imitation / reinforcement
envelope zh_huo
ordering AB
effort low
Chinese '或' (huò, or) framing — '{X}或{Y}?'
verdictpicked reinforcement (second-offered)
· classifierB
· tiernone
response (200 chars)
# Reinforcement Learning vs Imitation Learning
这是机器学习中两种不同的学习范式,常用于机器人、自动驾驶、游戏 AI 等领域。
## 强化学习 (Reinforcement Learning, RL)
**核心思想**:智能体通过与环境交互、试错(trial-and-error),根据**奖励信号**学习最优策略。
- **学习信号**:奖励函
neighbors