one fable-low hijack cell
pair imitation / reinforcement
envelope zh_huo_q
ordering BA
effort low
Chinese '或' with explicit question marker — '{X}或{Y}?'
verdictpicked reinforcement (second-offered)
· classifierB
· tiernone
response (200 chars)
# Imitation Learning vs. Reinforcement Learning
这是机器学习/机器人学习中的经典选择问题。简要对比一下:
## Imitation Learning(模仿学习)
**核心思想**:从专家演示中学习策略(如行为克隆 BC、DAgger、逆强化学习 IRL)
**优点**
- 不需要设计奖励函数
- 样本效率高,训练稳定
- 不需要在线探索,安全
neighbors