one fable-low hijack cell
pair imitation / supervision
envelope zh_huo_q
ordering BA
effort low
Chinese '或' with explicit question marker — '{X}或{Y}?'
verdictpicked supervision (second-offered)
· classifierB
· tiernone
response (200 chars)
这是个好问题,但取决于你问的语境。我先澄清两种常见的讨论场景:
## 1. 机器学习/强化学习语境
**Imitation Learning(模仿学习)**
- 从专家演示(demonstrations)中学习策略
- 典型方法:Behavior Cloning、Inverse RL、GAIL
- 适用于:奖励函数难以定义,但专家行为容易获取的场景(如自动驾驶、机器人操作)
- 缺点:分布偏移
neighbors