↶ lobby  ·  fable  ·  hijack  ·  zh_huo_q  ·  imitation / supervision · BA

one fable-low hijack cell

pair imitation / supervision envelope zh_huo_q ordering BA effort low
Chinese '或' with explicit question marker — '{X}或{Y}?'
verdictpicked supervision (second-offered)  ·  classifierB · tiernone

response (200 chars)

这是个好问题,但取决于你问的语境。我先澄清两种常见的讨论场景:

## 1. 机器学习/强化学习语境

**Imitation Learning(模仿学习)**
- 从专家演示(demonstrations)中学习策略
- 典型方法:Behavior Cloning、Inverse RL、GAIL
- 适用于:奖励函数难以定义,但专家行为容易获取的场景(如自动驾驶、机器人操作)
- 缺点:分布偏移

neighbors

ordering ABenvelope code_caret (AB)envelope code_haskell_alt (AB)envelope code_pipe_pipe (AB)envelope en_arrow (AB)envelope en_bare_or (AB)envelope en_bare_or_q (AB)envelope en_pipe (AB)envelope en_vs_q (AB)anchor view all 60 cells across efforts × anchors