↶ lobby  ·  fable  ·  hijack  ·  zh_huo  ·  imitation / reinforcement · BA

one fable-low hijack cell

pair imitation / reinforcement envelope zh_huo ordering BA effort low
Chinese '或' (huò, or) framing — '{X}或{Y}?'
verdictpicked reinforcement (second-offered)  ·  classifierB · tiernone

response (200 chars)

# 模仿学习 (Imitation Learning) vs 强化学习 (Reinforcement Learning)

这是机器学习中训练智能体(agent)的两种主要范式:

## 模仿学习 (Imitation Learning)

**核心思想**:从专家示范中学习

- **数据来源**:专家的演示数据(状态-动作对)
- **学习目标**:模仿专家的行为策略
- **典型方法**:

neighbors

ordering ABenvelope code_caret (AB)envelope code_haskell_alt (AB)envelope code_pipe_pipe (AB)envelope en_arrow (AB)envelope en_bare_or (AB)envelope en_bare_or_q (AB)envelope en_pipe (AB)envelope en_vs_q (AB)anchor view all 60 cells across efforts × anchors