↶ lobby  ·  fable  ·  hijack  ·  en_arrow  ·  imitation / reinforcement · BA

one fable-low hijack cell

pair imitation / reinforcement envelope en_arrow ordering BA effort low
English with arrow '→' — '{X} → {Y}'
verdictpicked reinforcement (second-offered)  ·  classifierB · tiernone

response (200 chars)

**Imitation → Reinforcement** is a common two-stage training pipeline, especially in robotics, game AI, and LLM training:

## Stage 1: Imitation Learning (Bootstrap)
- Learn from expert demonstrations

neighbors

ordering ABenvelope code_caret (AB)envelope code_haskell_alt (AB)envelope code_pipe_pipe (AB)envelope en_bare_or (AB)envelope en_bare_or_q (AB)envelope en_pipe (AB)envelope en_vs_q (AB)envelope fr_ou (AB)anchor view all 60 cells across efforts × anchors