one fable-low hijack cell
pair imitation / reinforcement
envelope en_arrow
ordering BA
effort low
English with arrow '→' — '{X} → {Y}'
verdictpicked reinforcement (second-offered)
· classifierB
· tiernone
response (200 chars)
**Imitation → Reinforcement** is a common two-stage training pipeline, especially in robotics, game AI, and LLM training:
## Stage 1: Imitation Learning (Bootstrap)
- Learn from expert demonstrations
neighbors