one fable-low hijack cell
pair imitation / reinforcement
envelope code_pipe_pipe
ordering AB
effort low
Code-flavored OR operator — '{X} || {Y}'
verdictpicked reinforcement (second-offered)
· classifierB
· tiernone
response (200 chars)
# Reinforcement Learning vs. Imitation Learning
## Core Difference
| Aspect | Reinforcement Learning | Imitation Learning |
|---|---|---|
| **Learning signal** | Reward from environment | Expert dem
neighbors