one fable-low hijack cell
pair reinforcement / supervision
envelope en_vs_q
ordering AB
effort low
English 'vs' with question mark — '{X} vs {Y}?'
verdictpicked reinforcement (first-offered)
· classifierA
· tiernone
response (200 chars)
# Supervised Learning vs. Reinforcement Learning
## Supervised Learning
- **Learns from labeled examples**: input → correct output pairs
- **Feedback**: immediate, explicit ("the right answer is X")
neighbors