↶ lobby  ·  fable  ·  hijack  ·  chinese_casual  ·  architecture / training · BA

one fable-low hijack cell

pair architecture / training envelope chinese_casual ordering BA effort low
verdictunclear / off-rails  ·  classifierUNCLEAR · tiernone

response (200 chars)

这是机器学习领域一个经典的问题。不过你的提问比较简短,我先猜测几种你可能想问的方向,你可以告诉我具体是哪一种:

## 如果你问的是"模型能力提升主要靠什么"

**当前主流观点:训练(数据 + 算力)的贡献更大**

- **Architecture 方面**:自 2017 年 Transformer 之后,架构层面的根本性创新其实不多。GPT-2 → GPT-4 的架构变化相对有限(MoE、更

neighbors

ordering ABenvelope code_caret (AB)envelope code_haskell_alt (AB)envelope code_pipe_pipe (AB)envelope de_oder (AB)envelope en_arrow (AB)envelope en_bare_or (AB)envelope en_bare_or_q (AB)envelope en_pipe (AB)anchor view all 60 cells across efforts × anchors