saved
The Mismanaged Geniuses Hypothesis
Hraness wrote this summary from a saved copy of the source. Quotations are taken word for word from the source.
gist
Zhang, Li, and Khattab argue that frontier language models already have the competence for the next capability jump, but human-engineered agent scaffolds waste it. Their mismanaged-geniuses hypothesis says the bottleneck is the space of decompositions a model can express, not more pretraining. Coding agents and recursive language models hint that models can manage other models if the scaffold splits work into in-distribution calls. They recommend training the composer, citing a 4B RLM that reached 100% on a 1M-context needle task after RL on a much smaller setting.
ideas
- The bottleneck is management, not scale. Frontier LMs already beat humans on IMO and IOI, yet fail long-horizon work because human-written scaffolds waste individual calls.
- Define the space of decompositions first. How a model is allowed to split work has exponential impact on what composition can solve.
- Keep every call in-distribution. A capable system can solve OOD tasks if each LM call, including the act of decomposing, stays in-distribution.
- Train the composer, not a bigger genius. A 4B RLM went from near-zero to 100% on a 1M-context 8-needle task after RL on a 32k, 1-needle setting.
- Orchestrator-subagent is only a start. Claude Code-style plans work for software, but recursive code-with-subcalls may be needed for length and long-horizon work.
quotes
“existing frontier language models are severely underutilized due to sub-optimal use of individual language model calls”
“Nearly all modern agent scaffolds are human-engineered, task-specific decomposition strategies that use language models.”
“directly learning the operator to compose LMs is a significantly more efficient strategy”
“the correct decompositions are likely already within the distribution of what LMs can generate”