Lengthy-horizon execution in Massive Language Fashions (LLMs) stays unstable even when high-level methods are supplied. Evaluating on managed algorithmic puzzles, we reveal that whereas decomposition is crucial for stability, excessive decomposition creates a “no-recovery bottleneck”. We present that this bottleneck turns into essential on account of extremely non-uniform error distribution, the place constant errors on a couple of “onerous” steps grow to be irreversible. To handle this, we suggest Lookahead-Enhanced Atomic Decomposition (LEAD). By incorporating short-horizon future validation and aggregating overlapping rollouts, LEAD gives sufficient isolation to keep up stability whereas retaining sufficient native context to appropriate errors. This permits the o4-mini mannequin to unravel Checkers Leaping as much as complexity n = 13, whereas excessive decomposition fails past n = 11.







