A study published September 25 describes "Anchored Planning," a technique that improves action selection in frozen visual world models by aiming at intermediate targets rather than final goals Aim Short to Reach Far: Your Frozen World Model Can Plan Better Than You Think.

Standard planners built on visual world models typically evaluate predicted trajectories against a fixed end-state image, a strategy that can stall control when optimal paths temporarily move away from the target Aim Short to Reach Far: Your Frozen World Model Can Plan Better Than You Think. The reported technique bypasses this limitation using an existing, untrained model, offering gains in simulation environments including Cube, PushT, Reacher, and TwoRoom without requiring further optimization steps.

Automatica Press reported in mid-September that recent studies highlighted discrepancies between benchmark scores and real-world AI performance, noting that evaluation metrics often fail to capture practical utility AI Benchmarks Are Broken: Six New Studies Expose the Gap Between Test Scores and Real-World Performance. The new work reinforces these concerns by demonstrating that minimizing successor-prediction error does not guarantee effective control, underscoring limitations in how planning algorithms are currently optimized and scored.

The authors propose "Anchored Planning," which retrieves a recorded trajectory segment where the start and end points resemble the agent's current and desired observations. The system then directs actions toward an observation taken shortly after the segment's start, scoring these moves against the frozen model Aim Short to Reach Far: Your Frozen World Model Can Plan Better Than You Think. Evaluations showed this approach improved action synthesis and recorded-action ranking across the four tasks, surpassing the default LeWM planner in every long-range scenario tested. Performance gains required shrinking the retrieval span as execution progressed.

The analysis comes from a Cornell University-hosted preprint that has not undergone peer review and notes potential commercial conflicts among the authors. The report includes no independent verification of results or details on deployment outside the listed simulation environments.

The paper argues that adjusting the intermediate target allows the same frozen model to solve scenarios that final-goal scoring previously missed.