NavGPT-3 Metacognition

NavGPT-3

Harnessing Context in a
Hierarchical Navigation Runtime

Connecting language-model reasoning with physical action,
through a runtime built for navigation.

Gengze Zhou1,2†, Yicong Hong4, Jiazhao Zhang5, Xunyi Zhao1, Jian Zhou1,
Zixing Lei6, Zun Wang7, Chongyang Zhao8, Xionghui Chen5,
Stephen Gould2,3, Anton van den Hengel1,2, Qi Wu1,2†

1AIML, Adelaide University   2Metacognition   3ANU
4Roblox   5PKU   6SJTU   7UNC Chapel Hill   8UNSW

†Correspondence to: {gengze.zhou, qi.wu01}@adelaide.edu.au

Reason about the task. Act in the world.
Keep the two connected.

Language models can interpret complex instructions and reason over long tasks. Navigation policies provide the low-level actions that move an agent through its environment. NavGPT-3 connects them through a navigation harness and a runtime that coordinates execution.

The NavGPT Planner chooses the next step and decides when the task is complete. The harness provides observations, spatial tools, persistent route state, and execution feedback. The runtime controls logical threads, their permissions, and motion authority. The Planner can act through tools or delegate extended navigation to NavGPT VLA, then inspect the returned evidence and revise the route.

Reasoning with context

Observations, route history, and spatial references make the agent’s progress inspectable.

Action with feedback

Delegation, return, and local repair let the Planner revise physical execution.

From a navigation policy
to a complete system.

NavGPT-3 achieves higher success rates than its standalone 8B navigation policy on R2R-CE and RxR-CE. The gains below use GPT-6 Astra as the Planner.

R2R-CE success rate
74.5181.51%

NavGPT VLA 8B → NavGPT-3 +7.00 percentage points

RxR-CE success rate
78.1990.43%

NavGPT VLA 8B → NavGPT-3 +12.24 percentage points

Inside a navigation run.

Follow recorded conversations with Claude Opus 5 and GPT-6 Astra on R2R-CE and RxR-CE. Inspect each tool call, its route, and the observations behind the next decision. These selected earlier recordings illustrate system behavior; their sample statistics are separate from the full benchmark results above.

Loading recorded navigation examples…

Full recorded trajectory

Select a numbered tool marker to inspect its call and recorded observations.

Methodology.

The harness connects the Planner to spatial tools and an optional navigation policy, maintaining context, state, and execution feedback. The runtime coordinates logical threads, permissions, and motion authority. NavGPT VLA provides low-level navigation and allocates visual tokens using scene change, recency, and camera view.

NavGPT-3 overview: navigation components, synchronous tool use and asynchronous thread coordination, and benchmark performance
The NavGPT-3 methodology: navigation components, execution coordination, and benchmark performance.

Thread management

The runtime coordinates concurrent reasoning and execution while controlling which thread can move the agent. The diagram illustrates route review interrupting a rollout for correction and a safety interrupt revoking motion authority until fresh state and runtime authorization allow execution to resume.

Figure 2: concurrent runtime threads and motion authority during a route-review correction and a safety interrupt
Figure 2. Thread coordination in the NavGPT-3 Runtime. Thin bars show concurrent activity; the teal path shows motion authority.

Codec allocation

NavGPT VLA allocates visual context across a four-view history using scene change, recency, and camera view. The codec chooses whole-image resolutions to distribute the available visual-token budget.

Figure 3: codec allocation across four camera views, optical-flow diagnostics and token grids, and measured allocation changes across 41,216 replayed image/context pairs
Figure 3. Codec allocation of visual context. Scene change, recency, and view determine image resolution; diagnostics and replay measurements show the resulting allocation.
VLA training mixture and codec comparison

NavGPT VLA is trained on approximately 19.28M effective records covering navigation, embodied question answering, and visual grounding. This weighted total includes reused and augmented examples; it is not a count of unique trajectories. The codec study compares 4B and 8B policies across three cumulative training scopes.

NavGPT VLA training mixture: action and next-token supervision shares and effective records by eleven data families, including four-view and single-view navigation
NavGPT VLA training mixture: supervision split and effective records by data family. Instruction-following navigation sources are grouped by camera views.
Success rate and SPL for 4B and 8B policies with and without codec allocation on R2R-CE and RxR-CE across three cumulative training scopes
With-codec and without-codec results across VLN, + Nav. QA, and + Cross-embod. training scopes, evaluated on full R2R-CE and RxR-CE Val-Unseen splits.

Citation

BibTeX
@article{zhou2026navgpt3,
  title = {NavGPT-3: Harnessing Context in a Hierarchical Navigation Runtime},
  author = {Zhou, Gengze and Hong, Yicong and Zhang, Jiazhao and
            Zhao, Xunyi and Zhou, Jian and Lei, Zixing and Wang, Zun and
            Zhao, Chongyang and Chen, Xionghui and Gould, Stephen and
            van den Hengel, Anton and Wu, Qi},
  journal = {arXiv preprint arXiv:2610.10787},
  year = {2026},
  eprint = {2610.10787},
  archivePrefix = {arXiv},
  primaryClass = {cs.RO},
  url = {https://arxiv.org/abs/2610.10787}
}

@inproceedings{zhou2024navgpt,
  title={{NavGPT}: Explicit Reasoning in Vision-and-Language Navigation with Large Language Models},
  author={Zhou, Gengze and Hong, Yicong and Wu, Qi},
  booktitle={Proceedings of the AAAI Conference on Artificial Intelligence},
  volume={38},
  pages={7641--7649},
  year={2024},
  doi={10.1609/aaai.v38i7.28597}
}

@inproceedings{zhou2024navgpt2,
  title={{NavGPT-2}: Unleashing Navigational Reasoning Capability for Large Vision-Language Models},
  author={Zhou, Gengze and Hong, Yicong and Wang, Zun and Wang, Xin Eric and Wu, Qi},
  booktitle={European Conference on Computer Vision (ECCV)},
  pages={260--278},
  organization={Springer},
  year={2024}
}