NavGPT-3
Harnessing Context in a
Hierarchical Navigation Runtime
Connecting language-model reasoning with physical action,
through a runtime built for navigation.
1AIML, Adelaide University 2Metacognition 3ANU
4Roblox 5PKU 6SJTU 7UNC Chapel Hill 8UNSW
About the benchmarks
Success rate measures how often the agent reaches its goal. SPL (success weighted by path length) also rewards efficient routes. Both use a 0–100 scale; higher is better.
NavGPT-3 variants appear first; other methods are ordered by the selected metric. The trajectory explorer contains selected earlier recordings, with sample statistics separate from these full benchmark results.
Reason about the task. Act in the world.
Keep the two connected.
Language models can interpret complex instructions and reason over long tasks. Navigation policies provide the low-level actions that move an agent through its environment. NavGPT-3 connects them through a navigation harness and a runtime that coordinates execution.
The NavGPT Planner chooses the next step and decides when the task is complete. The harness provides observations, spatial tools, persistent route state, and execution feedback. The runtime controls logical threads, their permissions, and motion authority. The Planner can act through tools or delegate extended navigation to NavGPT VLA, then inspect the returned evidence and revise the route.
Reasoning with context
Observations, route history, and spatial references make the agent’s progress inspectable.
Action with feedback
Delegation, return, and local repair let the Planner revise physical execution.
From a navigation policy
to a complete system.
NavGPT-3 achieves higher success rates than its standalone 8B navigation policy on R2R-CE and RxR-CE. The gains below use GPT-6 Astra as the Planner.
NavGPT VLA 8B → NavGPT-3 +7.00 percentage points
NavGPT VLA 8B → NavGPT-3 +12.24 percentage points
Inside a navigation run.
Follow recorded conversations with Claude Opus 5 and GPT-6 Astra on R2R-CE and RxR-CE. Inspect each tool call, its route, and the observations behind the next decision. These selected earlier recordings illustrate system behavior; their sample statistics are separate from the full benchmark results above.
Loading recorded navigation examples…
Conversation & tools
Loading recorded conversation…
No episodes match these filters.
Browse trajectories
Each map is fitted independently. Select a tile to open its conversation and full trajectory.
Methodology.
The harness connects the Planner to spatial tools and an optional navigation policy, maintaining context, state, and execution feedback. The runtime coordinates logical threads, permissions, and motion authority. NavGPT VLA provides low-level navigation and allocates visual tokens using scene change, recency, and camera view.
Thread management
The runtime coordinates concurrent reasoning and execution while controlling which thread can move the agent. The diagram illustrates route review interrupting a rollout for correction and a safety interrupt revoking motion authority until fresh state and runtime authorization allow execution to resume.
Codec allocation
NavGPT VLA allocates visual context across a four-view history using scene change, recency, and camera view. The codec chooses whole-image resolutions to distribute the available visual-token budget.
VLA training mixture and codec comparison
NavGPT VLA is trained on approximately 19.28M effective records covering navigation, embodied question answering, and visual grounding. This weighted total includes reused and augmented examples; it is not a count of unique trajectories. The codec study compares 4B and 8B policies across three cumulative training scopes.


Citation
@article{zhou2026navgpt3,
title = {NavGPT-3: Harnessing Context in a Hierarchical Navigation Runtime},
author = {Zhou, Gengze and Hong, Yicong and Zhang, Jiazhao and
Zhao, Xunyi and Zhou, Jian and Lei, Zixing and Wang, Zun and
Zhao, Chongyang and Chen, Xionghui and Gould, Stephen and
van den Hengel, Anton and Wu, Qi},
journal = {arXiv preprint arXiv:2610.10787},
year = {2026},
eprint = {2610.10787},
archivePrefix = {arXiv},
primaryClass = {cs.RO},
url = {https://arxiv.org/abs/2610.10787}
}
@inproceedings{zhou2024navgpt,
title={{NavGPT}: Explicit Reasoning in Vision-and-Language Navigation with Large Language Models},
author={Zhou, Gengze and Hong, Yicong and Wu, Qi},
booktitle={Proceedings of the AAAI Conference on Artificial Intelligence},
volume={38},
pages={7641--7649},
year={2024},
doi={10.1609/aaai.v38i7.28597}
}
@inproceedings{zhou2024navgpt2,
title={{NavGPT-2}: Unleashing Navigational Reasoning Capability for Large Vision-Language Models},
author={Zhou, Gengze and Hong, Yicong and Wang, Zun and Wang, Xin Eric and Wu, Qi},
booktitle={European Conference on Computer Vision (ECCV)},
pages={260--278},
organization={Springer},
year={2024}
}