PRISM: Planning with Belief under Incomplete World Models
Embodied task planning requires reasoning under uncertainty about world states that are partially observed and incrementally revealed through interaction. However, mostvision–language model (VLM) planners operate in a reactive sense–plan–act loop without maintaining a persistent representation of object locations. As a result, they cannot accumulate negative evidence across time (e.g., “the object is not here”), leading to redundant exploration and repeated manipulation failures. We introduce PRISM, Planning with Belief under Incomplete World Models, a lightweight inference-time framework that augments VLM planners with an explicit probabilistic belief over candidate object locations. PRISM maintains a categorical distribution over locations for each object and updates it online via Bayesian conditioning from visual observations and structured action feedback. To reduce initial uncertainty, the belief is initialized from instruction-aware affordance priors and integrated into the planning loop via prompt conditioning and belief-guided action filtering. We evaluate PRISM on EmbodiedBench across EB-Habitat (Habitat) and EB-ALFRED (AI2-THOR) using InternVL2.5-8B and GPT-4o-mini backbones with 50 episodes per task category. On EB-Habitat, average success improves from 21.6% to 42.8% with InternVL2.5-8B and from 30.4% to 47.2% with GPT-4omini. On EB-ALFRED, the belief pipeline enables InternVL2.5-8B, which achieves near-zero success under ReAct, to reach up to 30% success across multiple task categories. Beyond performance gains, the belief state provides a transparent uncertainty representation that exposes the agent’s location hypotheses at each step, supporting interpretability for human supervisors in shared-space deployments.
