NeurIPS 2026 Main Track

SphMind:Towards Robust, Training-Free VLM-based Spatial Reasoning with a 360° Camera

Shriram Damodaran1, Soumyaratna Debnath1, Cheston Tan2,
Addison Lin Wang1†

1 EmPACT Lab, Nanyang Technological University, Singapore2 Institute of High Performance Computing, A*STAR, Singapore

TL;DR SphMind separates visual semantics from spherical geometry, helping frozen vision-language models answer spatial questions about 360° scenes without retraining.

† Corresponding author

SphMind corrects a VLM's spatial answer on a 360-degree panorama and compares accuracy across reasoning categories
SphMind grounds VLM answers in spherical scene geometry while preserving the model’s visual understanding.

Abstract

Omnidirectional or 360° cameras give embodied agents a holistic view of their surroundings, but vision-language models (VLMs) trained on perspective images struggle with spherical distortion and the wrap-around seam. We present SphMind, a training-free, plug-and-play framework that decouples semantic perception from geometric reasoning. Its Spherical Harmonics-based Spatial Graph (SHSG) represents object relationships directly on the sphere using rotation-equivariant transformations. Inference-Time Geometric Grounding (IGG) then steers a frozen VLM’s internal representation toward geometrically valid answers, without changing model weights. Across three benchmark datasets, SphMind improves directional reasoning by an average of 21.4% on MP3D and Stanford2D-3D, exceeds prompt-engineering strategies by 8.7% on real-world ODI-Bench, and delivers 5.9× the full rotational consistency of the pixel-space baseline. It also resolves directional queries in an in-the-wild 360° capture where baseline VLMs fail.

Motivation

Core Question

How can frozen VLMs reason reliably about spatial relationships in a 360° scene without learning spherical geometry from new training data?

SphMind correctly answers which object is closer in a 360-degree scene, while the baseline VLM gives the opposite answer
Geodesic reasoning on the sphere identifies the board as closer to the observer; the current VLM reverses the distance relationship.

VLMs bring strong semantic understanding, but their image representations assume a flat grid. A 360° panorama maps a sphere onto that grid, stretching regions near the poles and splitting continuous scenes at the seam. As a result, models can recognize objects yet misjudge which object is left, right, nearer, or farther. SphMind leaves semantic perception with the VLM and handles spatial geometry explicitly on the sphere.

Method

Framework Summary

Given a 360° image and a spatial question, SphMind detects scene objects, maps their locations onto the sphere, and builds a query-aware geometric representation. It uses that evidence to guide the frozen VLM toward a consistent answer at inference time.

Component 01

Spherical Harmonics-based Spatial Graph (SHSG)

SHSG lifts detected objects to the sphere and encodes their locations and relationships as continuous spherical-harmonic fields. Rotation-equivariant scoring handles directional queries while preserving the panorama’s global geometry.

Component 02

Inference-Time Geometric Grounding (IGG)

IGG optimizes the VLM’s final-token hidden state against a differentiable geometric cost. A norm-preserving update steers the answer toward a geometrically valid candidate without updating model weights.

SphMind Pipeline

SphMind framework overview showing spherical scene graph reasoning and inference-time geometric grounding
SHSG builds a spherical scene representation; IGG uses its geometric constraints to guide the frozen VLM’s answer.

The two components work together: SHSG computes spherical evidence for candidate answers, and IGG injects that evidence into the VLM’s inference process. The model’s semantic knowledge remains available while its spatial prediction is corrected.

Spherical Directional Evidence

t-SNE projection showing distinct clusters for six spherical spatial directions

Key Finding

The plot shows SHSG representations separating into distinct groups for the six directions. Left and right form separate arcs; front and behind occupy different regions; above and below appear as isolated groups. The measured silhouette score of 0.532 supports this directional separation. The visualization projects 450 uniformly sampled points on the sphere. SHSG uses these separated patterns to score candidate answers by direction.

Results

Evaluation Highlights

SphMind improves spatial reasoning across synthetic and real-world panoramas, transfers to viewpoint changes, and remains consistent as the panorama rotates.

MP3D + Stanford2D-3D

+21.4%

Average gain in directional reasoning across the two panoramic benchmarks.

ODI-Bench

+8.7%

Improvement over prompt-engineering strategies on real-world panoramas.

Rotational Consistency

5.9×

Full rotational consistency score compared with ERP-pixel reasoning.

Real-world transfer

ODI-Bench directional accuracy

%
StrategyRelative directionEgo-view orientation
Baseline46.1345.71
Viewpoint guidance45.8045.10
Crop grounding44.1844.15
Response refinement47.6545.76
Chain of thought46.3343.20
SphMind56.3552.60

SphMind leads all four prompt-only strategies on both ODI-Bench measures; the paper reports an average gain of 8.7%.

Viewpoint generalization

OpenView-VQA accuracy

%
Question typeBaseSphMindTrained
Overall21.039.746.0
Pitch22.536.244.3
Yaw25.045.648.7
Pitch + yaw23.142.145.0

The training-free method transfers to pitch and yaw changes, approaching the trained specialist.

Rotation robustness

Stanford 2D-3D-S consistency

Score
MethodFull90°180°270°
ERP-pixel0.0430.4370.2120.451
SphMind0.2530.6530.5950.648

Full score requires a correct answer that remains consistent across all tested rotations.

Component study

Ablation on MP3D

Accuracy %
ConfigurationOverallDirectionDistance
Full SphMind58.453.164.1
No IGG gradient50.546.654.9
SHSG only45.142.548.5
Cartesian encoding47.343.152.6
ERP centroid encoding43.638.949.8

Both spherical encoding and inference-time grounding contribute to the gains.

Contributions

01

We identify the spatial-semantic gap that causes VLMs to misread spherical geometry in 360° images.

02

We introduce SHSG, a spherical-harmonic scene graph for continuous, rotation-equivariant spatial reasoning.

03

We introduce IGG, a model-agnostic inference-time optimization that grounds VLM answers without weight updates.

04

We demonstrate gains across panoramic VQA, real-world ODI-Bench, viewpoint generalization, and rotational consistency.

Citation

@misc{debnath2026sphmind,
  title={SphMind: Towards Robust, Training-Free VLM-based Spatial Reasoning with a 360 Camera},
  author={Damodaran, Shriram and Debnath, Soumyaratna and Tan, Cheston and Wang, Addison Lin},
  year={2026},
  note={Manuscript}
}