EMNLP 2026 Findings

SCoRE spatial consistency reasoning

Spatial Reasoning in Large Vision-Language Models via Egocentric-to-Allocentric Transformation

Yoonji Kim*·Jieun Kim*·Sung-Bae Cho

Yonsei University · *Equal contribution

SCoRE overview: stage 1 learns egocentric spatial grounding from single camera views; stage 2 learns perspective-taking reasoning with egocentric observation, decentering, mental rotation and allocentric conclusion, then GRPO with 2 to 8 views and reward 0.3 accuracy + 0.6 consistency + 0.1 format.
Stage 1 builds camera-centric spatial perception from single views. Stage 2 teaches egocentric-to-allocentric transformation with structured reasoning traces, then refines the policy with GRPO and a cross-view consistency reward.

motivation

Same scene, different views, different answers

An LVLM can answer a spatial question correctly from one view of a scene and wrongly from another. Asked where the road is when sitting on the bench, Qwen2.5-VL answers front, left and behind for three views of the same park.

Existing training supervises only the final answer, so models learn view-specific shortcuts. People instead transform what they see (egocentric) into a frame tied to the scene (allocentric), where the answer no longer depends on the viewpoint. SCoRE trains an LVLM to do this transformation explicitly.

A park with a bench and a road seen from three cameras. Qwen answers front, left and behind for the three views; SCoRE answers front for all. Below: egocentric relations per view, then decentering, mental rotation and reference frame coordination leading to the answer front.

abstract

Reasoning in a shared, scene-centered frame

Spatial reasoning in large vision-language models involves inferring view-specific spatial relations and integrating them into a consistent scene-centered representation. However, existing approaches often learn view-dependent relations tied to individual camera frames, which limits their ability to form viewpoint-invariant spatial representations. Motivated by cognitive models of spatial perspective-taking, we propose SCoRE (Spatial Consistency Reasoning), a method that realizes multi-view spatial reasoning with egocentric-to-allocentric transformation, which first trains the model to infer egocentric object relations from single-view observations, and then applies reinforcement learning to encourage the model to estimate viewpoint orientation, transform egocentric relations into a shared scene-centered reference frame, and predict allocentric object relations. To enforce viewpoint-invariant reasoning, we optimize the model with a multi-view consistency reward that aligns predictions from different views in the common allocentric space. Experiments show that SCoRE improves spatial reasoning performance by +3.69%p on multi-view benchmarks and +4.72%p on general spatial reasoning benchmarks over the state-of-the-art methods, demonstrating the effectiveness of explicit reference-frame transformation for robust multi-view spatial understanding.

method

From what each camera sees to what holds for the scene

SCoRE writes its reasoning in four parts: the egocentric observation rego, the facing direction of the reference subject rhead, the transformation rtrans, and the allocentric conclusion rallo.

QWhen I’m sitting on the bench, where is the road relative to me?

View 1 of the park, bench on the right and road on the left.V1
bench
right
road
left
Qwenfront ✓
SCoREfront ✓
View 2 of the park, bench on the right and road on the left.V2
bench
right
road
left
Qwenleft ✗
SCoREfront ✓
View 3 of the park, bench in front and road behind.V3
bench
front
road
behind
Qwenbehind ✗
SCoREfront ✓

01Decentering

In view 1, the bench faces left.

02Mental rotation

The road is on the camera’s left, which is in front of the bench.

03Reference-frame coordination

From the bench, the road is in front.

V1, V2, V3 → front R_cons: every view agrees with the answer

Example from Figure 1 of the paper.

Why train it, not prompt it

Prompting five LVLMs to reason from egocentric to allocentric view helps on the multi-view data, but the gains are uneven across models and sensitive to how the prompt is written. SCoRE therefore builds the transformation into training.

Accuracy of five LVLMs on multi-view data with and without egocentric-to-allocentric chain-of-thought prompting.

Stage 1: egocentric spatial grounding

Supervised fine-tuning on 4,240 single-view questions such as “From the camera’s perspective, where is the chair relative to the black thing?” The model answers with one of eight directions. No reasoning trace is needed at this stage.

Stage 2: perspective-taking reasoning

Transformation fine-tuning on 788 four-view samples teaches the trace format: for each view, where the target appears, which way the subject faces, and the transformed relation, then one final answer. Traces are generated automatically from SCOPE scene metadata and camera poses.

GRPO then trains on 5,516 samples with 2, 4, 6 or 8 views and G = 4 rollouts, without a reference trace.

Consistency reward

R = 0.3·Racc + 0.6·Rcons + 0.1·Rfmt

From each rollout, SCoRE reads the conclusion after “Therefore” in every view block, giving the set of view-wise answers 𝒱.

Rcons = Rmatch − 0.3·Runiq
Rmatch = share of views that match the answer
Runiq = (distinct answers − 1) / |𝒱|

A rollout with no view conclusions gets −0.5. Racc is +1 or −1 for the final answer, and Rfmt gives 0.25 for each required part of the trace. Consistency has the largest weight because agreement across views is the goal.

results

More accurate and more consistent across views

Qwen2.5-VL-7B trained with SCoRE, compared with GRPO-based spatial reasoning models: VILASR, SpatialReasoner, SpatialThinker and Spatial-SSRL. In-domain: MindCube, MMSI-Bench, ViewSpatial-Bench. Out-of-domain: What’sUp, SPAR-Bench, BLINK.

+3.69
multi-view accuracy over the best prior model, averaged per benchmark
+4.72
general spatial accuracy over the best prior model
34.0
global consistency on MindCube-Hard, up from 23.0

Multi-view spatial reasoning

MindCubeMMSI-Bench
ModelRotationAmongAroundAvg.PositionalAttributeMotionMSRAvg.Avg.Δ
Qwen2.5-VL-7B31.0030.1727.2529.4728.8725.5320.6119.7025.4027.44–
Baselines
VILASR-7B33.0028.8325.0028.9433.3628.6018.6224.7528.8028.87+1.43
SpatialReasoner32.5038.6729.2534.5020.8727.8219.3324.7522.2028.35+0.91
SpatialThinker35.0030.5025.2530.2528.3727.7526.5726.2628.1429.20+1.76
Spatial-SSRL39.5029.0025.7531.4228.2221.7425.3027.7826.4328.93+1.49
Ours
SCoRE (w/o RL)34.0039.3342.5038.6128.6223.9119.3524.7525.8032.21+4.77
SCoRE (full)34.5040.8342.5039.2829.0627.6526.6439.3930.5034.89+7.45

QA accuracy (%). Bold: best; underlined: second best. Δ: gain over Qwen2.5-VL-7B.

SCoRE has the best average on both benchmarks. Its gains are largest on tasks that need a consistent layout across views: Among and Around on MindCube and multi-step reasoning (MSR) on MMSI-Bench.

Egocentric vs. allocentric questions

On a 1,000-question subset of ViewSpatial-Bench, SCoRE reaches 42.65%, 4.60 points above Spatial-SSRL. The gain comes mostly from allocentric questions: 46.31%, 11.14 points above the base model, against 2.83 points on egocentric questions.

ViewSpatial-Bench accuracy for egocentric, allocentric and average settings; SCoRE is highest at 37.15, 46.31 and 42.65.

Beyond multi-view data

Training on multi-view consistency also helps single-image and general spatial tasks. SCoRE reaches 98.00% on What’sUp and raises the average over three out-of-domain benchmarks from 64.32% to 68.57%.

ModelWhat’sUpSPARBLINKAvg.Δ
Qwen2.5-VL-7B90.1033.8569.0064.32–
SpatialReasoner84.2033.2149.3455.58−8.74
SpatialThinker92.3034.0465.2063.85−0.47
SCoRE (w/o RL)97.6034.1071.7567.82+3.50
SCoRE (full)98.0035.4772.2568.57+4.25

MindCube categories

Across the 11 MindCube subcategories, the biggest gains are on agent–object relations (A-O: 44.73% against 23.65% for the base model) and on viewpoint sequences (50.17% against 24.42%). Both require relating objects to an agent and updating those relations as the view changes.

Accuracy of SCoRE, the base model and baselines across 11 MindCube subcategories grouped into object arrangement, perspective taking, relation pattern and viewpoint dynamics.

Relational consistency

We turn MindCube-Hard into relational questions and check whether all predicted pairwise relations form one coherent layout. SCoRE is best on every metric, and global consistency, the share of samples where every relation is right, rises by 11 points.

ModelRelation acc.Pair cov.Global cons.
Qwen2.5-VL-7B48.5045.9323.00
SpatialReasoner36.7440.3319.00
SpatialThinker50.6749.4027.00
Spatial-SSRL51.2346.3732.00
SCoRE52.4750.3334.00
An office seen from three views, and the object layouts reconstructed by each model; SCoRE's layout of the blue bin, monitors and black sofa matches the ground truth, while baselines distort it.

Ablations

TrainingMindCubeMMSIAvg.
Single stage
Perception SFT only34.0322.9028.47
RL only30.2526.2328.24
Transformation SFT only31.0927.8429.47
Remove one stage
w/o perception SFT36.8731.3234.10
w/o transformation SFT35.7630.4933.13
w/o RL38.6125.8032.21
SCoRE (full)39.2830.5034.89

No single stage comes close to the full pipeline. Removing RL costs the most, and removing transformation SFT costs more than removing perception SFT.

RewardAcc.Cons.Fmt.MindCubeMMSI
w/o consistency✓✓37.6725.90
w/o accuracy✓✓38.8326.90
Full✓✓✓39.2830.50

Dropping the consistency reward hurts more than dropping the accuracy reward, especially on MMSI-Bench (−4.60 points). Correct answers alone do not teach the model to agree with itself across views.

Checking each reasoning step

StepAccuracy (%)
Egocentric relation (rego)84.5
Heading (rhead)81.2
Transformation (rtrans)98.9
Allocentric relation (rallo)68.1

Given its own egocentric relation and heading, the model almost always composes the right transformed relation (98.9%). The gains come from the transformation itself, not from shortcuts to the final answer.

Without viewpoint labels

ModelNo V&OWith V&ONoisy V&O
Qwen2.5-VL-7B24.0025.0025.33
SpatialReasoner32.3329.3330.67
SpatialThinker26.3326.0028.00
VILASR-7B22.0022.0020.00
Spatial-SSRL29.0029.6730.00
SCoRE39.6735.3334.33

On MindCube, SCoRE stays best whether viewpoint and orientation (V&O) labels are given, removed or corrupted. It does best with no labels at all, so it does not depend on them.

Case study

Asked where the toilet is while using the handwashing area, SCoRE takes the perspective of the person at the sink, links the towel across both images, and answers to my left. Qwen reads the local cues correctly but stays in the camera’s frame and answers to my right.

Two bathroom images with a towel, sink and toilet. SCoRE reasons from the person's facing direction and answers that the toilet is to the left; Qwen answers right.

cite

BibTeX

@inproceedings{kim2026score,
  title     = {Spatial Reasoning in Large Vision-Language Models via
               Egocentric-to-Allocentric Transformation},
  author    = {Kim, Yoonji and Kim, Jieun and Cho, Sung-Bae},
  booktitle = {Findings of the Association for Computational Linguistics:
               EMNLP 2026},
  year      = {2026}
}