motivation
Same scene, different views, different answers
An LVLM can answer a spatial question correctly from one view of a scene and wrongly from another. Asked where the road is when sitting on the bench, Qwen2.5-VL answers front, left and behind for three views of the same park.
Existing training supervises only the final answer, so models learn view-specific shortcuts. People instead transform what they see (egocentric) into a frame tied to the scene (allocentric), where the answer no longer depends on the viewpoint. SCoRE trains an LVLM to do this transformation explicitly.


V1
V2
V3



