A proposed AAAI-27 workshop

Geometric & Spatial Reasoning in Vision-Language Models

FORMATFull-day workshopWHEN & WHERETo be announced
SPATIAL CONSTRUCTION INTERACTIVE
zxy
Separate components

Illustration of a 3 × 3 × 3 assembly.

Overview

Vision-language models struggle to maintain geometric consistency across fragments, viewpoints, and reasoning steps. This workshop examines these limitations and methods for improving spatial reasoning as task complexity increases.

The program brings together researchers in multimodal learning, embodied AI, and constraint reasoning through invited talks, contributed papers, posters, a competition, and a panel.

01

Spatial perception

Recognize relations, depth, and orientation already present in an image.

02

Spatial construction

Construct and maintain globally consistent structures under geometric constraints.

Topics of interest

01Spatial perception & grounding+

Spatial relations, depth, size, orientation, perspective-taking, and generalization across viewpoints.

02Construction under constraints+

Combining appearance evidence with geometric constraints in reassembly, multi-step assembly, and grid reasoning.

03Architectures & mechanisms+

Visual tokenization, resolution, positional encodings, spatial memory, combinatorial decoding, and locating where geometric information is lost.

04Neuro-symbolic reasoning+

Pairing VLM proposals with constraint solvers, search, and classical methods for geometric compatibility.

05Training & representations+

Spatial pretext tasks, exactly verifiable construction tasks, and reinforcement learning with verifiable rewards.

06Evaluation methodology+

Uniqueness guarantees, machine-checkable answers, contamination resistance, and metrics that distinguish local from global correctness.

07Embodied AI & applications+

Manipulation, assembly, navigation, and world models—and whether benchmark gains transfer to physical competence.

Invited speakers

Mohit Bansal

Mohit Bansal

University of North Carolina
at Chapel Hill

Ming-Hsuan Yang

Ming-Hsuan Yang

University of California,
Merced

Manling Li

Manling Li

Northwestern
University

Bistra Dilkina

Bistra Dilkina

University of
Southern California

Call for papers

We invite research and position papers on the topics above.

Full papers

Original research with substantive contributions.

Up to 8 pages

Short & position papers

Early results, focused studies, and new arguments.

Up to 4 pages

Extended abstracts

Recently published work, for presentation only.

Presentation

Every accepted paper will be presented as a poster. Selected papers will also receive an oral presentation slot.

Spatial reconstruction competition

Recover the original positions of shuffled, interlocking image pieces using visual evidence and geometric constraints.

2,000Held-out instances
16–256Pieces per instance
4Difficulty levels

Each instance is verified by a constraint solver to have exactly one solution.

Opens 4 months before the workshopCloses 3 weeks before the workshop

Planned competition · Dates, data, baselines, and evaluation code to be announced.

COMPETITION TRACKS

Zero-shot prompting

Use API-only models with zero-shot prompting.

Program

Tentative schedule
WELCOME

Opening remarks & problem framing

INVITED TALK I

Multimodal reasoning & evaluation

CONTRIBUTED RESEARCH

Oral presentations I

Four talks · 15 minutes each

POSTER SESSION

Coffee break & poster session I

INVITED TALK II

Embodied & multimodal agents

CHALLENGE

Competition session & results

Organizing committee

Contact organizers
Emilio Ferrara

Emilio Ferrara

University of Southern California

Contact ↗
SUPPORTING TEAM
Tianyu Shi

Tianyu Shi

McGill University

Ruolin Li

Ruolin Li

University of Southern California