Spatial audio-visual reasoning

FloorSAV

Elucidating Spatial Audio-Visual Context
with 2D Floormap for AV-LLMs

Korea Advanced Institute of Science and Technology

† Corresponding authors

Ground what models see and hear
in a shared spatial context.

A dynamic 2D floormap brings geometry, motion, sound, and object landmarks into one view—so audio-visual models can reason about viewpoints, places, and paths.

Training-freeTask-agnosticSingle inference
One scene. Two complementary views.20 fps · source footage
Egocentric videoWhat the camera sees
RGB + Audio
FloorSAV mapSpatial context over time
Geometry + Motion + Sound
01:00 / 01:20
Original RGB frames + original FloorSAV map rendering at 20 fps.
CameraSound estimateTrajectory
The idea

Give the model a map.
Let it connect the evidence.

A first-person video shows only part of a changing scene. FloorSAV adds a top-down 2D floormap that brings geometry, camera motion, spatial sound, and object landmarks into one shared frame.

The audio-visual language model (AV-LLM) reads both synchronized videos, with a floormap interpretation guide, to answer spatial questions.

Training-free

Use an existing AV-LLM without fine-tuning its weights.

Task-agnostic

Use the same spatial representation for viewpoints, places, and paths.

Single inference

After preparing the scene map, the AV-LLM reasons over both streams in one call.

Read the abstract +

While 3D spatial reasoning in dynamic egocentric environments is crucial for embodied intelligence, audio-visual large language models (AV-LLMs) lack explicit mechanisms to process and internalize global geometry directly from raw sensory streams. Existing approaches either require costly fine-tuning or underutilize the model’s cross-modal reasoning capacities. In this paper, we propose FloorSAV, a novel framework that explicitly grounds spatial audio-visual context by rendering a dynamic 2D floormap. By integrating 3D point clouds, camera trajectories, spatial audio cues, and semantically grounded object landmarks, we inject this floormap into the AV-LLM as a synchronized stream with an egocentric video. AV-LLMs utilize their multi-modal capabilities to jointly reason over visual, auditory, and geometric cues in a single inference with floormap interpretation guidance. We further introduce SAVED-Bench (Spatial Audio-Visual Egocentric Benchmark with Dynamic Agents), constructing essential tasks of spatial capability in real-world scenarios: dynamic relativity, regional, and path reasoning QAs. FloorSAV improves AV-LLMs’ spatial reasoning on various tasks from both SAVED-Bench and SAVVY-Bench. Studies with ground-truth floormaps demonstrate the substantial potential of FloorSAV with accurate spatial information.

01 / See it in action

See the scene. Follow the spatial reasoning.

Three spatial questions, explained one moment at a time.
Watch the full scene with brief pauses at key moments. Pause longer or jump to the next step whenever you like.

Regional reasoning

You hear them. But where are they?

When “I mean, I like beef rare” is spoken, which area of the home is the other person in?

Paused · after departure01:42.00
First-person videoSource 01:00

2D floormapMatched source time

00:00 / 00:20

Guided summaries of selected model explanations. Numbered markers connect the text to the evidence; they are presentation aids. Green dots in the original maps are sound estimates, not tracked people.

02 / Method

From sensory streams
to a shared spatial reference.

Project the scene into 2D. Add motion, sound, and semantics.
Feed both synchronized videos and a floormap interpretation guide to the AV-LLM.

Sensor inputs → Floormap videoInteractive illustration
3D Point CloudRoom and object structure
Camera + IMUPosition, direction, and motion
Microphone ArrayEstimated sound position
RGB Frames + VLMObject names and locations
3D Points
Geometry provided by AEA
3D POINT CLOUD01 / 05
CameraField of viewTrajectorySound
1

Project the 3D points

Flatten the scene onto a top-down metric grid.

00:00 / 00:20

Illustration of scene preparation. The qualitative videos above use recorded scene data.

Two synchronized inputsFirst-person video + audio
2D floormap video
+ Question & floormap interpretation guide
One inferenceAV-LLMJoint audio-visual-spatial reasoning
Spatial answerViewpoints.
Places. Paths.
Explore the method figure ↗
03 / SAVED-Bench

Spatial questions.
Dynamic people.

A benchmark for reasoning about viewpoints,
places, and movement in real egocentric scenes. We introduce SAVED-Bench to test reasoning with dynamic agents.

1,988Questions
55Recordings
9Tasks
800 questions · 4 tasks

Dynamic relativity

Reason about direction and distance across the camera wearer’s and another person’s viewpoints.

“From their viewpoint, where is the object?”
592 questions · 2 tasks

Regional reasoning

Identify a person’s location and the places they visited during a sound event.

“Which area was the speaker in?”
596 questions · 3 tasks

Path reasoning

Find objects near a finite path and estimate how far a person traveled.

“What would I pass on the way to the couch?”
Task definitions and benchmark distributions ↗
04 / Results

Better spatial context.
Stronger spatial reasoning.

SAVED-Bench · Gemini-3.6-Flash · No fine-tuning

Across spatial skills

A wider view of spatial reasoning.

Compare eight skill groups across SAVED-Bench and SAVVY-Bench. Add ground-truth maps to explore the role of spatial accuracy.

Ego / Exo: SAVVY-Bench · Gemini-2.5-Pro. Other axes: SAVED-Bench · Gemini-3.6-Flash. Paired tasks are averaged. Each axis starts at 0 and uses its labeled maximum; scales stay fixed when GT maps are toggled. Sources: paper Tables 1 & 4.

Overall · 1,988 questions58.3 65.3

+7.0 points with FloorSAV.
All three category averages improve.

SAVED-Bench scores · 0–100 · higher is better

Current SAVED-Bench release: accuracy or AbsMRA scores on a 0–100 scale. Higher is better.
Task QAs Baseline FloorSAV Δ points
Ego-to-Exo Direction 200 56.0 65.0 +9.0
Ego-to-Exo Distance · AbsMRA 200 34.3 34.6 +0.3
Exo-to-Ego Direction 200 32.0 50.5 +18.5
Exo-to-Ego Distance · AbsMRA 200 46.4 44.7 -1.7
Dynamic Relativity Overall 800 42.2 48.7 +6.5
Location Awareness 200 92.0 94.5 +2.5
Visit History 392 86.7 92.3 +5.6
Regional Overall 592 89.4 93.4 +4.0
Line-Path Search (static) 196 60.2 77.0 +16.8
Line-Path Search (dynamic) 200 73.0 78.0 +5.0
Trajectory Distance · AbsMRA 200 16.9 25.0 +8.1
Path Reasoning Overall 596 50.0 60.0 +10.0
Overall · QA-weighted average 1988 58.3 65.3 +7.0

A clearer spatial reference helps across viewpoints, regions, and paths.

Accuracy / AbsMRA · 0–100, higher is better. Category averages weight tasks equally; the headline overall score weights QAs.

Download reported results ↓
Evaluation details +

SAVED-Bench values match the released Table 1 results at the paper’s one-decimal precision. Category and overall averages are preserved as reported. Displayed gains are differences between these reported values.

Direction, region, and line-path tasks use accuracy; distance tasks use AbsMRA. All category averages improve; Exo-to-Ego Distance decreases from 46.4 to 44.7.

The radar averages direction and distance for each viewpoint pair, and static and dynamic tasks for Line-Path Search. It groups different benchmarks and models for the overview; it is not a single-model aggregate score.

Looking ahead

Better maps help.
Spatial reasoning is still hard.

FloorSAV improves performance on SAVED-Bench and SAVVY-Bench. It is not error-free: sound localization, long video contexts, and map reading remain challenging. Ground-truth map studies show further gains when spatial information is more accurate.

Build on this work

Citation

FloorSAV · Kim, Kim & Oh · 2026

@misc{kim2026floorsav,
  title = {FloorSAV: Elucidating Spatial Audio-Visual
           Context with 2D Floormap for AV-LLMs},
  author = {Kim, Kyeong-Rae and Kim, Sungnyun and Oh, Tae-Hyun},
  year = {2026},
  url = {https://byulharang.github.io/FloorSAV/}
}

Paper figure

Original manuscript figure. The current benchmark release is reported in the results table.

Explore the tasks

You · camera wearer Other person Top-down views
Schematic examples to explain each task. Not recorded predictions.