Training-free
Use an existing AV-LLM without fine-tuning its weights.
Elucidating Spatial Audio-Visual Context
with 2D Floormap for AV-LLMs
Korea Advanced Institute of Science and Technology
Ground what models see and hear
in a shared spatial context.
A dynamic 2D floormap brings geometry, motion, sound, and object landmarks into one view—so audio-visual models can reason about viewpoints, places, and paths.
A first-person video shows only part of a changing scene. FloorSAV adds a top-down 2D floormap that brings geometry, camera motion, spatial sound, and object landmarks into one shared frame.
The audio-visual language model (AV-LLM) reads both synchronized videos, with a floormap interpretation guide, to answer spatial questions.
Use an existing AV-LLM without fine-tuning its weights.
Use the same spatial representation for viewpoints, places, and paths.
After preparing the scene map, the AV-LLM reasons over both streams in one call.
While 3D spatial reasoning in dynamic egocentric environments is crucial for embodied intelligence, audio-visual large language models (AV-LLMs) lack explicit mechanisms to process and internalize global geometry directly from raw sensory streams. Existing approaches either require costly fine-tuning or underutilize the model’s cross-modal reasoning capacities. In this paper, we propose FloorSAV, a novel framework that explicitly grounds spatial audio-visual context by rendering a dynamic 2D floormap. By integrating 3D point clouds, camera trajectories, spatial audio cues, and semantically grounded object landmarks, we inject this floormap into the AV-LLM as a synchronized stream with an egocentric video. AV-LLMs utilize their multi-modal capabilities to jointly reason over visual, auditory, and geometric cues in a single inference with floormap interpretation guidance. We further introduce SAVED-Bench (Spatial Audio-Visual Egocentric Benchmark with Dynamic Agents), constructing essential tasks of spatial capability in real-world scenarios: dynamic relativity, regional, and path reasoning QAs. FloorSAV improves AV-LLMs’ spatial reasoning on various tasks from both SAVED-Bench and SAVVY-Bench. Studies with ground-truth floormaps demonstrate the substantial potential of FloorSAV with accurate spatial information.
Three spatial questions, explained one moment at a time.
Watch the full scene with brief pauses at key moments. Pause
longer or jump to the next step whenever you like.
When “I mean, I like beef rare” is spoken, which area of the home is the other person in?
Guided summaries of selected model explanations. Numbered markers connect the text to the evidence; they are presentation aids. Green dots in the original maps are sound estimates, not tracked people.
Project the scene into 2D. Add motion, sound, and semantics.
Feed both synchronized videos and a floormap interpretation guide
to the AV-LLM.
Flatten the scene onto a top-down metric grid.
Illustration of scene preparation. The qualitative videos above use recorded scene data.
A benchmark for reasoning about viewpoints,
places, and movement in real egocentric scenes. We introduce
SAVED-Bench to test reasoning with dynamic agents.
Reason about direction and distance across the camera wearer’s and another person’s viewpoints.
“From their viewpoint, where is the object?”
Identify a person’s location and the places they visited during a sound event.
“Which area was the speaker in?”
Find objects near a finite path and estimate how far a person traveled.
“What would I pass on the way to the couch?”
SAVED-Bench · Gemini-3.6-Flash · No fine-tuning
Compare eight skill groups across SAVED-Bench and SAVVY-Bench. Add ground-truth maps to explore the role of spatial accuracy.
Ego / Exo: SAVVY-Bench · Gemini-2.5-Pro. Other axes: SAVED-Bench · Gemini-3.6-Flash. Paired tasks are averaged. Each axis starts at 0 and uses its labeled maximum; scales stay fixed when GT maps are toggled. Sources: paper Tables 1 & 4.
Partial GT map: ground-truth other-person positions. Full GT map: additionally uses ground-truth object labels. The same switch updates the graph and table below.
+7.0 points with FloorSAV.
All three category averages
improve.
| Task | QAs | Baseline | FloorSAV | Δ points | Partial GT map | Full GT map |
|---|---|---|---|---|---|---|
| Ego-to-Exo Direction | 200 | 56.0 | 65.0 | +9.0 | 76.5 | 72.0 |
| Ego-to-Exo Distance · AbsMRA | 200 | 34.3 | 34.6 | +0.3 | 47.9 | 45.9 |
| Exo-to-Ego Direction | 200 | 32.0 | 50.5 | +18.5 | 50.5 | 53.5 |
| Exo-to-Ego Distance · AbsMRA | 200 | 46.4 | 44.7 | -1.7 | 57.5 | 55.0 |
| Dynamic Relativity Overall | 800 | 42.2 | 48.7 | +6.5 | 58.1 | 56.6 |
| Location Awareness | 200 | 92.0 | 94.5 | +2.5 | 96.0 | 96.5 |
| Visit History | 392 | 86.7 | 92.3 | +5.6 | 89.8 | 97.2 |
| Regional Overall | 592 | 89.4 | 93.4 | +4.0 | 92.9 | 96.8 |
| Line-Path Search (static) | 196 | 60.2 | 77.0 | +16.8 | 78.1 | 77.6 |
| Line-Path Search (dynamic) | 200 | 73.0 | 78.0 | +5.0 | 87.5 | 91.0 |
| Trajectory Distance · AbsMRA | 200 | 16.9 | 25.0 | +8.1 | 25.0 | 28.2 |
| Path Reasoning Overall | 596 | 50.0 | 60.0 | +10.0 | 63.5 | 65.6 |
| Overall · QA-weighted average | 1988 | 58.3 | 65.3 | +7.0 | 69.8 | 71.3 |
A clearer spatial reference helps across viewpoints, regions, and paths.
Accuracy / AbsMRA · 0–100, higher is better. Category averages weight tasks equally; the headline overall score weights QAs.
Download reported results ↓SAVED-Bench values match the released Table 1 results at the paper’s one-decimal precision. Category and overall averages are preserved as reported. Displayed gains are differences between these reported values.
Direction, region, and line-path tasks use accuracy; distance tasks use AbsMRA. All category averages improve; Exo-to-Ego Distance decreases from 46.4 to 44.7.
The radar averages direction and distance for each viewpoint pair, and static and dynamic tasks for Line-Path Search. It groups different benchmarks and models for the overview; it is not a single-model aggregate score.
FloorSAV improves performance on SAVED-Bench and SAVVY-Bench. It is not error-free: sound localization, long video contexts, and map reading remain challenging. Ground-truth map studies show further gains when spatial information is more accurate.
FloorSAV · Kim, Kim & Oh · 2026
@misc{kim2026floorsav,
title = {FloorSAV: Elucidating Spatial Audio-Visual
Context with 2D Floormap for AV-LLMs},
author = {Kim, Kyeong-Rae and Kim, Sungnyun and Oh, Tae-Hyun},
year = {2026},
url = {https://byulharang.github.io/FloorSAV/}
}