UniFusion: Sparse-View 4D Reconstruction via Unified Spatio-temporal Depth Alignment

* Equal contribution   Corresponding authors

Video

UniFusion jointly aligns monocular depths across sparse views and time, providing reliable initialization and supervision for 4D Gaussian splatting.

Abstract

We address sparse-view 4D reconstruction, where limited cross-view overlap and temporal variation make monocular depth predictions inconsistent across both views and time. Existing methods align these dimensions in separate stages and depend on foreground masks and tracking models.

UniFusion introduces a unified spatio-temporal depth alignment framework that represents all depth maps as spatio-temporal neural fields. It jointly resolves cross-view and cross-time inconsistencies without distinguishing foreground from background. A multi-view depth-order loss and a scale-and-shift-invariant loss further improve depth quality.

The aligned depths initialize and supervise Gaussian splatting models. Experiments on Ego-Exo4D and EgoHuman show consistent improvements in novel-view / time synthesis and geometry, including PSNR gains of 4–6 dB and absolute relative depth-error reductions of 14–40%.

Method

All temporal frames from one view share the same spatial representation, while compact residual weights condition the alignment network on time-dependent variation. The aligned depths then provide geometric initialization and supervision for 4D Gaussian splatting.

Overview of unified spatio-temporal depth alignment and 4D Gaussian splatting reconstruction.

Shared spatial chart

Per-view encodings capture geometry shared across all frames.

Residual temporal weights

Low-rank residuals efficiently model time-dependent depth variation.

Geometry-aware losses

Multi-view ordering and scale-shift invariance refine aligned depth.

Results

UniFusion consistently improves two Gaussian-splatting backends across novel-time and novel-view synthesis. Browse the video and image comparisons below.

Video Comparisons

Static Comparisons

Drag the divider to compare each backend with its UniFusion-enhanced result. Use the arrows or swipe to browse scenes.

BibTeX

@article{lyu2026unifusion,
  author = {Lyu, Yongzhe and Wang, Shaofei and Chen, Yixin and Huang, Siyuan},
  title  = {UniFusion: Sparse-View 4D Reconstruction via Unified Spatio-temporal Depth Alignment},
  year   = {2026}
}