GenIA: Generative Reconstruction
with Test-Time Input Alignment

Stefano Esposito1 Naama Pearl1 Polina Karpikova1 Samuel Rota Bulò2
Lorenzo Porzi2 Peter Kontschieder2 Andreas Geiger1 Jonathon Luiten2

1Tübingen AI Center, University of Tübingen, Germany 2Meta Reality Labs

Click on images to open at full resolution.

GenIA reconstructs complete, detailed, and input-aligned 3D objects from monocular images, multi-view images, and monocular videos. Our inference-time framework aligns the frozen SAM3D model to the inputs with geometric and photometric constraints, extending it beyond its single-view static training setting to multi-view and dynamic reconstruction. Rows: ours, a baseline (badged), and the input observations. Hover a tile to magnify the same point in all three.


Abstract

Reconstructing complete 3D object assets from monocular or sparse multi-view observations remains challenging. Generative 3D foundation models can complete object geometry beyond the observed views, but their predictions may not faithfully reproduce the observed geometry, appearance, or pose. We introduce GenIA, a framework for test-time input-aligned generation that grounds SAM3D's generative prior in geometric and photometric observations without retraining the foundation model.

We improve object pose by deriving translation and scale from geometry while retaining the learned rotation prior, and align appearance through visibility-biased attention, cross-observation fusion, and differentiable rendering guidance during denoising. An optional post-denoising refinement further adapts the appearance latent, lightweight decoder adapters, and object placement to the observations. Our framework also supports externally supplied geometry; when given temporal shapes of dynamic objects, it recovers a shared, input-aligned canonical appearance and stable world-space placement. Across synthetic and real benchmarks, GenIA improves pose prediction and object reconstruction from monocular, multi-view, and dynamic inputs, outperforming recent optimization-based, per-frame image-to-3D, and video-to-4D methods.


3D Gallery

Pick a scene by its input views, then load its interactive 3D reconstruction: the input cameras and the reconstructed objects. Drag to orbit, scroll to zoom, right-drag to pan, R to reset the view, and E to cycle between splats, ellipsoids and points.


Method

An inference-time framework: no retraining. All four stages run per scene on the frozen SAM3D model. Test-time refinement optimizes the appearance latent, a low-rank decoder adapter, and the object placement.

We (1) obtain the coarse shape from SAM3D’s first pass or (optionally) inject an external one into SAM3D’s shape tokens, (2) denoise each object’s per-frame pose over that fixed shape, (3) predict observation-aligned appearance as SLAT features, and (4) do test-time refinement of reconstruction against the input views. Dotted arrows denote gradient flow. We illustrate the single-view setting; multi-view and dynamic extensions are described in the paper.


Completion of unobserved regions

Drag the render to rotate the object; the method buttons switch reconstructions while keeping the viewing angle.

A real object reconstructed from a single photograph. SAM3D's completion looks smooth and synthetic; TripoSplat and CUPID match the observed side but wash out or drift on the unobserved surface. GenIA with TTR completes it with colors consistent with the input.


Qualitative comparisons

For each scene we compare: ours without test-time refinement, ours with it, and a baseline. The supplementary's static all-methods-at-once figures are on their own page: additional qualitative results.

Views picks the input-view count, Baseline which method to compare against. The arrows change scenes.

Quantitative results

PSNR is foreground-masked on training views; LPIPS and CLIP-I are held-out, and co-PSNR restricts held-out pixels to those visible from an input view. For fairness, every method's pose is refined against the inputs with shape and appearance fixed; per-scene optimization methods are exempt.


Runtime vs. quality

GenIA adds minimal overhead over SAM3D. Median runtime (log axis, as a multiple of SAM3D's 25.3 s on an A100) versus mean held-out CLIP-I; the static panels score on CO3D, the dynamic panel on DAVIS with ActionMesh included in our runtime. TRELLISv2 and Pixal3D appear as dashed rules: cost only, no asserted quality. On dynamic sequences GenIA costs 13× SAM3D's single-image runtime including ActionMesh (12× without TTR), versus 160× for Lift4D and 250× for HiMoR, and with TTR achieves the highest CLIP-I.


Ablations

Each row is cumulative, as described in the paper. The supplementary's full appearance ladders, on every benchmark, are on their own page: appearance ablations.

Our appearance prediction better aligns with the input observations. One GSO object from a single view: attention bias injects pose into SAM3D's pose-blind appearance predictor, rendering guidance improves color and detail matching, and optional TTR adds further gains, including on the held-out view. Rows: one training and one held-out view.

Hover a tile to magnify the same point across stacked contributions.


Limitations and failure cases

Dynamic geometry currently relies on an external predictor (ActionMesh) rather than being directly grounded in the observations. Rotation stays prior-driven, corrected only by noise-sensitive ICP registration. Both failure modes surface on dynamic scenes below: incorrect world-space placement from registration, non-rigid deformation errors inherited from ActionMesh, or a combination of the two, can strongly impact the final reconstruction quality.


BibTeX

@article{esposito2026genia,
  title   = {GenIA: Generative Reconstruction with Test-Time Input Alignment},
  author  = {Esposito, Stefano and Pearl, Naama and Karpikova, Polina and
             Rota Bul{\`o}, Samuel and Porzi, Lorenzo and Kontschieder, Peter and
             Geiger, Andreas and Luiten, Jonathon},
  journal = {arXiv preprint arXiv:2610.12388},
  year    = {2026}
}