DensiTok

DensiTok Making Feed-Forward 3D Gaussian Splatting
See More Views Than It Is Given

  • 1Yonsei University
  • 2NAVER AI Lab
  • 3Korea Institute of Science and Technology (KIST)

DensiTok in three minutes

TL;DR

  • DensiTok is a plug-in for pretrained feed-forward 3DGS models. It generates the geometry tokens of views that were never captured, so a frozen backbone reconstructs as if it had been given a dense capture.
  • All token levels are compressed into a 4-channel latent, and the unobserved views are completed in a single flow-matching step. No image is synthesized and the encoder is not run again.
  • On AnySplat, WorldMirror and Depth Anything 3, rendering quality improves at every sparse-view setting. With 3 input views, AnySplat with DensiTok renders better than AnySplat with 16, and faster.

What the model lacks is not pixels, but tokens

Overview figure. Three sparse input images pass through a geometry encoder. Backbone only: the reconstruction heads receive tokens for the three observed views and render a blurry novel view and depth map. Backbone plus DensiTok: DensiTok fills in tokens for the views in between, and the same heads render a sharp novel view and depth map. At the top right, a rendering labelled Generated Gaussians shows the reconstructed scene; below it, a plot shows PSNR on RE10K against the number of input views for AnySplat and AnySplat plus DensiTok, with a dashed line for AnySplat with 16 views.
Overview. DensiTok sits between the frozen geometry encoder and the frozen reconstruction heads, and adds geometry tokens for views that were never observed.
  1. 1

    Sparse views leave gaps in the tokens

    A geometry encoder turns the input images into per-view geometry tokens, and camera, depth and Gaussian heads decode them. From a few unposed views, the tokens cover only the surfaces that were observed. The heads turn everything else into holes, floaters and blur.

  2. 2

    Pixels are a detour

    The common remedy synthesizes extra RGB views with an image or video generator and encodes them again. That is costly, and the generated views are not 3D-consistent by construction. The heads only need the tokens those views would have produced.

  3. 3

    DensiTok densifies the tokens directly

    Placed between the frozen encoder and the heads, it adds token sets for virtual viewpoints along the camera path, in exactly the form the heads already consume. The same design plugs into different backbones.

Method

Compress the geometry tokens into a compact latent, complete the latents of the unobserved views, and decode them back into tokens for the original heads. The geometry encoder stays frozen throughout; the heads stay frozen or are minimally adapted.

Two-stage training diagram. Stage 1, VAE with Level-Adaptive Modulation: N-view images pass through the frozen geometry encoder; level-adaptive fusion and a VAE encoder, with a level-wise residual, compress the geometry tokens into latents, and a VAE decoder reconstructs tokens that the frozen heads decode into cameras, depths and Gaussians. Stage 2, DiT for geometry latent augmentation: n-view images pass through the frozen encoder and the frozen LAM-VAE encoder; the observed latents plus noise enter a camera-conditioned DiT that outputs completed latents for all N views, which the frozen VAE decoder and heads turn into cameras, depths and Gaussians.
Model. Stage 1 trains LAM-VAE to compress the encoder's multi-level geometry tokens into compact latents. Stage 2 freezes LAM-VAE and trains a camera-conditioned DiT that completes the latents of unobserved views.
  1. Stage 1

    LAM-VAE: one latent for all levels

    The heads read tokens from four encoder layers, each more than a thousand channels wide. Level-adaptive fusion lets every token decide how much each level contributes, and a VAE compresses the result into a 4-channel latent. A level-wise residual into the posterior mean keeps the detail that belongs to a single level.

  2. Stage 2

    DiT: complete the unobserved views

    A camera-conditioned DiT lays out all 16 views along the camera path. Observed latents stay clean and act as the condition. Unobserved views start from noise and are generated with rectified flow matching, guided by Plücker-ray embeddings of their cameras.

  3. Stage 3

    One-step sampling

    The DiT is then finetuned so that a single Euler step maps noise to the completed latents. The VAE decoder expands them back into tokens at every level, and the pretrained heads decode cameras, depth and Gaussians.

No input poses. The backbone's own camera head predicts the cameras of the observed views, and the cameras of the new views are interpolated along the path through them. DensiTok therefore densifies the captured trajectory; it does not extrapolate beyond it.

Results

DensiTok is trained separately for AnySplat, WorldMirror and Depth Anything 3, and evaluated on RealEstate10K and DL3DV with 2, 3, 5 and 7 unposed input views.

Quantitative results

DensiTok improves PSNR, SSIM and LPIPS for every backbone, dataset and number of input views. Depth improves in every setting as well, and camera pose in most, which suggests that completing the shared tokens benefits geometry as much as appearance. With three inputs, AnySplat with DensiTok surpasses 16-view AnySplat in all three appearance metrics on both datasets, in 0.27 s instead of 0.43 s.

Backbone
Dataset
Views Method Appearance Camera Depth Time (s) ↓
PSNR ↑ SSIM ↑ LPIPS ↓ AUC@3° ↑ AUC@30° ↑ δ1.25 ↑ AbsRel ↓

Each pair is a backbone with and without DensiTok on the same input views; the first row is the dense-view reference, the backbone given all 16 views. Bold marks the better value of a pair. AnySplat and WorldMirror are from Table 1 of the paper, Depth Anything 3 from Table 6. Depth is scored against MoGe-2 predictions, and runtime is reported once per setting rather than per dataset.

Comparison with backbones

Drag the divider to compare the backbone alone (left) with the same backbone plus DensiTok (right). Both use the input views shown underneath.

Backbone
Input views

Comparison with sparse-view methods

Feed-forward 3DGS methods grouped by the input camera parameters they need. DensiTok runs on the AnySplat backbone and needs none.

Scene
Input views
Inputs

Comparison with diffusion-based novel-view synthesis methods

All methods receive the same input views. Yellow boxes mark the enlarged regions; hover over an image to move them.

Scene
Input views
Inputs

Runtime per scene

seconds, 512 × 512

Explore the Gaussians

The reconstructed Gaussian scenes themselves, from the same sparse input views, with AnySplat alone and with DensiTok added. The two views share one camera: move either one and the other follows.

Scene
AnySplat + DensiTok

Drag to orbit, scroll to zoom, right-drag to pan; on a touch screen, pinch to zoom and drag with two fingers to pan. The shared camera starts at the first input camera.

Abstract

Feed-forward 3D Gaussian Splatting (3DGS) reconstructs a scene in a single forward pass, replacing per-scene optimization with a network trained across many scenes. Its quality, however, degrades sharply as the number of input images drops. The bottleneck is upstream of the reconstruction heads: from a few unposed views, the internal representation they read carries no evidence for unobserved regions, leaving holes, floaters, and blur. The common remedy supplies that evidence as pixels, synthesizing extra views with an image or video generator and re-encoding them, which is costly and not 3D-consistent by construction. We instead densify the evidence itself.

We present DensiTok, a plug-in module for pretrained feed-forward 3DGS models that densifies their internal geometry tokens directly, making a frozen backbone behave as though it had observed many more views than it was given. DensiTok compresses those tokens into a compact latent space, completes the latents of the unobserved viewpoints in a single flow-matching step conditioned on camera geometry, and decodes them back into tokens that the original reconstruction heads consume. The same module design can be integrated into different pretrained predictors while keeping each backbone and its reconstruction heads frozen. Completion in a low-dimensional latent space requires no image synthesis or additional encoder passes. Across three pretrained backbones and two benchmarks, DensiTok consistently improves sparse-view reconstruction and recovers much of the gap to dense-view reconstruction.

BibTeX

@article{lee2026densitok,
  title   = {{DensiTok}: Making Feed-Forward {3D} {Gaussian}
             Splatting See More Views Than It Is Given},
  author  = {Lee, Minhyeok and Lee, Jungho and Kang, Minseok and
             Choi, Heeseung and Kim, Ig-Jae and Lee, Sangyoun},
  journal = {arXiv preprint arXiv:2610.07958},
  year    = {2026}
}