DensiTok in three minutes
TL;DR
- DensiTok is a plug-in for pretrained feed-forward 3DGS models. It generates the geometry tokens of views that were never captured, so a frozen backbone reconstructs as if it had been given a dense capture.
- All token levels are compressed into a 4-channel latent, and the unobserved views are completed in a single flow-matching step. No image is synthesized and the encoder is not run again.
- On AnySplat, WorldMirror and Depth Anything 3, rendering quality improves at every sparse-view setting. With 3 input views, AnySplat with DensiTok renders better than AnySplat with 16, and faster.
What the model lacks is not pixels, but tokens
-
1
Sparse views leave gaps in the tokens
A geometry encoder turns the input images into per-view geometry tokens, and camera, depth and Gaussian heads decode them. From a few unposed views, the tokens cover only the surfaces that were observed. The heads turn everything else into holes, floaters and blur.
-
2
Pixels are a detour
The common remedy synthesizes extra RGB views with an image or video generator and encodes them again. That is costly, and the generated views are not 3D-consistent by construction. The heads only need the tokens those views would have produced.
-
3
DensiTok densifies the tokens directly
Placed between the frozen encoder and the heads, it adds token sets for virtual viewpoints along the camera path, in exactly the form the heads already consume. The same design plugs into different backbones.
Method
Compress the geometry tokens into a compact latent, complete the latents of the unobserved views, and decode them back into tokens for the original heads. The geometry encoder stays frozen throughout; the heads stay frozen or are minimally adapted.
-
Stage 1
LAM-VAE: one latent for all levels
The heads read tokens from four encoder layers, each more than a thousand channels wide. Level-adaptive fusion lets every token decide how much each level contributes, and a VAE compresses the result into a 4-channel latent. A level-wise residual into the posterior mean keeps the detail that belongs to a single level.
-
Stage 2
DiT: complete the unobserved views
A camera-conditioned DiT lays out all 16 views along the camera path. Observed latents stay clean and act as the condition. Unobserved views start from noise and are generated with rectified flow matching, guided by Plücker-ray embeddings of their cameras.
-
Stage 3
One-step sampling
The DiT is then finetuned so that a single Euler step maps noise to the completed latents. The VAE decoder expands them back into tokens at every level, and the pretrained heads decode cameras, depth and Gaussians.
No input poses. The backbone's own camera head predicts the cameras of the observed views, and the cameras of the new views are interpolated along the path through them. DensiTok therefore densifies the captured trajectory; it does not extrapolate beyond it.
Results
DensiTok is trained separately for AnySplat, WorldMirror and Depth Anything 3, and evaluated on RealEstate10K and DL3DV with 2, 3, 5 and 7 unposed input views.
Quantitative results
DensiTok improves PSNR, SSIM and LPIPS for every backbone, dataset and number of input views. Depth improves in every setting as well, and camera pose in most, which suggests that completing the shared tokens benefits geometry as much as appearance. With three inputs, AnySplat with DensiTok surpasses 16-view AnySplat in all three appearance metrics on both datasets, in 0.27 s instead of 0.43 s.
| Views | Method | Appearance | Camera | Depth | Time (s) ↓ | ||||
|---|---|---|---|---|---|---|---|---|---|
| PSNR ↑ | SSIM ↑ | LPIPS ↓ | AUC@3° ↑ | AUC@30° ↑ | δ1.25 ↑ | AbsRel ↓ | |||
Each pair is a backbone with and without DensiTok on the same input views; the first row is the dense-view reference, the backbone given all 16 views. Bold marks the better value of a pair. AnySplat and WorldMirror are from Table 1 of the paper, Depth Anything 3 from Table 6. Depth is scored against MoGe-2 predictions, and runtime is reported once per setting rather than per dataset.
Comparison with backbones
Drag the divider to compare the backbone alone (left) with the same backbone plus DensiTok (right). Both use the input views shown underneath.
Comparison with sparse-view methods
Feed-forward 3DGS methods grouped by the input camera parameters they need. DensiTok runs on the AnySplat backbone and needs none.
Comparison with diffusion-based novel-view synthesis methods
All methods receive the same input views. Yellow boxes mark the enlarged regions; hover over an image to move them.
Runtime per scene
seconds, 512 × 512
Explore the Gaussians
The reconstructed Gaussian scenes themselves, from the same sparse input views, with AnySplat alone and with DensiTok added. The two views share one camera: move either one and the other follows.
Drag to orbit, scroll to zoom, right-drag to pan; on a touch screen, pinch to zoom and drag with two fingers to pan. The shared camera starts at the first input camera.
Abstract
Feed-forward 3D Gaussian Splatting (3DGS) reconstructs a scene in a single forward pass, replacing per-scene optimization with a network trained across many scenes. Its quality, however, degrades sharply as the number of input images drops. The bottleneck is upstream of the reconstruction heads: from a few unposed views, the internal representation they read carries no evidence for unobserved regions, leaving holes, floaters, and blur. The common remedy supplies that evidence as pixels, synthesizing extra views with an image or video generator and re-encoding them, which is costly and not 3D-consistent by construction. We instead densify the evidence itself.
We present DensiTok, a plug-in module for pretrained feed-forward 3DGS models that densifies their internal geometry tokens directly, making a frozen backbone behave as though it had observed many more views than it was given. DensiTok compresses those tokens into a compact latent space, completes the latents of the unobserved viewpoints in a single flow-matching step conditioned on camera geometry, and decodes them back into tokens that the original reconstruction heads consume. The same module design can be integrated into different pretrained predictors while keeping each backbone and its reconstruction heads frozen. Completion in a low-dimensional latent space requires no image synthesis or additional encoder passes. Across three pretrained backbones and two benchmarks, DensiTok consistently improves sparse-view reconstruction and recovers much of the gap to dense-view reconstruction.
BibTeX
@article{lee2026densitok,
title = {{DensiTok}: Making Feed-Forward {3D} {Gaussian}
Splatting See More Views Than It Is Given},
author = {Lee, Minhyeok and Lee, Jungho and Kang, Minseok and
Choi, Heeseung and Kim, Ig-Jae and Lee, Sangyoun},
journal = {arXiv preprint arXiv:2610.07958},
year = {2026}
}