ISPRS Geospatial Week 2025 · Oral
AI R&D Team, Kakao Mobility
3D Gaussian Splatting (3DGS) is a state-of-the-art technique for 3D scene rendering, offering high efficiency and excellent visual quality. However, because 3DGS relies on an initial sparse point set from Structure-from-Motion (SfM) and view-dependent properties, it can suffer from geometric inaccuracies and visual artifacts, particularly in complex scenes. To address these challenges, we propose an improved 3DGS approach that regularizes the optimization process by integrating geometric priors, including surface normals and dense depth information. Surface normal regularization improves geometric consistency by aligning Gaussian covariance with local surface structures, while dense depth priors combined with an initial points from SfM enhance per-pixel depth estimation, increasing accuracy and reducing ambiguities. These enhancements enable robust handling of diverse and complex real-world scenarios, minimizing visual distortions and improving reconstruction quality across various environments. To validate our method, we evaluate it on challenging datasets, including street-view scenes and highly reflective environments, while testing it across multiple SfM pipelines. Our results demonstrate compatibility across diverse environments and highlight the robustness of our approach. Experimental findings further show that our method enhances geometric accuracy and visual quality, establishing a reliable solution for real-time 3D scene rendering in complex environments.
The same images feed two branches. An SfM pipeline (COLMAP, SuperPoint + SuperGlue, or LoFTR) estimates the camera poses and the sparse points that initialize the Gaussians. A monocular network (Metric3D) predicts a surface normal and a depth for every pixel. Both priors enter the 3DGS optimization, which otherwise stays unchanged.
Eq. (7). The first two terms are the original 3DGS losses, and the terms in red are ours.
Geometry-aware initialization aligns the orientation and scale of each initial Gaussian with the predicted surface normals. The normal-consistency loss then aligns the orientation and scale of every Gaussian's covariance with the surface, flattening it along the normal (after Hwang et al., 2024):
The predicted depth Ddense has an unknown scale. It is rescaled to match the sparse depth Dsparse, obtained by projecting the SfM points onto each image, which gives the guide Dguide. The depth D rendered by Gaussian splatting is supervised with an L1 loss:
Dense depth refines per-pixel geometry where the sparse SfM points leave it ambiguous, such as in repetitive patterns or on low-texture surfaces.
A 2D slice of one Gaussian lying on a surface with normal n. Rotate it and change its shape. The two terms follow Eqs. (5)–(6) for this Gaussian. Both reach their minimum when the Gaussian lies flat on the surface, with its thinnest axis along the normal.
In 3D, the normal can only be orthogonal to two of the three axes. The minimum therefore puts the smallest scale on the axis that points along the normal, and the Gaussian becomes a thin disk on the surface.
We use two Tanks and Temples scenes, Train and Horse, and two scenes we captured ourselves. Both of our scenes have repetitive patterns, strong reflections or low-texture surfaces, which make feature matching and pose estimation hard.
Underground parking lots, captured with the three front and side cameras of a Ladybug6. 2992×4096 images, 85.9° FOV, 1 FPS, undistorted. Low texture, reflections and varying illumination.
A six-camera MMS on a moving vehicle, captured at fixed distance intervals. Each equirectangular image gives seven 90° views at 45° steps (the rear view is excluded), each 1080×1080. Dynamic objects, road markings and similar facades.
| Dataset | Method | COLMAP | SP-SG | LoFTR | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| MRE↓ | PSNR↑ | SSIM↑ | LPIPS↓ | MRE↓ | PSNR↑ | SSIM↑ | LPIPS↓ | MRE↓ | PSNR↑ | SSIM↑ | LPIPS↓ | ||
| Train | 3DGS | 0.75 | 21.10 | 0.802 | 0.218 | 1.40 | 21.10 | 0.750 | 0.282 | 0.79 | 20.97 | 0.767 | 0.274 |
| Ours | 21.97 | 0.799 | 0.252 | 21.32 | 0.749 | 0.287 | 21.11 | 0.759 | 0.291 | ||||
| Horse | 3DGS | 0.71 | 24.18 | 0.889 | 0.239 | 1.31 | 21.01 | 0.802 | 0.239 | 0.80 | 23.39 | 0.870 | 0.174 |
| Ours | 25.50 | 0.903 | 0.153 | 21.12 | 0.801 | 0.246 | 24.80 | 0.881 | 0.165 | ||||
| Parking lots | 3DGS | invalid | – | – | – | 1.37 | 28.08 | 0.828 | 0.424 | 0.64 | 29.33 | 0.842 | 0.410 |
| Ours | – | – | – | 28.37 | 0.829 | 0.427 | 29.07 | 0.840 | 0.414 | ||||
| Street-view | 3DGS | invalid | – | – | – | 1.09 | 18.33 | 0.468 | 0.410 | 0.56 | 22.28 | 0.729 | 0.331 |
| Ours | – | – | – | 19.00 | 0.479 | 0.492 | 22.49 | 0.722 | 0.340 | ||||
Table 1 of the paper. MRE is the mean reprojection error of the SfM point cloud. COLMAP fails to reconstruct the Parking lots and Street-view scenes.
Trained Tanks and Temples models (LoFTR SfM), rendered live in the browser. For the web, each model keeps its 800k most significant Gaussians (about 25 MB per view). Full models are on Hugging Face. The views start from a held-out test camera.
Held-out test views of Tanks and Temples, rendered by both methods from the same LoFTR SfM initialization. The views were selected to show the difference clearly. Drag the divider, or use ← → to switch views.
@article{lee2025prior,
title = {Prior-Driven Enhancements in {3D} {Gaussian} Splatting:
Normals and Depths Regularization},
author = {Lee, Gyeonggwan and Hong, Seunghwan and Suh, Junghun},
journal = {The International Archives of the Photogrammetry,
Remote Sensing and Spatial Information Sciences},
volume = {XLVIII-G-2025},
pages = {891--897},
year = {2025},
doi = {10.5194/isprs-archives-XLVIII-G-2025-891-2025}
}