Prior-Driven Enhancements in 3D Gaussian Splatting:
Normals and Depths Regularization

ISPRS Geospatial Week 2025 · Oral

Gyeonggwan Lee Seunghwan Hong Junghun Suh

AI R&D Team, Kakao Mobility

Qualitative comparison of 3DGS and ours on Train, Horse, Parking lots and Street-view
TL;DR. 3D Gaussian Splatting starts from a sparse SfM point set, so reflective, low-texture and repetitive scenes leave it with misplaced Gaussians and artifacts. We regularize its optimization with surface normals and dense depth from a monocular network. The normals align each Gaussian's covariance with the local surface, and the depth refines per-pixel geometry. We validate the approach with three SfM pipelines, on benchmark scenes and on parking-lot and street-view data we captured. (Fig. 2 of the paper.)

Abstract

3D Gaussian Splatting (3DGS) is a state-of-the-art technique for 3D scene rendering, offering high efficiency and excellent visual quality. However, because 3DGS relies on an initial sparse point set from Structure-from-Motion (SfM) and view-dependent properties, it can suffer from geometric inaccuracies and visual artifacts, particularly in complex scenes. To address these challenges, we propose an improved 3DGS approach that regularizes the optimization process by integrating geometric priors, including surface normals and dense depth information. Surface normal regularization improves geometric consistency by aligning Gaussian covariance with local surface structures, while dense depth priors combined with an initial points from SfM enhance per-pixel depth estimation, increasing accuracy and reducing ambiguities. These enhancements enable robust handling of diverse and complex real-world scenarios, minimizing visual distortions and improving reconstruction quality across various environments. To validate our method, we evaluate it on challenging datasets, including street-view scenes and highly reflective environments, while testing it across multiple SfM pipelines. Our results demonstrate compatibility across diverse environments and highlight the robustness of our approach. Experimental findings further show that our method enhances geometric accuracy and visual quality, establishing a reliable solution for real-time 3D scene rendering in complex environments.

Method

The same images feed two branches. An SfM pipeline (COLMAP, SuperPoint + SuperGlue, or LoFTR) estimates the camera poses and the sparse points that initialize the Gaussians. A monocular network (Metric3D) predicts a surface normal and a depth for every pixel. Both priors enter the 3DGS optimization, which otherwise stays unchanged.

Images multi-view SfM COLMAP · SP-SG · LoFTR Monocular priors Metric3D Poses + sparse points Gaussian initialization Normals n · depth Ddense per pixel 3DGS optimization L_color + L_D-SSIM + geometry-aware init + L_normal + L_depth (D_guide from SfM scale) scale of Dsparse

Objective

Eq. (7). The first two terms are the original 3DGS losses, and the terms in red are ours.

Normal prior regularization

Geometry-aware initialization aligns the orientation and scale of each initial Gaussian with the predicted surface normals. The normal-consistency loss then aligns the orientation and scale of every Gaussian's covariance with the surface, flattening it along the normal (after Hwang et al., 2024):

Depth prior regularization

The predicted depth Ddense has an unknown scale. It is rescaled to match the sparse depth Dsparse, obtained by projecting the SfM points onto each image, which gives the guide Dguide. The depth D rendered by Gaussian splatting is supervised with an L1 loss:

Dense depth refines per-pixel geometry where the sparse SfM points leave it ambiguous, such as in repetitive patterns or on low-texture surfaces.

Interactive: the normal-consistency loss

A 2D slice of one Gaussian lying on a surface with normal n. Rotate it and change its shape. The two terms follow Eqs. (5)–(6) for this Gaussian. Both reach their minimum when the Gaussian lies flat on the surface, with its thinnest axis along the normal.

–Laxis
–Lscale

In 3D, the normal can only be orthogonal to two of the three axes. The minimum therefore puts the smallest scale on the axis that points along the normal, and the Gaussian becomes a thin disk on the surface.

Data

We use two Tanks and Temples scenes, Train and Horse, and two scenes we captured ourselves. Both of our scenes have repetitive patterns, strong reflections or low-texture surfaces, which make feature matching and pose estimation hard.

Ladybug6 camera

Parking lots · Ladybug6

Underground parking lots, captured with the three front and side cameras of a Ladybug6. 2992×4096 images, 85.9° FOV, 1 FPS, undistorted. Low texture, reflections and varying illumination.

Mobile mapping system on a vehicle

Street-view · Mobile Mapping System

A six-camera MMS on a moving vehicle, captured at fixed distance intervals. Each equirectangular image gives seven 90° views at 45° steps (the rear view is excluded), each 1080×1080. Dynamic objects, road markings and similar facades.

Results

DatasetMethodCOLMAPSP-SGLoFTR
MRE↓PSNR↑SSIM↑LPIPS↓MRE↓PSNR↑SSIM↑LPIPS↓MRE↓PSNR↑SSIM↑LPIPS↓
Train3DGS0.7521.100.8020.2181.4021.100.7500.2820.7920.970.7670.274
Ours21.970.7990.25221.320.7490.28721.110.7590.291
Horse3DGS0.7124.180.8890.2391.3121.010.8020.2390.8023.390.8700.174
Ours25.500.9030.15321.120.8010.24624.800.8810.165
Parking lots3DGSinvalid–––1.3728.080.8280.4240.6429.330.8420.410
Ours–––28.370.8290.42729.070.8400.414
Street-view3DGSinvalid–––1.0918.330.4680.4100.5622.280.7290.331
Ours–––19.000.4790.49222.490.7220.340

Table 1 of the paper. MRE is the mean reprojection error of the SfM point cloud. COLMAP fails to reconstruct the Parking lots and Street-view scenes.

PSNR gain over 3DGS (dB)

Interactive 3D Gaussians

drag to orbit · right-drag to pan · scroll to zoom · both views move together
3DGS
Ours

Trained Tanks and Temples models (LoFTR SfM), rendered live in the browser. For the web, each model keeps its 800k most significant Gaussians (about 25 MB per view). Full models are on Hugging Face. The views start from a held-out test camera.

BibTeX

@article{lee2025prior,
  title   = {Prior-Driven Enhancements in {3D} {Gaussian} Splatting:
             Normals and Depths Regularization},
  author  = {Lee, Gyeonggwan and Hong, Seunghwan and Suh, Junghun},
  journal = {The International Archives of the Photogrammetry,
             Remote Sensing and Spatial Information Sciences},
  volume  = {XLVIII-G-2025},
  pages   = {891--897},
  year    = {2025},
  doi     = {10.5194/isprs-archives-XLVIII-G-2025-891-2025}
}