Lens Flare Removal and Reconstruction

Anonymous

Composite RenderSceneFlare
Decomposing a captured scene into flares and geometry. Rendered along a smooth novel-view trajectory for the hat scene: the composite render, the flare-free scene alone, and the isolated flare — a view-consistent 3D decomposition, not limited to the training cameras.

The presence of lens flares in images can significantly reduce the quality of downstream application results for tasks such as 3D scene reconstruction. This is because lens flares are a property of the camera imaging system, and not a part of the underlying scene being modeled. There exist previous methods that tackle removing small flares that are focused around a light source. However, such existing methods cannot handle large flares such as those that take up a full image. In this work, we compile a novel dataset for large flare removal based on both publicly available real-world data, and a procedural generation pipeline. We train a diffusion-based model using our dataset to be able to remove complex and large lens flares.

On the other hand, lens flares remain effective storytelling tools, widely used in the media. While there are ways to simulate 2D flares, representing and reconstructing lens flares consistently in 3D has not yet been explored. To achieve this, we introduce a 3D flare representation model that leverages the symmetric property of lens flares around the principal point of the camera. We propose a computational pipeline for the joint optimization of this flare model alongside a Gaussian splatting model (3DGS). This enables the decomposition of a 3D scene into lens flares and the 3D scene itself, utilizing our flare removal model. Our model enables a view consistent 3D lens flare representation. We perform extensive evaluations to show the effectiveness of both our flare removal and flare representation models.

What are Lens Flares?

Lens flares arise from unintended light paths through a lens when it observes a sufficiently bright light source. We distinguish two subjective categories: scattering flares — glare, shimmer, and streaks from dust and lens defects — and reflective flares, caused by inter-lens reflections in a lens ensemble.

Categories of lens flares
Streaks, glare (both scattering), and reflective flares
Same light, different lenses
Same light, different lenses
Different lights, same lens
Different lights, same lens

The exact shape of a flare depends jointly on the lens ensemble and the light source — existing stock, procedural, and physically-based simulators can synthesize plausible flares, but none reconstruct the actual pattern present in a given capture. Our goal instead is to reconstruct the flare patterns actually present in footage, preserving lens- and capture-specific characteristics, so they can be removed, reconstructed, and even transferred.

Removal model & training data

We fine-tune a single-step latent diffusion image-to-image model (following Difix3D+) with LoRA adapters — including the VAE encoder, which we found necessary since flare-corrupted images are out of distribution for a pretrained VAE. Supervision includes the light source back in the target image (rather than a post-hoc add-back step), which we show consistently improves results. Fine-tuning takes 20 GPU-hours on a single GPU, versus >4 days for retrained baselines.

Diffusion-based flare removal model pipeline
Our flare removal pipeline. A flare image is composited onto a clean image in linear space to synthesize a corrupted input, which is encoded, denoised in a single step, and decoded by a latent diffusion model fine-tuned end-to-end with LoRA — including the VAE encoder and decoder, not just the UNet. We train two variants against different supervision targets: with the light source's glare restored (Ours‑wGlare) and without (Ours).

Existing flare datasets are dominated by small scattering flares; full-frame reflective flares are essentially absent. We compile a new dataset from public VFX footage of large reflective flares (1,866 training images from 90 clips, and a 128-image test benchmark built from 18 held-out clips) and augment it with a procedural generator for ring-shaped reflective flares with radially diminishing opacity — a common real-world flare shape that off-the-shelf datasets miss entirely.

State-of-the-art on scattering & reflective flares

We compare against LightsOut, FR, Flare7K++, ACL-FR, and FlareX — retrained where possible on our combined dataset — on the established Flare7K++ benchmark (scattering flares) and our new VFX benchmark (large reflective flares). Ours‑wGlare / Ours denote our model trained with and without light-source glare in the supervision target.

BenchmarkMetricLightsOutFRFlare7K++ACL-FRFlareXOurs‑wGlareOurs
Flare7K++PSNR ↑16.0424.4926.4224.4825.2626.6825.76
SSIM ↑0.6950.8800.8930.8690.8870.8990.893
LPIPS ↓0.2690.0980.0910.0930.0930.0780.085
VFX (ours)PSNR ↑13.0421.0421.3720.2024.1725.9327.06
SSIM ↑0.3360.9270.9190.8860.9380.9370.941
LPIPS ↓0.4890.0970.0850.0920.0700.0670.066

best best in gold, 2nd second-best in blue. Full ablation columns (w/o light supervision, w/o encoder fine-tuning, w/o VFX data) are in the paper.

InputLightsOutFRACL-FRFlare7K++FlareXOurs‑wGlareOursGround Truth
Qualitative flare removal comparison, Flare7K++ benchmark, example 1
19.5721.1322.6225.3421.9630.0827.24
Qualitative flare removal comparison, Flare7K++ benchmark, example 2
9.9829.4031.5833.5429.7738.1238.29
Flare7K++ benchmark. Per-column PSNR shown below each row (best per row in gold). Our models consistently remove residual flare and recover the light source more faithfully.
InputLightsOutFRACL-FRFlare7K++FlareXOurs‑wGlareOursGround Truth
Qualitative flare removal comparison, VFX benchmark, example 1
15.3818.7718.5815.7020.1120.4125.15
Qualitative flare removal comparison, VFX benchmark, example 2
13.9318.7817.8316.2722.0822.7027.11
VFX benchmark (large reflective flares). Same column order as above. Baselines mistakenly bake large flare corruption into the sky and scene; our models remove it and reconstruct plausible detail underneath.

Consistent removal across views improves 3DGS

We train 3DGS on flare-corrupted captures, feeding it images pre-processed by each removal baseline, and evaluate on held-out test views that are genuinely flare-free (not model-processed). Inconsistent per-view removal by baselines shows up directly as reconstruction artifacts; our models yield the cleanest geometry.

MethodPSNR ↑SSIM ↑LPIPS ↓
Vanilla 3DGS18.680.7410.318
+ BilateralGrid18.960.7870.268
+ Flare7K++22.400.8150.252
+ FR22.650.8090.262
+ ACL-FR21.110.7890.279
+ FlareX21.640.7990.262
+ Ours‑wGlare23.240.8340.226
+ Ours23.930.8370.220

Average over 5 captured multi-view scenes (amp, billboard, dog, guitar, plant), evaluated on flare-free held-out test views.

3D reconstruction quality with different flare-removal pre-processing
Qualitative 3DGS reconstructions trained on inputs pre-processed by BilateralGrid, Flare7K++, FR, ACL-FR, FlareX, and our two models, vs. ground truth (right). Our models leave far fewer flare-induced artifacts baked into the reconstructed geometry.

Why 3DGS isn't built to handle lens flares

3DGS assumes a static scene observed from multiple viewpoints. Lens flares violate this: they move with the relative motion between camera and light source, so they are fundamentally 3D-inconsistent. Standard and even deformable 3DGS variants end up baking flares into the scene geometry or ignoring them outright.

3DGS and deformable 3DGS fail to reconstruct lens flares
Reconstructions of a scene with lens flares using 3DGS and Deformable 3DGS. Because flares aren't 3D-consistent, both bake them into geometry or lose them (circled).

Our camera-anchored, line-constrained flare representation

Camera-anchored plane, not world-space geometry 1D Gaussians on the principal-point–light line Hash-grid deformation MLP per parameter Jointly rasterized with 3DGS — no renderer changes

We model each flare as a set of 1D canonical Gaussians constrained to lie on the line through the camera's principal point and the projected light-source location — a well-known geometric property of lens flares. A hash-grid MLP, conditioned on camera and light position, deforms these canonical Gaussians (radial position, scale, opacity, color) per view, plus a small learned 2D offset for off-axis asymmetries. The result is unprojected onto a plane just past the near plane and rasterized together with the scene Gaussians using an unmodified 3DGS rasterizer — so the flare is rigidly tied to the camera, not to scene geometry, and adds negligible rendering overhead.

Flare representation and reconstruction pipeline Geometric construction of a single flare component
Left: our lens-flare representation alongside a 3DGS scene model — flare Gaussians are deformed and placed along the line connecting the principal point to the light, then rendered together with the scene and supervised with the input images; the scene alone is supervised with our flare-removal model's output. Right: geometric construction of a single flare component — a canonical 1D Gaussian, parameterized by radial offset μ along the principal-point–light line, is deformed per view, mapped to a 2D point on that line, then unprojected to a depth just past the near plane to obtain a standard 3DGS primitive in the camera frame.

A learned light-position correction (δl) absorbs noise from the upstream light detector (Grounded-SAM2 + 3D clustering), so reconstruction quality is largely decoupled from detection quality — see Robustness below. For scenes with several light sources, we instantiate one flare model per source and merge them, including sources that are off-screen or occluded, as long as their projection has positive depth.

Moving the light along the principal-point line

Because our flare model is parameterized on camera and 2D light positions, we can directly control a flare's on-screen position by moving the light-position input, with no retraining. The video below, from the hat sequence, demonstrates our model's performance with the 2D light position changing. To the left we show a control signal where each spoke is a planned motion of the light source, and the red dot marks the light position driving the current render on the right.

Light position swept around the principal point, camera fixed.

In the second video, we additionally change the camera input location to one of the training views, demonstrating how the flare updates consistently as the camera moves too.

Light position and camera moved together. The bar under the spoke diagram is the second control signal, marking progress through the camera trajectory.

Note that our model learns to generalize to these locations just based on 400 training images. This allows us to use the flare model to transfer the flare to another image or scene.

Decomposition videos — all 9 scenes

Each video sweeps all captured cameras for that scene, showing Given, Composite Render, Scene, and Flare, left to right.

Novel-view videos — same 9 scenes

Same decomposition, but along each scene's smooth novel-view trajectory instead of a sweep across all captured cameras. There's no Given column here since novel views were never actually captured.

Composite RenderSceneFlare
chair
Composite RenderSceneFlare
decos
Composite RenderSceneFlare
intrunk
Composite RenderSceneFlare
metro
Composite RenderSceneFlare
outtree
Composite RenderSceneFlare
hat

What each design choice buys us

We ablate the flare representation along three axes: the 1D line-constrained parameterization (vs. an unconstrained 2D variant), the scene-supervision strategy (subtracting the reconstructed flare from the input, or no decomposition supervision at all), and the light-position handling (removing the learned correction δl, removing detected light positions and G-SAM2 entirely, or removing the finer off-axis deformation δf).

3DGS2D‑unconstr.In‑FlareUnsup.w/o δlw/o G‑SAM2w/o δfOursGT
Ablation comparison on the workshop scene
Workshop scene. Top to bottom: composite render, isolated flare component (blank where not applicable, e.g. 3DGS/GT), and scene-only component. The unconstrained 2D variant and the unsupervised/in-flare-supervised scenes visibly misplace or leak flare into the scene; removing δl or δf degrades flare shape more subtly. Our full model gives the cleanest decomposition.
Variant Flare (Decomposed) Scene (Decomposed) Scene + Flare
PSNR ↑SSIM ↑LPIPS ↓ PSNR ↑SSIM ↑LPIPS ↓ PSNR ↑SSIM ↑LPIPS ↓
Vanilla 3DGS32.370.9360.169
2D‑unconstr.27.770.7980.30031.780.9210.18327.760.9170.190
In‑Flare18.350.5370.44618.010.6970.28531.160.9230.206
Unsup.21.790.5420.38520.970.8130.26233.670.9390.166
w/o δl33.190.8720.27231.810.9230.18233.060.9380.170
w/o G‑SAM232.190.8590.27931.890.9240.18032.180.9360.174
w/o δf32.990.8720.26331.760.9220.18332.900.9370.170
Ours33.270.8740.27031.790.9220.18233.060.9370.168

Averaged over all 9 reconstruction scenes. best best in gold, 2nd second-best in blue. Vanilla 3DGS has no scene/flare decomposition, so it's only scored on the composite.

Robustness to noisy light-source detections

In-the-wild light detection (Grounded-SAM2) can be off by tens to hundreds of pixels under saturation, motion blur, or occlusion. We stress-test by injecting Gaussian noise into the projected light position, up to an average deviation of ∼262 px, across every combination of our light-position correction δl and finer deformations δf. Our full model degrades gracefully: composite PSNR falls only ∼1.5 dB (33.82→32.29 dB) and flare PSNR ∼1.4 dB (33.48→32.12 dB) across the full noise range. Removing either correction alone costs little — at the highest noise level composite PSNR still reaches 32.04 dB without δl and 31.86 dB without δf — but removing both breaks robustness, with composite PSNR falling to 28.24 dB and flare PSNR to 27.27 dB. Combined with the w/o G‑SAM2 ablation above (light positions from a noisy Colmap centroid instead of detection), this shows reconstruction quality is largely decoupled from the upstream light detector.

PSNR vs. injected light-position noise, across all combinations of light-position and finer-deformation corrections
Composite (solid) and flare (dotted) PSNR vs. injected light-position noise, across all four combinations of δl and δf. With at least one correction enabled the model degrades gracefully; with both removed (red), PSNR falls off much more steeply.

Flare transfer to new images and new scenes

Because flares are represented explicitly and independently of scene geometry, a flare learned from one capture can be transferred elsewhere. For a bare 2D target image with no known camera or geometry, we query the flare MLP with the source camera/light as during training and anchor its placement to a single point in the target image. For a full 3D target scene, we instead re-anchor the transferred flare to the target camera and light position, and run a per-light visibility test against the target scene's depth so the flare is correctly occluded by target geometry.

GivenComposite RenderSceneFlare
A new capture for one of two new outdoor sequences. We captured two new outdoor sequences for this application. This one serves as the source: we extract its flare model and use it to transfer the flare onto other scenes, shown here decomposed into the composite render, the scene alone, and the isolated flare.
Lens-flare transfer to a novel scene. We transfer flares reconstructed from three source scenes — two from our captured reconstruction dataset and one additional captured scene — onto a single target scene that does not appear in either of our captured datasets. One example plays at a time: its source flare and the target scene play together, then the transferred result. The second example additionally shows the transferred flare with its Gaussians uniformly scaled 2×, demonstrating simple post-hoc artistic control over the flare's apparent intensity and extent. The third example instead transfers entirely within our main reconstruction dataset, showing the technique doesn't depend on the additional captured sequences.
Lens flare transfer to images
Lens-flare transfer to images. Given multi-view images (left column) of scenes with lens flares, our model decomposes each scene into a flare-free scene and our 3DGS-based lens-flare representation. We use the reconstructed flares to transfer them onto a diverse set of target images (top row). Each following row composites the source flare reconstructed on the left onto every target image shown above it.

What doesn't work yet

  • Our flare Gaussians have peak opacity at the mean, decreasing radially outward — a single Gaussian cannot form the annular profile of ring-shaped flares. Our 2D removal model already handles these (via procedural training data), but their 3D representation is left for future work.
  • Light sources occluded by scene geometry within a captured view are currently handled only implicitly by the flare model, rather than through an explicit occlusion-aware mechanism.
GivenComposite RenderSceneFlare
Even without explicit occlusion handling, the model still responds correctly as the light source becomes occluded across this capture.
  • The removal model is bounded by the flare types present in our training distribution; flares far outside it are harder to remove convincingly.