Less Decoder is More Encoder: Extracting Robust 3D Geometric Representations via Novel View Synthesis
減少解碼器反而增強編碼器:從新視角合成中提煉強大三維幾何表徵
Novel View Synthesis (NVS) should theoretically teach models 3D geometry, but current encoders learn poor representations. The authors identify two main culprits: spatially expressive decoders that offload spatial reasoning from the encoder, and low-level pixel reconstruction targets. They introduce SNAP, a self-supervised transformer that employs a constrained 'pose-conditioned local decoder' and a latent-space reconstruction target. By bottlenecking the decoder, SNAP forces the encoder to learn rich geometric representations, achieving outstanding performance across visual localization, depth estimation, and robot manipulation tasks.
Key points
The Expressive Decoder Dilemma
Highly expressive decoders take over spatial reasoning, inadvertently diluting the geometric representations learned by the encoder.
SNAP Bottleneck Paradigm
Restricts decoder capacity using a pose-conditioned local decoder, forcing the encoder to fully capture the 3D geometry.
Latent-Space Reconstruction
Replaces pixel-space targets with latent-space reconstruction to prevent representation learning from getting bogged down in pixel-level details.
Emergent Viewpoint Invariance
Exhibits superior resilience and viewpoint invariance under camera shifts that cause standard 2D representations to collapse.
How it works
Why it matters
This work counters the intuition that all network components should be highly expressive. It proves that a bottlenecked decoder forces the encoder to step up and learn better 3D representations. This paradigm shift offers a highly efficient, self-supervised pre-training route for applications in autonomous driving and robotics where geometric annotations are scarce.
Who it affects
- AI Researcher
- AI Developer
- Student & Learner
How to use it
- 1Multi-view robot manipulation and trajectory planning
- 2Unsupervised 3D reconstruction, depth, and pose estimation
- 3Visual localization systems under complex indoor/outdoor scenarios
Limitations & caveats
- Because the local decoder is intentionally constrained, SNAP is not suitable for generating high-fidelity, photorealistic novel view images.
- Training remains heavily reliant on multi-view or sequential frames to provide the necessary pose-conditioned supervisory signals.
Related

Ai2 Open-Sources AstaBrief: An 8B Scientific Report Generator 3.5x Faster than Claude
艾倫人工智慧研究所開源 AstaBrief:比 Claude 快 3.5 倍的 8B 科學報告生成模型
Allen Institute for AI (Ai2) has open-sourced AstaBrief 8B, a specialized model for scientific report generation that achieves a 3.5x speedup over proprietary pipelines while maintaining high citation accuracy.
GALA: Distilling 3D Gaussian Avatars into Linear Blendshapes for Real-Time Animation
GALA:用線性混合變形蒸餾技術實現 3D Gaussian 虛擬化身即時動畫
GALA distills complex neural decoding of 3D Gaussian avatars into lightweight linear blendshapes, reducing CPU animation costs by up to 1000x and enabling 60fps real-time performance on mobile devices.
ScholarCatalyst: A Benchmark for Testing AI's Intuition in Retrieving Inspiring Research Papers
ScholarCatalyst:評估 AI 是否擁有「科學家直覺」的學術文獻檢索基準
ScholarCatalyst is a novel benchmark featuring annotations from 184 lead authors to evaluate whether AI can retrieve key inspiring papers from past literature based only on an initial research question.