SILSA: Topology-Preserving High-Resolution 3D Generation with Sliding-Window Slice Latents
SILSA:利用滑動視窗切片潛在特徵實現保持拓撲結構的高解析度 3D 生成
High-resolution 3D generation often struggles with fragmented surfaces and high computational costs caused by voxel representation. SILSA bypasses voxel tokens by utilizing a fixed set of overlapping sliding-window slices across three canonical axes. Integrating a Slice VAE, a Volumetric Anchor Lattice, and slice-level topology supervision (aligning Betti transitions), SILSA enables efficient single-stage rectified-flow 3D generation with outstanding structural fidelity and minimal resource overhead.
Key points
Sliding-Window Slice Latents
Uses overlapping sliding-window slices along three axes to project local depth into 2D latents, preserving cross-sectional continuity.
Sparse Decoder & Shared Workspace
Employs a Slice VAE and a Volumetric Anchor Lattice to coordinate multi-axis slice streams within a shared 3D workspace.
Slice-Level Topology Supervision
Integrates persistence diagram matching and Betti transition alignment to prevent thin or complex connections from breaking.
Unmatched Token Efficiency
Requires 70% fewer tokens than the next-most compact baseline, cutting training memory by 40.4% and inference time by 58.5%.
How it works
Why it matters
Generating topologically accurate thin structures (such as wire fences or chair legs) has been a bottleneck in 3D asset creation due to high voxel computational requirements. SILSA demonstrates that multi-axis 2D slice projection, paired with topology guidance, can reconstruct high-fidelity 3D shapes at a fraction of the cost. This opens up highly efficient, single-stage 3D generation on consumer-grade hardware for gaming, VR, and simulation.
Who it affects
- AI Researcher
- AI Developer
- Content Creator
How to use it
- 1High-resolution 3D asset generation for gaming, especially for complex shapes containing thin, repetitive, or porous elements.
- 2Creating geometric digital twins and robot simulation environments with minimal memory and compute overhead.
Limitations & caveats
- The model relies on predefined orthogonal slicing axes, which may limit performance on highly asymmetric, tilted, or non-manifold geometric shapes.
- Calculating persistent homology for slice-level topology supervision is computationally demanding and may cause CPU bottlenecks during massive-scale training.
Related

Ai2 Open-Sources AstaBrief: An 8B Scientific Report Generator 3.5x Faster than Claude
艾倫人工智慧研究所開源 AstaBrief:比 Claude 快 3.5 倍的 8B 科學報告生成模型
Allen Institute for AI (Ai2) has open-sourced AstaBrief 8B, a specialized model for scientific report generation that achieves a 3.5x speedup over proprietary pipelines while maintaining high citation accuracy.
GALA: Distilling 3D Gaussian Avatars into Linear Blendshapes for Real-Time Animation
GALA:用線性混合變形蒸餾技術實現 3D Gaussian 虛擬化身即時動畫
GALA distills complex neural decoding of 3D Gaussian avatars into lightweight linear blendshapes, reducing CPU animation costs by up to 1000x and enabling 60fps real-time performance on mobile devices.
ScholarCatalyst: A Benchmark for Testing AI's Intuition in Retrieving Inspiring Research Papers
ScholarCatalyst:評估 AI 是否擁有「科學家直覺」的學術文獻檢索基準
ScholarCatalyst is a novel benchmark featuring annotations from 184 lead authors to evaluate whether AI can retrieve key inspiring papers from past literature based only on an initial research question.