The goal of this thesis is to present my research contributions towards handling partially visible input in different visual tasks, comprising 360° RGB-D panorama image completion, 3D scene decomposition, and 3D amodal reconstruction. This thesis consists of three main pieces of work, each of which presents a novel learning-based approach for recovering complete visual representations from partially observed data. From different perspectives, each work demonstrates the novelty and superiority of the proposed method on utilizing occluded, inconsistent, and partially visible input for 3D generation. Chapter 2 provides a comprehensive literature review of 3D visual synthesis and generation, establishing the theoretical foundation for the subsequent contributions. This chapter systematically examines the evolution of 3D visual representations, followed by exploring three fundamental tasks in 3D computer vision: novel-view synthesis for generating photorealistic images from unseen viewpoints, 3D object generation, and 3D scene decomposition & generation. By analyzing the strengths and limitations of existing methodologies, we identify key research gaps and challenges that motivate the novel approaches presented in the following chapters. Chapter 3 presents a method for completing 360° panoramic scenes from limited narrow field-of-view (NFoV) input. We introduce a learning-based framework that treats panorama outpainting as a dual-modal latent diffusion process. During training, the model jointly encodes full resolution RGB and depth panoramas into a unified latent space and learns to denoise corrupted latents back into coherent RGB-D outputs. Two novel strategies - horizontal cyclic consistency through simulated camera rotations and explicit alignment of panorama boundaries during sampling - ensure that the generated panorama wraps seamlessly from left to right. At test time, given only a NFoV RGB crop, the system hallucinates both the missing color regions and the associated depth values, then decodes them into a full-resolution RGB-D panorama. Extensive experiments on public indoor datasets demonstrate that this approach produces semantically consistent, geometrically accurate 360° outputs, outperforming prior GAN-based and autoregressive diffusion baselines in both RGB fidelity and depth prediction. Chapter 4 focuses on decomposing complex indoor scenes into individual 3D object surfaces using only noisy multi-view 2D segmentation masks, without relying on any ground-truth 3D annotations. We present a neural implicit representation in which the network outputs multiple signed distance function (SDF) channels, each interpreted as the occupancy probability of a distinct object. A novel clustering-oriented loss combines three components: (1) a differentiation term that pushes different output channels apart, (2) a one-hot constraint that encourages each surface point to activate exactly a single channel, and (3) a regularization term to prevent trivial collapse. During training, rays sampled from multiple RGB-D viewpoints are rendered through the implicit network and matched against the corresponding 2D masks; the clustering loss then guides voxels from the same object to share one channel, while voxels from different objects occupy different channels. Experimental results on multiple benchmarks show that this weakly supervised pipeline yields high-quality per-object 3D decompositions, matching or exceeding the performance of fully supervised alternatives. Chapter 5 addresses 3D object reconstruction under occlusion. We describe a framework that integrates occlusion-aware attention mechanisms into a pretrained 3D generative diffusion backbone. Given one or more RGB views and their automatically extracted visibility and occlusion masks, the model first uses mask-weighted cross-attention to focus on observed regions during latent denoising, then employs a specialized occlusion attention layer to hallucinate the geometry of hidden surfaces. By performing diffusion directly in a 3D Gaussian implicit space, the system simultaneously recovers the geometry of visible parts and infers missing occluded regions. When multiple views are available, an adaptive fusion strategy orders features by apparent visibility to further enhance reconstruction fidelity. Comprehensive evaluations on multiple 3D datasets and real-scene occlusion benchmarks demonstrate that this end-to-end approach consistently outperforms two-stage baselines which first perform 2D amodal completion and then 3D reconstruction. Together, these three contributions form a cohesive research pipeline for handling partially visible inputs of different types: beginning with completing missing panorama content from NFoV images, moving to weakly supervised object-level decomposition of 3D scenes, and culminating in occlusion-aware single-stage 3D object generation. By integrating latent diffusion, neural implicit representations, and occlusion-aware attention, this thesis demonstrates how to robustly recover complete visual representations from limited or inconsistent observations, thereby advancing the state-of-the-art in real-world 3D vision tasks.
Read more