- Conference Instance
12
- 10.1145/2072572
Proceedings of the 2011 joint ACM workshop on Human gesture and behavior understanding
- Dec 01, 2011
Proceedings of the 2011 joint ACM workshop on Human gesture and behavior understanding
The goal of this thesis is to present my research contributions towards handling partially visible input in different visual tasks, comprising 360° RGB-D panorama image completion, 3D scene decomposition, and 3D amodal reconstruction. This thesis consists of three main pieces of work, each of which presents a novel learning-based approach for recovering complete visual representations from partially observed data. From different perspectives, each work demonstrates the novelty and superiority of the proposed method on utilizing occluded, inconsistent, and partially visible input for 3D generation. Chapter 2 provides a comprehensive literature review of 3D visual synthesis and generation, establishing the theoretical foundation for the subsequent contributions. This chapter systematically examines the evolution of 3D visual representations, followed by exploring three fundamental tasks in 3D computer vision: novel-view synthesis for generating photorealistic images from unseen viewpoints, 3D object generation, and 3D scene decomposition & generation. By analyzing the strengths and limitations of existing methodologies, we identify key research gaps and challenges that motivate the novel approaches presented in the following chapters. Chapter 3 presents a method for completing 360° panoramic scenes from limited narrow field-of-view (NFoV) input. We introduce a learning-based framework that treats panorama outpainting as a dual-modal latent diffusion process. During training, the model jointly encodes full resolution RGB and depth panoramas into a unified latent space and learns to denoise corrupted latents back into coherent RGB-D outputs. Two novel strategies - horizontal cyclic consistency through simulated camera rotations and explicit alignment of panorama boundaries during sampling - ensure that the generated panorama wraps seamlessly from left to right. At test time, given only a NFoV RGB crop, the system hallucinates both the missing color regions and the associated depth values, then decodes them into a full-resolution RGB-D panorama. Extensive experiments on public indoor datasets demonstrate that this approach produces semantically consistent, geometrically accurate 360° outputs, outperforming prior GAN-based and autoregressive diffusion baselines in both RGB fidelity and depth prediction. Chapter 4 focuses on decomposing complex indoor scenes into individual 3D object surfaces using only noisy multi-view 2D segmentation masks, without relying on any ground-truth 3D annotations. We present a neural implicit representation in which the network outputs multiple signed distance function (SDF) channels, each interpreted as the occupancy probability of a distinct object. A novel clustering-oriented loss combines three components: (1) a differentiation term that pushes different output channels apart, (2) a one-hot constraint that encourages each surface point to activate exactly a single channel, and (3) a regularization term to prevent trivial collapse. During training, rays sampled from multiple RGB-D viewpoints are rendered through the implicit network and matched against the corresponding 2D masks; the clustering loss then guides voxels from the same object to share one channel, while voxels from different objects occupy different channels. Experimental results on multiple benchmarks show that this weakly supervised pipeline yields high-quality per-object 3D decompositions, matching or exceeding the performance of fully supervised alternatives. Chapter 5 addresses 3D object reconstruction under occlusion. We describe a framework that integrates occlusion-aware attention mechanisms into a pretrained 3D generative diffusion backbone. Given one or more RGB views and their automatically extracted visibility and occlusion masks, the model first uses mask-weighted cross-attention to focus on observed regions during latent denoising, then employs a specialized occlusion attention layer to hallucinate the geometry of hidden surfaces. By performing diffusion directly in a 3D Gaussian implicit space, the system simultaneously recovers the geometry of visible parts and infers missing occluded regions. When multiple views are available, an adaptive fusion strategy orders features by apparent visibility to further enhance reconstruction fidelity. Comprehensive evaluations on multiple 3D datasets and real-scene occlusion benchmarks demonstrate that this end-to-end approach consistently outperforms two-stage baselines which first perform 2D amodal completion and then 3D reconstruction. Together, these three contributions form a cohesive research pipeline for handling partially visible inputs of different types: beginning with completing missing panorama content from NFoV images, moving to weakly supervised object-level decomposition of 3D scenes, and culminating in occlusion-aware single-stage 3D object generation. By integrating latent diffusion, neural implicit representations, and occlusion-aware attention, this thesis demonstrates how to robustly recover complete visual representations from limited or inconsistent observations, thereby advancing the state-of-the-art in real-world 3D vision tasks.
Proceedings of the 2011 joint ACM workshop on Human gesture and behavior understanding
Proceedings of the 2011 joint ACM workshop on Human gesture and behavior understanding
SDFDiff: Differentiable Rendering of Signed Distance Fields for 3D Shape Optimization
We propose SDFDiff, a novel approach for image-based shape optimization using differentiable rendering of 3D shapes represented by signed distance functions (SDFs). Compared to other representations, SDFs have the advantage that they can represent shapes with arbitrary topology, and that they guarantee watertight surfaces. We apply our approach to the problem of multi-view 3D reconstruction, where we achieve high reconstruction quality and can capture complex topology of 3D objects. In addition, we employ a multi-resolution strategy to obtain a robust optimization algorithm. We further demonstrate that our SDF-based differentiable renderer can be integrated with deep learning models, which opens up options for learning approaches on 3D objects without 3D supervision. In particular, we apply our method to single-view 3D reconstruction and achieve state-of-the-art results.
Read moreDeep Learning based Single-view 3D Reconstruction: A Survey
Reconstruction 3D objects from a single view is an important research direction in computer vision. With the development of deep learning technology, a single-view 3D reconstruction based on deep learning has made remarkable progress in recent years. In reviewing 3D object reconstruction based on deep learning, firstly, the research progress of 3D reconstruction methods based on deep learning is analyzed in detail according to different representations of 3D objects, i.e., point cloud, voxel, and mesh. Secondly, the standard 3D reconstruction datasets and evaluation criteria are summarized. Finally, the challenges of 3D object reconstruction based on deep learning are summarized.
Read moreReal-time visualization of Dynamic Unlimited Objects Instancing
In this paper, we propose a novel approach to an efficient rendering of an unlimited number of dynamic and unique\n3D objects in real-time. We present an extension to the Holistic Unlimited Object Instancing (UOI) rendering\npipeline and the holistic computer graphics paradigm. We called this extension Dynamic Unlimited Object Instancing\nrendering pipeline. Using Signed Distance Functions (SDF) for the virtual scene representation and the\nHolistic Scene Dynamics Function, we can control and render an unlimited number of dynamic 3D objects in\nreal-time. In order to solve some issues of the original UOI rendering pipeline, we developed two extensions:\nfirst, a collection of holistic Dynamic operators, and, second, the Multipass Depth-Based Ray Marching rendering\npipeline. The operators are used to apply affine transformations to an unlimited number of 3D objects and also to\nanimate their materials and other attributes. In order to solve the problem of the uniform object distribution within\nthe scene, we redefined the original definition of the scene SDF component. The virtual scene equation is divided\ninto independent SDF components, which are rendered separately using the Multipass Depth-Based Ray Marching\npipeline. Thanks to both extensions, the new version of the Holistic UOI rendering pipeline can handle 3D\nobjects intersections what significantly enhances the realism of SDF scenes. The presented extensions to the UOI\nrendering pipeline are fully compatible with the Holistic UOI rendering pipeline, SDF and Sparse Voxel Octree\n(SVO) based algorithms. The only hardware requirement for our approach is the support for multipass rendering\nwith compute shaders or any GPGPU API.
Read more3D objects reconstruction from frontal images: an example with guitars
This work deals with the automatic 3D reconstruction of objects from frontal RGB images. This aims at a better understanding of the reconstruction of 3D objects from RGB images and their use in immersive virtual environments. We propose a complete workflow that can be easily adapted to almost any other family of rigid objects. To explain and validate our method, we focus on guitars. First, we detect and segment the guitars present in the image using semantic segmentation methods based on convolutional neural networks. In a second step, we perform the final 3D reconstruction of the guitar by warping the rendered depth maps of a fitted 3D template in 2D image space to match the input silhouette. We validated our method by obtaining guitar reconstructions from real input images and renders of all guitar models available in the ShapeNet database. Numerical results for different object families were obtained by computing standard mesh evaluation metrics such as Intersection over Union, Chamfer Distance, and the F-score. The results of this study show that our method can automatically generate high-quality 3D object reconstructions from frontal images using various segmentation and 3D reconstruction techniques.
Read moreReal-Time Camera Tracking and 3D Reconstruction Using Signed Distance Functions
The ability to quickly acquire 3D models is an essential capability needed in many disciplines including robotics, computer vision, geodesy, and architecture. In this paper we present a novel method for real-time camera tracking and 3D reconstruction of static indoor environments using an RGB-D sensor. We show that by representing the geometry with a signed distance function (SDF), the camera pose can be efficiently estimated by directly minimizing the error of the depth images on the SDF. As the SDF contains the distances to the surface for each voxel, the pose optimization can be carried out extremely fast. By iteratively estimating the camera poses and integrating the RGB-D data in the voxel grid, a detailed reconstruction of an indoor environment can be achieved. We present reconstructions of several rooms using a hand-held sensor and from onboard an autonomous quadrocopter. Our extensive evaluation on publicly available benchmark data shows that our approach is more accurate and robust than the iterated closest point algorithm (ICP) used by KinectFusion, and yields often a comparable accuracy at much higher speed to feature-based bundle adjustment methods such as RGB-D SLAM for up to medium-sized scenes.
Read more3D Graphics Reconstruction, Compression and Animation
This work covers entire pipeline of a 3D immersive system, comprising acquistion and reconstruction of 3D objects, transforming the 3D objects into popular 3D format, compression of data with MPEG-4 compliant encoders and acquisition of animation motion data. For the acquisition part we created 3D human face object with the help of depth-field camera, in particular we used Microsoft Kinect. We process the raw data given by Kinect and transform it into a mesh, and texture it with photometric data. The reconstructed objects are quiet large in size and need to be compressed for an efficient network transmission. We encoded the objects using MPEG-4 encoder, and measured the performance of scalable mesh encoding techniques. For an interactive immersive application, motion data of player must be cap-tured, which is transmitted to remote client for playing the animation on the virtual character of player. For this purpose MPEG-4 standard defines a Bone Based An-imation to acquire and compress the motion data. We acquired the motion data of a player using a novel algorithm which is compliant to bone based animation. Our proposed approach of extracting motion data is computationally efficient
Read moreLearning in the dark: 3D integral imaging object recognition in very low illumination conditions using convolutional neural networks
We propose a framework for three-dimensional (3D) object recognition and classification in very low illumination environments using convolutional neural networks (CNNs). 3D images are reconstructed using 3D integral imaging (InIm) with conventional visible spectrum image sensors. After imaging the low light scene using 3D InIm, the 3D reconstructed image has a higher signal-to-noise ratio than a single 2D image, which is a result of 3D InIm being optimal in the maximum likelihood sense for read-noise dominant images. Once 3D reconstruction has been performed, the 3D image is denoised and regions of interest are extracted to detect 3D objects in a scene. The extracted regions are then inputted into a CNN, which was trained under low illumination conditions using 3D InIm reconstructed images, to perform object recognition. To the best of our knowledge, this is the first report of utilizing 3D InIm and convolutional neural networks for 3D training and 3D object classification under very low illumination conditions.
Read moreISHS-Net: Single-View 3D Reconstruction by Fusing Features of Image and Shape Hierarchical Structures
The reconstruction of 3D shapes from a single view has been a longstanding challenge. Previous methods have primarily focused on learning either geometric features that depict overall shape contours but are insufficient for occluded regions, local features that capture details but cannot represent the complete structure, or structural features that encode part relationships but require predefined semantics. However, the fusion of geometric, local, and structural features has been lacking, leading to inaccurate reconstruction of shapes with occlusions or novel compositions. To address this issue, we propose a two-stage approach for achieving 3D shape reconstruction. In the first stage, we encode the hierarchical structure features of the 3D shape using an encoder-decoder network. In the second stage, we enhance the hierarchical structure features by fusing them with global and point features and feed the enhanced features into a signed distance function (SDF) prediction network to obtain rough SDF values. Using the camera pose, we project arbitrary 3D points in space onto different depth feature maps of the CNN and obtain their corresponding positions. Then, we concatenate the features of these corresponding positions together to form local features. These local features are also fed into the SDF prediction network to obtain fine-grained SDF values. By fusing the two sets of SDF values, we improve the accuracy of the model and enable it to reconstruct other object types with higher quality. Comparative experiments demonstrate that the proposed method outperforms state-of-the-art approaches in terms of accuracy.
Read moreComputer Stereo Vision for Autonomous Driving: Theory and Algorithms
As an important component of autonomous systems, autonomous car perception has had a big leap with recent advances in parallel computing architectures. With the use of tiny but full-feature embedded supercomputers, computer stereo vision has been prevalently applied in autonomous cars for depth perception. The two key aspects of computer stereo vision are speed and accuracy. They are both desirable but conflicting properties, as the algorithms with better disparity accuracy usually have higher computational complexity. Therefore, the main aim of developing a computer stereo vision algorithm for resource-limited hardware is to improve the trade-off between speed and accuracy. In this chapter, we introduce both the hardware and software aspects of computer stereo vision for autonomous car systems. Then, we discuss four autonomous car perception tasks, including (1) visual feature detection, description and matching, (2) 3D information acquisition, (3) object detection/recognition and (4) semantic image segmentation. The principles of computer stereo vision and parallel computing on multi-threading CPU and GPU architectures are then detailed.
Read moreMulti-SLM holographic display system with planar configuration
Holographic display system that uses six phase-only spatial light modulators (SLMs) performs holographic reconstructions from the phase-hologram of a point cloud that is extracted from 3D object. The SLMs are tiled as a three by two matrix on a virtual planar surface. The alignment is successful and the display system generates large holographic reconstructions. The proposed system can be used either to obtain reconstructions of large objects with a narrow field of view or reconstructions of smaller objects with a broader field of view. Therefore, since field of view is broader for smaller objects, observer has the flexibility to move around the reconstruction within a larger angle. This flexibility increases the motion parallax and as a consequence it increases the quality of 3D perception. Results show that even with three SLMs in horizontal direction the 3D perception is significantly increased. Experimental results are satisfactory.
Read more3D reconstruction of aerodynamic airfoils using computer stereo vision
One of the most important components of a wind turbine are the blades, the evaluation of their manufacturing quality and aerodynamic capabilities can be very costly, for this reason a 3D reconstruction by stereo vision is proposed. This technique consists of projecting a laser line in each face of the blade. Using a linear stage, two cameras will scan simultaneously, considering bidirectional disparities and feature correspondences between the two pictures. Two symmetric airfoils of the NACA 0012 family are evaluated. The expected precision is 0.1mm.
Read more8 Exploring practical use-cases of augmented reality using photogrammetry and other 3D reconstruction tools in the Metaverse
In today’s world, people increasingly rely on mobile apps to do their day-today activities, like checking their Instagram feed and online shopping from websites like Amazon and Flipkart. People are depending on WhatsApp and Instagram stories to communicate with local businesses and to leverage the said platforms for online advertising. Using Google Maps to find their way when they travel and finding out the immediate road and traffic conditions with digital banners around the road, has obviously led to a boom in advertising and marketing. Recently, internet users have increasingly desired to immerse themselves in a Metaverse-like platform where they can interact and socialize. Meta’s Metaverse is a tightly connected network of 3D digital spaces that allow users to escape into a virtual world. It is designed to change the way you socialize, work, shop, and connect with the real and virtual world around you. These platforms are not fully submerged in the real world; they are inclined toward virtual spaces only, making it obvious to fill this gap. Thus, the proposed framework in this chapter would be a new kind of system that may develop a socio-meta platform, powered by augmented reality and other technologies like photogrammetry and LiDAR. Augmented reality provides an interactive way of experiencing the real world, where the objects of the natural world are enhanced by computer-generated perceptual vision. One of the significant problems with augmented reality is the process of building virtual 3D objects that can be augmented into real spaces, which could be solved with photogrammetry and LiDAR. Photogrammetry is the technique of producing 3D objects using 2D images of a physical object taken from different angles and orientations. LiDAR, on the other hand, is another 3D reconstruction technology used widely by Apple’s eco-system. The functioning of LIDAR is very similar to sonar and radar, and the detection and ranging part are where it stands out from the others. The idea behind this platform is to open tons of virtual dimensions in the real world using the principles of mixed reality and geographic mapping tools such as Google Maps, MapBox, and GeoJSON; it would consequently transition the way people spend their time on social media by opening a portal for generating 3D objects that can be augmented to the real-world location using photogrammetry and cloud anchors by just a few 2D digital photographs taken from their camera.
Read moreA high-quality voxel 3D reconstruction system for large scenes based on the branch and bound method
A high-quality voxel 3D reconstruction system for large scenes based on the branch and bound method
DEMO] Mobile augmented reality — 3D object selection and reconstruction with an RGBD sensor and scene understanding
In this proposal we show case two 3D reconstruction systems running in real-time on a tablet equipped with a depth sensor. We believe that the proposed set of demonstrations will engage ISMAR attendees both in terms of tracking technology and user experience. Both demos show state-of-the art 3D reconstruction technology and give attendees a chance to try hands-on our tracking with simple and interactive user interfaces.
Read more