- Research Article
22
- 10.1016/j.neucom.2021.01.110
Monocular 3D object detection using dual quadric for autonomous driving
- Feb 16, 2021
- Neurocomputing
- Peixuan Li + 1 more +1
Monocular 3D object detection using dual quadric for autonomous driving
For many automated driving functions, a highly accurate perception of the vehicle environment is a crucial prerequisite. Modern high-resolution radar sensors generate multiple radar targets per object, which makes these sensors particularly suitable for the 2D object detection task. This work presents an approach to detect 2D objects solely depending on sparse radar data using PointNets. In literature, only methods are presented so far which perform either object classification or bounding box estimation for objects. In contrast, this method facilitates a classification together with a bounding box estimation of objects using a single radar sensor. To this end, PointNets are adjusted for radar data performing 2D object classification with segmentation, and 2D bounding box regression in order to estimate an amodal 2D bounding box. The algorithm is evaluated using an automatically created dataset which consist of various realistic driving maneuvers. The results show the great potential of object detection in high-resolution radar data using PointNets.
Monocular 3D object detection using dual quadric for autonomous driving
Monocular 3D object detection using dual quadric for autonomous driving
Integrating Multimodal Large Language Models with Point Clouds for Urban Vehicle Detection
Road infrastructure assessment plays a vital role in urban planning, traffic management, and the development of autonomous systems. Recent advances in Multimodal Large Language Models (MLLMs) have shown promising capabilities for zero-shot object detection and classification in complex visual scenes. However, these models are primarily limited to 2D image analysis, lacking the spatial awareness required for precise geometric understanding in real-world environments. Meanwhile, conventional 3D object detection methods using point clouds face challenges in handling diverse, dynamic, and unstructured road environments. This work introduces a hybrid approach that combines the zero-shot visual reasoning capabilities of MLLMs with the geometric precision of 3D LiDAR point clouds collected from Mobile Mapping Systems (MMS). Specifically, the proposed pipeline uses an MLLM (Google Gemini 2.0 flash) to detect vehicles in road scene images based on natural language prompts, generating 2D bounding boxes without the need for task-specific training. These 2D detections are then projected into 3D space using calibration parameters from the KITTI-360 dataset, enabling localization and clustering of corresponding LiDAR points. Finally, 3D bounding boxes are generated, and temporal merging is performed to remove redundant detections across consecutive frames. Experiments on two sequences of the KITTI-360 dataset demonstrate that this fusion approach enables flexible and scalable vehicle detection in complex urban environments. The findings highlight the potential of integrating language-driven perception with spatially grounded 3D reasoning, offering a promising direction for automated road asset monitoring and maintenance strategies.
Read more3DYOLO: Real-time 3D Object Detection in 3D Point Clouds for Autonomous Driving
In the recent era, a lot of interest is attracted by the autonomous vehicles which can sense surroundings and navigate without human intervention. Object detection and recognition form a major part of autonomous driving systems. Lidar sensors can be used to capture point clouds of driving environment. Detecting multiple 3D objects in point clouds in real time and defining their boundaries with the help of 3D bounding boxes are critical in motion planning by self driving systems. This paper proposes a LiDAR-based 3D object detection system that operates in real-time, with emphasis on autonomous driving scenarios. A state-of-the-art 2D standard object detector for RGB images, YOLOv4, is used as the base for object detection. The multi-class 3D bounding boxes are generated using a complex regression approach. An Euler-Region-Proposal Network (E-RPN) is used to predict the pose of the object. The proposed model receives point cloud data as input and outputs 3D bounding boxes with classes in real-time. The experiments done on the KITTI benchmark dataset proves that the proposed system outperforms existing methods in terms of accuracy and performance.
Read moreReply on RC2
We present a high-resolution airborne radar data set (EGRIP-NOR-2018) for the onset region of the Northeast Greenland Ice Stream (NEGIS). The radar data were acquired in May 2018 with Alfred Wegener Institute’s multichannel ultra-wideband (UWB) radar mounted on the Polar6 aircraft. Radar profiles cover an area of ~24000 km2 and extend over the well-defined shear margins of the NEGIS. The survey area is centred at the location of the drill site of the East Greenland Ice-Core Project (EastGRIP) and several radar lines intersect at this location. The survey layout was designed to: (i) map the stratigraphic signature of the shear margins with radar profiles aligned perpendicular to ice flow, (ii) trace the radar stratigraphy along several flow lines and (iii) provide spatial coverage of ice thickness and basal properties. While we are able to resolve radar reflections in the deep stratigraphy, we can not fully resolve the steeply inclined reflections at the tightly folded shear margins in the lower part of the ice column. The NEGIS is causing the most significant discrepancies between numerically modelled and observed ice surface velocities. Given the high likelihood of future climate and ocean warming, this extensive data set of new high-resolution radar data in combination with the EastGRIP ice core will be a key contribution to understand the past and future dynamics of the NEGIS. The EGRIP-NOR-2018 radar data products can be obtained at the PANGAEA Data Publisher (https://doi.pangaea.de/10.1594/PANGAEA.928569; Franke et al. 2021a).
Read moreAirborne ultra-wideband radar sounding over the shear margins and along flow lines at the onset region of the Northeast Greenland Ice Stream
Abstract. We present a high-resolution airborne radar data set (EGRIP-NOR-2018) for the onset region of the Northeast Greenland Ice Stream (NEGIS). The radar data were acquired in May 2018 with the Alfred Wegener Institute's multichannel ultra-wideband (UWB) radar mounted on the Polar 6 aircraft. Radar profiles cover an area of ∼24 000 km2 and extend over the well-defined shear margins of the NEGIS. The survey area is centered at the location of the drill site of the East Greenland Ice-Core Project (EastGRIP), and several radar lines intersect at this location. The survey layout was designed to (i) map the stratigraphic signature of the shear margins with radar profiles aligned perpendicular to ice flow, (ii) trace the radar stratigraphy along several flow lines, and (iii) provide spatial coverage of ice thickness and basal properties. While we are able to resolve radar reflections in the deep stratigraphy, we cannot fully resolve the steeply inclined reflections at the tightly folded shear margins in the lower part of the ice column. The NEGIS is causing the most significant discrepancies between numerically modeled and observed ice surface velocities. Given the high likelihood of future climate and ocean warming, this extensive data set of new high-resolution radar data in combination with the EastGRIP ice core will be a key contribution to understand the past and future dynamics of the NEGIS. The EGRIP-NOR-2018 radar data products can be obtained from the PANGAEA data publisher (https://doi.pangaea.de/10.1594/PANGAEA.928569; Franke et al., 2021a).
Read moreA Robust Strategy for Roadside Cooperative Perception Based on Multi-Sensor Fusion
Roadside perception is a fundamental task for vehicle-to-road cooperative perception and traffic scheduling. However, most existing roadside perception strategies prefer to deploy sensors in a single perspective or test in a simulation environment. Due to the limited field of view covered by a single sensor, such methods usually cannot continuously detect the same object from different viewpoints or provide a wide sensing range in complex scenarios. To address these issues, a robust strategy for roadside cooperative perception based on multi-sensor fusion (RCP-MSF) is proposed in this paper. A 2D object detector is improved based on the NanoDet model to handle multiple images simultaneously. In addition, an ultra-fast 3D object detection strategy is suggested based on point cloud processing rather than relying on existing high-cost deep-learning models. Moreover, to match the 2D and 3D bounding boxes, a data association module for multi-modal sensor information fusion is presented. Any 2D and 3D object detector can follow this module. Furthermore, a roadside perception dataset named SCUT-V2R is constructed to verify the performance of the proposed method. Experiments on the dataset demonstrate that the RCP-MSF outperforms the camera-only and lidar-only strategies in object detection precision while maintaining real-time performance.
Read moreA Traffic Information Awareness Approach Based on Video Data and Millimeter Wave Radar Data Fusion
In order to solve the problems of using a single radar or video sensor in the traffic information detection process, such as susceptibility to environmental influences and non-intuitive target reflections, we propose a detection method using target correlation matching, target tracking, and target data fusion, and we show how to adjust the detection weights of the video sensor and radar sensor by adjusting the noise matrix parameters, which forms a flexible, simple, and nimble method. The efficient architecture allows accurate results in challenging environments. We show the effect on radar detection data and video detection data before and after adjustment. We test the proposed method by building a hardware verification platform and finally demonstrate it on video images. The experimental results show that the proposed method can significantly reduce the testing time while providing high detection rate in more environments.
Read moreLeveraging Pre-Trained 3D Object Detection Models for Fast Ground Truth Generation
Training 3D object detectors for autonomous driving has been limited to small datasets due to the effort required to generate annotations. Reducing both task complexity and the amount of task switching done by annotators is key to reducing the effort and time required to generate 3D bounding box annotations. This paper introduces a novel ground truth generation method that combines human supervision with pre-trained neural networks to generate per-instance 3D point cloud segmentation, 3D bounding boxes, and class annotations. The annotators provide object anchor clicks which behave as a seed to generate instance segmentation results in 3D. The points belonging to each instance are then used to regress object centroids, bounding box dimensions, and object orientation. Our proposed annotation scheme requires 30x lower human annotation time. We use the KITTI 3D object detection dataset [1] to evaluate the efficiency and the quality of our annotation scheme. We also test the the proposed scheme on previously unseen data from the Autonomoose self-driving vehicle to demonstrate generalization capabilities of the network.
Read moreSynergy of very high resolution optical and radar data for object-based olive grove mapping
This study investigates the potential of very high resolution (VHR) optical and radar data for olive grove landscape mapping. VHR data were fed into a four-step processing chain performing an object-based land-use classification. The four steps included (i) image segmentation, (ii) object feature calculation, (iii) object-based classification and (iv) land-use map evaluation. First, the optical (ADS40) and radar (RAMSES SAR and TerraSAR-X) data were applied to the processing chain separately. As supported by two segmentation evaluation measures, the stand purity index (PI) and the potential mapping accuracy (PMA), the optical data thereby led to a significantly better segmentation and a more accurate olive cover map (Kruskal–Wallis test, ). Second, synergy models were developed combining data from the different sensors at different stages of the object-based classification process, namely, (1) during the segmentation step, (2) during the feature calculation step and (3) after the object classification step. The combined use of features from the different sensors resulted in a considerable improvement in mapping accuracy, with correctly classified objects supported by high probabilities. The assessment of feature importance revealed that optical data were most important for successful object-based olive grove mapping; however, features related to object shape and texture of the radar imagery added to its success. Comparison of the object-based synergy model with a pixel-based synergy model indicated a limited classification improvement. This research showed that the integrated use of VHR optical and radar data is appropriate in an object-based classification framework, leading towards more accurate olive grove landscape mapping.
Read moreTemporally consistent caption detection in videos using a spatiotemporal 3D method
Captions are text or logos superimposed on videos during a postproduction process. Caption detection in videos is useful for a variety of applications. For many applications, temporal consistency and stability is very important. Most of the prior work adopts certain post-processing procedures to smooth detected caption bounding boxes over time. Although these approaches mitigate the effect of the temporal inconsistency problem, they are unable to eliminate the problem. In this paper, we present a new caption detection algorithm that detects the 3D bounding boxes of caption regions in spatiotemporal volume space. 2D bounding boxes are then created by slicing the 3D bounding boxes. Since all the 2D bounding boxes corresponding to a caption area are sliced from one 3D bounding box, they are identical over time, thus ensuring temporal consistency of the result. The experiment results show that our new approach not only generates temporally consistent results but also results in higher detection accuracy.
Read moreEstimation of 6D Object Pose Using a 2D Bounding Box
This paper provides an efficient way of addressing the problem of detecting or estimating the 6-Dimensional (6D) pose of objects from an RGB image. A quaternion is used to define an object′s three-dimensional pose, but the pose represented by q and the pose represented by -q are equivalent, and the L2 loss between them is very large. Therefore, we define a new quaternion pose loss function to solve this problem. Based on this, we designed a new convolutional neural network named Q-Net to estimate an object’s pose. Considering that the quaternion′s output is a unit vector, a normalization layer is added in Q-Net to hold the output of pose on a four-dimensional unit sphere. We propose a new algorithm, called the Bounding Box Equation, to obtain 3D translation quickly and effectively from 2D bounding boxes. The algorithm uses an entirely new way of assessing the 3D rotation (R) and 3D translation rotation (t) in only one RGB image. This method can upgrade any traditional 2D-box prediction algorithm to a 3D prediction model. We evaluated our model using the LineMod dataset, and experiments have shown that our methodology is more acceptable and efficient in terms of L2 loss and computational time.
Read moreDirect 3D Detection of Vehicles in Monocular Images with a CNN based 3D Decoder
In autonomous driving, the detection of objects like surrounding vehicles based on monocular RGB images is usually performed by 2D bounding box detectors. The resulting 2D objects can be used for a first coarse 3D position estimate but for a precise location, additional sensor data has to be taken into account. For further use in sensor fusion systems and environment maps it is preferable to detect objects, their orientation and dimensions directly in 3D coordinates. To address this 3D object detection task, we propose a direct 3D bounding box estimator which is realized as CNN decoder module and can be connected to most 2D object detectors like SSD[1], OverFeat[2], YOLO[3] and RetinaNet[4] or directly to CNN feature extractors like VGG [2] and ResNet [5]. The 3D parameters of the objects such as dimension and orientation are directly predicted by the CNN module. To successfully train this complex MultiNet architecture, a combination and modification of current loss functions is proposed. The fastest of the proposed network module combinations is capable of detecting objects in 3D camera coordinates at a frame rate of 28 fps.
Read moreColoRadar: The direct 3D millimeter wave radar dataset
This work presents two different forms of dense, high-resolution radar data from two frequency modulated continuous wave radar sensors, along sparse radar pointclouds produced by one of the radar sensors. In addition, all datasets include 3D lidar and inertial measurements, and a lidar-based simultaneous localization and mapping pose estimation. Over 2 h of 6D pose data was generated across 52 datasets collected in highly diverse 3D environments including lab spaces, outside and inside large buildings, urban walkways, and a mine. One dataset, from the ASPEN Lab, also includes precision groundtruth generated from a motion capture system. Intrinsic radar calibration and measured extrinsic sensor position calibrations are also provided along with python based development tools to interact with the various datasets. This data is designed to assist with generating radar based localization algorithms and calibrations between radar and other sensors.
Read moreSpatial Attention Frustum: A 3D Object Detection Method Focusing on Occluded Objects.
Achieving the accurate perception of occluded objects for autonomous vehicles is a challenging problem. Human vision can always quickly locate important object regions in complex external scenes, while other regions are only roughly analysed or ignored, defined as the visual attention mechanism. However, the perception system of autonomous vehicles cannot know which part of the point cloud is in the region of interest. Therefore, it is meaningful to explore how to use the visual attention mechanism in the perception system of autonomous driving. In this paper, we propose the model of the spatial attention frustum to solve object occlusion in 3D object detection. The spatial attention frustum can suppress unimportant features and allocate limited neural computing resources to critical parts of the scene, thereby providing greater relevance and easier processing for higher-level perceptual reasoning tasks. To ensure that our method maintains good reasoning ability when faced with occluded objects with only a partial structure, we propose a local feature aggregation module to capture more complex local features of the point cloud. Finally, we discuss the projection constraint relationship between the 3D bounding box and the 2D bounding box and propose a joint anchor box projection loss function, which will help to improve the overall performance of our method. The results of the KITTI dataset show that our proposed method can effectively improve the detection accuracy of occluded objects. Our method achieves 89.46%, 79.91% and 75.53% detection accuracy in the easy, moderate, and hard difficulty levels of the car category, and achieves a 6.97% performance improvement especially in the hard category with a high degree of occlusion. Our one-stage method does not need to rely on another refining stage, comparable to the accuracy of the two-stage method.
Read moreDrizzle Measurements Using High Spectral Resolution Lidar and Radar Data
\nThe ratio of millimeter radar and High Spectral Resolution Lidar (HSRL) backscatter are used to determine drizzle rates which are compared to conventional ground based measurements. The robustly calibrated HSRL backscatter cross section provides advantages over measurements made with traditional lidars.\n
Read more