• Home
  • Search
  • 3DWG: 3D Weakly Supervised Visual Grounding via Category and Instance-Level Alignment
  • Cite Icon2
  • https://doi.org/10.1109/icra55743.2025.11127295Copy DOI Icon

3DWG: 3D Weakly Supervised Visual Grounding via Category and Instance-Level Alignment

  • May 19, 2025
  • Xiaoqi Li +6 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

The 3D weakly-supervised visual grounding task aims to localize oriented 3D boxes in point clouds based on natural language descriptions without requiring annotations to guide model learning. This setting presents two primary challenges: category-level ambiguity and instance-level complexity. Category-level ambiguity arises from representing objects of fine-grained categories in a highly sparse point cloud format, making category distinction challenging. Instance-level complexity stems from multiple instances of the same category coexisting in a scene, leading to distractions during grounding. To address these challenges, we propose a novel weaklysupervised grounding approach that explicitly differentiates between categories and instances. In the category-level branch, we utilize extensive category knowledge from a pre-trained external detector to align object proposal features with sentencelevel category features, thereby enhancing category awareness. In the instance-level branch, we utilize spatial relationship descriptions from language queries to refine object proposal features, ensuring clear differentiation among objects. These designs enable our model to accurately identify target-category objects while distinguishing instances within the same category. Compared to previous methods, our approach achieves state-of-the-art performance on three widely used benchmarks: Nr3D, Sr3D, and ScanRef.

Similar Papers
  • Research Article
  • Citations3

A Masked Reference Token Supervision-Based Iterative Visual-Language Framework for Robust Visual Grounding

  • Jan 01, 2025
  • IEEE Transactions on Circuits and Systems for Video Technology
  • Chunlei Wang +6
  • Research Article
  • Citations5

Toward Visual Grounding: A Survey.

  • Mar 01, 2026
  • IEEE transactions on pattern analysis and machine intelligence
  • Linhui Xiao +4
  • Research Article
  • Citations74

TransVG++: End-to-End Visual Grounding With Language Conditioned Vision Transformer.

  • Nov 01, 2023
  • IEEE Transactions on Pattern Analysis and Machine Intelligence
  • Jiajun Deng +7
  • Research Article
  • Citations10

Language-guided Residual Graph Attention Network and Data Augmentation for Visual Grounding

  • Aug 24, 2023
  • ACM Transactions on Multimedia Computing, Communications, and Applications
  • Jia Wang +3
  • Conference Article
  • Citations1

PSAIR: A Neuro-Symbolic Approach to Zero-Shot Visual Grounding

  • Jun 30, 2024
  • Yi Pan +3
  • Research Article
  • Citations19

GroundVLP: Harnessing Zero-Shot Visual Grounding from Vision-Language Pre-training and Open-Vocabulary Object Detection

  • Mar 24, 2024
  • Proceedings of the AAAI Conference on Artificial Intelligence
  • Haozhan Shen +3
  • Video Transcripts

Seeing the advantage: visually grounding word embeddings to better capture human semantic knowledge

  • May 17, 2022
  • Underline Science Inc.
  • Danny Merkx
  • PDF
  • Research Article

Region Collaborative Network for Detection-Based Vision-Language Understanding

  • Aug 30, 2022
  • Mathematics
  • Linyan Li +4
  • Research Article

End-to-end Visual Grounding Based on Query Text Guidance and Multi-stage Reasoning

  • Feb 01, 2024
  • 電腦學刊
  • Chao Wang Chao Wang +5
  • Research Article

How direct is the link between words and images?

  • Dec 31, 2023
  • The Mental Lexicon
  • Hassan Shahmohammadi +4
  • Research Article
  • Citations3

Learning language to symbol and language to vision mapping for visual grounding

  • Apr 14, 2022
  • Image and Vision Computing
  • Su He +2
  • Book Chapter
  • Citations125

SeqTR: A Simple Yet Universal Network for Visual Grounding

  • Jan 01, 2022
  • Chaoyang Zhu +9
  • Research Article
  • Citations8

Cycle-Consistency Learning for Captioning and Grounding

  • Mar 24, 2024
  • Proceedings of the AAAI Conference on Artificial Intelligence
  • Ning Wang +2
  • Conference Article
  • Citations79

3DJCG: A Unified Framework for Joint Dense Captioning and Visual Grounding on 3D Point Clouds

  • Jun 01, 2022
  • Daigang Cai +4
  • Video Transcripts

Attending Self-Attention: A Case Study ofVisually Grounded Supervision in Vision-and-Language Transformers

  • Jul 22, 2021
  • Underline Science Inc.
  • Noa Garcia +4
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.