• Home
  • Search
  • SeqTR: A Simple Yet Universal Network for Visual Grounding
  • Open Access IconOpen Access
  • Cite Icon125
  • https://doi.org/10.1007/978-3-031-19833-5_35Copy DOI Icon

SeqTR: A Simple Yet Universal Network for Visual Grounding

  • Jan 1, 2022
  • Chaoyang Zhu +9 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

In this paper, we propose a simple yet universal network termed SeqTR for visual grounding tasks, e.g., phrase localization, referring expression comprehension (REC) and segmentation (RES). The canonical paradigms for visual grounding often require substantial expertise in designing network architectures and loss functions, making them hard to generalize across tasks. To simplify and unify the modeling, we cast visual grounding as a point prediction problem conditioned on image and text inputs, where either the bounding box or binary mask is represented as a sequence of discrete coordinate tokens. Under this paradigm, visual grounding tasks are unified in our SeqTR network without task-specific branches or heads, e.g., the convolutional mask decoder for RES, which greatly reduces the complexity of multi-task modeling. In addition, SeqTR also shares the same optimization objective for all tasks with a simple cross-entropy loss, further reducing the complexity of deploying hand-crafted loss functions. Experiments on five benchmark datasets demonstrate that the proposed SeqTR outperforms (or is on par with) the existing state-of-the-arts, proving that a simple yet universal approach for visual grounding is indeed feasible. Source code is available at https://github.com/sean-zhuh/SeqTR.

Similar Papers
  • Research Article
  • Citations5

Toward Visual Grounding: A Survey.

  • Mar 01, 2026
  • IEEE transactions on pattern analysis and machine intelligence
  • Linhui Xiao +4
  • Research Article
  • Citations44

Unambiguous Scene Text Segmentation with Referring Expression Comprehension.

  • Jul 26, 2019
  • IEEE Transactions on Image Processing
  • Xuejian Rong +2
  • Research Article
  • Citations19

GroundVLP: Harnessing Zero-Shot Visual Grounding from Vision-Language Pre-training and Open-Vocabulary Object Detection

  • Mar 24, 2024
  • Proceedings of the AAAI Conference on Artificial Intelligence
  • Haozhan Shen +3
  • PDF
  • Research Article
  • Citations103

ICIoU: Improved Loss Based on Complete Intersection Over Union for Bounding Box Regression

  • Jan 01, 2021
  • IEEE Access
  • Xufei Wang +1
  • PDF
  • Research Article
  • Citations15

NGIoU Loss: Generalized Intersection over Union Loss Based on a New Bounding Box Regression

  • Dec 13, 2022
  • Applied Sciences
  • Chenghao Tong +3
  • Research Article
  • Citations5

HE-YOLO: Aerial Target Detection Based On Improved YOLOv3

  • Sep 30, 2021
  • International Journal of Pattern Recognition and Artificial Intelligence
  • Yuqing Zhao +3
  • Research Article
  • Citations194

Practical Blind Image Denoising via Swin-Conv-UNet and Data Synthesis

  • Sep 15, 2023
  • Machine Intelligence Research
  • Kai Zhang +8
  • Conference Article
  • Citations111

An Improved Bounding Box Regression Loss Function Based on CIOU Loss for Multi-scale Object Detection

  • Jul 16, 2021
  • Shuangjiang Du +3
  • PDF
  • Research Article
  • Citations11

Object Detection of Flexible Objects with Arbitrary Orientation Based on Rotation-Adaptive YOLOv5

  • May 20, 2023
  • Sensors
  • Jiajun Wu +5
  • Conference Article
  • Citations6

Talk2Radar: Bridging Natural Language with 4D mmWave Radar for 3D Referring Expression Comprehension

  • May 19, 2025
  • Runwei Guan +10
  • Conference Article
  • Citations35

Word Discovery in Visually Grounded, Self-Supervised Speech Models

  • Sep 18, 2022
  • Puyuan Peng +1
  • Conference Article
  • Citations4

FlowVQA: Mapping Multimodal Logic in Visual Question Answering with Flowcharts

  • Jan 01, 2024
  • Shubhankar Singh +6
  • Research Article
  • Citations17

Hierarchical Regression and Classification for Accurate Object Detection.

  • May 01, 2023
  • IEEE Transactions on Neural Networks and Learning Systems
  • Jiale Cao +3
  • Conference Article
  • Citations238

Task2Vec: Task Embedding for Meta-Learning

  • Oct 01, 2019
  • Alessandro Achille +7
  • Book Chapter
  • Citations24

Visual-Relation Conscious Image Generation from Structured-Text

  • Jan 01, 2020
  • Duc Minh Vo +1
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.