- Research Article
6
- 10.1016/j.neucom.2024.129241
QT-TextSR: Enhancing scene text image super-resolution via efficient interaction with text recognition using a Query-aware Transformer
- Mar 01, 2025
- Neurocomputing
- Chongyu Liu + 8 more +8
Publications from 2021 to 2026
Showing 10 of 20 papers
QT-TextSR: Enhancing scene text image super-resolution via efficient interaction with text recognition using a Query-aware Transformer
PEneo: Unifying Line Extraction, Line Grouping, and Entity Linking for End-to-end Document Pair Extraction
Document pair extraction aims to identify key and value entities as well as their relationships from visually-rich documents. Most existing methods divide it into two separate tasks: semantic entity recognition (SER) and relation extraction (RE). However, simply concatenating SER and RE serially can lead to severe error propagation, and it fails to handle cases like multi-line entities in real scenarios. To address these issues, this paper introduces a novel framework, PEneo (Pair Extraction new decoder option), which performs document pair extraction in a unified pipeline, incorporating three concurrent sub-tasks: line extraction, line grouping, and entity linking. This approach alleviates the error accumulation problem and can handle the case of multi-line entities. Furthermore, to better evaluate the model's performance and to facilitate future research on pair extraction, we introduce RFUND, a re-annotated version of the commonly used FUNSD and XFUND datasets, to make them more accurate and cover realistic situations. Experiments on various benchmarks demonstrate PEneo's superiority over previous pipelines, boosting the performance by a large margin (e.g., 19.89%-22.91% F1 score on RFUND-EN) when combined with various backbones like LiLT and LayoutLMv3, showing its effectiveness and generality. Codes and the new annotations are available at https://github.com/ZeningLin/PEneo.
Read moreMulti-Scale Adaptive Feature Network Drainage Pipe Image Dehazing Method Based on Multiple Attention
Drainage pipes are a critical component of urban infrastructure, and their safety and proper functioning are vital. However, haze problems caused by humid environments and temperature differences seriously affect the quality and detection accuracy of drainage pipe images. Traditional repair methods are difficult to meet the requirements when dealing with complex underground environments. To solve this problem, we researched and proposed a dehazing method for drainage pipe images based on multi-attention multi-scale adaptive feature networks. By designing multiple attention and adaptive modules, the network is able to capture global features with multi-scale resolution in complex underground environments, thereby achieving end-to-end dehazing processing. In addition, we also constructed a large drainage pipe dataset containing tens of thousands of clear/hazy image pairs of drainage pipes for network training and testing. Experimental results show that our network exhibits excellent dehazing performance in various complex underground environments, especially in the real scene of urban underground drainage pipes. The contributions of this paper are mainly reflected in the following aspects: first, a novel multi-scale adaptive feature network based on multiple attention is proposed to effectively solve the problem of dehazing drainage pipe images; second, a large-scale drainage pipe data is constructed. The collection provides valuable resources for related research work; finally, the effectiveness and superiority of the proposed method are verified through experiments, and it provides an efficient solution for dehazing work in scenes such as urban underground drainage pipes.
Read moreHuman-AI Co-creation for Intangible Cultural Heritage Dance: Cultural Genes Retaining and Innovation
An Examination of the Impact of Financial Sharing on the Quality of Corporate Accounting Information in the Context of the Financial Shared Service Model
In the context of globalization intensification, enterprises face the challenge of managing increasingly complex business operations. This study aims to investigate how the financial shared services model, characterized by centralized processing and standardized operations, impacts the quality of corporate accounting information. By examining literature and conducting empirical data analysis, we explore how the model enhances the efficiency and accuracy of accounting information processing, strengthens enterprises’ internal controls through increasing financial information transparency, and bolsters financial monitoring. Furthermore, we analyze how this model positively impacts decision support, aiding enterprises in making superior economic decisions. Despite the challenges in system compatibility and employee training during implementation, this study underscores the immense potential of the shared services model in optimizing accounting information systems, reinforcing internal controls, and improving decision support. The profound impact of this model on the quality of corporate accounting information underscores its value and the importance of its further research and application.
Read moreObject Recognition with Class Conditional Gaussian Mixture Model - A Statistical Learning Approach
Object recognition is one of the key tasks in robot vision. In RoboCup SPL, the Nao Robot must identify objects of interest such as the ball, field features et al. These objects are critical for the robot players to successfully play soccer games. We propose a new statistical learning method, Class Conditional Gaussian Mixture Model (ccGMM), that can be used either as an object detector or a false positive discriminator. It is able to achieve a high recall rate and a low false positive rate. The proposed model has low computational cost on a mobile robot and the learning process takes a relatively short time, so that it is suitable for real robot competition play.
Read moreEnlightening Low-Light Images With Dynamic Guidance for Context Enrichment
Images acquired in low-light conditions suffer from a series of visual quality degradations, <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">e.g.</i> , low visibility, degraded contrast, and intensive noise. These complicated degradations based on various contexts ( <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">e.g</i> ., noise in smooth regions, over-exposure in well-exposed regions and low contrast around edges) cast major challenges to the low-light image enhancement. Herein, we propose a new methodology by imposing a learnable guidance map from the signal and deep priors, making the deep neural network adaptively enhance low-light images in a region-dependent manner. The enhancement capability of the learnable guidance map is further exploited with the multi-scale dilated context collaboration, leading to contextually enriched feature representations extracted by the model with various receptive fields. Through assimilating the intrinsic perceptual information from the learned guidance map, richer and more realistic textures are generated. Extensive experiments on real low-light images demonstrate the effectiveness of our method, which delivers superior results quantitatively and qualitatively. The code is available at <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><uri>https://github.com/lingyzhu0101/GEMSC</uri></i> to facilitate future research.
Read morePUGCQ: A Large Scale Dataset for Quality Assessment of Professional User-Generated Content
Recent years have witnessed a surge of professional user-generated content (PUGC) based video services, coinciding with the accelerated proliferation of video acquisition devices such as mobile phones, wearable cameras, and unmanned aerial vehicles. Different from traditional UGC videos by impromptu shooting, PUGC videos produced by professional users tend to be carefully designed and edited, receiving high popularity with a relatively satisfactory playing count. In this paper, we systematically conduct the comprehensive study on the perceptual quality of PUGC videos and introduce a database consisting of 10,000 PUGC videos with subjective ratings. In particular, during the subjective testing, we collect the human opinions based upon not only the MOS, but also the attributes that could potentially influence the visual quality including face, noise, blur, brightness, and color. We make the attempt to analyze the large-scale PUGC database with a series of video quality assessment (VQA) algorithms and a dedicated baseline model based on pretrained deep neural network is further presented. The cross-dataset experiments reveal a large domain gap between the PUGC and the traditional user-generated videos, which are critical in learning based VQA. These results shed light on developing next-generation PUGC quality assessment algorithms with desired properties including promising generalization capability, high accuracy, and effectiveness in perceptual optimization. The dataset and the codes are released at https://github.com/wlkdb/pugcq_create.
Read moreToward Joint Thing-and-Stuff Mining for Weakly Supervised Panoptic Segmentation
Panoptic segmentation aims to partition an image to object instances and semantic content for thing and stuff categories, respectively. To date, learning weakly supervised panoptic segmentation (WSPS) with only image-level labels remains unexplored. In this paper, we propose an efficient jointly thing-and-stuff mining (JTSM) framework for WSPS. To this end, we design a novel mask of interest pooling (MoIPool) to extract fixed-size pixel-accurate feature maps of arbitrary-shape segmentations. MoIPool enables a panoptic mining branch to leverage multiple instance learning (MIL) to recognize things and stuff segmentation in a unified manner. We further refine segmentation masks with parallel instance and semantic segmentation branches via self-training, which collaborates the mined masks from panoptic mining with bottom-up object evidence as pseudo-ground-truth labels to improve spatial coherence and contour localization. Experimental results demonstrate the effectiveness of JTSM on PASCAL VOC and MS COCO. As a by-product, we achieve competitive results for weakly supervised object detection and instance segmentation. This work is a first step towards tackling challenge panoptic segmentation task with only image-level labels.
Read moreSelf-Regulated Learning for Egocentric Video Activity Anticipation.
Future activity anticipation is a challenging problem in egocentric vision. As a standard future activity anticipation paradigm, recursive sequence prediction suffers from the accumulation of errors. To address this problem, we propose a simple and effective Self-Regulated Learning framework, which aims to regulate the intermediate representation consecutively to produce representation that (a) emphasizes the novel information in the frame of the current time-stamp in contrast to previously observed content, and (b) reflects its correlation with previously observed frames. The former is achieved by minimizing a contrastive loss, and the latter can be achieved by a dynamic reweighing mechanism to attend to informative frames in the observed content with a similarity comparison between feature of the current frame and observed frames. The learned final video representation can be further enhanced by multi-task learning which performs joint feature learning on the target activity labels and the automatically detected action and object class tokens. SRL sharply outperforms existing state-of-the-art in most cases on two egocentric video datasets and two third-person video datasets. Its effectiveness is also verified by the experimental fact that the action and object concepts that support the activity semantics can be accurately identified.
Read more