• Home
  • Search
  • Snippext: Semi-supervised Opinion Mining with Augmented Data
  • Open Access IconOpen Access
  • Cite Icon81
  • https://doi.org/10.1145/3366423.3380144Copy DOI Icon

Snippext: Semi-supervised Opinion Mining with Augmented Data

  • Apr 20, 2020
  • Zhengjie Miao +3 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Online services are interested in solutions to opinion mining, which is the problem of extracting aspects, opinions, and sentiments from text. One method to mine opinions is to leverage the recent success of pre-trained language models which can be fine-tuned to obtain high-quality extractions from reviews. However, fine-tuning language models still requires a non-trivial amount of training data. In this paper, we study the problem of how to significantly reduce the amount of labeled training data required in fine-tuning language models for opinion mining. We describe Snippext, an opinion mining system developed over a language model that is fine-tuned through semi-supervised learning with augmented data. A novelty of Snippext is its clever use of a two-prong approach to achieve state-of-the-art (SOTA) performance with little labeled training data through: (1) data augmentation to automatically generate more labeled training data from existing ones, and (2) a semi-supervised learning technique to leverage the massive amount of unlabeled data in addition to the (limited amount of) labeled data. We show with extensive experiments that Snippext performs comparably and can even exceed previous SOTA results on several opinion mining tasks with only half the training data required. Furthermore, it achieves new SOTA results when all training data are leveraged. By comparison to a baseline pipeline, we found that Snippext extracts significantly more fine-grained opinions which enable new opportunities of downstream applications.

Similar Papers
  • Research Article
  • Citations19

In-House Knowledge Management Using a Large Language Model: Focusing on Technical Specification Documents Review

  • Mar 02, 2024
  • Applied Sciences
  • Jooyeup Lee +2
  • Research Article
  • Citations2

Optimizing BERT Models with Fine-Tuning for Indonesian Twitter Sentiment Analysis

  • Jun 30, 2025
  • Journal of Wireless Mobile Networks, Ubiquitous Computing, and Dependable Applications
  • Lasmedi Afuan +5
  • Book Chapter
  • Citations1

Opinion Mining and Information Retrieval

  • Jan 01, 2011
  • Shishir K Shandilya +1
  • Conference Article
  • Citations5

A multi-granularity data augmentation based fusion neural network model for short text sentiment analysis

  • Oct 01, 2017
  • Xiao Sun +2
  • Research Article
  • Citations2

Toward Cross-Hospital Deployment of Natural Language Processing Systems: Model Development and Validation of Fine-Tuned Large Language Models for Disease Name Recognition in Japanese

  • Jul 08, 2025
  • JMIR Medical Informatics
  • Seiji Shimizu +4
  • Research Article
  • Citations77

Bangla-BERT: Transformer-Based Efficient Model for Transfer Learning and Language Understanding

  • Jan 01, 2022
  • IEEE Access
  • M Kowsher +5
  • Research Article

Semi-Supervised Learning vs. Few-Shot Learning: Which is Better for Sentiment Analysis on Hotel Reviews Towards a Small Labeled Training Data?

  • Jan 01, 2025
  • International Journal of Advanced Computer Science and Applications
  • Retno Kusumaningrum +5
  • Research Article
  • Citations8

A common convolutional neural network model to classify plain, rolled and latent fingerprints

  • Jan 01, 2019
  • International Journal of Biometrics
  • Asif Iqbal Khan +1
  • Conference Article

Secret Point Recognition Algorithm via Test-Time Augmentation Based on Large Language Models

  • Jul 16, 2025
  • Zhendong Wu +3
  • PDF
  • Research Article
  • Citations12

Leveraging Chain-of-Thought to Enhance Stance Detection with Prompt-Tuning

  • Feb 13, 2024
  • Mathematics
  • Daijun Ding +5
  • Conference Article
  • Citations10

Better Synthetic Data by Retrieving and Transforming Existing Datasets

  • Jan 01, 2024
  • Saumya Gandhi +4
  • Research Article
  • Citations9

PreCurious: How Innocent Pre-Trained Language Models Turn into Privacy Traps.

  • Dec 02, 2024
  • Conference on Computer and Communications Security : proceedings of the ... conference on computer and communications security. ACM Conference on Computer and Communications Security
  • Ruixuan Liu +3
  • Research Article
  • Citations14

The Comparison of Language Models with a Novel Text Filtering Approach for Turkish Sentiment Analysis

  • Dec 27, 2022
  • ACM Transactions on Asian and Low-Resource Language Information Processing
  • Zekeriya Anil Guven
  • Conference Article
  • Citations3

Chinese-Korean Weibo Sentiment Classification Based on Pre-trained Language Model and Transfer Learning

  • May 06, 2022
  • Hengxuan Wang +3
  • PDF
  • Research Article
  • Citations10

AGI-P: A Gender Identification Framework for Authorship Analysis Using Customized Fine-Tuning of Multilingual Language Model

  • Jan 01, 2024
  • IEEE Access
  • Raheem Sarwar +6
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.