- Dissertation
- 10.32657/10356/181934
Towards effective neural topic modeling
- Jan 01, 2024
- Xiaobao Wu
Over the past few decades, the world has witnessed an unprecedented explosion of information. Of these, a substantial portion consists of unlabeled textual data, such as tweets, news articles, product reviews, and web snippets. \nAs labeling is extremely expensive, time-consuming, and sometimes biased, \nhow to effectively analyze these data becomes an imperative. \nOwing to this, Neural Topic Models (NTMs) have emerged as a promising solution that attracts considerable research attention for their capability and interpretability. \nThey automatically discover latent topics from unlabeled textual data through neural networks, enabling unsupervised document understanding. \nThey have derived various downstream applications, such as content recommendation, trend analysis, and text summarization. \nCompared to conventional topic models like LDA, NTMs offer structural flexibility and also support gradient back-propagation, \nwhich avoids complicated model-specific derivations and well handle large-scale data. \n \nHowever, despite their promise, existing NTMs generally encounter several critical challenges. On the one hand, NTMs often produce low-quality topics, which are even incomparable to conventional models. \nThese topics tend to be incoherent or repetitive, significantly diminishing their informativeness. \nOn the other hand, NTMs struggle with low inference ability, leading to less accurate topic distributions for documents. \nThis limitation greatly hinders document understanding and undermines the subsequent analysis, reasoning, or decision making processes. \nDue to these challenges, existing NTMs are less useful and applicable for downstream tasks or applications. \nAs a result, it is necessary to enhance NTMs to deliver more reliable and informative topic modeling. \n \nThis thesis aims to advance neural topic modeling by addressing these key challenges. \nIn particular, we focus on four most popular scenarios: short-texts, cross-lingual, hierarchical, and basic neural topic modeling. \nFirst, we propose a novel neural topic model tailored for short texts. \nThis model leverages a topic-semantic contrastive learning method to captures the similarity relations among short text samples, which works regardless of data augmentation availability. \nThis refines short text representations, enriches learning signals, thus effectively alleviating the data sparsity issue of short texts and producing informative topics. \n \nSecond, we explore dynamic topic modeling. \nWe present a neural dynamic topic model to track the evolution of dynamic topics by building the contrastive relations among them, rather than relying on the conventional Markov Chains. This model further explicitly excludes unassociated words from dynamic topics to enhance the alignment to their respective time slices. \nThese improvements enable our model to reliably track topic evolution with high diversity. \n \nThird, we focus on the the affinity, rationality, and diversity of hierarchical topic modeling. \nOur proposed neural hierarchical topic model ensures the sparsity and balance of cross-level topic dependencies using a transport plan dependency method. \nMoreover, it distributes different semantic granularity to topics at different levels by disentangled decoding. \nWith these, our model produces affinitive, rational, and diverse topic hierarchies. \n \nFourth, we tackle the topic collapsing issue and propose a basic neural topic model. \nApart from the common reconstruction error, this model introduces a new embedding clustering regularization to force each topic embedding to be the center of a separately aggregated word embedding cluster in the semantic space. \nThis produces topics with distinct semantics and effectively resolves the topic collapsing issue. \n \nFinally, we develop a comprehensive topic modeling toolkit that includes both previous conventional and cutting-edge neural topic models. \nThis toolkit covers the complete pipelines of topic modeling, such as dataset preprocessing, model training, and evaluation. \nThese improvements position our toolkit as a valuable resource to accelerate the research and applications of topic models. \n \nIn conclusion, this thesis makes significant contributions towards advancing neural topic modeling through multiple models, various scenarios, and a rigorous and comprehensive benchmark toolkit. \nThese contributions pave the way for the utilization of neural topic modeling in diverse real-world applications.
Read more