• Home
  • Search
  • Iterative NLP Query Refinement for Enhancing Domain-Specific Information Retrieval: A Case Study in Career Services
  • https://doi.org/10.48550/arxiv.2412.17075Copy DOI Icon

Iterative NLP Query Refinement for Enhancing Domain-Specific Information Retrieval: A Case Study in Career Services

Show More
  • Abstract
  • Literature Map
  • Similar Papers
Abstract

Retrieving semantically relevant documents in niche domains poses significant challenges for traditional TF-IDF-based systems, often resulting in low similarity scores and suboptimal retrieval performance. This paper addresses these challenges by introducing an iterative and semi-automated query refinement methodology tailored to Humber College's career services webpages. Initially, generic queries related to interview preparation yield low top-document similarities (approximately 0.2--0.3). To enhance retrieval effectiveness, we implement a two-fold approach: first, domain-aware query refinement by incorporating specialized terms such as resources-online-learning, student-online-services, and career-advising; second, the integration of structured educational descriptors like "online resume and interview improvement tools." Additionally, we automate the extraction of domain-specific keywords from top-ranked documents to suggest relevant terms for query expansion. Through experiments conducted on five baseline queries, our semi-automated iterative refinement process elevates the average top similarity score from approximately 0.18 to 0.42, marking a substantial improvement in retrieval performance. The implementation details, including reproducible code and experimental setups, are made available in our GitHub repositories \url{https://github.com/Elipei88/HumberChatbotBackend} and \url{https://github.com/Nisarg851/HumberChatbot}. We also discuss the limitations of our approach and propose future directions, including the integration of advanced neural retrieval models.

Similar Papers
  • Research Article
  • Citations45

Retrieval, alignment, and clustering of computational models based on semantic annotations

  • Jan 01, 2011
  • Molecular Systems Biology
  • Marvin Schulz +4
  • Research Article
  • Citations47

Exploiting the category structure of Wikipedia for entity ranking

  • Jun 19, 2012
  • Artificial Intelligence
  • Rianne Kaptein +1
  • Conference Article
  • Citations1

Image indexing and Retrieval using Error Diffusion BTC with relevance feedback

  • Jan 01, 2016
  • Meharban M S +1
  • Conference Article
  • Citations2

Causality Inspired Retrieval of Human-object Interactions from Video

  • Sep 01, 2019
  • Liting Zhou +4
  • Research Article
  • Citations32

Bayesian belief networks for IR

  • Sep 04, 2003
  • International Journal of Approximate Reasoning
  • Marco Antônio Pinheiro De Cristo +5
  • Book Chapter
  • Citations5

A Method for Query Expansion Using a Hierarchy of Clusters

  • Jan 01, 2005
  • Masaki Aono +1
  • Research Article
  • Citations12

Query specific graph-based query reformulation using UMLS for clinical information access.

  • Jun 25, 2020
  • Journal of Biomedical Informatics
  • Jainisha Sankhavara +3
  • Book Chapter
  • Citations10

Shape-Based Retrieval of Articulated 3D Models Using Spectral Embedding

  • Jan 01, 2006
  • Varun Jain +1
  • Book Chapter
  • Citations28

On Improving Pseudo-Relevance Feedback Using Pseudo-Irrelevant Documents

  • Jan 01, 2010
  • Karthik Raman +3
  • Research Article
  • Citations5

SEMANTIC GRAPH KNOWLEDGE REPRESENTATION FOR AL-QURAN VERSES BASED ON WORD DEPENDENCIES

  • Dec 31, 2021
  • Malaysian Journal of Computer Science
  • Muhammad Muhtadi Mohamad Khazani +8
  • Book Chapter
  • Citations4

Rebuilding Visual Vocabulary via Spatial-temporal Context Similarity for Video Retrieval

  • Jan 01, 2014
  • Lei Wang +2
  • PDF
  • Research Article
  • Citations10

Developing Two Different Novel Techniques for Arabic Text Stemming

  • Jan 01, 2019
  • Intelligent Information Management
  • Mohammad Mustafa +4
  • Book Chapter
  • Citations1

A Simple WordNet-Ontology Based Email Retrieval System for Digital Forensics

  • Jan 01, 2008
  • Phan Thien Son +5
  • Book Chapter
  • Citations16

Exploiting User Comments for Audio-Visual Content Indexing and Retrieval

  • Jan 01, 2013
  • Carsten Eickhoff +2
  • Conference Article
  • Citations12

Exploiting redundancy in cross-channel video retrieval

  • Sep 24, 2007
  • Bouke Huurnink +1
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.