- Research Article
2
- 10.1142/s0129183118500365
Personalized query suggestion based on user behavior
- Apr 01, 2018
- International Journal of Modern Physics C
- Wanyu Chen + 3 more +3
Personalized query suggestion based on user behavior
Search systems in online content platforms are typically biased toward a minority of highly consumed items, reflecting the most common user behavior of navigating toward content that is already familiar and popular. Query suggestions are a powerful tool to support query formulation and to encourage exploratory search and content discovery. However, classic approaches for query suggestions typically rely either on semantic similarity, which lacks diversity and does not reflect user searching behavior, or on a collaborative similarity measure mined from search logs, which suffers from data sparsity and is biased by highly popular queries. In this work, we argue that the task of query suggestion can be modelled as a link prediction task on a heterogeneous graph including queries and documents, enabling Graph Learning methods to effectively generate query suggestions encompassing both semantic and collaborative information. We perform an offline evaluation on an internal Spotify dataset of search logs and on two public datasets, showing that node2vec leads to an accurate and diversified set of results, especially on the large scale real-world data. We then describe the implementation in an instant search scenario and discuss a set of additional challenges tied to the specific production environment. Finally, we report the results of a large scale A/B test involving millions of users and prove that node2vec query suggestions lead to an increase in online metrics such as coverage (+1.42% shown search results pages with suggestions) and engagement (+1.21% clicks), with a specifically notable boost in the number of clicks on exploratory search queries (+9.37%).
Personalized query suggestion based on user behavior
Personalized query suggestion based on user behavior
Attention-based Hierarchical Neural Query Suggestion
Query suggestions help users of a search engine to refine their queries. Previous work on query suggestion has mainly focused on incorporating directly observable features such as query co-occurrence and semantic similarity. The structure of such features is often set manually, as a result of which hidden dependencies between queries and users may be ignored. We propose an AHNQS model that combines a hierarchical structure with a session-level neural network and a user-level neural network to model the short- and long-term search history of a user. An attention mechanism is used to capture user preferences. We quantify the improvements of AHNQS over state-of-the-art RNN-based query suggestion baselines on the AOL query log dataset, with improvements of up to 21.86% and 22.99% in terms of MRR@10 and Recall@10, respectively, over the state-of-the-art; improvements are especially large for short sessions.
Read moreRDQS: A Relevant and Diverse Query Suggestion Generation Framework
Traditional query suggestion methods mainly leverage click-through information to find related queries as recommendations, without considering the semantic relateness between queries. In addition, few studies use click-through distribution in diversifying query suggestions. To address these issues, we propose a novel and effective framework to generate relevant and diversified query suggestions. We combine query semantics and click-through information together to generate query suggestion candidates which are highly relevant to original query , we use click-through distribution to diversify the candidates. We evaluate our method on a large-scale search log dataset of a commercial engine, experimental results indicate that our framework has significantly improved the relevance and diversity of suggested queries by comparing to four baseline methods.
Read moreExploratory Search: Beyond the Query-Response Paradigm
As information becomes more ubiquitous and the demands that searchers have on search systems grow, there is a need to support search behaviors beyond simple lookup. Information seeking is the process or activity of attempting to obtain information in both human and technological contexts. Exploratory search describes an information-seeking problem context that is open-ended, persistent, and multifaceted, and information-seeking processes that are opportunistic, iterative, and multitactical. Exploratory searchers aim to solve complex problems and develop enhanced mental capacities. Exploratory search systems support this through symbiotic human-machine relationships that provide guidance in exploring unfamiliar information landscapes. Exploratory search has gained prominence in recent years. There is an increased interest from the information retrieval, information science, and human-computer interaction communities in moving beyond the traditional turn-taking interaction model support d by major Web search engines, and toward support for human intelligence amplification and information use. In this lecture, we introduce exploratory search, relate it to relevant extant research, outline the features of exploratory search systems, discuss the evaluation of these systems, and suggest some future directions for supporting exploratory search. Exploratory search is a new frontier in the search domain and is becoming increasingly important in shaping our future world. Table of Contents: Introduction / Defining Exploratory Search / Related Work / Features of Exploratory Search Systems / Evaluation of Exploratory Search Systems / Future Directions and concluding Remarks
Read moreFindability → Known-Item Search, Discoverability → Exploratory Search?
I keep confusing findability and discoverability. It seems that findability is often equated to known-item search, and discoverability to exploratory search. Known-item search is compatible with “instant search”, aka search-as-you-type interfaces. Exploratory search is compatible with “autocomplete” (incl. re-spelling, infix matching, synonym substitution, etc.) interfaces.
Read moreAdaptive visualization for exploratory information retrieval
Adaptive visualization for exploratory information retrieval
HMNet: Hybrid Matching Network for Few-Shot Link Prediction
Knowledge graphs (KGs) are widely used in many real-world applications, such as information retrieval, question answering system, and personal recommendation. However, most KGs are suffering from the incompleteness problem. To deal with the task of link prediction, previous knowledge graph embedding methods require numerous reference instances for each relation. It is worth noting that most relations in KGs have only a few reference instances available. Existing works for few-shot link prediction evaluate the authenticity of triplets from a single relation perspective. In this paper, we propose Hybrid Matching Network (HMNet) for few-shot link prediction, evaluating triplets from entity and relation two perspectives. At the entity-aware matching network, HMNet uses attentive inductive embedding layer to aggregate entity features and relation-aware topology, and then provides entity-aware score to implement first perspective evaluation. At the relation-aware matching network, HMNet integrates feature attention mechanism to implement relation perspective evaluation. Experiments on two public datasets indicate that HMNet achieves promising performance in few-shot link prediction.
Read moreSupport and Centrality: Learning Weights for Knowledge Graph Embedding Models
Computing knowledge graph (KG) embeddings is a technique to learn distributional representations for components of a knowledge graph while preserving structural information. The learned embeddings can be used in multiple downstream tasks such as question answering, information extraction, query expansion, semantic similarity, and information retrieval. Over the past years, multiple embedding techniques have been proposed based on different underlying assumptions. The most actively researched models are translation-based which treat relations as translation operations in a shared (or relation-specific) space. Interestingly, almost all KG embedding models treat each triple equally, regardless of the fact that the contribution of each triple to the global information content differs substantially. Many triples can be inferred from others, while some triples are the foundational (basis) statements that constitute a knowledge graph, thereby supporting other triples. Hence, in order to learn a suitable embedding model, each triple should be treated differently with respect to its information content. Here, we propose a data-driven approach to measure the information content of each triple with respect to the whole knowledge graph by using rule mining and PageRank. We show how to compute triple-specific weights to improve the performance of three KG embedding models (TransE, TransR and HolE). Link prediction tasks on two standard datasets, FB15K and WN18, show the effectiveness of our weighted KG embedding model over other more complex models. In fact, for FB15K our TransE-RW embeddings model outperforms models such as TransE, TransM, TransH, and TransR by at least 12.98% for measuring the Mean Rank and at least 1.45% for HIT@10. Our HolE-RW model also outperforms HolE and ComplEx by at least 14.3% for MRR and about 30.4% for HIT@1 on FB15K. Finally, TransR-RW show an improvement over TransR by 3.90% for Mean Rank and 0.87% for HIT@10.
Read morePanacea: Enhancing Graph Learning with Multimodal Semantics for Drug Repositioning
Drug repositioning is a promising approach to discover new therapeutic applications with existing drugs. However, existing Graph Neural Networks (GNNs)-based methods suffer from two limitations: (1) the sparsity of labeled drug-disease associations and high-quality node representations, despite the continuous emergence of biomedical knowledge from multiple modalities; and (2) the over-smoothing issue in GNNs, which causes representations to become indistinguishable, thereby degrading model performance. To address these issues, we propose Panacea, a graph learning framework that leverages multimodal information to enhance drug repositioning performance. Panacea introduces an automated knowledge acquisition and refinement pipeline that collects and encodes multimodal information, including drug molecular structures (SMILES strings), clinical symptom descriptions of diseases, and both homogeneous and heterogeneous biomedical graphs. The resulting multimodal embeddings are aligned via a learnable projection layer, providing strong initial node representations for graph learning. In the subsequent hierarchical graph learning stage, we improve the Graph Isomorphism Network (GIN) with gated mechanisms and residual connections to enhance the network's representation capacity and avoid over-smoothing. The multimodal embeddings and the graph learning module are jointly optimized in an end-toend manner, enabling node representations to fuse graph structure. Experimental results demonstrate that Panacea outperforms state-of-the-art methods, achieving significant improvements in drug repositioning tasks, especially in scenarios with sparse data and insufficient information representation. Code is available at here.
Read moreA semantic similarity computation method for virtual resources in cloud manufacturing environment based on information content
A semantic similarity computation method for virtual resources in cloud manufacturing environment based on information content
Read moreИсследовательский поиск научных статей
It is intuitively clear that the search for scientific publications often has many characteristics of a research search. The purpose of this paper is to formalize this intuitive understanding, explore which research tasks of scientists can be attributed to research search, what approaches exist to solve a research search problem in general, and how they are implemented in specialized search engines for scientists. We researched existing works regarding information seeking behavior of scientists and the special variant of a search called exploratory search. There are several types of search typical for scientists, and we showed that most of them are exploratory. Exploratory search is different to information retrieval and demands special support from search systems. We analyzed seventeen actual search systems for academicians (from Google Scholar, Scopus and Web of Science to ResearchGate) from the exploratory search support aspect. We found that most of them didn’t go far from simple information retrieval and there is a room for further improvements especially in the collaborative search support.
Read moreSwitching sources: A study of people's exploratory search behavior on social media and the web
Searching the Web for information via search engines is a ubiquitous phenomenon and a well‐established field of study in Information Science. Social media sites also continue to evolve and by now have gained enough popularity and momentum to be used as not just vessels for communication with others, but also as important repositories of information. However, it is not clear if the information behavior of users of traditional search engines differ from those performing information searches strictly on social media sites. To address this, we examined data from two user studies on people's exploratory searching behavior: one group only used Web search engines, while the other exclusively used social media sites to search for information. Information search behaviors of both groups regarding exploratory tasks were observed and analyzed through search log and surveys. The results indicate that, while people using social media sites for exploratory search tasks find a smaller quantity and a less diverse set of documents than what they might discover when utilizing traditional Web search engines, they do perceive to end up with more relevant documents. They also report doing less work and feeling less challenged.
Read moreAn end-to-end deep learning model for human activity recognition from highly sparse body sensor data in Internet of Medical Things environment
Recognizing human activity from highly sparse body sensor data is becoming an important problem in Internet of Medical Things (IoMT) industry. As IoMT currently uses batteryless or passive wearable body sensors for activity recognition, the data signals of these sensors have a high sparsity, which means that the time intervals of sensors’ readings are irregular and the number of sensors’ readings in time unit is frequently limited. Therefore, learning activity recognition models from temporally sparse data is challenging. Traditional machine learning techniques cannot be applicable in this scenario due to their focus on a single sensing modality requiring regular sampling rate. In this work, we propose an effective end-to-end deep neural network (DNN) model to recognize human activities from temporally sparse data signals of passive wearable sensors that improve the accuracy rate of human activity recognition. A dropout technique is used in the developed model to deal with sparsity problem and avoiding overfitting problem. In addition, optimization of the proposed DNN model was performed by evaluating a different number of hidden layers. Various experiments were conducted on a public clinical room dataset of sparse data signals to compare the performance of the proposed DNN model with the conventional and other deep learning approaches. The experimental results demonstrate that the proposed DNN model outperforms the existing state-of-the-art methods in terms of lower inference delay and activity recognition accuracy.
Read moreOlio: A Semantic Search Interface for Data Repositories
Search and information retrieval systems are becoming more expressive in interpreting user queries beyond the traditional weighted bag-of-words model of document retrieval. For example, searching for a flight status or a game score returns a dynamically generated response along with supporting, pre-authored documents contextually relevant to the query. In this paper, we extend this hybrid search paradigm to data repositories that contain curated data sources and visualization content. We introduce a semantic search interface, Olio, that provides a hybrid set of results comprising both auto-generated visualization responses and pre-authored charts to blend analytical question-answering with content discovery search goals. We specifically explore three search scenarios - question-and-answering, exploratory search, and design search over data repositories. The interface also provides faceted search support for users to refine and filter the conventional best-first search results based on parameters such as author name, time, and chart type. A preliminary user evaluation of the system demonstrates that Olio’s interface and the hybrid search paradigm collectively afford greater expressivity in how users discover insights and visualization content in data repositories.
Read moreRetrofitting Soft Rules for Knowledge Representation Learning
Recently, a significant number of studies have focused on knowledge graph completion using rule-enhanced learning techniques, supported by the mined soft rules in addition to the hard logic rules. However, due to the difficulty in determining the confidences of the soft rules without the global semantics of knowledge graph such as the semantic relatedness between relations, the knowledge representation may not be optimal, leading to degraded effectiveness in its application to knowledge graph completion tasks. To address this challenge, this paper proposes a retrofit framework that iteratively enhances the knowledge representation and confidences of soft rules. Specifically, the soft rules guide the learning of knowledge representation, and the representation, in turn, provides global semantic of the knowledge graph to optimize the confidences of soft rules. Extensive evaluation shows that our method achieves new state-of-the-art results on link prediction and triple classification tasks, brought by the fine-tuned confidences of soft rules.
Read more