- Research Article
1
- 10.2139/ssrn.3791837
Evaluation Of Recommendation System On A Movie Dataset
- Jan 01, 2020
- SSRN Electronic Journal
- Ronit Malik
Evaluation Of Recommendation System On A Movie Dataset
Recommender systems are essential engines to deliver product recommendations for e-commerce businesses. Successful adoption of recommender systems could significantly influence the growth of marketing targets. Collaborative filtering is a type of recommender system model that uses customers' activities in the past, such as ratings. Unfortunately, the number of ratings collected from customers is sparse, amounting to less than 4%. The latent factor model is a kind of collaborative filtering that involves matrix factorization to generate rating predictions. However, using only matrix factorization would result in an inaccurate recommendation. Several models include product review documents to increase the effectiveness of their rating prediction. Most of them use methods such as TF-IDF and LDA to interpret product review documents. However, traditional models such as LDA and TF-IDF face some shortcomings, in that they show a less contextual understanding of the document. This research integrated matrix factorization and novel models to interpret and understand product review documents using LSTM and word embedding. According to the experiment report, this model significantly outperformed the traditional latent factor model by more than 16% on an average and achieved 1% on an average based on RMSE evaluation metrics, compared to the previous best performance. Contextual insight of the product review document is an important aspect to improve performance in a sparse rating matrix. In the future work, generating contextual insight using bidirectional word sequential is required to increase the performance of e-commerce recommender systems with sparse data issues.
Evaluation Of Recommendation System On A Movie Dataset
Evaluation Of Recommendation System On A Movie Dataset
Social Network Analysis for the Effective Adoption of Recommender Systems
Recommender system is the system which, by using automated information filtering technology, recommends products or services to the customers who are likely to be interested in. Those systems are widely used in many different Web retailers such as Amazon.com, Netfix.com, and CDNow.com. Various recommender systems have been developed. Among them, Collaborative Filtering (CF) has been known as the most successful and commonly used approach. CF identifies customers whose tastes are similar to those of a given customer, and recommends items those customers have liked in the past. Numerous CF algorithms have been developed to increase the performance of recommender systems. However, the relative performances of CF algorithms are known to be domain and data dependent. It is very time-consuming and expensive to implement and launce a CF recommender system, and also the system unsuited for the given domain provides customers with poor quality recommendations that make them easily annoyed. Therefore, predicting in advance whether the performance of CF recommender system is acceptable or not is practically important and needed. In this study, we propose a decision making guideline which helps decide whether CF is adoptable for a given application with certain transaction data characteristics. Several previous studies reported that sparsity, gray sheep, cold-start, coverage, and serendipity could affect the performance of CF, but the theoretical and empirical justification of such factors is lacking. Recently there are many studies paying attention to Social Network Analysis (SNA) as a method to analyze social relationships among people. SNA is a method to measure and visualize the linkage structure and status focusing on interaction among objects within communication group. CF analyzes the similarity among previous ratings or purchases of each customer, finds the relationships among the customers who have similarities, and then uses the relationships for recommendations. Thus CF can be modeled as a social network in which customers are nodes and purchase relationships between customers are links. Under the assumption that SNA could facilitate an exploration of the topological properties of the network structure that are implicit in transaction data for CF recommendations, we focus on density, clustering coefficient, and centralization which are ones of the most commonly used measures to capture topological properties of the social network structure. While network density, expressed as a proportion of the maximum possible number of links, captures the density of the whole network, the clustering coefficient captures the degree to which the overall network contains localized pockets of dense connectivity. Centralization reflects the extent to which connections are concentrated in a small number of nodes rather than distributed equally among all nodes. We explore how these SNA measures affect the performance of CF performance and how they interact to each other. Our experiments used sales transaction data from H department store, one of the well?known department stores in Korea. Total 396 data set were sampled to construct various types of social networks. The dependant variable measuring process consists of three steps; analysis of customer similarities, construction of a social network, and analysis of social network patterns. We used UCINET 6.0 for SNA. The experiments conducted the 3-way ANOVA which employs three SNA measures as dependant variables, and the recommendation accuracy measured by F1-measure as an independent variable. The experiments report that 1) each of three SNA measures affects the recommendation accuracy, 2) the density's effect to the performance overrides those of clustering coefficient and centralization (i.e., CF adoption is not a good decision if the density is low), and 3) however though the density is low, the performance of CF is comparatively good when the clustering coefficient is low. We expect that these experiment results help firms decide whether CF recommender system is adoptable for their business domain with certain transaction data characteristics.
Read moreMatrix Factorization With Rating Completion: An Enhanced SVD Model for Collaborative Filtering Recommender Systems
Collaborative filtering algorithms, such as matrix factorization techniques, are recently gaining momentum due to their promising performance on recommender systems. However, most collaborative filtering algorithms suffer from data sparsity. Active learning algorithms are effective in reducing the sparsity problem for recommender systems by requesting users to give ratings to some items when they enter the systems. In this paper, a new matrix factorization model, called Enhanced SVD (ESVD) is proposed, which incorporates the classic matrix factorization algorithms with ratings completion inspired by active learning. In addition, the connection between the prediction accuracy and the density of matrix is built to further explore its potentials. We also propose the Multi-layer ESVD, which learns the model iteratively to further improve the prediction accuracy. To handle the imbalanced data sets that contain far more users than items or more items than users, the Item-wise ESVD and User-wise ESVD are presented, respectively. The proposed methods are evaluated on the famous Netflix and Movielens data sets. Experimental results validate their effectiveness in terms of both accuracy and efficiency when compared with traditional matrix factorization methods and active learning methods.
Read moreJoint latent factors and attributes to discover interpretable preferences in recommendation
Joint latent factors and attributes to discover interpretable preferences in recommendation
Multi-criteria collaborative filtering recommender by fusing deep neural network and matrix factorization
Recommender systems have been an efficient strategy to deal with information overload by producing personalized predictions. Recommendation systems based on deep learning have accomplished magnificent results, but most of these systems are traditional recommender systems that use a single rating. In this work, we introduce a multi-criteria collaborative filtering recommender by combining deep neural network and matrix factorization. Our model consists of two parts: the first part uses a fused model of deep neural network and matrix factorization to predict the criteria ratings and the second one employs a deep neural network to predict the overall rating. The experimental results on two datasets, including a real-world dataset, show that the proposed model outperformed several state-of-the-art methods across different datasets and performance evaluation metrics.
Read moreProviding reliability in recommender systems through Bernoulli Matrix Factorization
Providing reliability in recommender systems through Bernoulli Matrix Factorization
A New Recommender System for 3D E-Commerce: An EEG Based Approach
This position paper discusses a novel recommender system for e-commerce in virtual reality environments. The system provides recommendations by taking into account prepurchase ratings in addition to traditional postpurchase ratings. Users' positive emotions are captured in the form of electroencephalogram (EEG) signals while interacting with 3D virtual products prior to purchase. The prepurchase ratings are calculated from the averaged relative power of the collected EEG signals. Prepurchase ratings are complementary to postpurchase ratings and help in alleviating two severe issues that traditional recommender systems suffer from: data sparsity and cold start. By making proper use of both pre- and postpurchase ratings, user preference can be modeled more accurately. This will improve the effectiveness of the current recommender systems and may change the traditional e- business applications.
Read moreDeep learning techniques for recommender systems based on collaborative filtering
In the Big Data Era, recommender systems perform a fundamental role in data management and information filtering. In this context, Collaborative Filtering (CF) persists as one of the most prominent strategies to effectively deal with large datasets and is capable of offering users interesting content in a recommendation fashion. Nevertheless, it is well‐known CF recommenders suffer from data sparsity, mainly in cold‐start scenarios, substantially reducing the quality of recommendations. In the vast literature about the aforementioned topic, there are numerous solutions, in which the state‐of‐the‐art contributions are, in some sense, conditioned or associated with traditional CF methods such as Matrix Factorization (MF), that is, they rely on linear optimization procedures to model users and items into low‐dimensional embeddings. To overcome the aforementioned challenges, there has been an increasing number of studies exploring deep learning techniques in the CF context for latent factor modelling. In this research, authors conduct a systematic review focusing on state‐of‐the‐art literature on deep learning techniques applied in collaborative filtering recommendation, and also featuring primary studies related to mitigating the cold start problem. Additionally, authors considered the diverse non‐linear modelling strategies to deal with rating data and side information, the combination of deep learning techniques with traditional CF‐based linear methods, and an overview of the most used public datasets and evaluation metrics concerning CF scenarios.
Read moreAN OPTIMAL SCALING FRAMEWORK FOR COLLABORATIVE FILTERING RECOMMENDATION SYSTEMS
Collaborative Filtering (CF) is a popular technique employed by Recommender Systems, a term used to describe intelligent methods that generate personalized recommendations. Some of the most efficient approaches to CF are based on latent factor models and nearest neighbor methods, and have received considerable attention in recent literature. Latent factor models can tackle some fundamental challenges of CF, such as data sparsity and scalability. In this work, we present an optimal scaling framework to address these problems using Categorical Principal Component Analysis (CatPCA) for the low-rank approximation of the user-item ratings matrix, followed by a neighborhood formation step. CatPCA is a versatile technique that utilizes an optimal scaling process where original data are transformed so that their overall variance is maximized. We considered both smooth and non-smooth transformations for the observed variables (items), such as numeric, (spline) ordinal, (spline) nominal and multiple nominal. The method was extended to handle missing data and incorporate differential weighting for items. Experiments were executed on three data sets of different sparsity and size, MovieLens 100k, 1M and Jester, aiming to evaluate the aforementioned options in terms of accuracy. A combined approach with a multiple nominal transformation and a "passive" missing data strategy clearly outperformed the other tested options for all three data sets. The results are comparable with those reported for single methods in the CF literature.
Read moreA Survey of Collaborative Filtering Recommender Algorithms and Their Evaluation Metrics
Abstract—Recommender systems are often used to provide useful recommendations for users. They use previous history of the users-items interactions, e.g. purchase history and/or users rating on items, to provide a suitable recommendation list for any target user. They may also use contextual information available about items and users. Collaborative filtering algorithm and its variants are the most successful recommendation algorithms that have been applied to many applications. Collaborative filtering method works by first finding the most similar users (or items) for a target user (or items), and then building the recommendation lists. There is no unique evaluation metric to assess the performance of recommendations systems, and one often choose the one most appropriate for the application in hand. This paper compares the performance of a number of well-known collaborative filtering algorithms on movie recommendation. To this end, a number of performance criteria are used to test the algorithms. The algorithms are ranked for each evaluation metric and a rank aggregation method is used to determine the wining algorithm. Our experiments show that the probabilistic matrix factorization has the top performance in this dataset, followed by item-based and user-based collaborative filtering. Non-negative matrix factorization and Slope 1 has the worst performance among the considered algorithms.
 Keywords—Social networks analysis and mining, big data, recommender systems, collaborative filtering.
Read moreINMO
Collaborative filtering is one of the most common scenarios and popular\nresearch topics in recommender systems. Among existing methods, latent factor\nmodels, i.e., learning a specific embedding for each user/item by\nreconstructing the observed interaction matrix, have shown excellent\nperformances. However, such user-specific and item-specific embeddings are\nintrinsically transductive, making it difficult to deal with new users and new\nitems unseen during training. Besides, the number of model parameters heavily\ndepends on the number of all users and items, restricting its scalability to\nreal-world applications. To solve the above challenges, in this paper, we\npropose a novel model-agnostic and scalable Inductive Embedding Module for\ncollaborative filtering, namely INMO. INMO generates the inductive embeddings\nfor users (items) by characterizing their interactions with some template items\n(template users), instead of employing an embedding lookup table. Under the\ntheoretical analysis, we further propose an effective indicator for the\nselection of template users/items. Our proposed INMO can be attached to\nexisting latent factor models as a pre-module, inheriting the expressiveness of\nbackbone models, while bringing the inductive ability and reducing model\nparameters. We validate the generality of INMO by attaching it to both Matrix\nFactorization (MF) and LightGCN, which are two representative latent factor\nmodels for collaborative filtering. Extensive experiments on three public\nbenchmarks demonstrate the effectiveness and efficiency of INMO in both\ntransductive and inductive recommendation scenarios.\n
Read moreAlign Reviews with Topics in Attention Network for Rating Prediction
Rating prediction has long been a hot research topic in recommendation systems. Latent factor models, in particular, matrix factorization (MF), are the most prevalent techniques for rating prediction. However, MF based methods suffer from the problem of data sparsity and lack of explanation. In this paper, we present a novel model to address these problems by integrating ratings and topic-level review information into a deep neural framework. Our model can capture the varying attentions that a review contributes to a user/item at the topic level. We conduct extensive experiments on three datasets from Amazon. Results demonstrate our proposed method consistently outperforms the state-of-the-art recommendation approaches.
Read moreBook Recommendation System Based on Collaborative Filtering: User-Based, Item-Based, and Singular Value Decomposition Analysis
Recommender systems have become essential in the digital era to help users navigate overwhelming content. This study develops a book recommendation system using three collaborative filtering methods: user-based, item-based, and matrix factorization using singular value decomposition. We evaluate the system on a real-world dataset of 1,149,780 book ratings from 278,858 users across 271,360 books. A subset of 500 active users is used for experimental evaluation. The models are assessed using root mean square error and mean absolute error to measure rating prediction accuracy. The results show that the item-based collaborative filtering method achieves the best accuracy (root mean square error 7.362; mean absolute error 6.761), slightly outperforming the user-based approach (7.365; 6.809) and the matrix factorization method (7.643; 7.413). We analyze the results to understand the performance differences, noting the stability of item similarity as a key factor and the need for optimal tuning in the matrix factorization model. In conclusion, item-based collaborative filtering proved most effective for this context. This work provides insights into the comparative performance of foundational recommendation techniques and highlights practical considerations for improving book recommender systems.
Read moreRecent Advances in Recommender Systems
Recommender systems are designed to identify the items that a user will like or find useful based on the user's prior preferences and activities. These systems have become ubiquitous and are an essential tool for information filtering and (e-)commerce. Over the years, collaborative filtering, which derive these recommendations by leveraging past activities of groups of users, has emerged as the most prominent approach for solving this problem. This talk will present some of our recent work towards improving the performance of collaborative filtering-based recommender systems and understanding some of their fundamental limitations and characteristics. It will start by analyzing how the ratings that users provide to a set of items relate to their ratings of the set's individual items and, using these insights, will present rating prediction approaches that utilize distant supervision. It will then discuss extensions to approaches based on sparse linear and latent factor models that postulate that users' preferences are a combination of global and local preferences, which are shown to lead to better user modeling and as such improved prediction performance. Finally, the talk will conclude by discussing what can be accurately predicted by latent factor approaches and by analyzing the estimation error of sparse linear and latent factor models and how its characteristics impacts the performance of top N recommendation algorithms.
Read moreA genetic algorithms-based hybrid recommender system of matrix factorization and neighborhood-based techniques
A genetic algorithms-based hybrid recommender system of matrix factorization and neighborhood-based techniques