- Research Article
1
- 10.2139/ssrn.3791837
Evaluation Of Recommendation System On A Movie Dataset
- Jan 01, 2020
- SSRN Electronic Journal
- Ronit Malik
Evaluation Of Recommendation System On A Movie Dataset
Recommendation system can predict the ratings of users to items by leveraging machine learning algorithms. The use of recommendation systems is common in e-commerce websites now-a-days. Since enormous amounts of data including users’ click streams, purchase history, demographics, social networking comments and user-item ratings are stored in e-commerce systems databases, the volume of the data is getting bigger at high speed, and the data is sparse. However, the recommendations and predictions must be made in real time, enabling to bring enormous benefits to human beings. Apache spark is well suited for applications which require high speed query of data, transformation and analytics results. Therefore, the recommendation system developed in this research is implemented on Apache Spark. Also, the matrix factorization using Alternating Least Squares (ALS) algorithm which is a type of collaborative filtering is used to solve overfitting issues in sparse data and increases prediction accuracy. The overfitting problem arises in the data as the user-item rating matrix is sparse. In this research a recommendation system for e-commerce using alternating least squares (ALS) matrix factorization method on Apache Spark MLlib is developed. The research shows that the RMSE value is significantly reduced using ALS matrix factorization method and the RMSE is 0.870. Consequently, it is shown that the ALS algorithm is suitable for training explicit feedback data set where users provide ratings for items.
Evaluation Of Recommendation System On A Movie Dataset
Evaluation Of Recommendation System On A Movie Dataset
The Technology of Using the Information - Recommending System to Establish the Point of Contact of the Audience with the Product
Many market areas use recommender systems. Based on the information about a customer, the system can recommend news, articles, concerts, shows, exhibitions, performances, videos, books, games, applications. The article dwells on the basic principles of using recommender systems. The paper presents a review of the main algorithm development: associative (content) rules, collaborative filtering, Singular value decomposition (SVD) and Alternating Least Squares (ALS) algorithms. The authors propose quality metrics, such as Prediction Accuracy for evaluating the prognosis accuracy, Decision support for evaluating the recommendation relevance, and Rank Accuracy - the ranking quality of recommendations issued. The authors took data on users' grocery shopping from 2017 to 2019 (1 million transactions) from one of the major online retailers to examine SVD and ALS algorithms. The computer program was designed in Python using scipy collections, statsmodels, surprise, sklearn, matplotlib, and included data pre-processing and clearing. As a result, it was found out that the ALS model was quicker than SVD when processing large data volumes. SVD is more favorable if the criterion is the accuracy of recommendations.
Read moreChapter 5 - A recommendation system for the prediction of drug–target associations
Chapter 5 - A recommendation system for the prediction of drug–target associations
A real-time recommendation engine using lambda architecture
In a data science theory, the recommended methodology is one of the most popular theories and has been deployed in many real industries. However, one of the most challenging problems these days is how to recommend items with massively streaming data. Therefore, this paper aims to do a real-time recommendation engine using the Lambda architecture. The Apache Hadoop and Apache Spark frameworks were used in this research to process the MovieLens dataset comprised 100 K and 20 M ratings from the GroupLens research. Using alternating least squares (ALS) and k-means algorithms, the top K recommendation movies and the top K trending movies for each user were shown as results. Additionally, the mean squared error (MSE) and within cluster sum of squared error (WCSS) had been computed to evaluate the performance of the ALS and k-means algorithms, sequentially. The results showed that they are acceptable since the MSE and WCSS values are low when comparing to the size of data. However, they can still be improved by tuning some parameters.
Read moreMatrix Factorization Based Recommendation System using Hybrid Optimization Technique
In this paper, a matrix factorization recommendation algorithm is used to recommend items to the user by inculcating a hybrid optimization technique that combines Alternating Least Squares (ALS) and Stochastic Gradient Descent (SGD) in the advanced stage and compares the two individual algorithms with the hybrid model. This hybrid optimization algorithm can be easily implemented in the real world as a cold start can be easily reduced. The hybrid technique proposed is set side-by-side with the ALS and SGD algorithms individually to assess the pros and cons and the requirements to be met to choose a specific technique in a specific domain. The metric used for comparison and evaluation of this technique is Mean Squared Error (MSE).
Read moreSocial trust model for rating prediction in recommender systems: Effects of similarity, centrality, and social ties
Social trust model for rating prediction in recommender systems: Effects of similarity, centrality, and social ties
User Location and Collaborative based Recommender System using Naive Bayes Classifier and UIR Matrix
The world is filled with information and getting the right information is a challenging task for internet users and online buyers. Recommender system helps internet users to get their information in a short span of time. It acts as an information extraction system that works behind users to perform their search easier. The recommender system comes under user’s content or item based search, similar users browsing behavior called collaborative and combination of both known as a hybrid. Here collaborative-based approach is adopted which recommends items to their users based on their past browsing behavior. In this article, the User-Item-Rating matrix is formulated concerning user personal profile, rating of the product, and reviews given by the users during their previous browsing history. In this research, user location is considered as an important attribute to group similar users. It also attempts to suppress the scalability and sparsity problems of the traditional collaborative filtering approach. So, the User-Item-Rating (UIR) matrix has considered the location, ratings and reviews for future recommendation. The Navie Bayes classifier algorithm is used to provide accurate topmost recommendations to internet users. The data set is taken from the MovieLens and IMDb database. The accuracy of the recommender system is measured based on the main metric f-measure. The experimental result has proven the improvement of the recommender system with the mentioned added attributes.
Read moreIncreasing prediction accuracy in collaborative filtering with initialized factor matrices
Recommender systems are useful tools to give personalized recommendations to users. One of the most popular techniques used in these systems is collaborative filtering. Recommender system algorithms get into trouble with data sparsity and scalability. These challenges cause lack of convergence in our algorithms. In this research, we propose a new method based on matrix factorization which alleviates data sparsity. We suggest a new method which can be performed as a preprocessing method for initial latent factor matrices of users and items. Initialized latent factors in matrix factorization lead to two advantages: (1) sparsity and scalability would be covered and (2) convergence of algorithms would be faster. We have shown that our method has improved the accuracy of optimization-based matrix factorization technique. Also it has increased the speed of matrix factorization convergence.
Read moreA genetic algorithms-based hybrid recommender system of matrix factorization and neighborhood-based techniques
A genetic algorithms-based hybrid recommender system of matrix factorization and neighborhood-based techniques
Heterogeneous Graph Embedding for Cross-Domain Recommendation Through Adversarial Learning
Cross-domain recommendation is critically important to construct a practical recommender system. The challenges of building a cross-domain recommender system lie in both the data sparsity issue and lacking of sufficient semantic information. Traditional approaches focus on using the user-item rating matrix or other feedback information, but the contents associated with the objects like reviews and the relationships among the objects are largely ignored. Although some works merge the content information and the user-item rating network structure, they only focus on using the attributes of the items but ignore user generated contents such as reviews. In this paper, we propose a novel cross-domain recommender framework called ECHCDR (Embedding content and heterogeneous network for cross-domain recommendation), which contains two major steps of content embedding and heterogeneous network embedding. By considering the contents of objects and their relationships, ECHCDR can effectively alleviate the data sparsity issue. To enrich the semantic information, we construct a weighted heterogeneous network whose nodes are users and items of different domains. The weight of link is defined by an adjacency matrix and represents the similarity between users, books and movies. We also propose to use adversarial training method to learn the embeddings of users and cross-domain items in the constructed heterogeneous graph. Experimental results on two real-world datasets collected from Amazon show the effectiveness of our approach compared with state-of-art recommender algorithms.
Read moreBlock-Term Tensor Decomposition Via Constrained Matrix Factorization
In this work, we consider the problem of factoring a third-order tensor into multilinear rank -($L_{r}, L_{r}$, 1) terms. This model, referred to as the rank-($L_{r}, L_{r}$, 1) block-term decomposition (BTD), finds many applications in signal processing, especially blind separation of smooth sources and unmixing spectral-spatial data (e.g., hyper-spectral image). On the other hand, finding latent factors of rank-($L_{r}, L_{r}$, 1) BTD poses a very challenging optimization problem. Some computational tools designed for canonical polyadic decomposition (CPD) (e.g., alternating least squares (ALS) and Levenberg-Marquardt (LM) based algorithms) can be modified to handle rank-($L_{r}, L_{r}$, 1) BTD. Nonetheless, these methods essentially treat rank-($L_{r}, L_{r}$, 1) BTD as a special CPD problem. This raises a number of challenges, since rank -($L_{r}, L_{r}$, 1) BTD can be viewed as a CPD problem with rank-deficient latent factors and high CP rank–and these are known as hard cases for CPD algorithms. In this work, we reformulate the rank-($L_{r}, L_{r}$, 1) BTD problem as a matrix rank-constrained matrix factorization problem. We propose a simple algorithm that combines alternating optimization and projected gradient. This way, the per-iteration complexity is much smaller than those of the ALS and LM-based algorithms. We also show that the algorithm converges to a stationary point of the problem of interest at a sublinear rate–although nonconvex constraints are involved. Numerical experiments are conducted to showcase the effectiveness of our algorithm.
Read moreFine-grained Sentiment-enhanced Collaborative Filtering-based Hybrid Recommender System
Developing online educational platforms necessitates the incorporation of new intelligent procedures in order to improve long-term student experience. Presently, e-learning recommender systems rely on deep learning methods to recommend appropriate e-learning materials to the students based on their learner profiles. Fine-grained sentiment analysis (FSA) can be leveraged to enrich the recommender system. User-posted reviews and rating data are vital in accurately directing the student to the appropriate e-learning resources based on posted comments by comparable learners. In this work, a new e-learning recommendation system is proposed based on individualization and FSA. A hybrid framework is provided by integrating alternating least square (ALS) based collaborative filtering (CF) with FSA to generate an effective e-content recommendation named HCFSAR. ALS attempts to capture the learner’s latent factors based on their selections of interest to build the learner profile. Three FSA models based on attention mechanisms and bidirectional long short-term memory (bi-LSTM) are suggested and used to train twelve models in order to predict new ratings from learner-posted book reviews based on the extracted learner profile. HCFSAR used multiplication word embeddings for stronger corpus representation that were trained on a dataset generated for an educational context and showed a better accuracy of 93.39% for the best model entitled MHAM based ABHR-2 with multiplication (MHAAM), which performed better than other models. A tailored dataset that has been created by scraping reviews of different e-learning resources is leveraged to train different proposed models and validate against public datasets.
Read moreMatrix Factorization With Rating Completion: An Enhanced SVD Model for Collaborative Filtering Recommender Systems
Collaborative filtering algorithms, such as matrix factorization techniques, are recently gaining momentum due to their promising performance on recommender systems. However, most collaborative filtering algorithms suffer from data sparsity. Active learning algorithms are effective in reducing the sparsity problem for recommender systems by requesting users to give ratings to some items when they enter the systems. In this paper, a new matrix factorization model, called Enhanced SVD (ESVD) is proposed, which incorporates the classic matrix factorization algorithms with ratings completion inspired by active learning. In addition, the connection between the prediction accuracy and the density of matrix is built to further explore its potentials. We also propose the Multi-layer ESVD, which learns the model iteratively to further improve the prediction accuracy. To handle the imbalanced data sets that contain far more users than items or more items than users, the Item-wise ESVD and User-wise ESVD are presented, respectively. The proposed methods are evaluated on the famous Netflix and Movielens data sets. Experimental results validate their effectiveness in terms of both accuracy and efficiency when compared with traditional matrix factorization methods and active learning methods.
Read moreCoupled Item-Based Matrix Factorization
The essence of the challenges cold start and sparsity in Recommender Systems (RS) is that the extant techniques, such as Collaborative Filtering (CF) and Matrix Factorization (MF), mainly rely on the user-item rating matrix, which sometimes is not informative enough for predicting recommendations. To solve these challenges, the objective item attributes are incorporated as complementary information. However, most of the existing methods for inferring the relationships between items assume that the attributes are “independently and identically distributed (iid)”, which does not always hold in reality. In fact, the attributes are more or less coupled with each other by some implicit relationships. Therefore, in this paper we propose an attribute-based coupled similarity measure to capture the implicit relationships between items. We then integrate the implicit item coupling into MF to form the Coupled Item-based Matrix Factorization (CIMF) model. Experimental results on two open data sets demonstrate that CIMF outperforms the benchmark methods.
Read moreDetect User’s Rating Characteristics by Separate Scores for Matrix Factorization Technique
A recommender system can effectively solve the problem of information overload in the era of big data. Recent research on recommender systems, specifically Collaborative Filtering, has focused on Matrix Factorization methods, which have been shown to have excellent performance. However, these methods do not pay attention to the influence of a user’s rating characteristics, which are especially important for the accuracy of prediction or recommendation. Therefore, in order to get better performance, we propose a novel method based on matrix factorization. We consider that the user’s rating score is composed of two parts: the real score, which is decided by the user’s preferences; and the bias score, which is decided by the user’s rating characteristics. We then analyze the user’s historical behavior to find his rating characteristics by using the matrix factorization technique and use them to adjust the final prediction results. Finally, by comparing with the latest algorithms on the open datasets, we verified that the proposed method can significantly improve the accuracy of recommender systems and achieve the best performance in terms of prediction accuracy criterion over other state-of-the-art methods.
Read more