- Research Article
1
- 10.2139/ssrn.3791837
Evaluation Of Recommendation System On A Movie Dataset
- Jan 01, 2020
- SSRN Electronic Journal
- Ronit Malik
Evaluation Of Recommendation System On A Movie Dataset
Along with the rapid rise of the internet, an e-commerce website brings enormous benefits for both customers and vendors. However, many choices are given at the same time makes customers have difficulty in choosing the most suitable products. A rising star solution for this is the recommender system which helps to narrow down the amount of suitable and relevant products for each customer. Matrix factorization is one of the most popular techniques used in recommender systems because of its effectiveness and simplicity. In this paper, we introduce a matrix factorization-based recommender system using Singular Value Decomposition (SVD) with some improvements in collaborative filtering and incremental learning. The SVD-based collaborative filtering methods can help generate personalized recommendations by combining user profiles. Moreover, the recommendation lists generated by the system are enhanced with diversity, which might attract more customer interests. Amazon’s Electronic data set is used to evaluate our proposed framework of the SVD-based recommender system. The experimental results show that our framework is promising.
Evaluation Of Recommendation System On A Movie Dataset
Evaluation Of Recommendation System On A Movie Dataset
A Tech Hybrid-Recommendation Engine and Personalized Notification: An integrated tool to assist users through Recommendations (Project ATHENA)
Purpose – Project ATHENA aims to develop an application to address information overload, primarily focused on Recommendation Systems (RSs) with the personalization and user experience design of a modern system. Method – Two machine learning (ML) algorithms were used: (1) TF-IDF for Content-based filtering (CBF); (2) Classification with Matrix Factorization- Singular Value Decomposition (SVD) applied with Collaborative filtering (CF) and mean (normalization) for prediction accuracy of the CF. Data sampling in academic Research and Development (R&D) of Philippine Council for Agriculture, Aquatic, and Natural Resources Research and Development (PCAARRD) e-Library and Project SARAI publications plus simulated data used as training sets to generate a recommendation of items that uses the three RS filtering (CF, CBF, and personalized version of item recommendations). Series of Testing and TAM performed and discussed. Results – Findings allow users to engage in online information and quickly evaluate retrieved items produced by the application. Compatibility-testing (CoT) shows the application is compatible with all major browsers and mobile-friendly. Performance-testing (PT) recommended v-parameter specs and TAM evaluations results indicate strongly associated with overall positive feedback, thoroughly enough to address the information-overload problem as the core of the paper. Conclusion – A modular architecture presented addressing the information overload, primarily focused on RSs with the personalization and design of modern systems. Developers utilized Two ML algorithms and prototyped a simplified version of the architecture. Series of testing (CoT and PT) and evaluations with TAM were performed and discussed. Project ATHENA added a UX feature design of a modern system. Recommendations – High-end hardware specs of v-Parameter are recommended, at least 8-cores of vCPU with 16GiB of memory and sufficient driver-type size to run the model and End-to-end jobs execution for continuously maintaining the application. Research Implications – Future Developers must use/integrate a large dataset. Other ML approaches can expand Hybrid-RS better. Implicit data and additional filtering methods may enhance the application in the future.
Read moreA Comparative Analysis of Matrix Factorization Models for Collaborative Filtering: From Basic SVD to SVD++
Recommendation systems have permeated society and everyones daily life with the expanding use of applications. It plays an important role in providing recommendations to active users based on either their historical interests or interests shown by similar users. This paper aims to introduce several components of a recommendation system, such as Collaborative Filtering (CF) and Matrix Factorization, and combining the dominant approach and main technique: Singular Value Decomposition (SVD). Additionally, with the advances in technology, various recently established SVD-based models are mentioned in this paper, attempting to give an answer to the question that though studies have shown the coexistence of advantages and disadvantages for each mode, the trade-offs between their predictive accuracy, computational cost, and robustness to data sparsity have still not been fully answered. Furthermore, this paper concentrates on the comparison and analysis of diverse matrix factorization variants, meanwhile evaluating their performances. For instance, basic SVD, BiasSVD, and SVD++. Finding indicates that SVD shows more efficiency in small-scale datasets, BiasSVD is suitable for cases with noticeable biases, and SVD++ gives high-quality results when handling both explicit and implicit feedback data. This study provides practical guidelines for practitioners to select the most desirable model based on their existing database.
Read moreCollaborative Filtering Based Food Recommendation System Using Matrix Factorization
A recommendation system is a method that provides suggestions of items that might users like. There are many domains that can be recommended, one of the most demanded domains by users today is food. In the era of big data, food choices from the large amount of data make it difficult for users to choose the right food for them. The collaborative filtering (CF) approach is considered capable of providing accurate and high quality item suggestions. One of the algorithms that can provide good performance results from the CF approach is Matrix Factorization (MF). This study aims to test a dataset that contains product ratings of food items using three MF algorithms, which are Singular Value Decomposition (SVD), SVD with Implicit Ratings (SVD++), and Non-Negative Matrix Factorization (NMF). Different latent factors are also used for the purpose of improving the performance of the proposed recommendation system algorithm. The dataset used is Amazon Fine Food Reviews. The study shows NMF and SVD++ as the best algorithm for generating user rating predictions for items. NMF has the smallest average prediction error as measured by MAE which is 0.7311. While SVD++ obtains the smallest prediction error value of 1.0607 as measured using RMSE. In addition to these results, the top-n evaluation also shows that the proposed algorithm performs quite well. The hit ratio value for each different n-item always increases proportionally to the number of recommended n-items. The highest hit ratio value is generated from the SVD++ algorithm of 0.0025 on n-item recommendations of 25 items. Overall it can be said that the proposed algorithm has good performance in providing item recommendations.
Read moreBook Recommendation System Based on Collaborative Filtering: User-Based, Item-Based, and Singular Value Decomposition Analysis
Recommender systems have become essential in the digital era to help users navigate overwhelming content. This study develops a book recommendation system using three collaborative filtering methods: user-based, item-based, and matrix factorization using singular value decomposition. We evaluate the system on a real-world dataset of 1,149,780 book ratings from 278,858 users across 271,360 books. A subset of 500 active users is used for experimental evaluation. The models are assessed using root mean square error and mean absolute error to measure rating prediction accuracy. The results show that the item-based collaborative filtering method achieves the best accuracy (root mean square error 7.362; mean absolute error 6.761), slightly outperforming the user-based approach (7.365; 6.809) and the matrix factorization method (7.643; 7.413). We analyze the results to understand the performance differences, noting the stability of item similarity as a key factor and the need for optimal tuning in the matrix factorization model. In conclusion, item-based collaborative filtering proved most effective for this context. This work provides insights into the comparative performance of foundational recommendation techniques and highlights practical considerations for improving book recommender systems.
Read moreA novel hybrid approach towards movie recommender systems
Collaborative and content-based filtering are the majorly used approaches towards implementing recommender systems which can successfully predict the item(s) to be recommended to users based upon their preferences. Each of the methods has its own advantages and disadvantages that are specific to the situation they are used in. The hybrid approach aims to combine both techniques in multiple ways to optimize results and overcome shortcomings. In this paper, we propose a hybrid approach based recommendation engine to improve accuracy. While content based recommendations are made on the similarity of items’ attributes, collaborative recommendations are based on similarity of users. Collaborative filtering is achieved by matrix factorization technique. One of the most effective algorithms for matrix factorization: Singular Value Decomposition (SVD) has been used to perform collaborative filtering and combined with content based model to predict the item ratings per user. The paper aims to achieve the final result with minimum possible errors as per user preferences and recommended items.
Read moreA Survey of Collaborative Filtering Recommender Algorithms and Their Evaluation Metrics
Abstract—Recommender systems are often used to provide useful recommendations for users. They use previous history of the users-items interactions, e.g. purchase history and/or users rating on items, to provide a suitable recommendation list for any target user. They may also use contextual information available about items and users. Collaborative filtering algorithm and its variants are the most successful recommendation algorithms that have been applied to many applications. Collaborative filtering method works by first finding the most similar users (or items) for a target user (or items), and then building the recommendation lists. There is no unique evaluation metric to assess the performance of recommendations systems, and one often choose the one most appropriate for the application in hand. This paper compares the performance of a number of well-known collaborative filtering algorithms on movie recommendation. To this end, a number of performance criteria are used to test the algorithms. The algorithms are ranked for each evaluation metric and a rank aggregation method is used to determine the wining algorithm. Our experiments show that the probabilistic matrix factorization has the top performance in this dataset, followed by item-based and user-based collaborative filtering. Non-negative matrix factorization and Slope 1 has the worst performance among the considered algorithms.
 Keywords—Social networks analysis and mining, big data, recommender systems, collaborative filtering.
Read moreA Collaborative Filtering Recommendation Algorithm Based on SVD Smoothing
Recommender system is one of the most important technologies in electronic commerce. And the collaborative filtering is almost the popular approach used in the recommender systems. With the development of electronic commerce systems, the magnitudes of users and items grow rapidly, resulted in the extreme sparsity of user rating data set. Traditional similarity measure methods work poor in this situation, make the quality of recommendation system decreased dramatically. Sparsity of users’ ratings is the major reason causing the poor quality. To address this issue, a collaborative filtering recommendation algorithm based on singular value decomposition (SVD) smoothing is presented. This approach predicts item ratings that users have not rated by the employ of SVD technology, and then uses Pearson correlation similarity measurement to find the target users’ neighbors, lastly produces the recommendations. The collaborative filtering recommendation algorithm based on SVD smoothing can alleviate the sparsity problems of the user item rating dataset, and can provide better recommendation than traditional collaborative filtering algorithms.
Read moreFood Recipe Recommendation System with Content-Based Filtering and Collaborative Filtering Methods
Cooking your own food at home is a good step toward reducing fast food consumption. Fast food increases the risk of dangerous diseases. The diversity of recipe information available on the internet makes it difficult to choose recipes that match user preferences. Mobile technology can help with this by recommending recipes that better suit users' eating habits. This makes the transition to a healthier diet easier. Therefore, in this study, a recommendation system was developed that can recommend recipes based on the preferences of Android users. Two main recommendation methods are used in this study: content-based filtering and collaborative filtering. Using cosine similarity, a content-based recommendation system identifies the proximity between a recipe for food and its related context. The history of user comments on recipes serves as implicit feedback for the collaborative recommendation algorithm. This eliminates the need for explicit evaluations, such as ratings. This recommendation system generates recommendations in the form of the top ten food recipes with an evaluation matrix, referred to as NDCG@k and Hit-Ratio@k. The tests revealed that a content-based filtering technique may produce helpful recommendations, with the highest similarity score of 0.41 for the entry "chocolate cake that you can easily make at home." Meanwhile, in the collaborative filtering method using the Neural Collaborative Filtering (NCF) approach, the system shows consistent performance improvements, with the MAP@10 value increasing from 0.705 to 0.767 and the NDCG@10 from 0.78 to 0.83 after 10 training epochs. Keywords: Recommendation systems; content-based filtering; neural collaborative filtering; cosine similarity; implicit feedback
Read moreSocial Network Analysis for the Effective Adoption of Recommender Systems
Recommender system is the system which, by using automated information filtering technology, recommends products or services to the customers who are likely to be interested in. Those systems are widely used in many different Web retailers such as Amazon.com, Netfix.com, and CDNow.com. Various recommender systems have been developed. Among them, Collaborative Filtering (CF) has been known as the most successful and commonly used approach. CF identifies customers whose tastes are similar to those of a given customer, and recommends items those customers have liked in the past. Numerous CF algorithms have been developed to increase the performance of recommender systems. However, the relative performances of CF algorithms are known to be domain and data dependent. It is very time-consuming and expensive to implement and launce a CF recommender system, and also the system unsuited for the given domain provides customers with poor quality recommendations that make them easily annoyed. Therefore, predicting in advance whether the performance of CF recommender system is acceptable or not is practically important and needed. In this study, we propose a decision making guideline which helps decide whether CF is adoptable for a given application with certain transaction data characteristics. Several previous studies reported that sparsity, gray sheep, cold-start, coverage, and serendipity could affect the performance of CF, but the theoretical and empirical justification of such factors is lacking. Recently there are many studies paying attention to Social Network Analysis (SNA) as a method to analyze social relationships among people. SNA is a method to measure and visualize the linkage structure and status focusing on interaction among objects within communication group. CF analyzes the similarity among previous ratings or purchases of each customer, finds the relationships among the customers who have similarities, and then uses the relationships for recommendations. Thus CF can be modeled as a social network in which customers are nodes and purchase relationships between customers are links. Under the assumption that SNA could facilitate an exploration of the topological properties of the network structure that are implicit in transaction data for CF recommendations, we focus on density, clustering coefficient, and centralization which are ones of the most commonly used measures to capture topological properties of the social network structure. While network density, expressed as a proportion of the maximum possible number of links, captures the density of the whole network, the clustering coefficient captures the degree to which the overall network contains localized pockets of dense connectivity. Centralization reflects the extent to which connections are concentrated in a small number of nodes rather than distributed equally among all nodes. We explore how these SNA measures affect the performance of CF performance and how they interact to each other. Our experiments used sales transaction data from H department store, one of the well?known department stores in Korea. Total 396 data set were sampled to construct various types of social networks. The dependant variable measuring process consists of three steps; analysis of customer similarities, construction of a social network, and analysis of social network patterns. We used UCINET 6.0 for SNA. The experiments conducted the 3-way ANOVA which employs three SNA measures as dependant variables, and the recommendation accuracy measured by F1-measure as an independent variable. The experiments report that 1) each of three SNA measures affects the recommendation accuracy, 2) the density's effect to the performance overrides those of clustering coefficient and centralization (i.e., CF adoption is not a good decision if the density is low), and 3) however though the density is low, the performance of CF is comparatively good when the clustering coefficient is low. We expect that these experiment results help firms decide whether CF recommender system is adoptable for their business domain with certain transaction data characteristics.
Read moreNovelty and diversity based recommendation systems
Traditional recommendation systems aim at generating recommendations that are relevant to the user’s interest. Thus, they are called relevance-based recommendation systems (RBRSs). The major drawback of this approach is that user soon becomes very familiar with the recommendations and loses interest in reading and exploring them. In other words, relevance-based recommendations cannot help users to expand their interest and keep the recommendations exciting to them. Discovery-oriented recommendation systems (DORSs) aim to solve this problem by introducing discover utilities (DUs) such as novelty and diversity to improve the attractiveness of the recommendations to the user. In this thesis, we investigate techniques for improving the effectiveness of DORSs. Since novelty and diversity are the most important and widely studied DUs, we focus on recommendation systems that aim to improve the novelty and diversity of the recommendations. We study two important aspects of DORSs, namely, novelty and diversity of the recommendations. Existing DORSs generate recommendations that are optimized to balance between the accuracy and DUs of the recommendations to make the recommendations relevant and yet interesting to the user. However, they disregard an important fact that different users’ appetites for DUs are different. For example, a curious user can accept highly novel and diversified recommendations but a conservative user tends to respond only to recommendations she is familiar with. Thus, we propose a framework for curiosity-based recommendation systems (CBRSs) which can produce recommendations with an amount of DUs personalized to fit an individual user’s curiosity level. As a result, the recommendations are neither too surprising nor too boring for a user because the recommendations are customized to fit her unique curiosity. In order to model and quantify human curiosity, we adopt the curiosity arousing model (CAM) developed in psychology research and propose a probabilistic curiosity model (PCM) to model the psychological model computationally. Extensive experiments have been performed to evaluate the performance of CBRS against the baselines using a music dataset from last.fm. The results show that compared to the baselines CBRS not only provides more personalized recommendations that adapt to the user’s curiosity level but also improves the recommendation accuracy. To improve the diversity of the recommendations, we propose a recommendation framework by the unification of two types of diversity, namely, intra-list and temporal diversity, of the recommendations. Traditional RBRSs recommend items which are very similar to the user’s interest. As a result, the recommended items are also very similar between each other, making the items in a recommendation list monotonous. We name this “intra-monotony problem” (IMP). Further, most existing recommendation systems make recommendations without considering what has been recommended before. Thus, they may make similar recommendations over and over again, making the recommended items across recommendation lists monotonous. We name this “temporal monotony problem” (TMP). To address these two problems, previous research has utilized intralist diversity (intraD) and temporal diversity (timeD) to improve, respectively, the diversity within a recommendation list and across recommendation lists. However, existing work studies these two diversity types separately. We propose an approach to unify the two diversity types into a single framework so that both intra-list and temporal diversity can be considered holistically. This is a challenging problem, since a high intra-list diversity does not guarantee a high temporal diversity, and vice versa. Rather than arbitrarily combining intraD and timeD, we propose a new diversity type called jointD and optimize it by formulating the problem as a constraint quadratic optimization problem. This approach allows both intraD and timeD to be jointly processed. We design a new performance metric called F-div to measure a recommendation system’s ability to improve the overall intraD and timeD. Experiment results show that optimizing jointD produces better F-div performance compared to optimize intraD or timeD alone.
Read moreA genetic algorithms-based hybrid recommender system of matrix factorization and neighborhood-based techniques
A genetic algorithms-based hybrid recommender system of matrix factorization and neighborhood-based techniques
Comparative study of recommender system approaches and movie recommendation using collaborative filtering
The increasing demand for personalized information has resulted in the development of the Recommender System (RS). RS has been widely utilized and broadly studied to suggest the interests of users and make an appropriate recommendation. This paper gives an overview of several types of recommendation approaches based on user preferences, ratings, domain knowledge, users demographic data, users context and also lists the advantages and disadvantages of each RS approach. In this paper, we also proposed the movie recommendation based on collaborative filtering and singular value decomposition plus-plus (SVD++). The proposed approach is compared with well-known machine learning approaches namely k nearest neighbor (K-NN), singular value decomposition (SVD) and Co-clustering. The proposed approach is experimentally verified using MovieLens 100 K datasets and error of the RS is measured using Root Mean Square Error (RMSE) and Mean Absolute Error (MAE). The result shows that the proposed approach gives a lesser error rate with RMSE (0.9201) and MAE (0.7219). This approach also overcomes cold-start, data sparsity problems and provides them relevant items and services.
Read moreEnhanced SVD for Collaborative Filtering
Matrix factorization is one of the most popular techniques for prediction problems in the fields of intelligent systems and data mining. It has shown its effectiveness in many real-world applications such as recommender systems. As a collaborative filtering method, it gives users recommendations based on their previous preferences or ratings. Due to the extreme sparseness of the ratings matrix, active learning is used for eliciting ratings for a user to get better recommendations. In this paper, we propose a new matrix factorization model called Enhanced SVD ESVD which combines the classic matrix factorization method with a specific rating elicitation strategy. We evaluate the proposed ESVD method on the Movielens data set, and the experimental results suggest its effectiveness in terms of both accuracy and efficiency, when compared with traditional matrix factorization methods and active learning methods.
Read moreAdvancing Book Recommendation Systems: A Comparative Analysis of Collaborative Filtering and Matrix Factorization Algorithms
Introduction: Recommendation systems (RS) are pivotal in enhancing personalized user experiences, particularly in the educational domain, where tailored book recommendation systems (BRS) support individualized learning. Despite their importance, selecting an optimal recommendation algorithm remains challenging, especially in scenarios involving large and sparse datasets. Objectives: This study aims to evaluate the performance of four recommendation algorithms—Random Baseline, Popular, User-Based Collaborative Filtering (UBCF), and Singular Value Decomposition (SVD)—to identify their strengths and tradeoffs regarding accuracy and computational efficiency. The goal is to provide actionable insights for selecting appropriate algorithms for educational platforms and similar applications. Methods: The evaluation employed a dataset containing over 10,000 books, 53,424 users, and 981,756 ratings. The algorithms were assessed using Root Mean Square Error (RMSE) to measure predictive accuracy and training time to indicate computational efficiency. The analysis addressed challenges such as data sparsity and rating biases. Results: The results demonstrated that SVD achieved the highest accuracy (RMSE: 0.950), effectively uncovering latent relationships in sparse datasets. UBCF, with a slightly lower accuracy (RMSE: 1.020), offered a balance between accuracy and computational efficiency, making it suitable for real-time applications. Conversely, simpler algorithms like Random Baseline and Popular exhibited faster training times but significantly lower predictive accuracy, highlighting the limitations of non-personalized methods. Conclusions: This study underscores the tradeoffs between accuracy and efficiency among recommendation algorithms. While SVD is ideal for accuracy-driven applications, UBCF provides a practical alternative for scenarios requiring computational efficiency. The findings have substantial implications for educational platforms and e-commerce, where personalized recommendations enhance user satisfaction and engagement. Future research should focus on integrating deep learning models and expanding evaluation criteria, including user satisfaction and diversity, to further improve the performance of recommendation systems.
Read more