- Research Article
1
- 10.2139/ssrn.3791837
Evaluation Of Recommendation System On A Movie Dataset
- Jan 01, 2020
- SSRN Electronic Journal
- Ronit Malik
Evaluation Of Recommendation System On A Movie Dataset
A recommendation system is a method that provides suggestions of items that might users like. There are many domains that can be recommended, one of the most demanded domains by users today is food. In the era of big data, food choices from the large amount of data make it difficult for users to choose the right food for them. The collaborative filtering (CF) approach is considered capable of providing accurate and high quality item suggestions. One of the algorithms that can provide good performance results from the CF approach is Matrix Factorization (MF). This study aims to test a dataset that contains product ratings of food items using three MF algorithms, which are Singular Value Decomposition (SVD), SVD with Implicit Ratings (SVD++), and Non-Negative Matrix Factorization (NMF). Different latent factors are also used for the purpose of improving the performance of the proposed recommendation system algorithm. The dataset used is Amazon Fine Food Reviews. The study shows NMF and SVD++ as the best algorithm for generating user rating predictions for items. NMF has the smallest average prediction error as measured by MAE which is 0.7311. While SVD++ obtains the smallest prediction error value of 1.0607 as measured using RMSE. In addition to these results, the top-n evaluation also shows that the proposed algorithm performs quite well. The hit ratio value for each different n-item always increases proportionally to the number of recommended n-items. The highest hit ratio value is generated from the SVD++ algorithm of 0.0025 on n-item recommendations of 25 items. Overall it can be said that the proposed algorithm has good performance in providing item recommendations.
Evaluation Of Recommendation System On A Movie Dataset
Evaluation Of Recommendation System On A Movie Dataset
Nonnegative singular value decomposition for microarray data analysis of spermatogenesis
Matrix factorization plays an important role in scientific computation. The widely used one is singular value decomposition (SVD) which approximates the original data matrix with three lower rank matrices with orthogonality constraints. Recently nonnegative matrix factorization (NMF) considering the nonnegativity of data makes the results more interpretable than those of SVD. However NMF finds only two factor matrices and there is no significant index as singular values of SVD which can be used for sorting learned basis vectors. In this paper we take into account the nonnegativity for SVD and propose nonnegative SVD (NNSVD). The preliminary results on the microarray data of spermatogenesis show that NNSVD has advantages of both SVD and NMF.
Read moreA Collaborative Filtering Recommendation Algorithm Based on SVD Smoothing
Recommender system is one of the most important technologies in electronic commerce. And the collaborative filtering is almost the popular approach used in the recommender systems. With the development of electronic commerce systems, the magnitudes of users and items grow rapidly, resulted in the extreme sparsity of user rating data set. Traditional similarity measure methods work poor in this situation, make the quality of recommendation system decreased dramatically. Sparsity of users’ ratings is the major reason causing the poor quality. To address this issue, a collaborative filtering recommendation algorithm based on singular value decomposition (SVD) smoothing is presented. This approach predicts item ratings that users have not rated by the employ of SVD technology, and then uses Pearson correlation similarity measurement to find the target users’ neighbors, lastly produces the recommendations. The collaborative filtering recommendation algorithm based on SVD smoothing can alleviate the sparsity problems of the user item rating dataset, and can provide better recommendation than traditional collaborative filtering algorithms.
Read moreA new collaborative approach to solve the gray-sheep users problem in recommender systems
Recommender systems aim to help users to find items that fit their requirements and preferences. In that field, the collaborative filtering (CF) approach is considered as a widely used one. There are two main approaches for CF: memory-based and model-based. Both of the two approaches are based on the use of users' ratings to predict the top-N recommendation for the active user. Despite its simplicity and efficiency, The CF approach stills suffer from many drawbacks including sparsity, gray sheep and scalability. The aim of this work is to deal with the gray sheep problem, by proposing a novel collaborative filtering approach. This novel approach aims to enhance the accuracy of prediction by turning the users whose preferences disagree with the target user, into new similar neighbors. For instance, if a user X is dissimilar to a user Y then the user $\lnot$X is similar to the user Y.To evaluate the performance of the proposed approach, we have used two datasets including MovieLens and FilmTrust. The Experimental results show that our approach outperforms many traditional recommendation techniques.
Read moreA Comparative Analysis of Matrix Factorization Models for Collaborative Filtering: From Basic SVD to SVD++
Recommendation systems have permeated society and everyones daily life with the expanding use of applications. It plays an important role in providing recommendations to active users based on either their historical interests or interests shown by similar users. This paper aims to introduce several components of a recommendation system, such as Collaborative Filtering (CF) and Matrix Factorization, and combining the dominant approach and main technique: Singular Value Decomposition (SVD). Additionally, with the advances in technology, various recently established SVD-based models are mentioned in this paper, attempting to give an answer to the question that though studies have shown the coexistence of advantages and disadvantages for each mode, the trade-offs between their predictive accuracy, computational cost, and robustness to data sparsity have still not been fully answered. Furthermore, this paper concentrates on the comparison and analysis of diverse matrix factorization variants, meanwhile evaluating their performances. For instance, basic SVD, BiasSVD, and SVD++. Finding indicates that SVD shows more efficiency in small-scale datasets, BiasSVD is suitable for cases with noticeable biases, and SVD++ gives high-quality results when handling both explicit and implicit feedback data. This study provides practical guidelines for practitioners to select the most desirable model based on their existing database.
Read moreAdvancing Book Recommendation Systems: A Comparative Analysis of Collaborative Filtering and Matrix Factorization Algorithms
Introduction: Recommendation systems (RS) are pivotal in enhancing personalized user experiences, particularly in the educational domain, where tailored book recommendation systems (BRS) support individualized learning. Despite their importance, selecting an optimal recommendation algorithm remains challenging, especially in scenarios involving large and sparse datasets. Objectives: This study aims to evaluate the performance of four recommendation algorithms—Random Baseline, Popular, User-Based Collaborative Filtering (UBCF), and Singular Value Decomposition (SVD)—to identify their strengths and tradeoffs regarding accuracy and computational efficiency. The goal is to provide actionable insights for selecting appropriate algorithms for educational platforms and similar applications. Methods: The evaluation employed a dataset containing over 10,000 books, 53,424 users, and 981,756 ratings. The algorithms were assessed using Root Mean Square Error (RMSE) to measure predictive accuracy and training time to indicate computational efficiency. The analysis addressed challenges such as data sparsity and rating biases. Results: The results demonstrated that SVD achieved the highest accuracy (RMSE: 0.950), effectively uncovering latent relationships in sparse datasets. UBCF, with a slightly lower accuracy (RMSE: 1.020), offered a balance between accuracy and computational efficiency, making it suitable for real-time applications. Conversely, simpler algorithms like Random Baseline and Popular exhibited faster training times but significantly lower predictive accuracy, highlighting the limitations of non-personalized methods. Conclusions: This study underscores the tradeoffs between accuracy and efficiency among recommendation algorithms. While SVD is ideal for accuracy-driven applications, UBCF provides a practical alternative for scenarios requiring computational efficiency. The findings have substantial implications for educational platforms and e-commerce, where personalized recommendations enhance user satisfaction and engagement. Future research should focus on integrating deep learning models and expanding evaluation criteria, including user satisfaction and diversity, to further improve the performance of recommendation systems.
Read moreA Tech Hybrid-Recommendation Engine and Personalized Notification: An integrated tool to assist users through Recommendations (Project ATHENA)
Purpose – Project ATHENA aims to develop an application to address information overload, primarily focused on Recommendation Systems (RSs) with the personalization and user experience design of a modern system. Method – Two machine learning (ML) algorithms were used: (1) TF-IDF for Content-based filtering (CBF); (2) Classification with Matrix Factorization- Singular Value Decomposition (SVD) applied with Collaborative filtering (CF) and mean (normalization) for prediction accuracy of the CF. Data sampling in academic Research and Development (R&D) of Philippine Council for Agriculture, Aquatic, and Natural Resources Research and Development (PCAARRD) e-Library and Project SARAI publications plus simulated data used as training sets to generate a recommendation of items that uses the three RS filtering (CF, CBF, and personalized version of item recommendations). Series of Testing and TAM performed and discussed. Results – Findings allow users to engage in online information and quickly evaluate retrieved items produced by the application. Compatibility-testing (CoT) shows the application is compatible with all major browsers and mobile-friendly. Performance-testing (PT) recommended v-parameter specs and TAM evaluations results indicate strongly associated with overall positive feedback, thoroughly enough to address the information-overload problem as the core of the paper. Conclusion – A modular architecture presented addressing the information overload, primarily focused on RSs with the personalization and design of modern systems. Developers utilized Two ML algorithms and prototyped a simplified version of the architecture. Series of testing (CoT and PT) and evaluations with TAM were performed and discussed. Project ATHENA added a UX feature design of a modern system. Recommendations – High-end hardware specs of v-Parameter are recommended, at least 8-cores of vCPU with 16GiB of memory and sufficient driver-type size to run the model and End-to-end jobs execution for continuously maintaining the application. Research Implications – Future Developers must use/integrate a large dataset. Other ML approaches can expand Hybrid-RS better. Implicit data and additional filtering methods may enhance the application in the future.
Read moreA novel hybrid approach towards movie recommender systems
Collaborative and content-based filtering are the majorly used approaches towards implementing recommender systems which can successfully predict the item(s) to be recommended to users based upon their preferences. Each of the methods has its own advantages and disadvantages that are specific to the situation they are used in. The hybrid approach aims to combine both techniques in multiple ways to optimize results and overcome shortcomings. In this paper, we propose a hybrid approach based recommendation engine to improve accuracy. While content based recommendations are made on the similarity of items’ attributes, collaborative recommendations are based on similarity of users. Collaborative filtering is achieved by matrix factorization technique. One of the most effective algorithms for matrix factorization: Singular Value Decomposition (SVD) has been used to perform collaborative filtering and combined with content based model to predict the item ratings per user. The paper aims to achieve the final result with minimum possible errors as per user preferences and recommended items.
Read moreA Hybrid Matrix Factorization Framework for Balancing Personalization, Diversity, and Coverage in Ad Recommendation Systems
This paper presents a comprehensive framework for addressing the inherent trade-offs among personalization, diversity, and coverage in hybrid advertisement recommendation systems. In response to the growing complexity of online advertising, the proposed system integrates collaborative filtering via Singular Value Decomposition (SVD) with content-based filtering using Non-negative Matrix Factorization (NMF), enhanced by a re-ranking mechanism based on coverage profiles. This hybrid design aims to deliver relevant, diverse, and widely distributed ads, thereby improving user engagement and fairness among content providers. To quantitatively evaluate these objectives, three novel metrics—Ad User Profile (AUP), Ad Candidate Set (ACS), and Ad Matching Degree (AMD)—are introduced. Through empirical analysis, including visualization and comparative experiments, the study demonstrates that the proposed system outperforms traditional models in achieving a more balanced recommendation outcome. Additionally, real-world case studies and optimization formulations are explored to support scalable deployment. This work contributes to the evolving landscape of recommender systems by proposing a scalable, user-centric approach that harmonizes personalization with fairness and discovery in digital advertising.
Read moreExtracting unrecognized gene relationships from the biomedical literature via matrix factorizations using a priori knowledge of gene relationships
The construction of literature-based networks of gene-gene interactions is one of the most important applications of text mining in bioinformatics. Extracting potential gene relationships from the biomedical literature may be helpful in building biological hypotheses that can be explored further experimentally. In this paper, we explore the utility of singular value decomposition (SVD) and non-negative matrix factorization (NMF) to extract unrecognized gene relationships from the biomedical literature by taking advantage of known gene relationships. We introduce a way to incorporate a priori knowledge of gene relationships into LSI/SVD and NMF. In addition, we propose a gene retrieval method based on NMF (GR/NMF), which shows comparable performance with latent semantic indexing based on SVD.
Read moreCollaborative filtering and deep learning based recommendation system for cold start items
Collaborative filtering and deep learning based recommendation system for cold start items
Optimizing deep learning architectures for SEM image classification using advanced dimensionality reduction techniques
Non-negative Matrix Factorization (NMF) and Singular Value Decomposition (SVD) are widely recognized as pivotal dimensionality reduction techniques in the literature, particularly for deep learning applications involving large and high-dimensional datasets like SEM images. This study systematically evaluates the impact of SVD and NMF on the performance, efficiency, and energy consumption of four deep learning architectures: GoogleNet, AlexNet, ResNet, and SqueezeNet. By applying these techniques to reduce dataset dimensions, we observed that SVD excelled in computational efficiency, achieving up to 35% faster processing times compared to raw datasets. NMF, on the other hand, provided superior feature interpretability, which proved beneficial for tasks requiring meaningful pattern extraction. Energy consumption analysis revealed that SVD led to a 28% reduction in computational energy cost on average, making it a practical choice for resource-constrained environments. Among the evaluated models, ResNet consistently delivered the highest classification accuracy after dimensionality reduction, showing an improvement of 4-6% over models trained on non-reduced data. These findings underscore the critical role of dimensionality reduction in enhancing the scalability, energy efficiency, and classification accuracy of deep learning models, offering valuable insights for optimizing high-dimensional data applications in both academic and industrial contexts.
Read moreJoint latent factors and attributes to discover interpretable preferences in recommendation
Joint latent factors and attributes to discover interpretable preferences in recommendation
Survey on Probabilistic Models of Low-Rank Matrix Factorizations
Low-rank matrix factorizations such as Principal Component Analysis (PCA), Singular Value Decomposition (SVD) and Non-negative Matrix Factorization (NMF) are a large class of methods for pursuing the low-rank approximation of a given data matrix. The conventional factorization models are based on the assumption that the data matrices are contaminated stochastically by some type of noise. Thus the point estimations of low-rank components can be obtained by Maximum Likelihood (ML) estimation or Maximum a posteriori (MAP). In the past decade, a variety of probabilistic models of low-rank matrix factorizations have emerged. The most significant difference between low-rank matrix factorizations and their corresponding probabilistic models is that the latter treat the low-rank components as random variables. This paper makes a survey of the probabilistic models of low-rank matrix factorizations. Firstly, we review some probability distributions commonly-used in probabilistic models of low-rank matrix factorizations and introduce the conjugate priors of some probability distributions to simplify the Bayesian inference. Then we provide two main inference methods for probabilistic low-rank matrix factorizations, i.e., Gibbs sampling and variational Bayesian inference. Next, we classify roughly the important probabilistic models of low-rank matrix factorizations into several categories and review them respectively. The categories are performed via different matrix factorizations formulations, which mainly include PCA, matrix factorizations, robust PCA, NMF and tensor factorizations. Finally, we discuss the research issues needed to be studied in the future.
Read moreCOMPUTING THERAPY FOR PRECISION MEDICINE: COLLABORATIVE FILTERING INTEGRATES AND PREDICTS MULTI-ENTITY INTERACTIONS
Biomedicine produces copious information it cannot fully exploit. Specifically, there is considerable need to integrate knowledge from disparate studies to discover connections across domains. Here, we used a Collaborative Filtering approach, inspired by online recommendation algorithms, in which non-negative matrix factorization (NMF) predicts interactions among chemicals, genes, and diseases only from pairwise information about their interactions. Our approach, applied to matrices derived from the Comparative Toxicogenomics Database, successfully recovered Chemical-Disease, Chemical-Gene, and Disease-Gene networks in 10-fold cross-validation experiments. Additionally, we could predict each of these interaction matrices from the other two. Integrating all three CTD interaction matrices with NMF led to good predictions of STRING, an independent, external network of protein-protein interactions. Finally, this approach could integrate the CTD and STRING interaction data to improve Chemical-Gene cross-validation performance significantly, and, in a time-stamped study, it predicted information added to CTD after a given date, using only data prior to that date. We conclude that collaborative filtering can integrate information across multiple types of biological entities, and that as a first step towards precision medicine it can compute drug repurposing hypotheses.
Read more