- Research Article
6
- 10.1016/j.pedhc.2012.04.008
Googling for Health Information
- Jun 21, 2012
- Journal of Pediatric Health Care
- Jennifer P D'Auria
Googling for Health Information
The World Wide Web (WWW) allows the people to share the information (data) from the large database repositories globally. The amount of information grows billions of databases. We need to search the information will specialize tools known generically search engine. There are many of search engines available today, retrieving meaningful information is difficult. However to overcome this problem in search engines to retrieve meaningful information intelligently, semantic web technologies are playing a major role. In this paper we present survey on the search engine generations and the role of search engines in intelligent web and semantic search technologies.
Googling for Health Information
Googling for Health Information
Avaliação e implementação de propostas de melhoria para o protocolo IRIS baseadas em tecnologias de web semântica.
The aim of this thesis is to evaluate whether Semantic Web technologies can contribute to Internet Registry Information Service Protocol (IRIS) protocol development. IRIS is a new protocol for providing an information service for Internet resources. It is currently still under development by an Internet Engineering Task Force (IETF) working group. The objective of the working group is to develop and standardize a new protocol to replace the Whois protocol. Whois is the standard protocol used today by information services for Internet resources, i.e. domain names, Internet Protocol (IP) addresses, autonomous systems, amongst others. The motivation to develop a new protocol was based on increasing concerns regarding the security of data stored in the Whois database as the Whois protocol does not provide any security mechanism. Another motivation was the absence of support for distributed databases as the Whois protocol was developed for a centralized database, hence it no longer meets the standard requirements for Internet protocols. So far, the working group has tackled and solved two main issues concerning the Whois protocol: (1) security and (2) support for distributed databases. However, the development of a new standard demands a great investment from the community, in particular with respect to consensus-based policies. Additionally, there is one major barrier against adopting the new protocol: the users adoption. The new protocol must have longevity without being updated or replaced by another protocol. To reach this goal, it is necessary to meet not only the current requirements, such as security issues, but to cater also for future requirements. This thesis is concerned with the following research activities: (1) comparative analysis of the current protocols used to provide information services on Internet resources, (2) the IRIS protocol analysis and (3) the evaluation of new technologies that could be incorporated in the new protocol, in particular Semantic technologies. The results demonstrate that Semantic Web technologies could provide the necessary flexibility and extensibility to meet the current and future requirements of IRIS. To validate the theoretical results a prototype based on the IRIS specification was implemented using Semantic Web technologies. Two types of experiments were conducted: (1) experiments comparing the Whois and the prototype performance and (2) performance evaluation of the prototype based on load tests. Finally, the prototype implementation and subsequent experiment results serve as a proof-of-concept that Semantic Web technologies could contribute towards the IRIS protocol success.
Read moreDesign of a Metacrawler for web document retrieval
Web Crawlers ‘browse’ the World Wide Web (WWW) on behalf of search engine, to collect web pages from numerous collections of billions of documents. Metacrawler is similar to that of a meta search engine that combines the top web search results from popular search engines. World Wide Web is growing rapidly. This possesses great challenges to general purpose crawlers. This paper introduces an architectural framework of a Metacrawler. This crawler enables the user to retrieve information that is relevant to the topic from more than one traditional web search engines. The crawler works in such a way that it fetches only the pages that are relevant to the topic. The PageRank algorithm is often used in ranking web pages. But, the ranking causes the problem of topic-drift. So, modified PageRank algorithm is used to rank the retrieved web pages in such a way that it reduces this problem. The clustering method is used to combine the search results so that the user can easily select web pages from the clustered results based upon the requirement. Experimental results show the effectiveness of the Metacrawler.
Read moreTHE USE OF SEMANTIC WEB APPLICATIONS IN A DIGITAL LIBRARY
ABSTRACT: The semantic technologies that will underlie the next generation of knowledge management systems. A key element of the project is to evaluate and assess the impact of semantic web technology in case study settings. The overall aim of the case study, described here, is to investigate how the semantic web technologies being researched and developed in the project can enhance the functionality of a digital library. Keywords: Semantics Web, World Wide Web, Digital Libraries
Read moreChapter 1 New Perspectives on Web Search Engine Research
Purpose — The purpose of this chapter is to give an overview of the context of Web search and search engine related research, as well as to introduce the reader to the sections and chapters of the book. Methodology/approach — We review literature dealing with various aspects of search engines, with special emphasis on emerging areas of Web searching, search engine evaluation going beyond traditional methods and new perspectives on Web searching. Findings — The approaches to studying Web search engines are manifold. Given the importance of Web search engines for knowledge acquisition, research from different perspectives needs to be integrated into a more cohesive perspective. Research limitations/implications — The chapter suggests a basis for research in the field and also introduces further research directions. Originality/value of paper — The chapter gives a concise overview of the topics dealt within the book and also shows directions for researchers interested in Web search engines.
Read moreCYBER NEWS
CYBER NEWS
Enhancing Efficiency of Web Search Engines through Ontology Learning from Unstructured Information Sources
With the fast growth rate of information availability through the World Wide Web, search engines' ranking become limited to deal with such enormous amount of information. Web search engines should be enriched with methodologies that enable it to understand the content of Web pages, then to align pages to the correct query category that highly match its content. In this paper, a proposed system is introduced to deal with the abundance of information by automatically understand the content of a Web page, and semantically model the ontological concepts that exist inside it. The semantic relations between ontological concepts are automatically given a score or weight based on its influence to the given query. The weighted semantic relations between ontological concepts can be viewed as a signature for the query, the highly similarity of an article to this signature, the more relevant to the query. A new relevancy measure is introduced to semantically re-rank or classify Web pages based on computing the semantic similarity of the weighted intersection ratio between ontological concepts extracted from retrieved Web pages, and ontological concepts that represents the query. Results shows that the proposed system has the highest Pearson correlation coefficient (0.890) to human judgments which outperforms semantic similarity state-of-the-art methods and Web-based methods. The proposed model, was tested to re-rank Web pages according to the semantic relevancy of the query, experiments shows that it has the highest convergence to expert ranking order of Web pages compared to other Web search engines.
Read moreSpecial Issue on Software Engineering for Web Intelligence
Software engineering, in applying engineering to software, advocates the application of systematic, disciplined, quantifiable approaches to development, operation, and maintenance of software. Software engineering is also considered a transformational process that models the real world on a corresponding software world. Recent advances in software engineering applied to Web intelligence have been recognized as a wave in scientific research and development, exploring basic roles and practical impact of advanced information technology on next-generation Web-empowered products, systems, and services. Various web systems and services currently provide many benefits to users, with Web intelligence becoming increasingly important in research and business. Such Web intelligent systems have been realized through related technologies as software engineering, interactive machine learning, etc. This special issue brings together researchers who have inspired software engineering in Web intelligence. Submissions have been reviewed for relevance, originality, significance, and presentation based on JACIII review criteria. This special issue presents four papers describing outstanding studies on project management, service-oriented architecture, semantic Web technologies, and service composition planning in Web intelligence. All papers introduce promising approaches and interesting results that readers will find highly relevant! We believe that software engineering in Web Intelligence has tremendous potential as a new, active research field, and we hope this issue will motivate researchers to expand their own studies on software engineering in Web Intelligence.
Read moreThe semantic web and language technology, its potential and practicalities: EUROLAN-2003
EUROLAN, which has been held biennially since 1993, is one of the most significant European summer schools in the area of natural language processing. Each of the EUROLAN sessions has focused on an area of timely interest to researchers in the field; this year's EUROLAN involved students in tutorials and hands-on sessions concerned with semantic web technologies as applied to language processing, ontology creation and use, and consideration of the semantic web's potential and limitations.
Read moreSimilarity based Automatic Web Search Engine Evaluation
Nowadays, as the usage of Internet has incredibly increased, web search engines become the common approach to find and retrieve needed information. Hence, evaluating search engine quality is a hot topic which attracts many researches' attention. In this paper, we propose a framework named Similarity based Automatic Web Search Engine Evaluation, SAWSEE, to evaluate web search engines. SAWSEE measures the information retrieval effectiveness of web search engines by comparing and voting their returned results, particularly using nDCG metric to rank search engines. SAWSEE compares the search engines' results based on their similarity which is calculated in two consecutive levels, the web page address level and the main content of web page level. To find the similarity of the main content of two search engines' results, SAWSEE utilizes Winnowing algorithm, a well-known and widely-used plagiarism detection method. We compared our method with the results acquired from human assessors' evaluations. The promising comparison shows that SAWSEE provides rankings that are consistent with the rankings resulted from human assessors' evaluations. Hence, the proposed method can be applied in real world environments for evaluation of web search engines.
Read moreEstimating the Prominence of AHP in a Selected Internet Search Engine
Internet has become an instant source of informatio n for almost anyone. By viewing the results of sear ch terms displayed by an Internet Search Engine (ISE), a person may decide his next course of action whether to continue using the Internet, abort or to combine it with other sources of information. The results produced by an ISE can be regarded as an in dex of relative availability of references. Given a list of suitable search terms, a user will firstly have to decide on the ISE to be used. Due to huge potent ial references available for given search terms the use r has to create heuristics to choose the entries sh own on the computer screen. As there are varying breadths and depths of information revealed by various Inter net Search Engines (ISE’s), this paper will not attempt to make a comparison among ISE’s, rather will focu s only on a particular search engine and ascertain th e results it produces given a list of search terms. Google is chosen as a proxy, being one of the most popular ISE’s. By confining to only a search engine, the s tudy affords to control variability among ISE’s should m ultiple ISE’s be used. With the search engine place d under control, it is easy to achieve the primary ob jective of the study, i.e., to ascertain availabili ty of relative breadth of sub-themes of Analytic Hierarch y Process (AHP) in Google. It is natural for user, especially researcher to be concerned with number, quantity. If there is seemingly abundant literature , one would be motivated to pursue research along the the me or sub-theme. In this study, the strength of presence of a sub-theme is measured by using two measures: (i) result of a sub-theme of AHP over the sub-theme itself, and (ii) result of the sub-theme over the total results generated for all of the thi rty six sub-themes used in the search. This study controlle d biasness in specifying the sub-themes of AHP by adopting the sub-themes or search terms specified b y the 2013 AHP conference organizers. This decision helps make the study efficient without with it has to distill the sub-themes by surveying the AHP literature. The data for analysis was gathered by s urfing Google on 26 Feb 2013 8.55 p.m. - 9.26 p.m. Peninsular Malaysian time. The search results were computed to generate two types of ratios specified earlier. A composite index was created using the re sulting two types of ratios which are used to class ify the efficiency, hence dominance of the original res ults (hits). Kendall’s correlation produced statist ically significant correlations between the composite inde x and Rank of AHP specific and area results. Using indices greater than 1.000 as the base, 6 AHP specific areas occupy the top positions with ratios ranging from 19.341 to 61.574; 15 AHP specific areas occupy the second top positions with ratios ranging from 1.119 to 9.602, and 14 AHP specific areas occupy the third and last position with indi ces below 1.000 ranging from 0.050 to 0.894. The paper includes dis cussion, implications, limitations, conclusions and suggestions for further research.
Read moreSemantic Web Search Engines : A Comparative Survey
Search engines play important role in the success of the Web. Search engine helps the users to find the relevant information on the internet. Due to many problems in traditional search engines has led to the development of semantic web. Semantic web technologies are playing a crucial role in enhancing traditional search, as it work to create machines readable data and focus on metadata. However, it will not replace traditional search engines. In the environment of semantic web, search engine should be more useful and efficient for searching the relevant web information. It is a way to increase the accuracy of information retrieval system. This is possible because semantic web uses software agents; these agents collect the information, perform relevant transactions and interact with physical devices. This paper includes the survey on the prevalent Semantic Search Engines based on their advantages, working and disadvantages and presents a comparative study based on techniques, type of results, crawling, and indexing.
Read moreLeveraging electronic healthcare record standards and semantic web technologies for the identification of patient cohorts.
The secondary use of electronic healthcare records (EHRs) often requires the identification of patient cohorts. In this context, an important problem is the heterogeneity of clinical data sources, which can be overcome with the combined use of standardized information models, virtual health records, and semantic technologies, since each of them contributes to solving aspects related to the semantic interoperability of EHR data. To develop methods allowing for a direct use of EHR data for the identification of patient cohorts leveraging current EHR standards and semantic web technologies. We propose to take advantage of the best features of working with EHR standards and ontologies. Our proposal is based on our previous results and experience working with both technological infrastructures. Our main principle is to perform each activity at the abstraction level with the most appropriate technology available. This means that part of the processing will be performed using archetypes (ie, data level) and the rest using ontologies (ie, knowledge level). Our approach will start working with EHR data in proprietary format, which will be first normalized and elaborated using EHR standards and then transformed into a semantic representation, which will be exploited by automated reasoning. We have applied our approach to protocols for colorectal cancer screening. The results comprise the archetypes, ontologies, and datasets developed for the standardization and semantic analysis of EHR data. Anonymized real data have been used and the patients have been successfully classified by the risk of developing colorectal cancer. This work provides new insights in how archetypes and ontologies can be effectively combined for EHR-driven phenotyping. The methodological approach can be applied to other problems provided that suitable archetypes, ontologies, and classification rules can be designed.
Read moreSemantic Web and Internet of Things: Challenges, Applications and Perspectives
The apparent growth of the internet of things (IoT) has allowed its deployment in many domains. The IoT devices sense their surroundings and transmit the data via the Web. According to statistics, due to the proliferation of smart devices, the number of active IoT devices is expected to exceed 25.4 billion by 2030.1 A large number of IoT objects gather an enormous amount of raw data. The data generated by various IoT objects and sensors are heterogeneous, with varying types and formats. Therefore, it is difficult for IoT systems to share and reuse raw IoT data, which causes the problem of lack of interoperability. The lack of interoperability in IoT systems creates a problematic issue that prevents IoT systems from performing well. To address this issue, data modeling and knowledge representation using semantic web technologies may be an appropriate solution to give meaning to raw IoT data and convert it to an enriched data format. The primary goal of this research section is to highlight the best outcomes for semantic interoperability among IoT systems, which can serve as a guideline for future studies via the presentation of a literature review on semantic interoperability for Internet of Things systems, including challenges, prospects, and recent work. The paper also provides an overview of the application of semantic web technologies in IoT systems, such as specific ontologies, frameworks, and application domains that use semantic technologies in the IoT areas to solve interoperability and heterogeneity problems.
Read moreDesign Issues for Search Engines and Web Crawlers: A Review
The World Wide Web is a huge source of hyperlinked information contained in hypertext documents. Search engines use web crawlers to collect these web documents from web for storage and indexing. The prompt growth of the World Wide Web has posed incomparable challenges for the designers of search engines and web crawlers; that help users to retrieve web pages in a reasonable amount of time. In this paper, a review on need and working of a search engine, and role of a web crawler is being presented.
Read more