- Supplementary Content
1
- 10.4225/28/5afa3eebb90f0
Mining people's semantic trajectory behaviours from geotagged photographs
- Jan 01, 2017
- Guochen Cai
Mining people's semantic trajectory behaviours from geotagged photographs
Events are ubiquitous in real-life. With the rapid rise of the popularity of social media channels, massive amounts of event data, such as information about festivals, concerts, or meetings, are increasingly created and shared by users on the Internet. Deriving insights or knowledge from such social media data provides a semantically rich basis for many applications, for instance, social media marketing, service recommendation, sales promotion, or enrichment of existing data sources. In spite of substantial research on discovering valuable knowledge from various types of social media data such as microblog data, check-in data, or GPS trajectories, interestingly there has been only little work on mining event data for useful patterns. \n \nIn this thesis, we focus on the discovery of interesting, useful patterns from datasets of events, where information about these events is shared by and spread across social media platforms. To deal with the existence of heterogeneous event data sources, we propose a comprehensive framework to model events for pattern mining purposes, where each event is described by three components: context, time, and location. This framework allows one to easily define how events are related in terms of conceptual, temporal, and spatial (geographic) relationships. Moreover, we also take into account hierarchies for contexts, time, and locations of events, which naturally exist as useful background knowledge to derive patterns at different levels of abstraction and granularity. Based on this framework, we focus on the following problems: (i) mining interval-based event sequence patterns, (ii) mining periodic event patterns, and (iii) extracting semantic annotations for locations of events. Generally, the first two problems consider correlations of events whereas the last one takes correlations of event components into account. \n \nIn particular, the first problem is a generalization of mining sequential patterns from traditional data, where patterns representing complex temporal relationships among events can be discovered at different levels of abstraction and granularity. The second problem is to find periodic event patterns, where a notion of relaxed periodicity is formulated for events as well as for groups of events that co-occur. The third~problem is to extract semantic annotations for locations on the basis of exploiting correlations of contexts, time, and locations of events. For the three problems above, we respectively propose novel and efficient approaches. Our experiments clearly indicate that extracted patterns and knowledge can be well utilized in various useful tasks, such as event prediction, semantic search for locations, or topic-based clustering of locations.
Mining people's semantic trajectory behaviours from geotagged photographs
Mining people's semantic trajectory behaviours from geotagged photographs
The Image of Bali Tourism in Social Networking Media
Tourism is growing very rapidly with the development of technology, especially information technology .Through information technology people can easily get access to information about various destinations and tourist attractions as well as hotel by online media communication. Therefore, this study was trying to use social media such as facebook and twitter, where both social media were widely used by companies engaged in tourism in the world as well as in Bali. Internet consists of several websites, but the most effective one to conduct the research on tourism such as webseries 2.0 or commonly called as social networking media. Social medias - Facebook and Twitter - both were wildly used by the companies running in the sector of tourism in Bali, such as Hotels, Villas, Restaurants, Spas, Travel Agents, Airlines, and Tourist Attraction Managements, to advertised their products and services in their respective field. Through those social media, tourists will be able to access the comments or reviews about the companies. The problem raised in this research is how Bali’s tourism data in social networking media were used and considered, and how was the image of Bali’s tourism in their point of viewbased on the research. The method used here was descriptive qualitative method which describes this phenomenon descriptively by analyzing the obtained comments on social networking media such as facebook and twitter, and then classified in the form of comments positive, negative, and unidentified comment. Afterwards comments were re-analyzed through 4A Approach which consist of Attraction, amenities, Ancillary, and Accessibility. Data which has been obtained re-analyzed via the data obtained from the homepage of (TripAdvisor and Agoda.com). In addition, to strengthen the research data obtained from the homepage of TripAdvisor and Agoda.com , then it has been elaborated on the point rating system which indicatedthat the company were registered in both of those social networking media such as Hotel, Villa, SPA, restaurants, travel agents, airlines, media, tourist attraction management has a good rating result, and it has represented that the image of Bali tourism has a good rank in the eyes of local and international tourists. This data was also strengthened by the existence of empirical studies where there were the data from a hundred tourists came to Bali both locally and internationally has filled out the questionnaire about Bali in order to obtain maximum results on the image of Bali tourism. The results showed that, social networking media were effectively used in this study. Many companies in Bali used facebook and twitter. In addition, Bali Tourism image placed in positive rank on both of social networking sites, and it was gained by the results of questionnaires spread towards the visitor.
Read moreBuilding a Social Media rating model
Social Media (SM) data are growing, and SM is becoming an acceptable part of daily life for billions of people around the world. Extracting information from Social Networking Sites (SNS) can provide great challenges as well as opportunities. Using SM data beyond day-to-day communication can provide additional values. There is much research and many products that are dedicated to take SNS beyond communication channels. In our research, we are going beyond specific tools inherent to the SM tools, such as Hashtag mentions and Like counts. Instead it will use text-based modeling, data mining techniques, natural process language, machine language, etc. to understand SM content to produce numeric ratings. The final contribution of this research is building a SM users' rating model for an event using SM data. At this point of our research, we are laying out a road map.
Read moreExploring digital remediation in support of personal reflection
We explored the transformation of social media content into three radically different formats; a book, a triptych of photographs and film.Using interviews, we investigated users' responses to their remediated data.Our findings were used to develop a digital curation process to assist the development of different tools for social media analysis and display.Our paper establishes design implications to aid reflection on social media use through remediation. Increasingly our digital traces are providing new opportunities for self-reflection. In particular, social media (SM) data can be used to support self-reflection, but to what extent is this affected by the form in which SM data is presented? Here, we present three studies where we work with individuals to transform or remediate their SM data into a physical book, a photographic triptych and a film. We describe the editorial decisions that take place as part of the remediation process and show how the transformations allow users to reflect on their digital identity in new ways. We discuss our findings in terms of the application of Goffman's (1959) self-presentation theories to the SM context, showing that a fluid rather than bounded interpretation of our social media spaces may be appropriate. We argue that remediation can contribute to the understanding of digital self and consider the design implications for new SM systems designed to support self-reflection.
Read moreIdentifying Different Types of Social Ties in Events from Publicly Available Social Media Data
Tie strength is an essential concept in identifying different kind of social ties - strong ties and weak ties. Most present studies that evaluated tie strength from social media were carried out in a controlled environment and used private/closed social media data. Even though social media has become a very important way of networking in professional events, access to such private social media data in those events is almost impossible. There is very limited research on how to facilitate networking between event participants and especially on how to automate this networking aspect in events using social media. Tie strength evaluated using social media will be key in automating this process of networking. To create such tie strength based event participant recommendation systems and tools in the future, first, we need to understand how to evaluate tie strength using publicly available social media data. The purpose of this study is to evaluate tie strength from publicly available social media data in the context of a professional event. Our case study environment is community managers’ online discussions in social media (Twitter and Facebook) about the CMAD2016 event in Finland. In this work, we analyzed social media data from that event to evaluate tie strength and compared the social media analysis-based findings with the individuals’ perceptions of the actual tie strengths of the event participants using a questionnaire. We present our findings and conclude with directions for future work.
Read moreFrom benchmark to bedside: transfer learning from social media to patient-provider text messages for suicide risk prediction.
Compared to natural language processing research investigating suicide risk prediction with social media (SM) data, research utilizing data from clinical settings are scarce. However, the utility of models trained on SM data in text from clinical settings remains unclear. In addition, commonly used performance metrics do not directly translate to operational value in a real-world deployment. The objectives of this study were to evaluate the utility of SM-derived training data for suicide risk prediction in a clinical setting and to develop a metric of the clinical utility of automated triage of patient messages for suicide risk. Using clinical data, we developed a Bidirectional Encoder Representations from Transformers-based suicide risk detection model to identify messages indicating potential suicide risk. We used both annotated and unlabeled suicide-related SM posts for multi-stage transfer learning, leveraging customized contemporary learning rate schedules. We also developed a novel metric estimating predictive models' potential to reduce follow-up delays with patients in distress and used it to assess model utility. Multi-stage transfer learning from SM data outperformed baseline approaches by traditional classification performance metrics, improving performance from 0.734 to a best F1 score of 0.797. Using this approach for automated triage could reduce response times by 15 minutes per urgent message. Despite differences in data characteristics and distribution, publicly available SM data benefit clinical suicide risk prediction when used in conjunction with contemporary transfer learning techniques. Estimates of time saved due to automated triage indicate the potential for the practical impact of such models when deployed as part of established suicide prevention interventions. This work demonstrates a pathway for leveraging publicly available SM data toward improving risk assessment, paving the way for better clinical care and improved clinical outcomes.
Read moreGaining historical and international relations insights from social media: spatio-temporal real-world news analysis using Twitter
The immense growth of the social Web, which has made a large amount of user data easily and publicly available, has opened a whole new spectrum for research in social behavioral sciences. However, as the volume of social media content increases at a very fast rate, it becomes extremely difficult to systematically obtain high-level information from this data. As a consequence, tasks related to the analysis of historical news events based on social media data have not been explored, which limits any type of comparative historical research, causality analysis, and discovery of knowledge from patterns extracted from aggregated social media event information.In this work, we target this issue by proposing a compact high-level representation of news events using social media information. This representation explicitly includes temporal information about the event and information about locations, in particular of geopolitical entities. We call this a spatio-temporal context-aware event representation. Our hypothesis is that by including social, temporal, and spatial information in the event representation, we are enabling the analysis of historical world news from a social and geopolitical perspective. This facilitates, new information retrieval tasks related to historical event information extraction and international relations analysis. We support our claims by presenting two applications of this idea: the first, a visual tool, named Galean, for retrieval and exploration of historical news events within their geopolitical and temporal context. The second, a quantitative analysis of a 2-year Twitter dataset of news events reported by U.S. and U.K. media, which we explore using data mining techniques on our event representations. We present two case studies of event exploration using Galean and user evaluation of this tool, as well as details of our data mining empirical results.
Read moreSentiment Computing for the News Event Based on the Social Media Big Data
The explosive increasing of the social media data on the Web has created and promoted the development of the social media big data mining area welcomed by researchers from both academia and industry. The sentiment computing of news event is a significant component of the social media big data. It has also attracted a lot of researches, which could support many real-world applications, such as public opinion monitoring for governments and news recommendation for Websites. However, existing sentiment computing methods are mainly based on the standard emotion thesaurus or supervised methods, which are not scalable to the social media big data. Therefore, we propose an innovative method to do the sentiment computing for news events. More specially, based on the social media data (i.e., words and emoticons) of a news event, a word emotion association network (WEAN) is built to jointly express its semantic and emotion, which lays the foundation for the news event sentiment computation. Based on WEAN, a word emotion computation algorithm is proposed to obtain the initial words emotion, which are further refined through the standard emotion thesaurus. With the words emotion in hand, we can compute every sentence’s sentiment. Experimental results on real-world data sets demonstrate the excellent performance of the proposed method on the emotion computing for news events.
Read moreSocial Media Data in the Big Data Environment
The article contains results of a study of social media data (SMD) which, being distinct from conventional data by their origin, require special methods for collection, processing and analysis. As shown by a literature review, in spite of great many research publications devoted to social media research and big data analysis, the SMD potential as a big data component still remains inadequately explored. Two approaches to research and analysis of SMD were highlighted in course of the study, in which SMD are addressed as an object of Internet statistics and an object of big data. When SMD are explored as an object of Internet statistics, collection of anonymized data is performed using the services that have network protocols for collection and analysis of data on social media customers using statistical methods. When SMD are explored as an object of big data, the collection is performed mostly by artificial intellect, whereas the storage and processing is operated by databases designed for large scopes of data and software with statistical data processing applications. The social media most popular with users in 2020 were identified in the study. Statistical indicators for assessment of users’ feedback, available now for statistical assessments of social media communities, are given. The study revealed several problems which solutions would require, apart from a multifaceted and complex approach to collection and processing, highly competent teams of specialists in various subject fields, including experts in computations, experts in machine learning and statisticians.
Read moreJournalists’ Use of Social Media to Infer Public Opinion: The citizens’ perspective
Journalists increasingly use social media data to infer and report public opinion by quoting social media posts, identifying trending topics, and reporting general sentiment. In contrast to traditional approaches of inferring public opinion, citizens are often unaware of how their publicly available social media data is being used and how public opinion is constructed using social media analytics. In this exploratory study based on a census-weighted online survey of Canadian adults (N=1,500), we examine citizens’ perceptions of journalistic use of social media data. We demonstrate that: (1) people find it more appropriate for journalists to use aggregate social media data rather than personally identifiable data; (2) people who use more social media are more likely to positively perceive journalistic use of social media data to infer public opinion; and (3) the frequency of political posting is positively related to acceptance of this emerging journalistic practice, which suggests some citizens want to be heard publicly on social media while others do not. We provide recommendations for journalists on the ethical use of social media data and social media platforms on opt-in functionality.
Read moreUtility of social media and crowd-intelligence data for pharmacovigilance: a scoping review
BackgroundA scoping review to characterize the literature on the use of conversations in social media as a potential source of data for detecting adverse events (AEs) related to health products.MethodsOur specific research questions were (1) What social media listening platforms exist to detect adverse events related to health products, and what are their capabilities and characteristics? (2) What is the validity and reliability of data from social media for detecting these adverse events? MEDLINE, EMBASE, Cochrane Library, and relevant websites were searched from inception to May 2016. Any type of document (e.g., manuscripts, reports) that described the use of social media data for detecting health product AEs was included. Two reviewers independently screened citations and full-texts, and one reviewer and one verifier performed data abstraction. Descriptive synthesis was conducted.ResultsAfter screening 3631 citations and 321 full-texts, 70 unique documents with 7 companion reports available from 2001 to 2016 were included. Forty-six documents (66%) described an automated or semi-automated information extraction system to detect health product AEs from social media conversations (in the developmental phase). Seven pre-existing information extraction systems to mine social media data were identified in eight documents. Nineteen documents compared AEs reported in social media data with validated data and found consistent AE discovery in all except two documents. None of the documents reported the validity and reliability of the overall system, but some reported on the performance of individual steps in processing the data. The validity and reliability results were found for the following steps in the data processing pipeline: data de-identification (n = 1), concept identification (n = 3), concept normalization (n = 2), and relation extraction (n = 8). The methods varied widely, and some approaches yielded better results than others.ConclusionsOur results suggest that the use of social media conversations for pharmacovigilance is in its infancy. Although social media data has the potential to supplement data from regulatory agency databases; is able to capture less frequently reported AEs; and can identify AEs earlier than official alerts or regulatory changes, the utility and validity of the data source remains under-studied.Trial registrationOpen Science Framework (https://osf.io/kv9hu/).
Read moreAt the Interface of Social Media Analytics, Big Data and Social Movements: Research Challenges
Social media is being employed to build support for social, economic, and political justice (Selander et al. 2016; Vaast et al. 2014) in movements such as occupy Wall Street, violence against women's movement and global sustainability movement. These movements have used social media in ways that goes beyond simple communications. As social media allows people to produce and share user generated content, they enable a certain set of affordances of these technologies (Bharati et al 2015; Bharati et al 2014). The social media affordances, available to both collective and individual actors, translate into capabilities afforded to social movements (Tufecki 2014). Research on connective action has examined the effects of digital action repertoires on interaction and engagement such as in the Tea Party and Occupy movements (Agarwal et al. 2014; Selander et al. 2016). Social media can also facilitate mobilization of movement and participation by new volunteers and oftentimes provides a transnational character by diffusing actions beyond the virtual (Van Laer and Van Aelst 2010). Conversation, an essential part of social movements, shapes “social life by altering individual and collective understandings, by creating and transforming social ties, by generating cultural materials that are then available for subsequent social interchange, and by establishing, obliterating, or shifting commitments on the part of participants&x201D; (Tilly 2002, p. 122). In a personal interaction that involves repeated organized interactions between individuals, typically, leads to shared values and trust. The role of social media technologies in furthering this conversation has to be studied and its' influence on social movement ascertained. A few scholars have started to investigate social media affordances and capabilities, especially focused on discourse, during contentious collective action. Still the research has been limited to studying mechanisms of participation, development of a sense of collective identity, creation of community, and framing of political discourse (Farrell 2012; Garrett 2006). Social media data on social movements can involve impersonal “like&x201D; and “share&x201D; to more engaged conversations. Twitter, Instagram, Facebook, and YouTube offer a wide communicative and discourse reach in networks with text, image, audio, and video data. This social media based big data, consisting mostly of unstructured data, comes in the form of social media posts, digital pictures and videos. The symposium will discuss how formal organizational structures and practices might be integrated with social media capabilities to reinforce and enhance social movement organizations as leaders in social change movements. It will explore the characteristics of social media discourse and assess the dynamics of social movements and, subsequent, impact on real-world protests. The symposium will also demonstrate how discourse analysis can be applied visually in order to understand communication patterns. The panel symposium will focus on theoretical and methodological challenges of social media analytics, big data and social movements. Panelists will engage the audience in an interactive discussion on: 1) Theoretical challenges: a. How and why are social movement recruitment and engagement mechanisms being impacted as a result of social media? b. How do we advance theory on social media and social movements when we are overwhelmed with social media based big data? c. What approach should we undertake if big data analysis contradicts most theories on social media and social movements? d. How do we address the issue of generalizability of social media and social movement research when data collection was limited to one social media platform, albeit involving big data? 2) Methodological challenges: a. What methodological approaches have worked in the analysis of social media, big data and social movements? b. How can discourse analysis be applied visually in order to understand communication patterns evident in social media-based big data? c. How can we employ social media analytics to investigate image and video data? d. What combinations of qualitative and quantitative methodologies be employed for big data and social media analytics in the context of social movements? e. What are the limitations of quantitative data analysis techniques, such as structural equation modeling, because of an extremely large sample size? f. What are the limitations of qualitative data analysis techniques as they become extremely labor intensive and, maybe even, impractical because of big data?
Read moreClinicians' Reports in Electronic Health Records Versus Patients' Concerns in Social Media: A Pilot Study of Adverse Drug Reactions of Aspirin and Atorvastatin.
Large databases of clinician reported (e.g., allergy repositories) and patient reported (e.g., social media) adverse drug reactions (ADRs) exist; however, whether patients and clinicians report the same concerns is not clear. Our objective was to compare electronic health record data and social media data to better understand differences and similarities between clinician-reported ADRs and patients' concerns regarding aspirin and atorvastatin. This pilot study explored a large repository of electronic health record data and social media data for clinician-reported ADRs and patients concerns for two common medications: aspirin (n=31,817 ADRs accessible in clinical data; n=19,186 potential ADRs accessible in social media data) and atorvastatin (n=15,047 ADRs accessible in clinical data; n=23,408 potential ADRs accessible in social media data). We found that the most frequently reported ADRs matched the most frequent patients' concerns. However, several less frequently reported reactions were more prevalent on social media (i.e., aspirin-induced hypoglycemia was discussed only on social media). Overall, we found a relatively strong positive and statistically significant correlation between the frequency ranking of reactions and patients' concerns for atorvastatin (Pearson's r=0.61, p<0.001) but not for aspirin (Pearson's r=0.1, p=0.69). Future studies should develop further natural language methods for a more detailed data analysis (i.e., identifying causality and temporal aspects in the social media data).
Read moreUsing social media data to study political science
Social media are now involved in many aspects of human life. People use social media to do business, find friends, have fun, and discuss social and political issues. As an important aspect of our lives, and just like many other topics, politics is widely discussed on social media. Politicians, activists, civilians, companies and terrorists have made politics a hot topic on social media. However, there is no consensus among researchers on how reliable social media data can be for political science research, and what would be the proper method of collecting and analyzing social media data. In this paper, various theoretical and empirical works concerning the relationship between social media data, specifically Twitter data, and politics are critically examinedto demonstrate how social media data affect politics and contribute to political research. The findings imply that social media data have significantly contributed to the field of political communication by offering inexpensive and easily-accessible information, empowering marginal social entities to participate in politics and internationalizing communication among political actors.
Read moreA Total Error Approach for Validating Event Data
Understanding how useful any particular set of event data might be for conflict research requires appropriate methods for assessing validity when ground truth data about the population of interest do not exist. We argue that a total error framework can provide better leverage on these critical questions than previous methods have been able to deliver. We first define a total event data error approach for identifying 19 types of error that can affect the validity of event data. We then address the challenge of applying a total error framework when authoritative ground truth about the actual distribution of relevant events is lacking. We argue that carefully constructed gold standard datasets can effectively benchmark validity problems even in the absence of ground truth data about event populations. To illustrate the limitations of conventional strategies for validating event data, we present a case study of Boko Haram activity in Nigeria over a 3-month offensive in 2015 that compares events generated by six prominent event extraction pipelines—ACLED, SCAD, ICEWS, GDELT, PETRARCH, and the Cline Center’s SPEED project. We conclude that conventional ways of assessing validity in event data using only published datasets offer little insight into potential sources of error or bias. Finally, we illustrate the benefits of validating event data using a total error approach by showing how the gold standard approach used to validate SPEED data offers a clear and robust method for detecting and evaluating the severity of temporal errors in event data.
Read more