- Book Chapter
6
- 10.1007/978-1-4614-7163-9_391-1
Stream Querying and Reasoning on Social Data
- Jan 01, 2017
- Jayanta Mondal + 1 more +1
2. INTRODUCTION Since the inception of online social networks, the amount of social data that is being published on a daily basis has been increasing at an unprecedented rate. Smart, GPS-enabled, always-connected personal devices have taken the data generation to a new level by making it tremendously easy to generate and share social content like check-in information, likes, microblogs (e.g., Twitter), multimedia data, and so on. There is an enormous value in reasoning about such streaming data and deriving meaningful insights from it in real-time. Examples of potential applications include advertising, sentiment analysis, detecting natural disasters, social recommendations, personalized trends, spam detection, to name a few. There is thus an increasing need to build scalable systems to support such applications. Complex nature of social networks and their rapid evolution, coupled with the huge volume of streaming social data and the need for real-time processing, raise many computational challenges that have not been addressed in prior work. Social network data comprises of two major components. First, there is a network (linkage) component that captures the underlying interconnection structure among the entities in the social network. Second, there is content data that is typically associated with the nodes and the edges in the social network. The social network data stream contains updates to both these components. The structure of the network may itself change rapidly in many cases, especially when things like webpages and user tags (e.g., Twitter hashtags) are treated as nodes of the network. However, most of the social network data stream consists of updates to the data associated with the nodes and the edges, e.g., status updates and other content uploaded by the users, communication among the users, and so on. There is interest in performing a wide variety of queries and analytics over such data streams in real time. The queries can range from simple publish-subscribe queries, where a user is interested in being notified when something happens in his friend circle, to complex anomaly detection queries, where the goal is to identify anomalous behavior as early as possible. In this paper, we present an introduction to this new research area of stream querying and reasoning over social data. This area combines aspects from several well-studied research areas, chief among them, social network analysis, graph databases, and data streams. We provide a formal definition of the problem, survey the related prior work, and discuss some of the key research challenges that need to be addressed (and some of the solutions that have been proposed). We note that we use the term stream reasoning in this paper to encompass a broad range of tasks including various types of analytics, probabilistic reasoning, statistical inference, and logical reasoning. We contrast our use of this term with the recent work by Valle et al. [104, 105] who define this term more specifically to refer to integration of logical reasoning systems with data streams in the context of the Semantic Web. Given the vast amount of work on this and related topics, it is not our intention to be comprehensive in this brief overview. Rather we aim to cover some of the key ideas and representative work.
Read more