- Research Article
25
- 10.1002/1944-2866.poi326
Addressing the policy challenges and opportunities of “Big data”
- Jun 01, 2013
- Policy & Internet
- Helen Margetts + 1 more +1
Addressing the policy challenges and opportunities of “Big data”
Extract Large and complex datasets collecting digital breadcrumbs of human behaviour have become non-negligible sources for social science researchers, students, and practitioners. ‘Big data’ has made a place for itself as a commodity in academic, commercial, and non-commercial settings. There have been numerous new descriptions for defining social scientists with quantitative expertise: empirical social scientist, computational social scientist, quantitative social scientists, social data scientist, etc. Nonetheless, a social scientist with solid big data skills is not yet a common profile in academia and research institutions. Neither is a computer scientist with broad interdisciplinary skills in social sciences. As the amount of published material increases exponentially in many disciplines, specialisation seems to be a reasonable response to keep up with the state of the art, and to produce technically and conceptually solid research. Interdisciplinary teams are vital in providing the critical synthesis from these separate specialisations, and this book is put together to help bridge the language gap between the different specialisations needed for migration analysis from big data sources.
Addressing the policy challenges and opportunities of “Big data”
Addressing the policy challenges and opportunities of “Big data”
Taking a byte out of big data
Taking a byte out of big data
Session details: Doctoral Consortium Abstracts
The Hypertext 2015 Doctorial Consortium session is held the first day of the 26th ACM Conference on Hypertext and Social Media. It offers Ph.D. students an opportunity to present their ongoing research towards obtaining a Ph.D. degree in the disciplines related to the conference, mostly in Computer and Information Sciences but not limited to them. The themes of the submissions are strongly connected to the research tracks of this year's conference: Digital Connectivity, Data Connectivity and Digital Humanities. The Digital Connectivity track targets developing insights into the mechanisms of information generation and dissemination, characterization of evolutionary processes on online social networks, and studies of models and systems that support these processes. The second track, Data Connectivity, deals with the methods, techniques and technologies that can be used to make data available on the Web, with a special focus on how heterogeneous data sources can be connected to each other. Finally, the track of Digital Humanities seeks to attract work from an interdisciplinary perspective, on the intersection between computer science on one hand, and the humanities and social sciences on the other. During the DC session, students receive constructive feedback from their peers and from a panel of mentors to support them in envisioning added-value contributions to the state-of-the-art research in Hypertext and Social Media. The consortium session is open to all doctorial students by application. Each submission was reviewed by three senior researchers in the topics the students made their submissions. The DC session is aimed in particular at students who have defined a dissertation topic but are still more than one year from graduating at the time of application in order to obtain benefits from feedback. This year, the committee has accepted three contributions to be presented during the Doctorial Consortium session: Automated Methods for Identity Resolution in Heterogeneous Social Platforms by Paridhi Jahin. In this work, connected to the digital and data connectivity tracks of the conference, the author proposes novel methods to search and link user identities scattered across heterogeneous social networks. The methods consider carefully users' privacy, so they are designed to access only public and historic data. The evaluation is proposed on large datasets over multiple platforms to prove their significance in identity resolution of an online user. Language Innovation and Change in On-line Social Networks by Daniel Kershaw. As a fundamental aspect for human communication, this research focuses on forecasting online language change through the use of predictive and descriptive methodologies. This work is framed within structuration theory which helps the researcher in structuring the analysis of the dynamics of language (re)production - i.e. by the agent (user), the social structure and their interplay. A Framework to Provide Customized Reuse of Open Corpus Content for Adaptive Systems by Mostafa Bayomi. This work deals with issues of reusability of contents on adaptive systems. Adaptive systems tailor content specific to user's needs, but they face two big problems: (a) they use a closed corpus content that has been prepared for them a priori, and (b) the content is tightly coupled with other parts of the system, which hinders its reusability. This work presents a proposal that leverages the semantic web by extending an existing content provision system, Slice-pedia.
Read moreComparison of Academic Procrastination in University Health and Social Science Students
Academic procrastination in university students of social and health sciences was compared according to socio-academic variables, such as: age, gender, occupation, area and year of study. The research was descriptive-comparative, quantitative, non-experimental; 1000 university social and health sciences students intentionally selected according to quotas participated. The information was collected with a duly validated instrument about academic procrastination. It was observed that: the majority of students perceive high academic procrastination (62%), low academic self-regulation (41%) and high procrastination of activities (74%). Concluding that the trend of academic procrastination is mostly present in students of social sciences, in male students, in the first years or academic cycles and in those who are studying while working. Therefore, most students who procrastinate avoid prioritizing the development of academic activities for others that are of particular interest such as: excessive use of technology, social networks, dependence on cell phones and work.
Read moreThe use of Big Data in Official Statistics
The thesis concerns the use of big data in Official Statistics; the aim is to bring some new experimental studies into the big data literature. In particular, the purpose is to evaluate if and how big data could be used in Official Statistics. Only a few experiments exists on the use of big data for statistical purposes; it is a challenge task and a lot of experimentation is needed to find out evidence and solutions to use big data for statistical purposes. The analysis performed in the thesis goes into two different directions: 1. combining a traditional data source with a big data source to verify the potential of the latter to replicate official results; 2. analyzing a big data source per se and then trying to combine with an Official Statistics source to identify common patterns. The thesis initially proposes a literature review of definitions of big data and experiments, in particular concerning the use of the new sources combined with traditional data sources. Then, three original studies have been performed: the first two concern mobility in Lombardy region using mobile phone data. They both refer to the same issue (mobility patterns), but they differ in the traditional data source used: Origin/Destination matrix in the first case, an integrated version of the O/D matrix in the second. The objective of these two studies is trying to put in a unique interpretative framework one traditional statistical source and one typical kind of big data in order to evaluate some informative potentialities of this approach. In particular, we wanted to check if the two sources show common patterns, to evaluate future uses of the big data source in Official Statistics. The third study shows the pilot that was carried out during the traineeship I had the opportunity to attend at Eurostat, in collaboration with the Task Force Big Data. It concerns the use of Wikipedia, free online encyclopedia, for Tourism Statistics. The aim is to evaluate the use of Wikipedia page views as a source of information for the identification of factors that drive tourism to an area and whether it is possible to predict tourism flows using these data. A final chapter proposes conclusions and future remarks on the use of big data in Official Statistics. Two of the studies (the first on mobility patterns and the one on Wikipedia) have been or are being published, in a shorter and revised version. The three experiments show some potential in the use of big data in Official Statistics. The study needs more in-depth analysis, many more experiments and considerations will be necessary before we can achieve some definitive and convincing approaches.
Read moreUtilizing big data analytics for information systems research: challenges, promises and guidelines
This essay discusses the use of big data analytics (BDA) as a strategy of enquiry for advancing information systems (IS) research. In broad terms, we understand BDA as the statistical modelling of large, diverse, and dynamic data sets of user-generated content and digital traces. BDA, as a new paradigm for utilising big data sources and advanced analytics, has already found its way into some social science disciplines. Sociology and economics are two examples that have successfully harnessed BDA for scientific enquiry. Often, BDA draws on methodologies and tools that are unfamiliar for some IS researchers (e.g., predictive modelling, natural language processing). Following the phases of a typical research process, this article is set out to dissect BDA’s challenges and promises for IS research, and illustrates them by means of an exemplary study about predicting the helpfulness of 1.3 million online customer reviews. In order to assist IS researchers in planning, executing, and interpreting their own studies, and evaluating the studies of others, we propose an initial set of guidelines for conducting rigorous BDA studies in IS.
Read moreUnlocking big data: at the crossroads of computer science and the social sciences
Digital data on human behaviour and social interactions are a seemingly abundant and valuable resource of the twenty-first century, promising deep insights into the social processes generating them. Such data - commonly referred to as digital trace data, process-generated data, or digital behavioural data - are increasingly available to us, but the goal of unlocking their potential remains elusive despite the emergence of specialized fields such as computational social science, which has produced an ample body of research based on this resource. Achieving that goal requires revisiting the foundations of research based on this data type, a thorough understanding of its unique characteristics, overcoming access barriers, understanding data generating processes, and rethinking the role of theory at the intersection between computer science and social science. This chapter discusses the origin and characteristics of such data, as well as the interdisciplinary challenges researchers face in accessing, understanding, and using it for research.
Read moreCollaboration Between Social Sciences and Computer Science: Toward a Cross-Disciplinary Methodology for Studying Big Social Data from Online Communities
Collaboration Between Social Sciences and Computer Science: Toward a Cross-Disciplinary Methodology for Studying Big Social Data from Online Communities
Read moreHandvatten voor een kwaliteitsbeoordeling van big data: de introductie van het Total Error raamwerk
Assessing the methodological quality of big data: an introduction to the Total Error framework The availability and use of big data sources is increasing exponentially. The variety of new and emerging data sources offers opportunities to complement, replace, improve or add to conventional data sources. Survey data are one kind of conventional data sources. In survey research, a framework to assess the accuracy of survey data already existed for quite some time. This framework is known as the Total Survey Error (TSE) framework. The philosophy behind this framework has only recently been universalized to (big) data in general in the form of the Total Error (TE) framework. This generic framework, which allows for assessing the accuracy of (big) data, is outlined in this article. Additionally, the TE framework is applied to big data sources that could be relevant for policing: police-registered crime data, Twitter data and mobile phone data.
Read moreThe perils of working with Big Data and a SMALL framework you can use to avoid them
The use of "Big Data" to explain fluctuations in the broader economy or guide the business decisions of a firm is now so commonplace that in some instances it has even begun to rival more traditional government statistics and business analytics. Big data sources can very often provide advantages when compared to these more traditional data sources, but with these advantages also comes the potential for pitfalls. We lay out a framework called SMALL that we have developed in order to help interested parties as they navigate the big data minefield. Based on a set of five questions, the SMALL framework should help users of big data spot concerns in their own work and that of others who rely on such data to draw conclusions with actionable public policy or business implications. To demonstrate, we provide several case studies that show a healthy dose of skepticism can be warranted when working with and interpreting these new big data sources.
Read moreKnowledge transfer in agent-based computational social science
Knowledge transfer in agent-based computational social science
Why Big Data?: Why Nursing?
Attention on “big data” spans nursing and the health sciences, and extends as well to engineering/computer sciences through to the liberal arts in professional literature. A current Google search (3 Nov 2016) of “big data” yields 288 million entries. A focused search of “big data and nursing” yields more than 3.9 million entries. Thus we ask, “Why big data? Why nursing?” The focus of this chapter is to provide an overview of why big data has emerged now and to make the case for how big data has the capacity to change health, healthcare systems, and nursing. This chapter lays a foundation for the chapters and case studies to follow that explore what data, knowledge, and transformation processes are needed to put information and knowledge into the hands of nursing wherever nurses are working. In this chapter we examine the big data sources within and beyond nursing and healthcare that can be collected and analyzed to improve nursing and patient, family and community health. This chapter entices the reader to examine “Why big data now?” and “Why big data in the future?” This chapter is meant to stir curiosity for “Why should I be knowledgeable?” Whether the reader’s role is in clinical practice, education, research, industry, or policy, the applied uses of big data analytics are empowering change at an exponential speed across all domains. Big data has the capacity to illuminate nursing’s discovery of new knowledge and best practices that are safe, effective and lead to improved outcomes including well-being of providers; it also can expand nursing’s vision and future possibilities through increasing awareness of what nursing doesn’t know. The importance of nursing’s lens on the new discoveries obtained through big data and data science is critical to the transformation of health and healthcare systems. This transformation completes the challenge of placing the person at the center of all care initiatives and actions.
Read moreA DECADE IN INTERNET TIME
This introductory article provides a critical assessment of the last decade of social research on the Internet and identifies directions for research over the next. Ten years is only a moment in the span of social research, but aeons in Internet time. Has social research across the disciplines been up to the challenges? Over more than 40 years, the unfolding development of the Internet and related information and communication technologies has been one of the most dynamic areas of technological and social innovation worldwide. In the first decade of the twenty-first century, its development was even more dramatic. While innovations in such areas as search, social media, big data, and the commercialization of the Internet became prominent only over the last decade, they are already taken for granted by most Internet users. It is becoming increasingly apparent to us that this interdisciplinary field must broaden even further to better connect with fields beyond the social sciences and information, communication, media and cultural studies to include stronger collaborative ties to law, ethics, and the sciences, engineering, and computer sciences, but also across the arts and humanities. This will be an almost certain requirement for interdisciplinary research over the coming decade with technology and society moving at Internet time.
Read moreNegotiating AI fairness: a call for rebalancing power relations
AI fairness is at the center of many debates, but there are different perspectives on what it entails. Currently, purely technical algorithmic fairness approaches dominate the scene, often neglecting a sufficiently well-rounded view of social implications and ignoring the voices of lay people. With the end goal of overcoming these issues, we investigate and synthesize the points of contact and differences among computer science, sociological and lay people perspectives and move towards a lay-socio-technical view of AI fairness. To this end, we conducted interviews with experts from the computer and social sciences, as well as a survey and co-creation workshops with lay people. Our results show that there are sometimes conflicting views on AI fairness across these perspectives. An integrated view requires two processes of negotiation: (a) between computer and social sciences, and (b) between experts from both disciplines and lay people. We identify strategies for supporting these processes, but we state that they are ultimately possible only if there is a rebalance of current power relations between the disciplines, and between experts and lay people.
Read moreDisciplinary dimensions and social relevance in the scientific communications on biofuels
The disciplinary structure of research on complex problems related to human activities is supported by the fundaments of the social, life, and hard sciences. In this work, we looked at the development of scientific research in the field of biofuels, as a sustainable source of energy, searching for references regarding its scientific roots and social relevance. Scientific communications on biofuels published between 1998 and 2007 were analyzed using a combination of bibliometric methods and text mining techniques. This field of research was characterized as interdisciplinary, with marked social relevance. Our bibliometric analysis shows that, in this research subject, 132 different, interacting fields of knowledge overlap, with dominance of Chemistry, Engineering and Agricultural Sciences. Through the use of text mining techniques, this field was configured into three groups of Disciplinary Dimensions. The first and most influential group includes the Agricultural Sciences, Social Sciences, and Environmental Sciences. The second group, which gives the field its technological basis, includes Chemistry, Engineering, and Microbiology. The third group includes disciplines with emerging involvement in the field of biofuels: Biology and Biochemistry, Animal and Plant Sciences, Molecular Biology and Genetics, Economics, Material Sciences, Nanosciences and Nanotechnology, Geosciences, Physics, Humanities, Multidisciplinary Sciences, Mathematics, and Computer Sciences. This study suggests that the first group of Disciplinary Dimensions conforms to the elements that socially validate the progress of research in the field of biofuels. This study also proposes a metric that can be used to measure the interdisciplinarity and the social framing of any other research field.
Read more