- Research Article
18
- 10.1016/s0169-023x(01)00031-3
Decomposing relationship types by pivoting and schema equivalence
- Oct 01, 2001
- Data & Knowledge Engineering
- Sven Hartmann
Decomposing relationship types by pivoting and schema equivalence
This paper presents on estimation method of input-output table construction with private business data offered by the Center for TDB Advanced Data Analysis and Modeling.1 We estimated part of Input-Output Table and its algebraic descriptions with Exchange Algebra, AADL, FALCONSEED. The paper presents algorithms and heuristics to classify and normalize private business data in order to construct economic statics.
Decomposing relationship types by pivoting and schema equivalence
Decomposing relationship types by pivoting and schema equivalence
Functional and multivalued dependencies in nested databases generated by record and list constructor
The impact of the list constructor on two important classes of relational dependencies is investigated. Lists represent an inevitable data structure whenever order matters and data is allowed to occur repeatedly. The list constructor is therefore supported by many advanced data models such as genomic sequence, deductive and object-oriented data models including XML. The article proposes finite axiomatisations of functional, multivalued and both functional and multivalued dependencies in nested databases supporting record and list constructor. In order to capture different data models at a time, an abstract algebraic approach based on nested attributes is taken. The presence of the list constructor calls for a new inference rule which allows to infer non-trivial functional dependencies from multivalued dependencies. Further differences to the relational theory become apparent when the independence of the inference rules is investigated. The extension of the relational theory to nested databases allows to specify more real-world constraints and increases therefore the number of application domains.
Read morePADL2Java: A Java code generator for process algebraic architectural descriptions
One of the main objectives of model-driven software engineering is to produce code automatically from high-level design models. This goal can be achieved by providing suitable models and model-to-code transformations that ensure the conformance of the produced code to its high-level specification. In this context we have developed PADL2Java, a software tool that translates PADL models into Java code. PADL is a process algebraic architectural description language equipped with a rigorous semantics and transformation rules into multithreaded object-oriented software, which is employed in the verification tool TwoTowers. This paper discusses the code generation approach underlying PADL2Java, the structure of the synthesized code, and the integration of the translator in TwoTowers. The effectiveness of PADL2Java is illustrated through the generation of a Java implementation of a cruise control system.
Read moreUsing the RAS Technique as a Test of Hybrid Methods of Regional Input–Output Table Updating
DEWHURST J. H. LL. (1992) Using the RAS technique as a test of hybrid of regional input–output table updating. Reg. Studies 26, 81–91. Hybrid methods of regional input–output table construction appear to offer cheaper ways of constructing such tables than using a full survey. In addition they are claimed to give more accurate results than short–cut methods of table construction. However empirical tests of the hybrid method of table construction are few. This paper presents an empirical test of such methods of updating tables using the 1973 and 1979 Scottish input–output tables. The results support the claim of advocates of hybrid tables that the accuracy of tables can be improved significantly by the use of a limited amount of superior information. DEWHURST J. H. LL. (1992) l'utilsation de la technique dite RAS afin de mettre à l'épreuve les méthodes hybrides qui permettent la mise à jour des tableaux d'échanges intersectoriels régionaux, Reg. Studies 26, 81–91. Les méthodes hybrides visant la construction des tableaux d'échanges intersectoriels fournissent, semble-t-il, des moyens moins chers de construire de tels tableaux par rapport à une enquête détaillée. En plus elles sont censées fournir des résultats plus précis que ne le font les méthodes abrégées. Toujours est-il que peu nombreux sont les tests empiriques de la méthode hybride de la construction des tableaux. A partir des tableaux d'échanges intersectoriels pour l'Écosse et datant de 1973 et de 1979, cet article présente un test empirique de telles méthodes qui mettent à jour les tableaux. Les résultats viennent à l'appui de l'affirmation deceux qui sont en faveur des tableaux hybrides: à savoir, la précision des tableaux s'améliore sensiblement au fur et à mesure que l'on se sert d'une quantité limitée d'information plus détaillée. DEWHURST J. H. LL. (1992) Zur Anwendung der RAS-Technik als Eignungsprüfung der gemischten Methoden, regionale Aufwand-Ertragstabellen auf den neuesten Stand zu bringen, Reg. Studies 26, 81–91. Gemischte Methoden der Aufstellung regionaler Aufwand–Ertragstabellen bieten anscheinend billigere Arten der Erstellung solcher Tabellen als die Anwendung einer erforschenden Aufnahme. Zusätzlich wird der Anspruch erhoben, dass sie zu genaueren Ergebnissen führen als Abkürzungsmethoden der Tabellenerstellung. Es gibt jedoch nur wenig empirische Überprüfungen der gemischten Methode der Tabellenerstellung. Dieser Aufsatz stellt eine empirische Eignungsprüfung solcher Methoden, Tabellen auf den neuesten Stand zu bringen, dar, wobei die schottischen Aufwand-Ertragstabellen der Jahre 1973 und 1979 benutzt werden. Die Ergebnisse bestätigen den Anspruch derer, die gemischte Tabellen befürworten, dass die Genauigkeit der Tabellen durch die Anwendung einer beschräankten Menge überlegener Daten bedeutend verbessert werden kann.
Read moreUnlocking Bottomhole Dynamics During Hydraulic Fracturing: Insights from Pressure and Temperature Measurements from Surface to Frac Face
This paper aims to analyse bottomhole pressure and temperature data acquired from multiple gauges during various hydraulic fracturing treatments. The study seeks to identify unusual bottomhole events, through analysis of large datasets, and correlate these events with key treatment parameters. Understanding of bottomhole pressure and temperature signatures provides deeper insights into downhole processes, enabling improved treatment design and execution. A data-driven approach is employed to extract meaningful insights from the extensive dataset of bottomhole pressure and temperature recordings from surface, middle of string, fracture entry point and below. Advanced data modelling and analytics techniques are applied to process and interpret the information effectively. The methodology includes identifying discrepancies between measured and calculated bottomhole pressure, evaluating temperature variations throughout the treatment process, and analysing fluid friction behaviour along the wellbore. Several sources of downhole pressure and temperature measurements were compiled, thus enabling a comprehensive assessment of hydraulic fracturing events. The analysis reveals several interesting observations regarding pressure and temperature behaviour during hydraulic fracturing treatments. Key findings include discrepancies between measured and calculated bottomhole pressure, indicating potential variations in fracture propagation or wellbore conditions. Additionally, the impact of slurry flow on different pipe diameters is explored, demonstrating how it can create misleading pressure interpretations. Temperature data analysis shows an initial rise at the start of the treatment, followed by a variable temperature evolution pattern during the job. Notably, a decrease in temperature away from fracture entrance point is observed, often dropping below the temperature at the fracturing port. Furthermore, fluid friction analysis highlights variations along the pipe, influenced by different injected fluid types. These insights contribute to a better pressure decline interpretation and better understanding of bottomhole processes, like fracture initiation and propagation, helping to optimize fracturing designs and operational strategies. This study describes the workflow of analysing large datasets of bottomhole pressure and temperature measurements from multiple hydraulic fracturing treatments. Field proven data handling and modelling tools applied in a computer programming language (Python), made possible the compilation and analysis of large amount of data, providing new perspectives on pressure and temperature events, potentially overlooked previously. The findings aim to assist with better understanding of bottomhole dynamics, thus provide additional inputs to improving the completion and treatment designs, as well as optimize operational planning and execution of hydraulic fracturing operations.
Read moreEstimation of SAM for India: An Application for India’s Energy Transition Targets
The social accounting matrix (SAM) for India was historically constructed based on the input–output table (IOT). However, since 2011–2012, the Government of India has been publishing the supply–use table, instead of the IOT. While the erstwhile IOT, published by the Government of India, had the same number of products and industries, the supply–use table provides one ‘supply matrix’ and one ‘use matrix’, each of which is a rectangular table with 140 products and 66 industries (for 2018–2019). Converting the supply–use table to a square IOT and subsequently extending it to SAM require the utilisation of various data sources, numerous steps and adjustments. Despite the usefulness of both IO and SAM matrices in macroeconomic policy design, not much literature is available on these. This study aims to bridge the gap by constructing IOT and SAM for India from the supply–use table incorporating information from many other sources and describing the method of construction of the matrices. Our IO and SAM also focus on various energy sectors, including different sources of power generation, biomass and so on, and disaggregate energy-intensive sectors like cement or aluminium, considering the immense usefulness of the energy-extended macro-structure to the research of energy and environment policies. The study focuses on the construction of a 59 × 59 SAM for India, with the base year of 2021–2022 incorporating three factors of production and ten categories of households. As an application of the newly constructed SAM, we have analysed the employment implication of India’s Nationally Determined Contribution emission commitments. JEL Codes: E16, C67, D57
Read more3D Water Saturation Estimation Using AVO Inversion Output and Artificial Neural Network; A Case Study, Offshore Angola
Thanks to high-resolution seismic enabled by hardware and software technology advances, 3D seismic surveying techniques have become particularly effective in hydrocarbon prospection and development. This paper, based on two seismic datasets, well data and production history from deep-offshore Angola, Block 4, highlights how seismic data may be used for both reservoir characterization (static model and in-place) and reservoir monitoring (saturation changes). In the following we shall present in detail the applied seismic characterization method, with particular emphasis on: (i) advanced data conditioning and modeling, (ii) sequential elastic inversion for the integration of two datasets of different vintages, (iii) utilization of Bayesian techniques for reservoir characterization and probability density functions and finally (iv) Neural Network application for the assessment of oil saturation. The high resolution seismic data contains strong AVO signature, but to exploit these fully, preserving the signal throughout the process is essential: (i) seismic inversion feasibility analysis and synthetic gathers confirmed the quality of the far offset up to 55 degrees; (ii) this allows producing reliable density inversion products, using data beyond 35 degrees angles with adequate angle bands (iii) to get best inversion results a tailored pre-stack data conditioning was applied using five partial stacks with particular attention to amplitude preserving and flattening at reservoir level; (iv) finally well and seismic results were integrated through inversion and litho-classification, which compared well with well data. The high-fidelity inversion results allowed trying a novel approach through an additional step: saturation estimation based on Neural Network (NN) technique: (i) a labeled data set of elastic well logs, namely, P-wave Impedance, P-wave and S-wave velocities ratio and Density were used to train a Neural Network engine to estimate the water saturation from petrophysics. A low level of cross-validation error was achieved and deemed acceptable. The trained NN model seemed to be able to estimate the Water Saturation logs accurately as confirmed by blind well tests. A saturation cube was generated by applying the trained NN model on the three properties established through pre-stack inversion. Owing to the excellent quality of the recent (2018) high resolution seismic data, and despite log quality issues (poor borehole condition at some wells), high-fidelity elastic inversion could be achieved for both datasets. This in-turn led to a significant reservoir characterization uplift (additional in-place, better segregation and distribution of reservoir facies for simulation). Finally, the comparison between inversion results from the older dataset (prior to production) and the recent one allowed highlighting saturation variations pointing to swept and unswept portions of the field.
Read moreEntropy-based Chinese city-level MRIO table framework
Cities are pivotal hubs of socioeconomic activities, and consumption in cities contributes to global environmental pressures. Compiling city-level multi-regional input-output (MRIO) tables is challenging due to the scarcity of city-level data. Here we propose an entropy-based framework to construct city-level MRIO tables. We demonstrate the new construction method and present an analysis of the carbon footprint of cities in China's Hebei province. A sensitivity analysis is conducted by introducing a weight reflecting the heterogeneity between city and province data, as an important source of uncertainty is the degree to which cities and provinces have an identical ratio of intermediate demand to total demand. We compare consumption-based emissions generated from the new MRIO to results of the MRIO based on individual city input-output tables. The findings reveal a large discrepancy in consumption-based emissions between the two MRIO tables but this is due to conflicting benchmark data used in the two tables.
Read moreTeaching data modeling
While competition for scarce space in a Database Systems course curriculum increases, the amount of time spent in many such courses on data modeling decreases. We instead recommend increasing the amount of time spent in the study of data modeling and encourage data model study beyond formalism syntax. We do this in an attempt to help computer science students better understand complex data domains and to help develop higher-level skills that serve them well in a job market threatened by the increased outsourcing of lower level programming jobs. We further recommend the study of process skills as part of data modeling, and develop the idea of data patterns to assist students in the development of advanced data modeling skills.
Read moreConstruction of an Input-Output Table Considering Business-to-Consumer Transactions by using Private Data
The current input-output table takes about 4 years from the start of the survey to the release, it is not suitable for analysis of real-time industries, and it takes a lot of labor to create. In order to solve this problem, the input-output table is being constructed using the corporate data possessed by Teikoku Databank (TDB) Co., Ltd. However, since the calculation of private consumption is based on the total sales of (BtoC) companies that are selling to consumers, it is not possible to record in the spent where it was consumed. Therefore, in this research, we estimate the sales ratio to consumers in the BtoC company and calculate the total sales amount to consumers in the sales of the company. Also, by dividing the total of the private consumption obtained proportionally according to the population size of each region, we construct an algorithm to calculate private consumption in each region.
Read moreWhat Can Be Learned from Billions of Invoices? The Construction and Application of China’s Multiregional Input-Output Table Based on Big Data from the Value-Added Tax
The big data on the value-added tax (VAT), which links upstream production and downstream consumption, can serve as a key basis for a structured macroeconomic analysis. This paper discusses the construction and aggregation method of the transaction matrix between enterprises based on this data and develops a multiregional input-output (MRIO) table. The intraregional structure of this table is consistent with the IO table in China’s national statistics and includes more detail on the domestic value chain. We use big data on the VAT in creating an analytical tool to track profit shifting and tax avoidance. Therefore, this new MRIO table can coordinate consumption-based measures with the “destination principle” in international VAT statistics.
Read moreApplication of the AI-Based Framework for Analyzing the Dynamics of Persistent Organic Pollutants (POPs) in Human Breast Milk
Human milk has been used for over 70 years to monitor pollutants such as polychlorinated biphenyls (PCBs) and organochlorine pesticides (OCPs). Despite the growing body of data, our understanding of the pollutant exposome, particularly co-exposure patterns and their interactions, remains limited. Artificial intelligence (AI) offers considerable potential to enhance biomonitoring efforts through advanced data modelling, yet its application to pollutant dynamics in complex biological matrices such as human milk remains underutilized. This study applied an AI-based framework, integrating machine learning, metaheuristic hyperparameter optimization, explainable AI, and postprocessing, to analyze PCB-170 levels in breast milk samples from 186 mothers in Zadar, Croatia. Among 24 analyzed POPs, the most influential predictors of PCB-170 concentrations were hexa- and hepta-chlorinated PCBs (PCB-180, -153, and -138), alongside p,p’-DDE. Maternal age and other POPs exhibited negligible global influence. SHAP-based interaction analysis revealed pronounced co-behavior among highly chlorinated congeners, especially PCB-138–PCB-153, PCB-138–PCB-180, and PCB-180–PCB-153. These findings highlight the importance of examining pollutant interactions rather than individual contributions alone. They also advocate for the revision of current monitoring strategies to prioritize multi-pollutant assessment and focus on toxicologically relevant PCB groups, improving risk evaluation in real-world exposure scenarios.
Read moreModeling Emission Reductions and Forest Carbon Sequestration in GTAP: Data Base and Model Improvements
Forest carbon sequestration (FCS) is considered a cost-effective method to mitigate climate change. This has attracted the attention of policy makers and academics in the last quarter century. Nevertheless, only few efforts have incorporated the FCS global supply in a global economic modeling framework. In our research, we developed the GTAP-BIO-FCS model, which offers a more comprehensive basis for climate change mitigation including alternatives such as FCS and biofuels. GTAP-BIO-FCS is a reliable model that satisfies accounting balances and provides useful tools such as detailed description of welfare variation sources. In this paper, we describe the improvements in the input-output table, the implementation of the emissions, forestry sequestration and the welfare decomposition. Likewise, we compare the simulation results with its predecessor, the GTAP-AEZ-GHG model, reducing emissions by 50% globally, under two scenarios: with and without induced climate change crop yields. Our model follows similar conclusions to its predecessor in terms of economic and environmental variables. Our model also shows consistency in its accounting balance, correct capture of the welfare variation as well as correct direction of price changes and input subsidy allocation.
Read moreA new version of the location quotient for estimating regional input output tables: the centred location quotient
The construction of regional input-output tables presents several challenges. Ideally, a detailed survey of the regional economy would be conducted to determine the strength of relationships among sectors. However, a major limitation arises due to the large volume of data required, which may not always be readily available and entails high costs and lengthy delays. One potential solution to this issue is the adaptation of national input-output coefficients using methods such as location quotients (LQs) or constrained matrix-balancing techniques. This paper focuses on location quotient methods and proposes a new formula for estimating regional input coefficients. This explicit formula enables the measurement of the impact of primary inputs (including imports) on the coefficient estimation. To validate the proposed method, the multi-regional input-output table for Korea in 2015 and the Figaro database were used. The results of the investigation confirm the effectiveness of the approach.
Read moreTechniques in the Compilation of Danish Input-Output Tables: A New Approach to the Treatment of Imports
National accounts work in Denmark has since its initiation in the 1930s been closely connected with the construction of input-output tables, and the commodity-flow method — or production statistical method — has been the core of the technique.
Read more