Deep learning techniques for geospatial data integration, knowledge mining, and real-world applications
Geospatial data is ubiquitous in today’s world, playing a crucial role in various aspects of modern life. As businesses increasingly establish and expand their online presence, they leverage location-based insights to better understand customer behavior and optimize services. Simultaneously, the proliferation of GPS-enabled devices enables users to continuously share their locations and activities online, contributing to an ever-growing pool of geospatial information. Social networks and web services, in particular, heavily rely on this data, integrating it into their platforms to provide more personalized and relevant experiences. This surge in geospatial data generation necessitates the development of novel algorithms to mine meaningful patterns and enhancing user experiences by delivering accurate information at the right place and time. This dissertation presents a pipeline of novel algorithms, spanning from the integration of geospatial data, to its effective application in real-world use cases. Specifically it introduces (1) an algorithm for deduplicating and consolidating geospatial data from multiple sources, a key step in creating a comprehensive geospatial database; it studies (2) the problem of Geospatial Knowledge Graph (GeoKG) construction, and presents an algorithm to automatically mine fine-grained spatial relationships between entities from a database, within an open-world setting; it presents (3) a framework for self-supervised pre-training of geospatial foundation models, capable of generating task-agnostic representations that are readily applicable to a wide range of downstream tasks. Finally, (4) this thesis explores the fine-tuning of pre-trained large language models (LLMs) on city-specific data, to build urban virtual assistants, capable of delivering accurate geospatial recommendations in a conversational manner while minimizing hallucinations. The study on geospatial Entity Resolution (ER) proposes a state-of-the-art approach to link records from different sources, that refer to the same real-world geospatial entity. This is a crucial step for deduplicating merged data. The algorithm, called Geo-ER, leverages textual, spatial, and contextual information, making a step forward towards automatic geospatial data integration. A novel neighborhood embedding component is designed, using Graph Attention Networks, to produce a more contextualized representation of an entity. Extensive experiments, on eight real-world datasets, are performed to demonstrate the advantages of Geo-ER compared to existing state-of-the-art solution for geospatial and traditional ER. Cross-city validation experiments showcase the algorithm's robustness in scenarios where training data for a specific city is unavailable. Subsequently, this thesis presents GTMiner, a novel framework for automatically constructing a Geospatial Knowledge Graph (GeoKG) from tabular data, derived from the merging and consolidation of multiple data sources. The architecture for geospatial relation prediction is an open-world Knowledge Graph Completion solution, consisting of a pre-trained language model, a geospatial encoder, and a Geo-Textual interaction component, which jointly model the textual and geospatial characteristics of an entity. GTMiner further includes a refinement module, specifically designed to improve both coverage and correctness of a GeoKG. The study further introduces four real-world datasets, collected from publicly available sources, and carries out extensive experiments to evaluate the performance of each of the system’s components and perform ablation studies. A study on general-purpose representation learning in the geospatial domain introduces CityFM, a new framework that leverages mutual information maximization and incorporates Nodes, Ways, and Relations into its learning frame-work to train geospatial Pre-trained Foundation Models (PFMs) for a selected area of interest. CityFM-trained models are task-agnostic and multimodal, capable of integrating an entity’s textual, visual, and spatial attributes. The study conducts experiments on road, building, and region-level downstream tasks to demonstrate the effectiveness of the embeddings produced by the foundation models in real-world applications, compared to task-specific algorithms. The last study in this dissertation presents LAMP, a virtual assistant for urban applications. It introduces a novel strategy to generate synthetic conversational data about geospatial entities, to fine-tune a pre-trained LLM. The generated data enables the language model to learn and memorize thousands of places within a city, while also equipping it with spatial (proximity) awareness specific to the urban area of interest.
Read more