Data-driven prediction and decision-making increasingly depend on large-scale data lakes and machine learning (ML) and deep learning (DL) analytics across a wide range of application domains. Nevertheless, although geospatial vector data constitutes a substantial fraction of contemporary datasets, it has remained primarily managed within traditional GIS and database systems and has benefited only to a limited extent from recent advances in artificial intelligence (AI). This limitation has been largely attributable to the intrinsic characteristics of geospatial information, in particular when represented in vector format, like multidimensional spatial coordinates, scale and projection effects, variable-size geometries, and a broad spectrum of spatial operations, which are difficult to reconcile with the fixeddimensional inputs typically required by DL models. This paper advanced DL-enabled geospatial analytics by enabling AI models and agents to provide native access, interpretation, and query processing over geospatial vector datasets. Rather than proposing novel architectures, we introduced a methodology for leveraging existing DL models by properly encoding both geospatial inputs and the spatial operations to be estimated. We investigated three representation/architecture families: (i) image-like, georeferenced histograms processed with ResNet and UNet, (ii) point- and graph-based variants of PointNet++, and (iii) fixed-length geometry embeddings obtained by extending Poly2Vec combined with a Transformer model. We instantiated these approaches on three representative geospatial operators, spatial synopsis, spatial clustering, and walkability estimation, and evaluated them using synthetic and real-world datasets. The results demonstrated that all three families could achieve high accuracy in specific configurations with distinct trade-offs: graph-based models were most effective for dense data and short-range interactions, whereas image- and vector-based approaches exhibited superior scalability due to fixed-size inputs, and the image-based approach was particularly effective for multi-dataset operations.
On the applicability of artificial intelligence models to geospatial vector data analysis and exploration
Belussi, Alberto;Migliorini, Sara;Eldawy, Ahmed
2026-01-01
Abstract
Data-driven prediction and decision-making increasingly depend on large-scale data lakes and machine learning (ML) and deep learning (DL) analytics across a wide range of application domains. Nevertheless, although geospatial vector data constitutes a substantial fraction of contemporary datasets, it has remained primarily managed within traditional GIS and database systems and has benefited only to a limited extent from recent advances in artificial intelligence (AI). This limitation has been largely attributable to the intrinsic characteristics of geospatial information, in particular when represented in vector format, like multidimensional spatial coordinates, scale and projection effects, variable-size geometries, and a broad spectrum of spatial operations, which are difficult to reconcile with the fixeddimensional inputs typically required by DL models. This paper advanced DL-enabled geospatial analytics by enabling AI models and agents to provide native access, interpretation, and query processing over geospatial vector datasets. Rather than proposing novel architectures, we introduced a methodology for leveraging existing DL models by properly encoding both geospatial inputs and the spatial operations to be estimated. We investigated three representation/architecture families: (i) image-like, georeferenced histograms processed with ResNet and UNet, (ii) point- and graph-based variants of PointNet++, and (iii) fixed-length geometry embeddings obtained by extending Poly2Vec combined with a Transformer model. We instantiated these approaches on three representative geospatial operators, spatial synopsis, spatial clustering, and walkability estimation, and evaluated them using synthetic and real-world datasets. The results demonstrated that all three families could achieve high accuracy in specific configurations with distinct trade-offs: graph-based models were most effective for dense data and short-range interactions, whereas image- and vector-based approaches exhibited superior scalability due to fixed-size inputs, and the image-based approach was particularly effective for multi-dataset operations.| File | Dimensione | Formato | |
|---|---|---|---|
|
Saeedan_et_al-2026-GeoInformatica.pdf
accesso aperto
Licenza:
Creative commons
Dimensione
4.48 MB
Formato
Adobe PDF
|
4.48 MB | Adobe PDF | Visualizza/Apri |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.



