Walk Extraction Strategies for Node Embeddings with RDF2Vec in Knowledge Graphs

2020-09-09 16:26:31

Gilles Vandewiele, Bram Steenwinckel, Pieter Bonte, Michael Weyns, Heiko Paulheim, Petar Ristoski, Filip De Turck, Femke Ongenae

arXiv_AI

arXiv_AI Classification Embedding Knowledge Knowledge_Graph Language_Model Unsupervised Pose Action

Abstract
Abstract (translated)
URL
PDF

Abstract

As KGs are symbolic constructs, specialized techniques have to be applied in order to make them compatible with data mining techniques. RDF2Vec is an unsupervised technique that can create task-agnostic numerical representations of the nodes in a KG by extending successful language modelling techniques. The original work proposed the Weisfeiler-Lehman (WL) kernel to improve the quality of the representations. However, in this work, we show both formally and empirically that the WL kernel does little to improve walk embeddings in the context of a single KG. As an alternative to the WL kernel, we propose five different strategies to extract information complementary to basic random walks. We compare these walks on several benchmark datasets to show that the \emph{n-gram} strategy performs best on average on node classification tasks and that tuning the walk strategy can result in improved predictive performances.

Abstract (translated)

URL

https://arxiv.org/abs/2009.04404

PDF

https://arxiv.org/pdf/2009.04404.pdf