A Distributed Approach for Parsing Large-scale OWL Datasets

Heba Mohamed, Heba Mohamed, Said Fathalla, Said Fathalla, Jens Lehmann, Jens Lehmann, Hajira Jabeen

2020

Abstract

Ontologies are widely used in many diverse disciplines, including but not limited to biology, geology, medicine, geography and scholarly communications. In order to understand the axiomatic structure of the ontologies in OWL/XML syntax, an OWL/XML parser is needed. Several research efforts offer such parsers; however, these parsers usually show severe limitations as the dataset size increases beyond a single machine’s capabilities. To meet increasing data requirements, we present a novel approach, i.e., DistOWL, for parsing large-scale OWL/XML datasets in a cost-effective and scalable manner. DistOWL is implemented using an in-memory and distributed framework, i.e., Apache Spark. While the application of the parser is rather generic, two use cases are presented for the usage of DistOWL. The Lehigh University Benchmark (LUBM) has been used for the evaluation of DistOWL. The preliminary results show that DistOWL provides a linear scale-up compared to prior centralized approaches.

Download


Paper Citation


in Harvard Style

Mohamed H., Fathalla S., Lehmann J. and Jabeen H. (2020). A Distributed Approach for Parsing Large-scale OWL Datasets. In Proceedings of the 12th International Joint Conference on Knowledge Discovery, Knowledge Engineering and Knowledge Management (IC3K 2020) - Volume 2: KEOD; ISBN 978-989-758-474-9, SciTePress, pages 227-234. DOI: 10.5220/0010138602270234


in Bibtex Style

@conference{keod20,
author={Heba Mohamed and Said Fathalla and Jens Lehmann and Hajira Jabeen},
title={A Distributed Approach for Parsing Large-scale OWL Datasets},
booktitle={Proceedings of the 12th International Joint Conference on Knowledge Discovery, Knowledge Engineering and Knowledge Management (IC3K 2020) - Volume 2: KEOD},
year={2020},
pages={227-234},
publisher={SciTePress},
organization={INSTICC},
doi={10.5220/0010138602270234},
isbn={978-989-758-474-9},
}


in EndNote Style

TY - CONF

JO - Proceedings of the 12th International Joint Conference on Knowledge Discovery, Knowledge Engineering and Knowledge Management (IC3K 2020) - Volume 2: KEOD
TI - A Distributed Approach for Parsing Large-scale OWL Datasets
SN - 978-989-758-474-9
AU - Mohamed H.
AU - Fathalla S.
AU - Lehmann J.
AU - Jabeen H.
PY - 2020
SP - 227
EP - 234
DO - 10.5220/0010138602270234
PB - SciTePress