loading
Documents

Research.Publish.Connect.

Paper

Authors: Oliver Schmidts 1 ; Bodo Kraft 1 ; Marvin Winkens 1 and Albert Zündorf 2

Affiliations: 1 FH Aachen, University of Applied Sciences, Germany ; 2 University of Kassel, Germany

ISBN: 978-989-758-440-4

Keyword(s): Catalog Integration, Data Integration, Data Quality, Label Prediction, Machine Learning, Neural Network Applications, Public Data.

Abstract: The integration of product data from heterogeneous sources and manufacturers into a single catalog is often still a laborious, manual task. Especially small- and medium-sized enterprises face the challenge of timely integrating the data their business relies on to have an up-to-date product catalog, due to format specifications, low quality of data and the requirement of expert knowledge. Additionally, modern approaches to simplify catalog integration demand experience in machine learning, word vectorization, or semantic similarity that such enterprises do not have. Furthermore, most approaches struggle with low-quality data. We propose Attribute Label Ranking (ALR), an easy to understand and simple to adapt learning approach. ALR leverages a model trained on real-world integration data to identify the best possible schema mapping of previously unknown, proprietary, tabular format into a standardized catalog schema. Our approach predicts multiple labels for every attribute of an input column. The whole column is taken into consideration to rank among these labels. We evaluate ALR regarding the correctness of predictions and compare the results on real-world data to state-of-the-art approaches. Additionally, we report findings during experiments and limitations of our approach. (More)

PDF ImageFull Text

Download
CC BY-NC-ND 4.0

Sign In Guest: Register as new SciTePress user now for free.

Sign In SciTePress user: please login.

PDF ImageMy Papers

You are not signed in, therefore limits apply to your IP address 100.24.125.162

In the current month:
Recent papers: 100 available of 100 total
2+ years older papers: 200 available of 200 total

Paper citation in several formats:
Schmidts, Oliver; Kraft, B.; Winkens, M. and Zündorf, A. (2020). Catalog Integration of Low-quality Product Data by Attribute Label Ranking.In Proceedings of the 9th International Conference on Data Science, Technology and Applications - Volume 1: DATA, ISBN 978-989-758-440-4, pages 90-101. DOI: 10.5220/0009831000900101

@conference{data20,
author={Schmidts, Oliver and Bodo Kraft. and Marvin Winkens. and Albert Zündorf.},
title={Catalog Integration of Low-quality Product Data by Attribute Label Ranking},
booktitle={Proceedings of the 9th International Conference on Data Science, Technology and Applications - Volume 1: DATA,},
year={2020},
pages={90-101},
publisher={SciTePress},
organization={INSTICC},
doi={10.5220/0009831000900101},
isbn={978-989-758-440-4},
}

TY - CONF

JO - Proceedings of the 9th International Conference on Data Science, Technology and Applications - Volume 1: DATA,
TI - Catalog Integration of Low-quality Product Data by Attribute Label Ranking
SN - 978-989-758-440-4
AU - Schmidts, Oliver
AU - Kraft, B.
AU - Winkens, M.
AU - Zündorf, A.
PY - 2020
SP - 90
EP - 101
DO - 10.5220/0009831000900101

Login or register to post comments.

Comments on this Paper: Be the first to review this paper.