loading
Papers

Research.Publish.Connect.

Paper

Paper Unlock

Authors: Nuno Moniz and Fátima Rodrigues

Affiliation: Institute of Engineering and Polytechnic of Porto, Portugal

ISBN: 978-989-8565-29-7

Keyword(s): Information Retrieval, Text Extraction, PDF.

Related Ontology Subjects/Areas/Topics: Artificial Intelligence ; Information Extraction ; Knowledge Discovery and Information Retrieval ; Knowledge-Based Systems ; Structured Data Analysis and Statistical Methods ; Symbolic Systems

Abstract: This paper presents an approach for text processing of PDF documents with well-defined layout structure. The scope of the approach is to explore the font’s structure of PDF documents, using perceptual grouping. It consists on the extraction of text objects from the content stream of the documents and its grouping according to a set criterion, making also use of geometric-based regions in order to achieve the correct reading order. The developed approach processes the PDF documents using logical and structural rules to extract the entities present in them, and returns an optimized XML representation of the PDF document, useful for re-use, for example in text categorization. The system was trained and tested with Portuguese Legislation PDF documents extracted from the electronic Republic’s Diary. Evaluation results show that our approach presents good results.

PDF ImageFull Text

Download
CC BY-NC-ND 4.0

Sign In Guest: Register as new SciTePress user now for free.

Sign In SciTePress user: please login.

PDF ImageMy Papers

You are not signed in, therefore limits apply to your IP address 18.207.249.15

In the current month:
Recent papers: 100 available of 100 total
2+ years older papers: 200 available of 200 total

Paper citation in several formats:
Moniz, N. and Rodrigues, F. (2012). Extracting Structure, Text and Entities from PDF Documents of the Portuguese Legislation.In Proceedings of the International Conference on Knowledge Discovery and Information Retrieval - Volume 1: KDIR, (IC3K 2012) ISBN 978-989-8565-29-7, pages 123-131. DOI: 10.5220/0004103501230131

@conference{kdir12,
author={Nuno Moniz. and Fátima Rodrigues.},
title={Extracting Structure, Text and Entities from PDF Documents of the Portuguese Legislation},
booktitle={Proceedings of the International Conference on Knowledge Discovery and Information Retrieval - Volume 1: KDIR, (IC3K 2012)},
year={2012},
pages={123-131},
publisher={SciTePress},
organization={INSTICC},
doi={10.5220/0004103501230131},
isbn={978-989-8565-29-7},
}

TY - CONF

JO - Proceedings of the International Conference on Knowledge Discovery and Information Retrieval - Volume 1: KDIR, (IC3K 2012)
TI - Extracting Structure, Text and Entities from PDF Documents of the Portuguese Legislation
SN - 978-989-8565-29-7
AU - Moniz, N.
AU - Rodrigues, F.
PY - 2012
SP - 123
EP - 131
DO - 10.5220/0004103501230131

Login or register to post comments.

Comments on this Paper: Be the first to review this paper.