Feature Selection for Automatic Categorization of Patent Documents

S  Don   and Dugki Min

doi:10.17485/ijst/2016/v9i37/98974

Article

Feature Selection for Automatic Categorization of Patent Documents

VIEWS 961
PDF 200

Abstract
Full-Text HTML
Full-Text PDF
How to Cite

Indian Journal of Science and Technology

DOI: 10.17485/ijst/2016/v9i37/98974

Year: 2016, Volume: 9, Issue: 37, Pages: 1-17

Original Article

Feature Selection for Automatic Categorization of Patent Documents

S. Don^1* and Dugki Min²

¹ Department of Analytics, School of Computer Science and Engineering, VIT University, Vellore - 632014, Tamil Nadu, India; [email protected]
² Department of Computer Science and Engineering, Konkuk University, Seoul, South Korea; [email protected]
*Author for correspondence
Don
Department of Analytics
Email:[email protected]

This work is licensed under a Creative Commons Attribution 4.0 International License.

Abstract

Objective: With the rapid increase in the number of patent documents worldwide, demand for their automatic categorization has grown significantly. The automatic categorization of patent documents is the organization of such documents in digital form, thus replacing the manual time-consuming process. In this work, we proposed a system that can automatically categorize patent document by considering the structural information of the patents. Methods: We propose a three-stage mechanism for automatic categorization. In the first stage, we apply a pre-processing mechanism to reduce unwanted noise that can influence the categorization process. Such noise includes terms that have less structural meaning in the document. In the second stage, feature selection is conducted based on the term frequencies. Feature vectors are constructed from the structural information of the patent. In the third stage, classifications are conducted using a Random Forest (RF), Support Vector Machine (SVM), and Naïve Bayes (NB) classifier. Findings: It was found that the semantic structural information of a patent document is an important feature set in constructing the terms of a document for the categorization. The experimental results also show that feature reduction using Information Gain (IG) is beneficial for obtaining a higher accuracy rate in a reduced dimensional space. Applications: The results reveal the importance of the proposed method for automatic categorization of patent documents.
Keywords: Classification, Feature Selection, Patent categorization, Structural information