首页 | 本学科首页   官方微博 | 高级检索  
     检索      

基于训练集裁剪的加权K近邻文本分类算法
引用本文:孙新,欧阳童,严西敏,尚煜茗,郭文浩.基于训练集裁剪的加权K近邻文本分类算法[J].情报工程,2016,2(6):008-016.
作者姓名:孙新  欧阳童  严西敏  尚煜茗  郭文浩
作者单位:北京理工大学计算机学院北京市海量语言信息处理与云计算应用工程技术研究中心 北京 100081
基金项目:: 本文受国家 973 课题(2013CB329605)的资助
摘    要:文本分类是信息检索领域的重要应用之一,由于采用统一特征向量形式表示所有文档,导致针对每个文档的特征向量具有高维性和稀疏性,从而影响文档分类的性能和精度。为有效提升文本特征选择的准确度,本文首先提出基于信息增益的特征选择函数改进方法,提高特征选择的精度。KNN(K-Nearest Neighbor)算法是文本分类中广泛应用的算法,本文针对经典KNN计算量大、类别标定函数精度不高的问题,提出基于训练集裁剪的加权KNN算法。该算法通过对训练集进行裁剪提升了分类算法的计算效率,通过模糊集的隶属度函数提升分类算法的准确性。在公开数据上的实验结果及实验分析证明了算法的有效性。

关 键 词:文本分类  特征选择  信息增益  最近邻分类算法

The Weighted KNN Text Categorization Algorithm Based on Training Set Cutting
Authors:SUN Xin  OUYANG Tong  YAN XiMin  SHANG YuMing and GUO WenHao
Institution:Beijing Engineering Research Center of Massive Language Information Processing and Cloud Computing Application, School of Computer Science and Technology,Beijing Engineering Research Center of Massive Language Information Processing and Cloud Computing Application, School of Computer Science and Technology,Beijing Engineering Research Center of Massive Language Information Processing and Cloud Computing Application, School of Computer Science and Technology,Beijing Engineering Research Center of Massive Language Information Processing and Cloud Computing Application, School of Computer Science and Technology and Beijing Engineering Research Center of Massive Language Information Processing and Cloud Computing Application, School of Computer Science and Technology
Abstract:Text categorization is one of the key research fields in the information retrieval. Feature selection is an important part in the document processing, and imposes great influence on the document classification. In this paper, an improved feature selection algorithm based on information gain was proposed to improve the accuracy of text feature selection effectively. Moreover, K-Nearest Neighbor (KNN) algorithm is used widely in text categorization, and the advantages of this method are high accuracy and stability.However, the number of training samples and their position may influence the classification performance of the KNN algorithm, thus we proposed the weighted KNN classification algorithm based on training set cutting, and the accuracy of the classification algorithm can be improved by the rough sets and the concept of membership function. Finally, this research tested the new algorithm based on the text categorization experiment, and the results indicated that the effectiveness of the proposed algorithm.
Keywords:Text categorization  feature selection  information gain  KNN algorithm
本文献已被 万方数据 等数据库收录!
点击此处可从《情报工程》浏览原始摘要信息
点击此处可从《情报工程》下载免费的PDF全文
设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司  京ICP备09084417号