Abstract:The size of text and weight of elements in feature vectors may affect text classification rule. In order to improve the classification accuracy
new concepts of the weighted frequent items and a weighted frequent item set mining algorithm to highlight great weight items were proposed. A pre-processing method for feature vectors was proposed to eliminate ill effects of the size of text on generating classification rules. Experiments demonstrated utility and feasibility of the method.
关键词
关联规则文本分类加权频繁项集
Keywords
association ruletext classificationweighted frequent itemsets
references
Li Xiaoli,Learning to classify texts using positive and unlabeled data,2003.
Liu B,Hsu W,Ma Y,Integrating classification and association rule mining,New York,1998.
Li W,Han J,Pe iJ,CMAR:accurate and efficient classification based on multiple class-association rules,2001.
Yin X,Han J,CPAR:Classification based on predictive association rules,2003.
陈晓云.陈袆.王雷.李荣陆.胡运发 基于分类规则树的频繁模式文本分类 [J].
Wang W,Yang J,Yu P,Efficient mining of Weighted Association Rules (WAR)[IBM,RC 21692(97734)],2000.
Zhang Z,Chen E,Wang J,Enabling personalization recommendation with weighted FP for text information retrieval based on user-focus,2004.
Salton G,Wong A,Yang C S,A vector space model for automatic indexing,Communications of the ACM,1995(1).
Cavnar W B,Trenkle J M,N-gram-based text categorization,1994.
Han J,Pei J,Yin Y,Mining frequent patterns without candidate generation,Dallas,TX,2000.