Invention Grant
US08650136B2 Text classification with confidence grading 有权
具有置信分级的文本分类

Text classification with confidence grading
Abstract:
A computer implemented method and system is provided for classifying a document. A classifier is trained using training documents. A list of first words is obtained from the training documents. A prior probability is determined for each class of multiple classes. Conditional probabilities are calculated for the first words for each class. Confidence thresholds are determined. Confidence grades are defined for the classes using the confidence thresholds. A list of second words is obtained from the document. Conditional probabilities for the list of second words are determined from the calculated conditional probabilities for the list of first words. A posterior probability is calculated for each of the classes and compared with the determined confidence thresholds. Each class is assigned to one of the defined confidence grades based on the comparison. The document is assigned to one of the classes based on the posterior probability and the assigned confidence grades.
Public/Granted literature
Information query
Patent Agency Ranking
0/0