helpfn
A WordNet-based Semantic Model for Enhancing Text Clustering
Shady Shehata University of Waterloo
Waterloo, Ontario, Canada N2L 3G1 [email protected]
Abstract—Most of text mining techniques are based on word and/or phrase analysis of the text. The statistical analysis of a term (word or phrase) frequency captures the importance of the term within a document. However, to achieve a more accurate analysis, the underlying mining technique should indicate terms that capture the semantics of the text from which the importance of a term in a sentence and in the document can be derived. Incorporating semantic features from the WordNet lexical database is one of many approaches that have been tried to improve the accuracy of text clustering techniques.
A new semantic-based model that analyzes documents based on their meaning is introduced. The proposed model analyzes terms and their corresponding synonyms and/or hypernyms on the sentence and document levels. In this model, if two docu- ments contain different words and these words are semantically related, the proposed model can measure the semantic-based similarity between the two documents. The similarity between documents relies on a new semantic-based similarity measure which is applied to the matching concepts between documents.
Experiments using the proposed semantic-based model in text clustering are conducted. Experimental results demon- strate that the newly developed semantic-based model enhances the clustering quality of sets of documents substantially.
Keywords-Semantic; WordNet; Clustering; Concepts;
I. INTRODUCTION
Text mining attempts to discover new, previously un- known information by applying techniques from natural language processing and data mining.
Clustering, one of the traditional data mining techniques, is an unsupervised learning paradigm where clustering meth- ods try to identify inherent groupings of the text documents so that a set of clusters are produced in which clusters exhibit high intra-cluster similarity and low inter-cluster similarity. Generally, text document clustering methods at- tempt to segregate the documents into groups where each group represents some topic that is different to those topics represented by the other groups[1]
Most current document clustering methods are based on the Vector Space Model (VSM)[1] which is a widely used data representation for text classification and clustering. The VSM represents each document as a feature vector of the terms (words or phrases) in the document. Each feature vector contains term-weights (usually term-frequencies) of the terms in the document. The similarity between the documents is measured by one of several similarity measures
that are based on such a feature vector. Examples include the cosine measure and the Jaccard measure.
In text clustering, it is important to note that selecting important features, which present the text data properly, has critical effect on the output of the clustering algorithm [2]. Moreover, weighting these features accurately also affect the result of the clustering algorithm substantially [3]. In- corporating semantic features from the WordNet [4] lexical database is one of many approaches that have been tried to improve the accuracy of text clustering techniques.
In this paper, a new semantic-based model is proposed. The proposed model captures the semantic structure of each term within a sentence and document rather than the frequency of the term within a document only. Each sentence in a document is labeled by a semantic role labeler. The role labeling task determines terms which contribute to the sentence semantics in a sentence. The labeled term that has a semantic role in the sentence, can be either word or phrase which is totally dependent on the semantic structure of the sentence. Based on the semantic-based analysis, each term is assigned a weight. The terms that have maximum weights are extracted as top terms. Synonyms and/or hypernyms of each word (in the top terms) are added to the term vector. These concepts are analyzed on the sentence and document levels as well as top terms and used in text document clustering. When a new document is introduced to the system, the proposed model can detect a concept match from this document to all the previously processed documents in the data set by scanning the new document and extracting the matching concepts.
A new semantic-based similarity measure which makes use of the synonyms and/or hypernyms is proposed. This similarity measure outperforms other similarity measures that are based on term analysis models of the document only. The similarity between documents is based on a combination of term and concept analysis on the sentence and document levels. The clustering results produced by the semantic-based model have higher quality than those produced by a single- term analysis similarity only. The results are evaluated using two quality measures, the F-measure and the Entropy. Both of these quality measures showed improvement versus the use of the single-term method when the semantic-based similarity measure is used to cluster sets of documents.
The explanations of the important terminologies, which
2009 IEEE International Conference on Data Mining Workshops
978-0-7695-3902-7/09 $26.00 © 2009 IEEE
DOI 10.1109/ICDMW.2009.86
471
2009 IEEE International Conference on Data Mining Workshops
978-0-7695-3902-7/09 $26.00 © 2009 IEEE
DOI 10.1109/ICDMW.2009.86
477
Authorized licensed use limited to: University of the Cumberlands. Downloaded on July 24,2021 at 04:14:27 UTC from IEEE Xplore. Restrictions apply.
are used in this paper, are listed as follows: * Verb-argument structure: (i.e. John hits the ball). ”hits” is the verb. ”John” and ”the ball” are the arguments of the verb ”hits”, * Label: A label is assigned to an argument. i.e: ”John” has subject (or Agent) label. ”the ball” has object (or theme) label, * Term: is either an argument or a verb. Term is also either a word or a phrase (which is a sequence of words), * Concept: is a WordNet concept either synonym or hyper- nym.
The rest of this paper is organized as follows. The semantic-based model which includes semantic-based term analysis and semantic-based similarity measure is presented in section 2. Experimental results are presented in section 3. The last section summarizes the conclusions and suggests future work.
II. SEMANTIC-BASED MODEL
The proposed semantic-based model aims to cluster doc- uments by meaning. The introduced model consists of semantic-based term analysis and semantic-based similarity measure.
This work presents an ongoing project for semantic-based analysis of terms to enhance the quality of the text mining process. The work introduced in [5] presents concept-based analysis to enhance text categorization. This work introduces deep semantic analysis for concepts extracted from WordNet to enhance the quality of text clustering. WordNet has not been used in the work of [5] as the definition of concept in [5] was a labeled term and has no relation to WordNet concepts.
The contribution of this paper relies on the following: * Extracting the synonyms and/or hypernyms for the ana- lyzed terms, * Proposing a new semantic-based analysis that analyzes terms and their corresponding synonyms and/or hypernyms on the sentence and document levels, * Applying the semantic-based analysis to text document clustering with sets of experiments that compare between clustering techniques when the extracted features are terms only, terms and synonyms, and terms and hypernyms. * Proposing a new semantic-based similarity measure that takes the advantage of the WordNet concepts to measure the similarity between documents based on the meaning of their words rather than the words themselves.
A raw text document is the input to the proposed model. Each document has well defined sentence boundaries. Each sentence in the document is labeled automatically based on the PropBank notations [6]. After running the semantic role labeler, each sentence in the document might have one or more labeled verb-argument structures. The number of generated labeled verb-argument structures is entirely dependent on the amount of information in the sentence.
The sentence that has many labeled verb-argument structures includes many verbs associated with their arguments. The labeled verb-argument structures, the output of the role labeling task, are captured and analyzed by the semantic- based model on sentence and document levels.
In this model, both the verb and the argument are con- sidered terms. One term can be an argument to more than one verb in the same sentence. This means that this term can have more than one semantic role in the same sentence. In such cases, this term plays important semantic roles that contribute to the meaning of the sentence. The entire terms are sorted based on their weights which are assigned by the semantic-based analysis. Terms that have maximum weights are considered top terms. Based on the part of speech of each word in a top term, synonyms and/or hypernyms are extracted from WordNet and added to the term vector. In the semantic-based model, the synonym and hypernym of a labeled term are considered concepts.
A. Semantic-based Term Analysis
The objective of this tasks is to analyze terms on the sentence and document levels. Then, top terms (which have maximum weights assigned by the semantic-based analysis) and their corresponding synonyms and/or hypernyms are used for text clustering. For instance, terms like beef and lamb are found to be similar, because they both are sub- concepts of meat in WordNet. If these words are added to the term vector, the clustering technique will cluster documents that are related based on the meaning of their words rather than the words themselves. The following phases depicts the process of the semantic-based term analysis: * Each sentence is labeled by semantic role labeler, * For each labeled term, stop-words that have no significance are removed, * Labeled terms are analyzed on the sentence and document levels as discussed in the next subsection, * Due to words with the same meaning appear in various morphological forms, words (in a labeled term) are normal- ized into a common root-form to capture their similarity, * Terms are sorted based on their weights (assigned by the semantic-based analysis) descendingly and top terms are extracted, * For each word (in a top term), nouns, verbs, adjectives and adverbs1 are looked up in WordNet and a global list of the first2 synonyms and hypernym synsets is assembled. Synonyms are extracted for nouns, verbs, adjectives and adverbs. Hypernyms are extracted for nouns and verbs. Terms that have no corresponding concept in WordNet are still used in text clustering. Extracting the synonyms and hypernyms from the top terms only is a kind of pruning to reduce the dimension of the term vector, * Term vector is extended by adding the corresponding
472478
Authorized licensed use limited to: University of the Cumberlands. Downloaded on July 24,2021 at 04:14:27 UTC from IEEE Xplore. Restrictions apply.
synonyms or hypernyms of the top terms. Term and Concept Analysis
To analyze each term at the sentence-level, a sentence-based frequency measure, called the conceptual term frequency ctf is utilized. The ctf is the number of occurrences of term t in verb-argument structures of sentence s. The term t, which frequently appears in different verb-argument structures of the same sentence s, has the principal role of contributing to the meaning of s.
To analyze each concept at the sentence-level, each WordNet concept is assigned the same (ctf ) value of its corresponding top term.
To analyze each concept at the document-level, the con- cept frequency cf is proposed, the number of occurrences of a WordNet concept c in the document, is calculated.
At this point, each term and it’s corresponding concepts have the same measures which are the ctf and cf (for concept and term) on the sentence and document levels respectively.
Semantic-based Analyzer Algorithm ————————————————————–
1. ddoci is a new Document where doci = {1, 2, .., N} and N is a total number of documents 2. L is an empty List (L is a synonyms or hypernyms list) 3. M is an empty List (M is a matched concepts list) 4. T is an empty List (T is a terms list) 5. for each sentence s in d do 6. for each labeled term t in d do 7. compute tfi of ti in d 8. compute ctfi of ti in s in d 9. compute the weighti = tfi + ctfi 10. add term t with weighti to T 11. end for 12. end for 13. sort T descendingly based on weight 14. output the max(weight) from list T 15. for each term t in T that has max(weight) do 16. extract synonyms and/or hypernyms of t 17. add concept c (synonyms and/or hypernyms) to L 18. end for 19. for each concept ci in L do 20. compute cfi of ci in d 21. assign ctfi to ci that corresponds to ti 22. if (ci == cj ) where j = {1, 2, .., doci} then 23. compute tf weight = avg(tfi, tfj ) 24. compute ctf weight = avg(ctfi, ctfj ) 25. add new concept matches to M
1Part of the task of the semantic role labeling is to assign a part of speech tag to each word in the corpus using the Brill tagger [7].
2The word sense disambiguation strategy used in this paper is to extract the first synonyms or hypernyms. Wordnet returns an ordered list of concepts. The ordering is supposed to reflect how common it is that a term reflects a concept in standard English language. More common term meanings are listed before less common ones. This approach extracts only the most common meaning of each term as shown in [8]
26. end if 27. end for 28. output the matched concepts list M ———————————————————————– The semantic-based analyzer algorithm describes the process of calculating the cf and the ctf of the matched concepts in the documents. The procedure begins with processing a new document (at line 1) which has well defined sentence boundaries. Each sentence is semantically labeled according to [6]. The lengths of the matched concepts and their verb-argument structures are stored for the concept-based similarity calculations in section II-B.
For each sentence (in the for loop at line 5) the concepts (synonyms and/or hypernyms) of the corresponding terms in the verb-argument structures which represent the semantic structures of the sentence are processed sequentially. Each concept in the current document is matched with the other concepts in the previously processed documents. To match the concepts in previous documents is accomplished by keeping a concept list M that holds the entry for each of the previous documents that shares a concept with the current document.
After the document is processed, M contains all the matching concepts between the current document and any previous document that shares at least one concept with the new document. Finally, M is output as the list of documents with the matching concepts and the necessary information about them. The concept-based term analyzer algorithm is capable of matching each concept in a new document (d) with all the previously processed documents in O(m) time, where m is the number of concepts in d.
Performance Issues For implementation and performance purposes, it is im- perative to note that the semantic-based analysis maintains the identification number of each term and corresponding concepts. There is a hash table that includes the unique terms (and their synonyms and/or hypernyms) that appeared in each verb-argument structure. Thus, the semantic analyzer uses the identification numbers of the terms and their con- cepts to look up in the hash table. This is the source of the efficiency of the analysis. It is also important to note that extracting the synonyms or hypernyms from the top terms only is a kind of pruning to reduce the dimension of the term vector. Moreover, the morphological analysis provided by looking up at another hash table of English language dictionary of words and their corresponding root-form.
B. Semantic-based Similarity Measure
Concepts convey local context information, which is es- sential in determining an accurate similarity between docu- ments.
A semantic-based similarity measure, based on matching concepts at the sentence and document levels rather than on individual words only, is devised. The semantic-based
473479
Authorized licensed use limited to: University of the Cumberlands. Downloaded on July 24,2021 at 04:14:27 UTC from IEEE Xplore. Restrictions apply.
similarity measure relies on two critical aspects. First, the analyzed top terms capture the semantic structure of each sentence. Secondly, the frequency of concepts, which are corresponding to top terms, is used to measure the contribution of the concept to the meaning of the sentence, as well as to the main topics of the document. These two aspects are measured by the proposed semantic-based similarity which measures the importance of each concept at the sentence-level by the ctf (same value of the its corresponding term) and document-level by the cf measure. This similarity measure is a function of the following factors: * The number of matching concepts, m, in each document (d), * The total number of the labeled verb-argument structures, v, in each sentence s, * The ctfi of each concept ci in s for each document d where (i = 1, 2, ..., m) * The cfi of each concept ci in each document d where (i = 1, 2, ..., m),
The semantic-based similarity between two documents d1 and d2 is calculated by:
sims(d1, d2) = m∑
i=1
weighti1 ∗ weighti2 , (1)
weighti = cf weighti + ctf weighti, (2)
cf weighti = cfij√∑cn
j=1 (cfij ) 2 , (3)
ctf weighti = ctfij√∑cn
j=1 (ctfij ) 2 , (4)
where cn is the total number of concepts which has a conceptual term frequency value in document d.
* The weight of concept i in document d is calculated by equation 2. * In equation 2, the cf weighti value presents the weight of concept i in document d at the document-level. * In equation 2, the ctf weighti value presents the weight of the concept i in document d at the sentence-level based on the contribution of concept i to the semantics of the sentences in d through its corresponding term. * The sum between the two values of cf weighti and ctf weighti in equation 2 presents a semantic-based measure of the contribution of each concept to the meaning of the sentences and to the topics mentioned in a document. * In equation 3, the cfij value is normalized by the length of the document vector of the concept frequency cfij in document d, where j = 1, 2, ..., cn.
* cn is the total number of the concepts which has a concept frequency value in document d. * In equation 4, the ctfij value is normalized by the length of the document vector of the conceptual term frequency ctfij in document d where j = 1, 2, ..., cn
A concept i can have many ctf values in different sentences in the same document d. Thus, the ctf value of concept i in document d is calculated by ctfi =
∑ sn n=1 ctfn
sn where sn is the total number of sentences that contain concept c in document d.
For the single-term similarity measure, the cosine corre- lation similarity measure in is adopted with the popular TF- IDF (Term Frequency/Inverse Document Frequency) term weighting. The cosine measure is chosen due to its wide use in the document clustering literature. Recall that the cosine measure calculates the cosine of the angle between the two document vectors. Accordingly, the single-term similarity measure (sims) is:
simc(d1, d2) = cos(x, y) = d1 · d2
‖ d1 ‖ ‖ d2 ‖ , (5)
The vectors d1 and d2 are represented as single-term weights calculated by using the TF-IDF weighting scheme.
III. EXPERIMENTAL RESULTS
To test the effectiveness of concept matching in deter- mining an accurate measure of the similarity between doc- uments, extensive sets of experiments using the semantic- based term analysis and similarity measure are conducted.
The experimental setup consisted of three datasets. The first data set contains 23,115 ACM abstract articles collected from the ACM digital library. The ACM articles are classi- fied according to the ACM computing classification system into five main categories: general literature, hardware, com- puter systems organization, software, and data. The second data set has 12,902 documents from the Reuters 21578 dataset. There are 9,603 documents in the training set, 3,299 documents in the test set, and 8,676 documents are unused. Out of the 5 category sets, the topic category set contains 135 categories, but only 90 categories have at least one document in the training set. These 90 categories were used in the experiment. The third dataset consisted of 361 samples from the Brown corpus [9]. Each sample has 2000+ words. The Brown corpus main categories used in the experiment were: press: reportage, press: reviews, religion, skills and hobbies, popular lore, belles-letters, learned, fiction: science, fiction: romance, and humor.
The similarities which are calculated by using the syn- onyms and/or hypernyms on the sentence and document lev- els are used to compute similarity matrix among documents.
Three standard document clustering techniques are chosen for testing the effect of the concept-based similarity on clustering [10]: (1) Hierarchical Agglomerative Clustering
474480
Authorized licensed use limited to: University of the Cumberlands. Downloaded on July 24,2021 at 04:14:27 UTC from IEEE Xplore. Restrictions apply.
Table I: Clustering Improvement using Terms and Corresponding Synonyms DataSet Single-Term Terms and Synonyms ImprovementF-measure Entropy F-measure Entropy
Reuters HAC (ward) 0.723 0.251 0.928 0.011 +28.35%F, -95.61%E Single Pass 0.411 0.523 0.852 0.062 >+100%F, -88.14%E
k-NN 0.511 0.348 0.921 0.012 +80.23%F, -96.55%E
ACM HAC (ward) 0.697 0.317 0.923 0.039 +32.42%F, -87.69%E Single Pass 0.398 0.608 0.811 0.148 >+100%F, -75.65%E
k-NN 0.491 0.402 0.896 0.102 +82.48%F, -74.62%E
Brown HAC (ward) 0.581 0.385 0.908 0.015 +56.28%F, -96.1%E Single Pass 0.437 0.551 0.807 0.039 +84.66%F, -92.9%E
k-NN 0.462 0.316 0.907 0.021 +96.32%F, -93.35%E
Table II: Clustering Improvement using Terms and Corresponding Hypernyms DataSet Single-Term Terms and Hypernyms ImprovementF-measure Entropy F-measure Entropy
Reuters HAC (ward) 0.723 0.251 0.931 0.0087 +28.76%F, -96.53%E Single Pass 0.411 0.523 0.856 0.059 >+100%F, -88.71%E
k-NN 0.511 0.348 0.926 0.011 +81.21%F, -96.83%E
ACM HAC (ward) 0.697 0.317 0.921 0.041 +32.13%F, -87.06%E Single Pass 0.398 0.608 0.805 0.149 >+100%F, -75.49%E
k-NN 0.491 0.402 0.894 0.108 +82.07%F, -73.13%E
Brown HAC (ward) 0.581 0.385 0.912 0.013 +56.97%F, -96.62%E Single Pass 0.437 0.551 0.813 0.032 +86.04%F, -94.19%E
k-NN 0.462 0.316 0.911 0.017 +97.18%F, -94.62%E
(HAC), (2) Single Pass Clustering, and (3) k-Nearest Neigh- bor (k-NN)2.
The semantic-based term analysis is one of the main factors that captures the importance of synonyms and/or hypernyms in text clustering. Thus, to study the effect of the semantic-based analysis on the clustering techniques, the entire set of the experiments is repeated using the following schemes for the different clustering techniques as shown in Tables [I and II]: * The weights of terms and their corresponding synonyms * The weights of terms and their corresponding hypernyms
In order to evaluate the quality of the clustering, two quality measures widely used in the text mining literature for the purpose of document clustering [12] are adopted. The first is the F-measure, which combines the Precision and Recall measures from the Information Retrieval literature. The precision P and recall R of a cluster j with respect to a class i are defined as: P = P recision(i, j) = Mij
Mj and
R = Recall(i, j) = Mij Mi
where Mij : is the number of members of class i in cluster j, Mj : is the number of members of cluster j, and Mi: is the number of members of class i.
The F-measure of a class i is defined as: F (i) = 2P R P +R
With respect to class i, the cluster with the highest F- measure is considered to be the cluster that maps to class i, and that F-measure becomes the score for class i. The overall
2Though k-NN is mostly known to be used for classification, it has also been used for clustering (example could be found in [11]).
F-measure for the clustering result C is the weighted average of the F-measure for each class i:
FC = ∑
i(|i| × F (i))∑ i |i|
, (6)
where |i| is the number of objects in class i. The higher the overall F-measure, the better the clustering, due to the higher accuracy of the clusters mapping to the original classes.
The second measure is the Entropy, which provides a measure of quality for unnested clusters or for the clusters at one level of a hierarchical clustering. Entropy measures how homogeneous a cluster is. The higher the homogeneity of a cluster, the lower the entropy is, and vice versa. The entropy of a cluster containing only one object (perfect homogeneity) is zero.
For every cluster j in the clustering result C, the prob- ability pij that a member of cluster j belongs to class i is computed. The entropy of each cluster j is calculated using the standard formula Ej = −
∑ i pij log(pij ), where
the sum is taken over all classes. The total entropy for a set of clusters is calculated as the sum of entropies of each cluster weighted by the size of that cluster:
EC = n∑
j=1
( Mj M
× Ej ), (7)
where Mj is the size of cluster j, and M is the total number of data objects.
Basically, the aim is to maximize the F-measure, and minimize the Entropy of clusters to achieve high quality clustering. As shown in Tables (I and II), the ward linkage
475481
Authorized licensed use limited to: University of the Cumberlands. Downloaded on July 24,2021 at 04:14:27 UTC from IEEE Xplore. Restrictions apply.
was used as the cluster distance measures for the HAC method since they tend to produce tight clusters with small diameter. A document-to-cluster similarity threshold of 0.3 was used in the single pass clustering method. A k of 5 and a cluster similarity threshold of 0.35 were used in the k-NN method.
The parameters chosen for the different algorithms were the ones that produced best results.
For the terms and synonyms weighting, the percentage of improvement ranges from +28.35% to (+49.67% above 100%) increase in the F-measure quality, and -63.81% to -96.55% drop in Entropy.
For the terms and hypernyms weighting, the percentage of improvement ranges from +28.76% to (+48.71% above 100%) increase in the F-measure quality, and -63.25% to -96.62% drop in Entropy.
It is shown that the HAC clustering with the ward linkage has the best performance. It is known that Single Pass clustering is very sensitive to noise; that is why it has the worst performance. However, when the semantic-based similarity was introduced the quality of clusters produced was pushed close to that produced by k-NN.
It is illustrated that (terms and synonyms) has better performance than (terms and hypernyms) for some datasets and the situation is revered in the others. However, both of them have better performance than the terms weighting only. The combination between the terms and their corresponding synonyms or hypernyms weighting schemes has the best performance as it can measure the importance of a concept at the sentence and document levels which leads to enhance the clustering quality substantially.
IV. CONCLUSIONS
This work bridges the gap between natural language processing and text mining disciplines. A new semantic- based model composed of two components, is proposed to improve the text clustering quality. By exploiting the semantic structure of the sentences and adding synonyms and/or hypernyms to documents, a better text clustering result is achieved. The first component is the new semantic- based term analysis. First, it analyzes the semantic structure of each sentence to capture the top sentence terms. Secondly, words in each top term are looked up in WordNet to get their corresponding synonyms and/or hypernyms. Lastly, the component analyzes each concept at the sentence and docu- ment levels. These two levels of semantic-based analysis are achieved by using two semantic-based measures: the conceptual term frequency ctf and the concept frequency cf . The second component is the semantic-based similarity measure which allows measuring the importance of each concept with respect to the semantics of the sentence, and the topic of the document.
By combining the factors affecting the weights of con- cepts on the sentences and documents, a semantic-based
similarity measure that is capable of the accurate calculation of pair-wise documents is devised. This allows performing concept matching and semantic-based similarity calculations among documents in a very robust and accurate way. The quality of text clustering achieved by this model significantly surpasses the traditional single-term based approaches.
There are a number of possibilities for extending this work. one direction is to combine between synonyms and hypernyms into one term vector. The intention is to inves- tigate the usage of both synonyms and hypernyms together on other corpora and its effect on clustering, compared to that of traditional methods. Another direction is to link this work to web document clustering or text classification.
REFERENCES
[1] G. Salton and M. J. McGill, Introduction to Modern Infor- mation Retrieval. McGraw-Hill, 1983.
[2] P. Mitra, C. Murthy, and S. K. Pal, “Unsupervised feature selection using feature similarity,” IEEE Transactions on Pattern Analysis and Machine, vol. 24, no. 3, pp. 301–312, 2002.
[3] R. Nock and F. Nielsen, “Unsupervised feature selection using feature similarity,” IEEE Transactions on Pattern Analysis and Machine, vol. 28, no. 8, pp. 1223–1235, 2006.
[4] G. A. Miller, “Wordnet: a lexical database for english,” Commun. ACM, vol. 38, no. 11, pp. 39–41, 1995.
[5] S. Shehata, F. Karray, and M. Kamel, “Enhancing text cluster- ing using concept-based mining model,” in Proceedings of the 6th IEEE International Conference on Data Mining (ICDM), 2006.
[6] P. Kingsbury and M. Palmer, “Propbank: the next level of treebank,” in Proceedings of Treebanks and Lexical Theories, 2003.
[7] E. Brill, “A simple rule-based part-of-speech tagger,” in Proceedings of ANLP-92, 3rd Conference on Applied Natural Language Processing, Trento, IT, 1992, pp. 152–155.
[8] A. Hotho, S. Staab, and G. Stumme, “Wordnet improves text document clustering,” in Proceedings of the Semantic Web Workshop at SIGIR-2003, 26th Annual International ACM SIGIR Conference, 2003.
[9] W. Francis and H. Kucera, Manual of information to ac- company a standard corpus of present-day edited american english, for use with digital computers, 1964.
[10] A. K. Jain and R. C. Dubes, Algorithms for Clustering Data. Prentice Hall, Englewood Cliffs, N.J., 1988.
[11] S. Y. Lu and K. S. Fu, “A sentence to sentence clustering pro- cedure for pattern analysis,” IEEE Transactions on Systems Mans and Cybernetics, vol. 8, pp. 381–389, 1978.
[12] M. Steinbach, G. Karypis, and V. Kumar, “A comparison of document clustering techniques,” in Knowledge Discovery and Data Mining (KDD) Workshop on TextMining, August 2000.
476482
Authorized licensed use limited to: University of the Cumberlands. Downloaded on July 24,2021 at 04:14:27 UTC from IEEE Xplore. Restrictions apply.