Scalable feature selection, classification and signature generation for organizing large text databases into hierarchical topic taxonomies
S. Chakrabarti, B. Dom, R. Agrawal, und P. Raghavan.
The VLDB Journal (August 1998)

We explore how to organize large text databases hierarchically by topic to aid better searching, browsing and filtering. Many corpora, such as internet directories, digital libraries, and patent databases are manually organized into topic hierarchies, also called taxonomies. Similar to indices for relational data, taxonomies make search and access more efficient. However, the exponential growth in the volume of on-line textual information makes it nearly impossible to maintain such taxonomic organization for large, fast-changing corpora by hand. We describe an automatic system that starts with a small sample of the corpus in which topics have been assigned by hand, and then updates the database with new documents as the corpus grows, assigning topics to these new documents with high speed and accuracy. To do this, we use techniques from statistical pattern recognition to efficiently separate the feature words, or discriminants, from thenoise words at each node of the taxonomy. Using these, we build a multilevel classifier. At each node, this classifier can ignore the large number of “noise” words in a document. Thus, the classifier has a small model size and is very fast. Owing to the use of context-sensitive features, the classifier is very accurate. As a by-product, we can compute for each document a set of terms that occur significantly more often in it than in the classes to which it belongs. We describe the design and implementation of our system, stressing how to exploit standard, efficient relational operations like sorts and joins. We report on experiences with the Reuters newswire benchmark, the US patent database, and web document samples from Yahoo!. We discuss applications where our system can improve searching and filtering capabilities.

URL

http://dx.doi.org/10.1007/s007780050061

Suchen auf

Diese Publikation wurde noch nicht bewertet.

Bewertungsverteilung

Durchschnittliche Benutzerbewertung0,0 von 5.0 auf Grundlage von 0 Rezensionen

Bitte melden Sie sich an um selbst Rezensionen oder Kommentare zu erstellen.

@article{Chakrabarti:1998:SFS:765529.765533,
 abstract = {We explore how to organize large text databases hierarchically by topic to aid better searching, browsing and filtering. Many corpora, such as internet directories, digital libraries, and patent databases are manually organized into topic hierarchies, also called taxonomies. Similar to indices for relational data, taxonomies make search and access more efficient. However, the exponential growth in the volume of on-line textual information makes it nearly impossible to maintain such taxonomic organization for large, fast-changing corpora by hand. We describe an automatic system that starts with a small sample of the corpus in which topics have been assigned by hand, and then updates the database with new documents as the corpus grows, assigning topics to these new documents with high speed and accuracy. To do this, we use techniques from statistical pattern recognition to efficiently separate the feature words, or discriminants, from thenoise words at each node of the taxonomy. Using these, we build a multilevel classifier. At each node, this classifier can ignore the large number of &ldquo;noise&rdquo; words in a document. Thus, the classifier has a small model size and is very fast. Owing to the use of context-sensitive features, the classifier is very accurate. As a by-product, we can compute for each document a set of terms that occur significantly more often in it than in the classes to which it belongs. We describe the design and implementation of our system, stressing how to exploit standard, efficient relational operations like sorts and joins. We report on experiences with the Reuters newswire benchmark, the US patent database, and web document samples from Yahoo!. We discuss applications where our system can improve searching and filtering capabilities.},
 acmid = {765533},
 added-at = {2011-11-21T11:08:21.000+0100},
 address = {Secaucus, NJ, USA},
 author = {Chakrabarti, Soumen and Dom, Byron and Agrawal, Rakesh and Raghavan, Prabhakar},
 biburl = {https://puma.uni-kassel.de/bibtex/2606e518c01399c1e0569a00a81719343/benz},
 description = {Scalable feature selection, classification and signature generation for organizing large text databases into hierarchical topic taxonomies},
 doi = {10.1007/s007780050061},
 interhash = {5969518d9723da79c33437322c06474a},
 intrahash = {606e518c01399c1e0569a00a81719343},
 issn = {1066-8888},
 issue = {3},
 journal = {The VLDB Journal},
 keywords = {bachelor:2011:bachmann organization web},
 month = {August},
 numpages = {16},
 pages = {163--178},
 publisher = {Springer-Verlag New York, Inc.},
 timestamp = {2011-11-21T11:08:21.000+0100},
 title = {Scalable feature selection, classification and signature generation for organizing large text databases into hierarchical topic taxonomies},
 url = {http://dx.doi.org/10.1007/s007780050061},
 volume = 7,
 year = 1998
}

%0 Journal Article
%1 Chakrabarti:1998:SFS:765529.765533
%A Chakrabarti, Soumen
%A Dom, Byron
%A Agrawal, Rakesh
%A Raghavan, Prabhakar
%C Secaucus, NJ, USA
%D 1998
%I Springer-Verlag New York, Inc.
%J The VLDB Journal
%K bachelor:2011:bachmann organization web
%P 163--178
%R 10.1007/s007780050061
%T Scalable feature selection, classification and signature generation for organizing large text databases into hierarchical topic taxonomies
%U http://dx.doi.org/10.1007/s007780050061
%V 7
%X We explore how to organize large text databases hierarchically by topic to aid better searching, browsing and filtering. Many corpora, such as internet directories, digital libraries, and patent databases are manually organized into topic hierarchies, also called <i>taxonomies</i>. Similar to indices for relational data, taxonomies make search and access more efficient. However, the exponential growth in the volume of on-line textual information makes it nearly impossible to maintain such taxonomic organization for large, fast-changing corpora by hand. We describe an automatic system that starts with a small sample of the corpus in which topics have been assigned by hand, and then updates the database with new documents as the corpus grows, assigning topics to these new documents with high speed and accuracy. To do this, we use techniques from statistical pattern recognition to efficiently separate the <i>feature</i> words, or <i>discriminants</i>, from the<i>noise</i> words at each node of the taxonomy. Using these, we build a multilevel classifier. At each node, this classifier can ignore the large number of &ldquo;noise&rdquo; words in a document. Thus, the classifier has a small model size and is very fast. Owing to the use of context-sensitive features, the classifier is very accurate. As a by-product, we can compute for each document a set of terms that occur significantly more often in it than in the classes to which it belongs. We describe the design and implementation of our system, stressing how to exploit standard, efficient relational operations like sorts and joins. We report on experiences with the Reuters newswire benchmark, the US patent database, and web document samples from Yahoo!. We discuss applications where our system can improve searching and filtering capabilities.

PUMA

Scalable feature selection, classification and signature generation for organizing large text databases into hierarchical topic taxonomies
S. Chakrabarti, B. Dom, R. Agrawal, und P. Raghavan.
The VLDB Journal (August 1998)

Tags

Nutzer

Kommentare und Rezensionen

Zitieren Sie diese Publikation

PUMA

Scalable feature selection, classification and signature generation for organizing large text databases into hierarchical topic taxonomiesS. Chakrabarti, B. Dom, R. Agrawal, und P. Raghavan. The VLDB Journal (August 1998)

Tags

Nutzer

Kommentare und Rezensionen

Zitieren Sie diese Publikation

Scalable feature selection, classification and signature generation for organizing large text databases into hierarchical topic taxonomies
S. Chakrabarti, B. Dom, R. Agrawal, und P. Raghavan.
The VLDB Journal (August 1998)