Automatic classification of scientific records using the German Subject Heading Authority File (SWD)

The following paper deals with an automatic text classification method which does not require training documents. For this method the German Subject Heading Authority File (SWD), provided by the linked data service of the German National Library is used. Recently the SWD was enriched with notations of the Dewey Decimal Classification (DDC). In consequence it became possible to utilize the subject headings as textual representations for the notations of the DDC. Basically, we we derive the classification of a text from the classification of the words in the text given by the thesaurus. The method was tested by classifying 3826 OAI-Records from 7 different repositories. Mean reciprocal rank and recall were chosen as evaluation measure. Direct comparison to a machine learning method has shown that this method is definitely competitive. Thus we can conclude that the enriched version of the SWD provides high quality information with a broad coverage for classification of German scientific articles.

Download full text files

Export metadata

  • Export Bibtex
  • Export RIS

Additional Services

    Share in Twitter Search Google Scholar
Author:Christian Wartena, Maike Sommer
Document Type:Conference Proceeding
Year of Completion:2012
Release Date:2012/10/29
SWD-Keyword:Dewey-Dezimalklassifikation; Notation <Klassifikation>; Schlagwortkatalog; Schlagwortnormdatei; Text Mining
Source:Erschinenen in: Proceedings of the 2nd International Workshop on Semantic Digital Archives (SDA 2012), 2012, S. 37-48,
To order the print edition:729343464
Institutes:Fakult├Ąt III - Medien, Information und Design
Dewey Decimal Classification:020 Bibliotheks- und Informationswissenschaften
Licence (German):License LogoHinweis zum Urheberrecht