Volltext-Downloads (blau) und Frontdoor-Views (grau)

Automatic classification of scientific records using the German Subject Heading Authority File (SWD)

  • The following paper deals with an automatic text classification method which does not require training documents. For this method the German Subject Heading Authority File (SWD), provided by the linked data service of the German National Library is used. Recently the SWD was enriched with notations of the Dewey Decimal Classification (DDC). In consequence it became possible to utilize the subject headings as textual representations for the notations of the DDC. Basically, we we derive the classification of a text from the classification of the words in the text given by the thesaurus. The method was tested by classifying 3826 OAI-Records from 7 different repositories. Mean reciprocal rank and recall were chosen as evaluation measure. Direct comparison to a machine learning method has shown that this method is definitely competitive. Thus we can conclude that the enriched version of the SWD provides high quality information with a broad coverage for classification of German scientific articles.

Download full text files

Export metadata

Additional Services

Search Google Scholar


Author:Christian WartenaORCiDGND, Maike Sommer
Parent Title (German):Proceedings of the 2nd International Workshop on Semantic Digital Archives (SDA 2012)
Document Type:Conference Proceeding
Year of Completion:2012
Publishing Institution:Hochschule Hannover
Release Date:2012/10/29
GND Keyword:Notation <Klassifikation>; Schlagwortnormdatei; Schlagwortkatalog; Dewey-Dezimalklassifikation; Text Mining
First Page:37
Last Page:48
Link to catalogue:729343464
Institutes:Fakultät III - Medien, Information und Design
DDC classes:020 Bibliotheks- und Informationswissenschaft
Licence (German):License LogoUrheberrechtlich geschützt