Automatic classification of scientific records using the German Subject Heading Authority File (SWD)
- The following paper deals with an automatic text classification method which does not require training documents. For this method the German Subject Heading Authority File (SWD), provided by the linked data service of the German National Library is used. Recently the SWD was enriched with notations of the Dewey Decimal Classification (DDC). In consequence it became possible to utilize the subject headings as textual representations for the notations of the DDC. Basically, we we derive the classification of a text from the classification of the words in the text given by the thesaurus. The method was tested by classifying 3826 OAI-Records from 7 different repositories. Mean reciprocal rank and recall were chosen as evaluation measure. Direct comparison to a machine learning method has shown that this method is definitely competitive. Thus we can conclude that the enriched version of the SWD provides high quality information with a broad coverage for classification of German scientific articles.
Author: | Christian WartenaORCiDGND, Maike Sommer |
---|---|
URN: | urn:nbn:de:bsz:960-opus-4008 |
URL: | http://ceur-ws.org/Vol-912/proceedings.pdf#page=37 |
DOI: | https://doi.org/10.25968/opus-328 |
Parent Title (German): | Proceedings of the 2nd International Workshop on Semantic Digital Archives (SDA 2012) |
Document Type: | Conference Proceeding |
Language: | English |
Year of Completion: | 2012 |
Publishing Institution: | Hochschule Hannover |
Release Date: | 2012/10/29 |
GND Keyword: | Notation <Klassifikation>; Schlagwortnormdatei; Schlagwortkatalog; Dewey-Dezimalklassifikation; Text Mining |
First Page: | 37 |
Last Page: | 48 |
Link to catalogue: | 729343464 |
Institutes: | Fakultät III - Medien, Information und Design |
DDC classes: | 020 Bibliotheks- und Informationswissenschaft |
Licence (German): | ![]() |