Document (#32319)

Author
Kaufmann, E.
Title
¬Das Indexieren von natürlichsprachlichen Dokumenten und die inverse Seitenhäufigkeit
Source
http://www.ifi.unizh.ch/cl/study/lizarbeiten/lizkaufmann.pdf
Year
2001
Abstract
Die Lizentiatsarbeit gibt im ersten theoretischen Teil einen Überblick über das Indexieren von Dokumenten. Sie zeigt die verschiedenen Typen von Indexen sowie die wichtigsten Aspekte bezüglich einer Indexsprache auf. Diverse manuelle und automatische Indexierungsverfahren werden präsentiert. Spezielle Aufmerksamkeit innerhalb des ersten Teils gilt den Schlagwortregistern, deren charakteristische Merkmale und Eigenheiten erörtert werden. Zusätzlich werden die gängigen Kriterien zur Bewertung von Indexen sowie die Masse zur Evaluation von Indexierungsverfahren und Indexierungsergebnissen vorgestellt. Im zweiten Teil der Arbeit werden fünf reale Bücher einer statistischen Untersuchung unterzogen. Zum einen werden die lexikalischen und syntaktischen Bestandteile der fünf Buchregister ermittelt, um den Inhalt von Schlagwortregistern zu erschliessen. Andererseits werden aus den Textausschnitten der Bücher Indexterme maschinell extrahiert und mit den Schlagworteinträgen in den Buchregistern verglichen. Das Hauptziel der Untersuchungen besteht darin, eine Indexierungsmethode, die auf linguistikorientierter Extraktion der Indexterme und Termhäufigkeitsgewichtung basiert, im Hinblick auf ihren Gebrauchswert für eine automatische Indexierung zu testen. Die Gewichtungsmethode ist die inverse Seitenhäufigkeit, eine Methode, welche von der inversen Dokumentfrequenz abgeleitet wurde, zur automatischen Erstellung von Schlagwortregistern für deutschsprachige Texte. Die Prüfung der Methode im statistischen Teil führte nicht zu zufriedenstellenden Resultaten.
Content
Lizentiatsarbeit der Philosphischen Fakultät der Universität Zürich, - Vgl. auch: http://www.ifi.unizh.ch/cl/study/lizarbeiten/lizkaufmann.pdf.
Theme
Automatisches Indexieren
Register

Similar documents (author)

  1. Kaufmann, N.C.: Kommt das Domainsterben? : Rechtsprechung gibt beschreibende Internet-Adressen zum Abschuss frei (2000) 6.01
    6.0137663 = sum of:
      6.0137663 = weight(author_txt:kaufmann in 5278) [ClassicSimilarity], result of:
        6.0137663 = fieldWeight in 5278, product of:
          1.0 = tf(freq=1.0), with freq of:
            1.0 = termFreq=1.0
          9.622026 = idf(docFreq=7, maxDocs=44421)
          0.625 = fieldNorm(doc=5278)
    
  2. Kaufmann, T.: Googeln wie die Profis : Perfekte Suche (2004) 6.01
    6.0137663 = sum of:
      6.0137663 = weight(author_txt:kaufmann in 2926) [ClassicSimilarity], result of:
        6.0137663 = fieldWeight in 2926, product of:
          1.0 = tf(freq=1.0), with freq of:
            1.0 = termFreq=1.0
          9.622026 = idf(docFreq=7, maxDocs=44421)
          0.625 = fieldNorm(doc=2926)
    
  3. Havemann, F.; Kaufmann, A.: ¬Der Wandel des Benutzerverhaltens in Zeiten des Internet : Ergebnisse von Befragungen an 13 Bibliotheken (2006) 4.81
    4.811013 = sum of:
      4.811013 = weight(author_txt:kaufmann in 149) [ClassicSimilarity], result of:
        4.811013 = fieldWeight in 149, product of:
          1.0 = tf(freq=1.0), with freq of:
            1.0 = termFreq=1.0
          9.622026 = idf(docFreq=7, maxDocs=44421)
          0.5 = fieldNorm(doc=149)
    
  4. Kaufmann, J.-C.: Wenn ICH ein anderer ist (2010) 4.81
    4.811013 = sum of:
      4.811013 = weight(author_txt:kaufmann in 625) [ClassicSimilarity], result of:
        4.811013 = fieldWeight in 625, product of:
          1.0 = tf(freq=1.0), with freq of:
            1.0 = termFreq=1.0
          9.622026 = idf(docFreq=7, maxDocs=44421)
          0.5 = fieldNorm(doc=625)
    
  5. Kaufmann, J.-C.: ¬Die Erfindung des Ich : eine Theorie der Identität (2005) 4.81
    4.811013 = sum of:
      4.811013 = weight(author_txt:kaufmann in 626) [ClassicSimilarity], result of:
        4.811013 = fieldWeight in 626, product of:
          1.0 = tf(freq=1.0), with freq of:
            1.0 = termFreq=1.0
          9.622026 = idf(docFreq=7, maxDocs=44421)
          0.5 = fieldNorm(doc=626)
    

Similar documents (content)

  1. Lepsky, K.: Automatisches Indexieren (2023) 0.28
    0.28376195 = sum of:
      0.28376195 = product of:
        1.1823415 = sum of:
          0.024151627 = weight(abstract_txt:eine in 1782) [ClassicSimilarity], result of:
            0.024151627 = score(doc=1782,freq=1.0), product of:
              0.07383915 = queryWeight, product of:
                1.1878996 = boost
                3.4888992 = idf(docFreq=3686, maxDocs=44421)
                0.01781634 = queryNorm
              0.3270843 = fieldWeight in 1782, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                3.4888992 = idf(docFreq=3686, maxDocs=44421)
                0.09375 = fieldNorm(doc=1782)
          0.101449706 = weight(abstract_txt:dokumenten in 1782) [ClassicSimilarity], result of:
            0.101449706 = score(doc=1782,freq=1.0), product of:
              0.16792905 = queryWeight, product of:
                1.4626948 = boost
                6.443972 = idf(docFreq=191, maxDocs=44421)
                0.01781634 = queryNorm
              0.6041224 = fieldWeight in 1782, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.443972 = idf(docFreq=191, maxDocs=44421)
                0.09375 = fieldNorm(doc=1782)
          0.125762 = weight(abstract_txt:automatische in 1782) [ClassicSimilarity], result of:
            0.125762 = score(doc=1782,freq=1.0), product of:
              0.19378714 = queryWeight, product of:
                1.5712788 = boost
                6.922344 = idf(docFreq=118, maxDocs=44421)
                0.01781634 = queryNorm
              0.64896977 = fieldWeight in 1782, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.922344 = idf(docFreq=118, maxDocs=44421)
                0.09375 = fieldNorm(doc=1782)
          0.2821766 = weight(abstract_txt:indexierungsverfahren in 1782) [ClassicSimilarity], result of:
            0.2821766 = score(doc=1782,freq=1.0), product of:
              0.3321284 = queryWeight, product of:
                2.057045 = boost
                9.06241 = idf(docFreq=13, maxDocs=44421)
                0.01781634 = queryNorm
              0.849601 = fieldWeight in 1782, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                9.06241 = idf(docFreq=13, maxDocs=44421)
                0.09375 = fieldNorm(doc=1782)
          0.5637714 = weight(abstract_txt:indexterme in 1782) [ClassicSimilarity], result of:
            0.5637714 = score(doc=1782,freq=3.0), product of:
              0.3653033 = queryWeight, product of:
                2.157335 = boost
                9.504243 = idf(docFreq=8, maxDocs=44421)
                0.01781634 = queryNorm
              1.5432968 = fieldWeight in 1782, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                9.504243 = idf(docFreq=8, maxDocs=44421)
                0.09375 = fieldNorm(doc=1782)
          0.08503012 = weight(abstract_txt:werden in 1782) [ClassicSimilarity], result of:
            0.08503012 = score(doc=1782,freq=3.0), product of:
              0.14928192 = queryWeight, product of:
                2.3886638 = boost
                3.507791 = idf(docFreq=3617, maxDocs=44421)
                0.01781634 = queryNorm
              0.56959426 = fieldWeight in 1782, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                3.507791 = idf(docFreq=3617, maxDocs=44421)
                0.09375 = fieldNorm(doc=1782)
        0.24 = coord(6/25)
    
  2. Halip, I.: Automatische Extrahierung von Schlagworten aus unstrukturierten Texten (2005) 0.19
    0.18685178 = sum of:
      0.18685178 = product of:
        0.58391184 = sum of:
          0.0756426 = weight(abstract_txt:manuelle in 986) [ClassicSimilarity], result of:
            0.0756426 = score(doc=986,freq=1.0), product of:
              0.15698148 = queryWeight, product of:
                8.811096 = idf(docFreq=17, maxDocs=44421)
                0.01781634 = queryNorm
              0.48185682 = fieldWeight in 986, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                8.811096 = idf(docFreq=17, maxDocs=44421)
                0.0546875 = fieldNorm(doc=986)
          0.021829253 = weight(abstract_txt:sowie in 986) [ClassicSimilarity], result of:
            0.021829253 = score(doc=986,freq=1.0), product of:
              0.086372085 = queryWeight, product of:
                1.0490048 = boost
                4.621441 = idf(docFreq=1187, maxDocs=44421)
                0.01781634 = queryNorm
              0.25273505 = fieldWeight in 986, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.621441 = idf(docFreq=1187, maxDocs=44421)
                0.0546875 = fieldNorm(doc=986)
          0.019924076 = weight(abstract_txt:eine in 986) [ClassicSimilarity], result of:
            0.019924076 = score(doc=986,freq=2.0), product of:
              0.07383915 = queryWeight, product of:
                1.1878996 = boost
                3.4888992 = idf(docFreq=3686, maxDocs=44421)
                0.01781634 = queryNorm
              0.2698308 = fieldWeight in 986, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                3.4888992 = idf(docFreq=3686, maxDocs=44421)
                0.0546875 = fieldNorm(doc=986)
          0.10250103 = weight(abstract_txt:dokumenten in 986) [ClassicSimilarity], result of:
            0.10250103 = score(doc=986,freq=3.0), product of:
              0.16792905 = queryWeight, product of:
                1.4626948 = boost
                6.443972 = idf(docFreq=191, maxDocs=44421)
                0.01781634 = queryNorm
              0.6103829 = fieldWeight in 986, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                6.443972 = idf(docFreq=191, maxDocs=44421)
                0.0546875 = fieldNorm(doc=986)
          0.073361166 = weight(abstract_txt:automatische in 986) [ClassicSimilarity], result of:
            0.073361166 = score(doc=986,freq=1.0), product of:
              0.19378714 = queryWeight, product of:
                1.5712788 = boost
                6.922344 = idf(docFreq=118, maxDocs=44421)
                0.01781634 = queryNorm
              0.3785657 = fieldWeight in 986, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.922344 = idf(docFreq=118, maxDocs=44421)
                0.0546875 = fieldNorm(doc=986)
          0.05590446 = weight(abstract_txt:teil in 986) [ClassicSimilarity], result of:
            0.05590446 = score(doc=986,freq=1.0), product of:
              0.18507262 = queryWeight, product of:
                1.8806479 = boost
                5.5235233 = idf(docFreq=481, maxDocs=44421)
                0.01781634 = queryNorm
              0.3020677 = fieldWeight in 986, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.5235233 = idf(docFreq=481, maxDocs=44421)
                0.0546875 = fieldNorm(doc=986)
          0.16460302 = weight(abstract_txt:indexierungsverfahren in 986) [ClassicSimilarity], result of:
            0.16460302 = score(doc=986,freq=1.0), product of:
              0.3321284 = queryWeight, product of:
                2.057045 = boost
                9.06241 = idf(docFreq=13, maxDocs=44421)
                0.01781634 = queryNorm
              0.49560058 = fieldWeight in 986, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                9.06241 = idf(docFreq=13, maxDocs=44421)
                0.0546875 = fieldNorm(doc=986)
          0.07014628 = weight(abstract_txt:werden in 986) [ClassicSimilarity], result of:
            0.07014628 = score(doc=986,freq=6.0), product of:
              0.14928192 = queryWeight, product of:
                2.3886638 = boost
                3.507791 = idf(docFreq=3617, maxDocs=44421)
                0.01781634 = queryNorm
              0.4698913 = fieldWeight in 986, product of:
                2.4494898 = tf(freq=6.0), with freq of:
                  6.0 = termFreq=6.0
                3.507791 = idf(docFreq=3617, maxDocs=44421)
                0.0546875 = fieldNorm(doc=986)
        0.32 = coord(8/25)
    
  3. Bredack, J.: Automatische Extraktion fachterminologischer Mehrwortbegriffe : ein Verfahrensvergleich (2016) 0.15
    0.15068533 = sum of:
      0.15068533 = product of:
        0.62785554 = sum of:
          0.09405887 = weight(abstract_txt:extrahiert in 4194) [ClassicSimilarity], result of:
            0.09405887 = score(doc=4194,freq=1.0), product of:
              0.1660642 = queryWeight, product of:
                1.0285225 = boost
                9.06241 = idf(docFreq=13, maxDocs=44421)
                0.01781634 = queryNorm
              0.56640065 = fieldWeight in 4194, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                9.06241 = idf(docFreq=13, maxDocs=44421)
                0.0625 = fieldNorm(doc=4194)
          0.022770373 = weight(abstract_txt:eine in 4194) [ClassicSimilarity], result of:
            0.022770373 = score(doc=4194,freq=2.0), product of:
              0.07383915 = queryWeight, product of:
                1.1878996 = boost
                3.4888992 = idf(docFreq=3686, maxDocs=44421)
                0.01781634 = queryNorm
              0.30837804 = fieldWeight in 4194, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                3.4888992 = idf(docFreq=3686, maxDocs=44421)
                0.0625 = fieldNorm(doc=4194)
          0.08384133 = weight(abstract_txt:automatische in 4194) [ClassicSimilarity], result of:
            0.08384133 = score(doc=4194,freq=1.0), product of:
              0.19378714 = queryWeight, product of:
                1.5712788 = boost
                6.922344 = idf(docFreq=118, maxDocs=44421)
                0.01781634 = queryNorm
              0.4326465 = fieldWeight in 4194, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.922344 = idf(docFreq=118, maxDocs=44421)
                0.0625 = fieldNorm(doc=4194)
          0.13002208 = weight(abstract_txt:statistischen in 4194) [ClassicSimilarity], result of:
            0.13002208 = score(doc=4194,freq=1.0), product of:
              0.2596356 = queryWeight, product of:
                1.8187495 = boost
                8.0125885 = idf(docFreq=39, maxDocs=44421)
                0.01781634 = queryNorm
              0.5007868 = fieldWeight in 4194, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                8.0125885 = idf(docFreq=39, maxDocs=44421)
                0.0625 = fieldNorm(doc=4194)
          0.21699572 = weight(abstract_txt:indexterme in 4194) [ClassicSimilarity], result of:
            0.21699572 = score(doc=4194,freq=1.0), product of:
              0.3653033 = queryWeight, product of:
                2.157335 = boost
                9.504243 = idf(docFreq=8, maxDocs=44421)
                0.01781634 = queryNorm
              0.5940152 = fieldWeight in 4194, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                9.504243 = idf(docFreq=8, maxDocs=44421)
                0.0625 = fieldNorm(doc=4194)
          0.080167174 = weight(abstract_txt:werden in 4194) [ClassicSimilarity], result of:
            0.080167174 = score(doc=4194,freq=6.0), product of:
              0.14928192 = queryWeight, product of:
                2.3886638 = boost
                3.507791 = idf(docFreq=3617, maxDocs=44421)
                0.01781634 = queryNorm
              0.53701866 = fieldWeight in 4194, product of:
                2.4494898 = tf(freq=6.0), with freq of:
                  6.0 = termFreq=6.0
                3.507791 = idf(docFreq=3617, maxDocs=44421)
                0.0625 = fieldNorm(doc=4194)
        0.24 = coord(6/25)
    
  4. Leonhardt, H.A.: Systematik "Ästhetische Kulturwissenschaft" an der Universitätsbibliothek Hildesheim : ein Innovationsbericht (2018) 0.13
    0.13194902 = sum of:
      0.13194902 = product of:
        0.6597451 = sum of:
          0.03220217 = weight(abstract_txt:eine in 490) [ClassicSimilarity], result of:
            0.03220217 = score(doc=490,freq=1.0), product of:
              0.07383915 = queryWeight, product of:
                1.1878996 = boost
                3.4888992 = idf(docFreq=3686, maxDocs=44421)
                0.01781634 = queryNorm
              0.4361124 = fieldWeight in 490, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                3.4888992 = idf(docFreq=3686, maxDocs=44421)
                0.125 = fieldNorm(doc=490)
          0.15186961 = weight(abstract_txt:bücher in 490) [ClassicSimilarity], result of:
            0.15186961 = score(doc=490,freq=1.0), product of:
              0.18140396 = queryWeight, product of:
                1.520247 = boost
                6.697521 = idf(docFreq=148, maxDocs=44421)
                0.01781634 = queryNorm
              0.83719015 = fieldWeight in 490, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.697521 = idf(docFreq=148, maxDocs=44421)
                0.125 = fieldNorm(doc=490)
          0.25532264 = weight(abstract_txt:indexieren in 490) [ClassicSimilarity], result of:
            0.25532264 = score(doc=490,freq=1.0), product of:
              0.2564833 = queryWeight, product of:
                1.8076748 = boost
                7.963798 = idf(docFreq=41, maxDocs=44421)
                0.01781634 = queryNorm
              0.99547476 = fieldWeight in 490, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                7.963798 = idf(docFreq=41, maxDocs=44421)
                0.125 = fieldNorm(doc=490)
          0.12778161 = weight(abstract_txt:teil in 490) [ClassicSimilarity], result of:
            0.12778161 = score(doc=490,freq=1.0), product of:
              0.18507262 = queryWeight, product of:
                1.8806479 = boost
                5.5235233 = idf(docFreq=481, maxDocs=44421)
                0.01781634 = queryNorm
              0.6904404 = fieldWeight in 490, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.5235233 = idf(docFreq=481, maxDocs=44421)
                0.125 = fieldNorm(doc=490)
          0.09256907 = weight(abstract_txt:werden in 490) [ClassicSimilarity], result of:
            0.09256907 = score(doc=490,freq=2.0), product of:
              0.14928192 = queryWeight, product of:
                2.3886638 = boost
                3.507791 = idf(docFreq=3617, maxDocs=44421)
                0.01781634 = queryNorm
              0.6200957 = fieldWeight in 490, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                3.507791 = idf(docFreq=3617, maxDocs=44421)
                0.125 = fieldNorm(doc=490)
        0.2 = coord(5/25)
    
  5. Larroche-Boutet, V.; Pöhl, K.: ¬Das Nominalsyntagna : über die Nutzbarmachung eines logico-semantischen Konzeptes für dokumentarische Fragestellungen (1993) 0.11
    0.11492395 = sum of:
      0.11492395 = product of:
        0.5746197 = sum of:
          0.09405887 = weight(abstract_txt:extrahiert in 6282) [ClassicSimilarity], result of:
            0.09405887 = score(doc=6282,freq=1.0), product of:
              0.1660642 = queryWeight, product of:
                1.0285225 = boost
                9.06241 = idf(docFreq=13, maxDocs=44421)
                0.01781634 = queryNorm
              0.56640065 = fieldWeight in 6282, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                9.06241 = idf(docFreq=13, maxDocs=44421)
                0.0625 = fieldNorm(doc=6282)
          0.022770373 = weight(abstract_txt:eine in 6282) [ClassicSimilarity], result of:
            0.022770373 = score(doc=6282,freq=2.0), product of:
              0.07383915 = queryWeight, product of:
                1.1878996 = boost
                3.4888992 = idf(docFreq=3686, maxDocs=44421)
                0.01781634 = queryNorm
              0.30837804 = fieldWeight in 6282, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                3.4888992 = idf(docFreq=3686, maxDocs=44421)
                0.0625 = fieldNorm(doc=6282)
          0.118569545 = weight(abstract_txt:automatische in 6282) [ClassicSimilarity], result of:
            0.118569545 = score(doc=6282,freq=2.0), product of:
              0.19378714 = queryWeight, product of:
                1.5712788 = boost
                6.922344 = idf(docFreq=118, maxDocs=44421)
                0.01781634 = queryNorm
              0.61185455 = fieldWeight in 6282, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                6.922344 = idf(docFreq=118, maxDocs=44421)
                0.0625 = fieldNorm(doc=6282)
          0.26603866 = weight(abstract_txt:indexierungsverfahren in 6282) [ClassicSimilarity], result of:
            0.26603866 = score(doc=6282,freq=2.0), product of:
              0.3321284 = queryWeight, product of:
                2.057045 = boost
                9.06241 = idf(docFreq=13, maxDocs=44421)
                0.01781634 = queryNorm
              0.80101144 = fieldWeight in 6282, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                9.06241 = idf(docFreq=13, maxDocs=44421)
                0.0625 = fieldNorm(doc=6282)
          0.073182285 = weight(abstract_txt:werden in 6282) [ClassicSimilarity], result of:
            0.073182285 = score(doc=6282,freq=5.0), product of:
              0.14928192 = queryWeight, product of:
                2.3886638 = boost
                3.507791 = idf(docFreq=3617, maxDocs=44421)
                0.01781634 = queryNorm
              0.4902287 = fieldWeight in 6282, product of:
                2.236068 = tf(freq=5.0), with freq of:
                  5.0 = termFreq=5.0
                3.507791 = idf(docFreq=3617, maxDocs=44421)
                0.0625 = fieldNorm(doc=6282)
        0.2 = coord(5/25)