Document (#28460)

Author
Ding, C.H.Q.
Title
¬A probabilistic model for Latent Semantic Indexing
Source
Journal of the American Society for Information Science and Technology. 56(2005) no.6, S.597-608
Year
2005
Abstract
Latent Semantic Indexing (LSI), when applied to semantic space built an text collections, improves information retrieval, information filtering, and word sense disambiguation. A new dual probability model based an the similarity concepts is introduced to provide deeper understanding of LSI. Semantic associations can be quantitatively characterized by their statistical significance, the likelihood. Semantic dimensions containing redundant and noisy information can be separated out and should be ignored because their negative contribution to the overall statistical significance. LSI is the optimal solution of the model. The peak in the likelihood curve indicates the existence of an intrinsic semantic dimension. The importance of LSI dimensions follows the Zipf-distribution, indicating that LSI dimensions represent latent concepts. Document frequency of words follows the Zipf distribution, and the number of distinct words follows log-normal distribution. Experiments an five standard document collections confirm and illustrate the analysis.
Theme
Retrievalstudien
Object
Latent Semantic Indexing

Similar documents (author)

  1. Ding, Y.: Visualization of intellectual structure in information retrieval : author cocitation analysis (1998) 4.76
    4.7649565 = sum of:
      4.7649565 = weight(author_txt:ding in 3792) [ClassicSimilarity], result of:
        4.7649565 = fieldWeight in 3792, product of:
          1.0 = tf(freq=1.0), with freq of:
            1.0 = termFreq=1.0
          7.62393 = idf(docFreq=58, maxDocs=44421)
          0.625 = fieldNorm(doc=3792)
    
  2. Ding, Y.: Scholarly communication and bibliometrics : Part 1: The scholarly communication model: literature review (1998) 4.76
    4.7649565 = sum of:
      4.7649565 = weight(author_txt:ding in 4995) [ClassicSimilarity], result of:
        4.7649565 = fieldWeight in 4995, product of:
          1.0 = tf(freq=1.0), with freq of:
            1.0 = termFreq=1.0
          7.62393 = idf(docFreq=58, maxDocs=44421)
          0.625 = fieldNorm(doc=4995)
    
  3. Ding, Y.: ¬A review of ontologies with the Semantic Web in view (2001) 4.76
    4.7649565 = sum of:
      4.7649565 = weight(author_txt:ding in 5152) [ClassicSimilarity], result of:
        4.7649565 = fieldWeight in 5152, product of:
          1.0 = tf(freq=1.0), with freq of:
            1.0 = termFreq=1.0
          7.62393 = idf(docFreq=58, maxDocs=44421)
          0.625 = fieldNorm(doc=5152)
    
  4. Ding, Y.: Applying weighted PageRank to author citation networks (2011) 4.76
    4.7649565 = sum of:
      4.7649565 = weight(author_txt:ding in 188) [ClassicSimilarity], result of:
        4.7649565 = fieldWeight in 188, product of:
          1.0 = tf(freq=1.0), with freq of:
            1.0 = termFreq=1.0
          7.62393 = idf(docFreq=58, maxDocs=44421)
          0.625 = fieldNorm(doc=188)
    
  5. Ding, Y.: Topic-based PageRank on author cocitation networks (2011) 4.76
    4.7649565 = sum of:
      4.7649565 = weight(author_txt:ding in 348) [ClassicSimilarity], result of:
        4.7649565 = fieldWeight in 348, product of:
          1.0 = tf(freq=1.0), with freq of:
            1.0 = termFreq=1.0
          7.62393 = idf(docFreq=58, maxDocs=44421)
          0.625 = fieldNorm(doc=348)
    

Similar documents (content)

  1. Zhu, W.Z.; Allen, R.B.: Document clustering using the LSI subspace signature model (2013) 0.31
    0.30636734 = sum of:
      0.30636734 = product of:
        0.95739794 = sum of:
          0.05173103 = weight(abstract_txt:document in 1690) [ClassicSimilarity], result of:
            0.05173103 = score(doc=1690,freq=3.0), product of:
              0.08902732 = queryWeight, product of:
                1.1501118 = boost
                4.29415 = idf(docFreq=1647, maxDocs=44421)
                0.018026277 = queryNorm
              0.5810692 = fieldWeight in 1690, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                4.29415 = idf(docFreq=1647, maxDocs=44421)
                0.078125 = fieldNorm(doc=1690)
          0.04391717 = weight(abstract_txt:indexing in 1690) [ClassicSimilarity], result of:
            0.04391717 = score(doc=1690,freq=2.0), product of:
              0.09137117 = queryWeight, product of:
                1.1651531 = boost
                4.3503094 = idf(docFreq=1557, maxDocs=44421)
                0.018026277 = queryNorm
              0.48064584 = fieldWeight in 1690, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                4.3503094 = idf(docFreq=1557, maxDocs=44421)
                0.078125 = fieldNorm(doc=1690)
          0.0643638 = weight(abstract_txt:statistical in 1690) [ClassicSimilarity], result of:
            0.0643638 = score(doc=1690,freq=1.0), product of:
              0.14853337 = queryWeight, product of:
                1.4855608 = boost
                5.5466094 = idf(docFreq=470, maxDocs=44421)
                0.018026277 = queryNorm
              0.43332887 = fieldWeight in 1690, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.5466094 = idf(docFreq=470, maxDocs=44421)
                0.078125 = fieldNorm(doc=1690)
          0.0714891 = weight(abstract_txt:model in 1690) [ClassicSimilarity], result of:
            0.0714891 = score(doc=1690,freq=4.0), product of:
              0.11487705 = queryWeight, product of:
                1.6000762 = boost
                3.9827821 = idf(docFreq=2249, maxDocs=44421)
                0.018026277 = queryNorm
              0.6223097 = fieldWeight in 1690, product of:
                2.0 = tf(freq=4.0), with freq of:
                  4.0 = termFreq=4.0
                3.9827821 = idf(docFreq=2249, maxDocs=44421)
                0.078125 = fieldNorm(doc=1690)
          0.09918371 = weight(abstract_txt:distribution in 1690) [ClassicSimilarity], result of:
            0.09918371 = score(doc=1690,freq=1.0), product of:
              0.22684033 = queryWeight, product of:
                2.2484548 = boost
                5.5966744 = idf(docFreq=447, maxDocs=44421)
                0.018026277 = queryNorm
              0.43724018 = fieldWeight in 1690, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.5966744 = idf(docFreq=447, maxDocs=44421)
                0.078125 = fieldNorm(doc=1690)
          0.116161585 = weight(abstract_txt:dimensions in 1690) [ClassicSimilarity], result of:
            0.116161585 = score(doc=1690,freq=1.0), product of:
              0.25203937 = queryWeight, product of:
                2.3700538 = boost
                5.899349 = idf(docFreq=330, maxDocs=44421)
                0.018026277 = queryNorm
              0.46088666 = fieldWeight in 1690, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.899349 = idf(docFreq=330, maxDocs=44421)
                0.078125 = fieldNorm(doc=1690)
          0.33496812 = weight(abstract_txt:latent in 1690) [ClassicSimilarity], result of:
            0.33496812 = score(doc=1690,freq=3.0), product of:
              0.3540424 = queryWeight, product of:
                2.8089995 = boost
                6.9919376 = idf(docFreq=110, maxDocs=44421)
                0.018026277 = queryNorm
              0.94612426 = fieldWeight in 1690, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                6.9919376 = idf(docFreq=110, maxDocs=44421)
                0.078125 = fieldNorm(doc=1690)
          0.17558344 = weight(abstract_txt:semantic in 1690) [ClassicSimilarity], result of:
            0.17558344 = score(doc=1690,freq=3.0), product of:
              0.28999153 = queryWeight, product of:
                3.5952716 = boost
                4.4745317 = idf(docFreq=1375, maxDocs=44421)
                0.018026277 = queryNorm
              0.6054778 = fieldWeight in 1690, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                4.4745317 = idf(docFreq=1375, maxDocs=44421)
                0.078125 = fieldNorm(doc=1690)
        0.32 = coord(8/25)
    
  2. Li, D.; Kwong, C.-P.; Lee, D.L.: Unified linear subspace approach to semantic analysis (2009) 0.20
    0.19539018 = sum of:
      0.19539018 = product of:
        0.6978221 = sum of:
          0.10084477 = weight(abstract_txt:dual in 308) [ClassicSimilarity], result of:
            0.10084477 = score(doc=308,freq=2.0), product of:
              0.14647108 = queryWeight, product of:
                1.0431322 = boost
                7.7894444 = idf(docFreq=49, maxDocs=44421)
                0.018026277 = queryNorm
              0.6884961 = fieldWeight in 308, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                7.7894444 = idf(docFreq=49, maxDocs=44421)
                0.0625 = fieldNorm(doc=308)
          0.03379057 = weight(abstract_txt:document in 308) [ClassicSimilarity], result of:
            0.03379057 = score(doc=308,freq=2.0), product of:
              0.08902732 = queryWeight, product of:
                1.1501118 = boost
                4.29415 = idf(docFreq=1647, maxDocs=44421)
                0.018026277 = queryNorm
              0.3795528 = fieldWeight in 308, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                4.29415 = idf(docFreq=1647, maxDocs=44421)
                0.0625 = fieldNorm(doc=308)
          0.024843303 = weight(abstract_txt:indexing in 308) [ClassicSimilarity], result of:
            0.024843303 = score(doc=308,freq=1.0), product of:
              0.09137117 = queryWeight, product of:
                1.1651531 = boost
                4.3503094 = idf(docFreq=1557, maxDocs=44421)
                0.018026277 = queryNorm
              0.27189434 = fieldWeight in 308, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.3503094 = idf(docFreq=1557, maxDocs=44421)
                0.0625 = fieldNorm(doc=308)
          0.031278614 = weight(abstract_txt:collections in 308) [ClassicSimilarity], result of:
            0.031278614 = score(doc=308,freq=1.0), product of:
              0.10653721 = queryWeight, product of:
                1.2581402 = boost
                4.6974936 = idf(docFreq=1100, maxDocs=44421)
                0.018026277 = queryNorm
              0.29359335 = fieldWeight in 308, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.6974936 = idf(docFreq=1100, maxDocs=44421)
                0.0625 = fieldNorm(doc=308)
          0.040440343 = weight(abstract_txt:model in 308) [ClassicSimilarity], result of:
            0.040440343 = score(doc=308,freq=2.0), product of:
              0.11487705 = queryWeight, product of:
                1.6000762 = boost
                3.9827821 = idf(docFreq=2249, maxDocs=44421)
                0.018026277 = queryNorm
              0.35203153 = fieldWeight in 308, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                3.9827821 = idf(docFreq=2249, maxDocs=44421)
                0.0625 = fieldNorm(doc=308)
          0.2679745 = weight(abstract_txt:latent in 308) [ClassicSimilarity], result of:
            0.2679745 = score(doc=308,freq=3.0), product of:
              0.3540424 = queryWeight, product of:
                2.8089995 = boost
                6.9919376 = idf(docFreq=110, maxDocs=44421)
                0.018026277 = queryNorm
              0.7568994 = fieldWeight in 308, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                6.9919376 = idf(docFreq=110, maxDocs=44421)
                0.0625 = fieldNorm(doc=308)
          0.19864999 = weight(abstract_txt:semantic in 308) [ClassicSimilarity], result of:
            0.19864999 = score(doc=308,freq=6.0), product of:
              0.28999153 = queryWeight, product of:
                3.5952716 = boost
                4.4745317 = idf(docFreq=1375, maxDocs=44421)
                0.018026277 = queryNorm
              0.68501997 = fieldWeight in 308, product of:
                2.4494898 = tf(freq=6.0), with freq of:
                  6.0 = termFreq=6.0
                4.4745317 = idf(docFreq=1375, maxDocs=44421)
                0.0625 = fieldNorm(doc=308)
        0.28 = coord(7/25)
    
  3. Choi, Y.: ¬A complete assessment of tagging quality : a consolidated methodology (2015) 0.19
    0.1932622 = sum of:
      0.1932622 = product of:
        0.69022214 = sum of:
          0.029866926 = weight(abstract_txt:document in 2730) [ClassicSimilarity], result of:
            0.029866926 = score(doc=2730,freq=1.0), product of:
              0.08902732 = queryWeight, product of:
                1.1501118 = boost
                4.29415 = idf(docFreq=1647, maxDocs=44421)
                0.018026277 = queryNorm
              0.33548045 = fieldWeight in 2730, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.29415 = idf(docFreq=1647, maxDocs=44421)
                0.078125 = fieldNorm(doc=2730)
          0.06943915 = weight(abstract_txt:indexing in 2730) [ClassicSimilarity], result of:
            0.06943915 = score(doc=2730,freq=5.0), product of:
              0.09137117 = queryWeight, product of:
                1.1651531 = boost
                4.3503094 = idf(docFreq=1557, maxDocs=44421)
                0.018026277 = queryNorm
              0.7599678 = fieldWeight in 2730, product of:
                2.236068 = tf(freq=5.0), with freq of:
                  5.0 = termFreq=5.0
                4.3503094 = idf(docFreq=1557, maxDocs=44421)
                0.078125 = fieldNorm(doc=2730)
          0.0643638 = weight(abstract_txt:statistical in 2730) [ClassicSimilarity], result of:
            0.0643638 = score(doc=2730,freq=1.0), product of:
              0.14853337 = queryWeight, product of:
                1.4855608 = boost
                5.5466094 = idf(docFreq=470, maxDocs=44421)
                0.018026277 = queryNorm
              0.43332887 = fieldWeight in 2730, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.5466094 = idf(docFreq=470, maxDocs=44421)
                0.078125 = fieldNorm(doc=2730)
          0.03574455 = weight(abstract_txt:model in 2730) [ClassicSimilarity], result of:
            0.03574455 = score(doc=2730,freq=1.0), product of:
              0.11487705 = queryWeight, product of:
                1.6000762 = boost
                3.9827821 = idf(docFreq=2249, maxDocs=44421)
                0.018026277 = queryNorm
              0.31115484 = fieldWeight in 2730, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                3.9827821 = idf(docFreq=2249, maxDocs=44421)
                0.078125 = fieldNorm(doc=2730)
          0.0946675 = weight(abstract_txt:significance in 2730) [ClassicSimilarity], result of:
            0.0946675 = score(doc=2730,freq=1.0), product of:
              0.19210127 = queryWeight, product of:
                1.6894429 = boost
                6.30784 = idf(docFreq=219, maxDocs=44421)
                0.018026277 = queryNorm
              0.4928 = fieldWeight in 2730, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.30784 = idf(docFreq=219, maxDocs=44421)
                0.078125 = fieldNorm(doc=2730)
          0.19339393 = weight(abstract_txt:latent in 2730) [ClassicSimilarity], result of:
            0.19339393 = score(doc=2730,freq=1.0), product of:
              0.3540424 = queryWeight, product of:
                2.8089995 = boost
                6.9919376 = idf(docFreq=110, maxDocs=44421)
                0.018026277 = queryNorm
              0.5462451 = fieldWeight in 2730, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.9919376 = idf(docFreq=110, maxDocs=44421)
                0.078125 = fieldNorm(doc=2730)
          0.20274629 = weight(abstract_txt:semantic in 2730) [ClassicSimilarity], result of:
            0.20274629 = score(doc=2730,freq=4.0), product of:
              0.28999153 = queryWeight, product of:
                3.5952716 = boost
                4.4745317 = idf(docFreq=1375, maxDocs=44421)
                0.018026277 = queryNorm
              0.69914556 = fieldWeight in 2730, product of:
                2.0 = tf(freq=4.0), with freq of:
                  4.0 = termFreq=4.0
                4.4745317 = idf(docFreq=1375, maxDocs=44421)
                0.078125 = fieldNorm(doc=2730)
        0.28 = coord(7/25)
    
  4. He, X.; Cai, D.; Liu, H.; Ma, W.Y.: Locality preserving indexing for document representation (2004) 0.16
    0.16256842 = sum of:
      0.16256842 = product of:
        1.3547369 = sum of:
          0.17566869 = weight(abstract_txt:indexing in 5079) [ClassicSimilarity], result of:
            0.17566869 = score(doc=5079,freq=2.0), product of:
              0.09137117 = queryWeight, product of:
                1.1651531 = boost
                4.3503094 = idf(docFreq=1557, maxDocs=44421)
                0.018026277 = queryNorm
              1.9225833 = fieldWeight in 5079, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                4.3503094 = idf(docFreq=1557, maxDocs=44421)
                0.3125 = fieldNorm(doc=5079)
          0.7735757 = weight(abstract_txt:latent in 5079) [ClassicSimilarity], result of:
            0.7735757 = score(doc=5079,freq=1.0), product of:
              0.3540424 = queryWeight, product of:
                2.8089995 = boost
                6.9919376 = idf(docFreq=110, maxDocs=44421)
                0.018026277 = queryNorm
              2.1849804 = fieldWeight in 5079, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.9919376 = idf(docFreq=110, maxDocs=44421)
                0.3125 = fieldNorm(doc=5079)
          0.40549257 = weight(abstract_txt:semantic in 5079) [ClassicSimilarity], result of:
            0.40549257 = score(doc=5079,freq=1.0), product of:
              0.28999153 = queryWeight, product of:
                3.5952716 = boost
                4.4745317 = idf(docFreq=1375, maxDocs=44421)
                0.018026277 = queryNorm
              1.3982911 = fieldWeight in 5079, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.4745317 = idf(docFreq=1375, maxDocs=44421)
                0.3125 = fieldNorm(doc=5079)
        0.12 = coord(3/25)
    
  5. Story, R.E.: ¬An explanation of the effectiveness of latent semantic indexing by means of a Baysian regression model (1996) 0.16
    0.16060802 = sum of:
      0.16060802 = product of:
        0.6692001 = sum of:
          0.041813698 = weight(abstract_txt:document in 2011) [ClassicSimilarity], result of:
            0.041813698 = score(doc=2011,freq=1.0), product of:
              0.08902732 = queryWeight, product of:
                1.1501118 = boost
                4.29415 = idf(docFreq=1647, maxDocs=44421)
                0.018026277 = queryNorm
              0.46967265 = fieldWeight in 2011, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.29415 = idf(docFreq=1647, maxDocs=44421)
                0.109375 = fieldNorm(doc=2011)
          0.04347578 = weight(abstract_txt:indexing in 2011) [ClassicSimilarity], result of:
            0.04347578 = score(doc=2011,freq=1.0), product of:
              0.09137117 = queryWeight, product of:
                1.1651531 = boost
                4.3503094 = idf(docFreq=1557, maxDocs=44421)
                0.018026277 = queryNorm
              0.4758151 = fieldWeight in 2011, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.3503094 = idf(docFreq=1557, maxDocs=44421)
                0.109375 = fieldNorm(doc=2011)
          0.081127405 = weight(abstract_txt:words in 2011) [ClassicSimilarity], result of:
            0.081127405 = score(doc=2011,freq=1.0), product of:
              0.13849135 = queryWeight, product of:
                1.4344642 = boost
                5.355831 = idf(docFreq=569, maxDocs=44421)
                0.018026277 = queryNorm
              0.58579403 = fieldWeight in 2011, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.355831 = idf(docFreq=569, maxDocs=44421)
                0.109375 = fieldNorm(doc=2011)
          0.09010932 = weight(abstract_txt:statistical in 2011) [ClassicSimilarity], result of:
            0.09010932 = score(doc=2011,freq=1.0), product of:
              0.14853337 = queryWeight, product of:
                1.4855608 = boost
                5.5466094 = idf(docFreq=470, maxDocs=44421)
                0.018026277 = queryNorm
              0.6066604 = fieldWeight in 2011, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.5466094 = idf(docFreq=470, maxDocs=44421)
                0.109375 = fieldNorm(doc=2011)
          0.27075154 = weight(abstract_txt:latent in 2011) [ClassicSimilarity], result of:
            0.27075154 = score(doc=2011,freq=1.0), product of:
              0.3540424 = queryWeight, product of:
                2.8089995 = boost
                6.9919376 = idf(docFreq=110, maxDocs=44421)
                0.018026277 = queryNorm
              0.7647432 = fieldWeight in 2011, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.9919376 = idf(docFreq=110, maxDocs=44421)
                0.109375 = fieldNorm(doc=2011)
          0.1419224 = weight(abstract_txt:semantic in 2011) [ClassicSimilarity], result of:
            0.1419224 = score(doc=2011,freq=1.0), product of:
              0.28999153 = queryWeight, product of:
                3.5952716 = boost
                4.4745317 = idf(docFreq=1375, maxDocs=44421)
                0.018026277 = queryNorm
              0.4894019 = fieldWeight in 2011, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.4745317 = idf(docFreq=1375, maxDocs=44421)
                0.109375 = fieldNorm(doc=2011)
        0.24 = coord(6/25)