Document (#41369)

Author
Zubiaga, A.
Title
¬A longitudinal assessment of the persistence of twitter datasets
Source
Journal of the Association for Information Science and Technology. 69(2018) no.8, S.974-984
Year
2018
Abstract
Social media datasets are not always completely replicable. Having to adhere to requirements of platforms such as Twitter, researchers can only release a list of unique identifiers, which others can then use to recollect the data themselves. This leads to subsets of the data no longer being available, as content can be deleted or user accounts deactivated. To quantify the long-term impact of this in the replicability of datasets, we perform a longitudinal analysis of the persistence of 30 Twitter datasets, which include more than 147 million tweets. By recollecting Twitter datasets ranging from 0 to 4 years old by using the tweet IDs, we look at four different factors quantifying the extent to which recollected datasets resemble original ones: completeness, representativity, similarity, and changingness. Although the ratio of available tweets keeps decreasing as the dataset gets older, we find that the textual content of the recollected subset is still largely representative of the original dataset. The representativity of the metadata, however, keeps fading over time, both because the dataset shrinks and because certain metadata, such as the users' number of followers, keeps changing. Our study has important implications for researchers sharing and using publicly shared Twitter datasets in their research.
Content
Vgl.: https://onlinelibrary.wiley.com/doi/abs/10.1002/asi.24026.
Theme
Informetrie
Object
Twitter

Similar documents (content)

  1. Sedhai, S.; Sun, A.: ¬An analysis of 14 Million tweets on hashtag-oriented spamming* (2017) 0.25
    0.25116807 = sum of:
      0.25116807 = product of:
        0.8970288 = sum of:
          0.015441311 = weight(abstract_txt:content in 3683) [ClassicSimilarity], result of:
            0.015441311 = score(doc=3683,freq=1.0), product of:
              0.059106763 = queryWeight, product of:
                1.0738556 = boost
                4.17991 = idf(docFreq=1838, maxDocs=44218)
                0.013168137 = queryNorm
              0.2612444 = fieldWeight in 3683, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.17991 = idf(docFreq=1838, maxDocs=44218)
                0.0625 = fieldNorm(doc=3683)
          0.11320205 = weight(abstract_txt:tweet in 3683) [ClassicSimilarity], result of:
            0.11320205 = score(doc=3683,freq=3.0), product of:
              0.122753404 = queryWeight, product of:
                1.0942816 = boost
                8.518833 = idf(docFreq=23, maxDocs=44218)
                0.013168137 = queryNorm
              0.9221907 = fieldWeight in 3683, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                8.518833 = idf(docFreq=23, maxDocs=44218)
                0.0625 = fieldNorm(doc=3683)
          0.0111293895 = weight(abstract_txt:which in 3683) [ClassicSimilarity], result of:
            0.0111293895 = score(doc=3683,freq=2.0), product of:
              0.043170035 = queryWeight, product of:
                1.1239944 = boost
                2.9167147 = idf(docFreq=6503, maxDocs=44218)
                0.013168137 = queryNorm
              0.2578036 = fieldWeight in 3683, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                2.9167147 = idf(docFreq=6503, maxDocs=44218)
                0.0625 = fieldNorm(doc=3683)
          0.024590971 = weight(abstract_txt:metadata in 3683) [ClassicSimilarity], result of:
            0.024590971 = score(doc=3683,freq=1.0), product of:
              0.08060554 = queryWeight, product of:
                1.2540352 = boost
                4.881247 = idf(docFreq=911, maxDocs=44218)
                0.013168137 = queryNorm
              0.30507794 = fieldWeight in 3683, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.881247 = idf(docFreq=911, maxDocs=44218)
                0.0625 = fieldNorm(doc=3683)
          0.24421541 = weight(abstract_txt:tweets in 3683) [ClassicSimilarity], result of:
            0.24421541 = score(doc=3683,freq=7.0), product of:
              0.19468409 = queryWeight, product of:
                1.9489135 = boost
                7.5860133 = idf(docFreq=60, maxDocs=44218)
                0.013168137 = queryNorm
              1.254419 = fieldWeight in 3683, product of:
                2.6457512 = tf(freq=7.0), with freq of:
                  7.0 = termFreq=7.0
                7.5860133 = idf(docFreq=60, maxDocs=44218)
                0.0625 = fieldNorm(doc=3683)
          0.09715366 = weight(abstract_txt:dataset in 3683) [ClassicSimilarity], result of:
            0.09715366 = score(doc=3683,freq=1.0), product of:
              0.23059557 = queryWeight, product of:
                2.5977564 = boost
                6.7410603 = idf(docFreq=141, maxDocs=44218)
                0.013168137 = queryNorm
              0.42131627 = fieldWeight in 3683, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.7410603 = idf(docFreq=141, maxDocs=44218)
                0.0625 = fieldNorm(doc=3683)
          0.39129603 = weight(abstract_txt:twitter in 3683) [ClassicSimilarity], result of:
            0.39129603 = score(doc=3683,freq=5.0), product of:
              0.40473866 = queryWeight, product of:
                4.443085 = boost
                6.9177637 = idf(docFreq=118, maxDocs=44218)
                0.013168137 = queryNorm
              0.96678686 = fieldWeight in 3683, product of:
                2.236068 = tf(freq=5.0), with freq of:
                  5.0 = termFreq=5.0
                6.9177637 = idf(docFreq=118, maxDocs=44218)
                0.0625 = fieldNorm(doc=3683)
        0.28 = coord(7/25)
    
  2. Saif, H.; He, Y.; Fernandez, M.; Alani, H.: Contextual semantics for sentiment analysis of Twitter (2016) 0.21
    0.21380632 = sum of:
      0.21380632 = product of:
        0.89085966 = sum of:
          0.09242909 = weight(abstract_txt:tweet in 2667) [ClassicSimilarity], result of:
            0.09242909 = score(doc=2667,freq=2.0), product of:
              0.122753404 = queryWeight, product of:
                1.0942816 = boost
                8.518833 = idf(docFreq=23, maxDocs=44218)
                0.013168137 = queryNorm
              0.75296557 = fieldWeight in 2667, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                8.518833 = idf(docFreq=23, maxDocs=44218)
                0.0625 = fieldNorm(doc=2667)
          0.007869667 = weight(abstract_txt:which in 2667) [ClassicSimilarity], result of:
            0.007869667 = score(doc=2667,freq=1.0), product of:
              0.043170035 = queryWeight, product of:
                1.1239944 = boost
                2.9167147 = idf(docFreq=6503, maxDocs=44218)
                0.013168137 = queryNorm
              0.18229467 = fieldWeight in 2667, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                2.9167147 = idf(docFreq=6503, maxDocs=44218)
                0.0625 = fieldNorm(doc=2667)
          0.09230476 = weight(abstract_txt:tweets in 2667) [ClassicSimilarity], result of:
            0.09230476 = score(doc=2667,freq=1.0), product of:
              0.19468409 = queryWeight, product of:
                1.9489135 = boost
                7.5860133 = idf(docFreq=60, maxDocs=44218)
                0.013168137 = queryNorm
              0.47412583 = fieldWeight in 2667, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                7.5860133 = idf(docFreq=60, maxDocs=44218)
                0.0625 = fieldNorm(doc=2667)
          0.09715366 = weight(abstract_txt:dataset in 2667) [ClassicSimilarity], result of:
            0.09715366 = score(doc=2667,freq=1.0), product of:
              0.23059557 = queryWeight, product of:
                2.5977564 = boost
                6.7410603 = idf(docFreq=141, maxDocs=44218)
                0.013168137 = queryNorm
              0.42131627 = fieldWeight in 2667, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.7410603 = idf(docFreq=141, maxDocs=44218)
                0.0625 = fieldNorm(doc=2667)
          0.3030966 = weight(abstract_txt:twitter in 2667) [ClassicSimilarity], result of:
            0.3030966 = score(doc=2667,freq=3.0), product of:
              0.40473866 = queryWeight, product of:
                4.443085 = boost
                6.9177637 = idf(docFreq=118, maxDocs=44218)
                0.013168137 = queryNorm
              0.7488699 = fieldWeight in 2667, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                6.9177637 = idf(docFreq=118, maxDocs=44218)
                0.0625 = fieldNorm(doc=2667)
          0.29800588 = weight(abstract_txt:datasets in 2667) [ClassicSimilarity], result of:
            0.29800588 = score(doc=2667,freq=2.0), product of:
              0.5124801 = queryWeight, product of:
                5.915614 = boost
                6.578893 = idf(docFreq=166, maxDocs=44218)
                0.013168137 = queryNorm
              0.5814975 = fieldWeight in 2667, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                6.578893 = idf(docFreq=166, maxDocs=44218)
                0.0625 = fieldNorm(doc=2667)
        0.24 = coord(6/25)
    
  3. Fang, Z.; Dudek, J.; Costas, R.: ¬The stability of Twitter metrics : a study on unavailable Twitter mentions of scientific publications (2020) 0.16
    0.16085377 = sum of:
      0.16085377 = product of:
        0.80426884 = sum of:
          0.06535724 = weight(abstract_txt:tweet in 35) [ClassicSimilarity], result of:
            0.06535724 = score(doc=35,freq=1.0), product of:
              0.122753404 = queryWeight, product of:
                1.0942816 = boost
                8.518833 = idf(docFreq=23, maxDocs=44218)
                0.013168137 = queryNorm
              0.5324271 = fieldWeight in 35, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                8.518833 = idf(docFreq=23, maxDocs=44218)
                0.0625 = fieldNorm(doc=35)
          0.007869667 = weight(abstract_txt:which in 35) [ClassicSimilarity], result of:
            0.007869667 = score(doc=35,freq=1.0), product of:
              0.043170035 = queryWeight, product of:
                1.1239944 = boost
                2.9167147 = idf(docFreq=6503, maxDocs=44218)
                0.013168137 = queryNorm
              0.18229467 = fieldWeight in 35, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                2.9167147 = idf(docFreq=6503, maxDocs=44218)
                0.0625 = fieldNorm(doc=35)
          0.047127098 = weight(abstract_txt:original in 35) [ClassicSimilarity], result of:
            0.047127098 = score(doc=35,freq=2.0), product of:
              0.09870782 = queryWeight, product of:
                1.3877239 = boost
                5.4016213 = idf(docFreq=541, maxDocs=44218)
                0.013168137 = queryNorm
              0.4774404 = fieldWeight in 35, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                5.4016213 = idf(docFreq=541, maxDocs=44218)
                0.0625 = fieldNorm(doc=35)
          0.13053864 = weight(abstract_txt:tweets in 35) [ClassicSimilarity], result of:
            0.13053864 = score(doc=35,freq=2.0), product of:
              0.19468409 = queryWeight, product of:
                1.9489135 = boost
                7.5860133 = idf(docFreq=60, maxDocs=44218)
                0.013168137 = queryNorm
              0.6705152 = fieldWeight in 35, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                7.5860133 = idf(docFreq=60, maxDocs=44218)
                0.0625 = fieldNorm(doc=35)
          0.5533762 = weight(abstract_txt:twitter in 35) [ClassicSimilarity], result of:
            0.5533762 = score(doc=35,freq=10.0), product of:
              0.40473866 = queryWeight, product of:
                4.443085 = boost
                6.9177637 = idf(docFreq=118, maxDocs=44218)
                0.013168137 = queryNorm
              1.3672432 = fieldWeight in 35, product of:
                3.1622777 = tf(freq=10.0), with freq of:
                  10.0 = termFreq=10.0
                6.9177637 = idf(docFreq=118, maxDocs=44218)
                0.0625 = fieldNorm(doc=35)
        0.2 = coord(5/25)
    
  4. Ortega, J.L.: ¬The presence of academic journals on Twitter and its relationship with dissemination (tweets) and research impact (citations) (2017) 0.15
    0.15096734 = sum of:
      0.15096734 = product of:
        0.7548367 = sum of:
          0.015441311 = weight(abstract_txt:content in 4410) [ClassicSimilarity], result of:
            0.015441311 = score(doc=4410,freq=1.0), product of:
              0.059106763 = queryWeight, product of:
                1.0738556 = boost
                4.17991 = idf(docFreq=1838, maxDocs=44218)
                0.013168137 = queryNorm
              0.2612444 = fieldWeight in 4410, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.17991 = idf(docFreq=1838, maxDocs=44218)
                0.0625 = fieldNorm(doc=4410)
          0.007869667 = weight(abstract_txt:which in 4410) [ClassicSimilarity], result of:
            0.007869667 = score(doc=4410,freq=1.0), product of:
              0.043170035 = queryWeight, product of:
                1.1239944 = boost
                2.9167147 = idf(docFreq=6503, maxDocs=44218)
                0.013168137 = queryNorm
              0.18229467 = fieldWeight in 4410, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                2.9167147 = idf(docFreq=6503, maxDocs=44218)
                0.0625 = fieldNorm(doc=4410)
          0.076782785 = weight(abstract_txt:followers in 4410) [ClassicSimilarity], result of:
            0.076782785 = score(doc=4410,freq=1.0), product of:
              0.13667224 = queryWeight, product of:
                1.1546556 = boost
                8.988837 = idf(docFreq=14, maxDocs=44218)
                0.013168137 = queryNorm
              0.5618023 = fieldWeight in 4410, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                8.988837 = idf(docFreq=14, maxDocs=44218)
                0.0625 = fieldNorm(doc=4410)
          0.22609957 = weight(abstract_txt:tweets in 4410) [ClassicSimilarity], result of:
            0.22609957 = score(doc=4410,freq=6.0), product of:
              0.19468409 = queryWeight, product of:
                1.9489135 = boost
                7.5860133 = idf(docFreq=60, maxDocs=44218)
                0.013168137 = queryNorm
              1.1613665 = fieldWeight in 4410, product of:
                2.4494898 = tf(freq=6.0), with freq of:
                  6.0 = termFreq=6.0
                7.5860133 = idf(docFreq=60, maxDocs=44218)
                0.0625 = fieldNorm(doc=4410)
          0.42864335 = weight(abstract_txt:twitter in 4410) [ClassicSimilarity], result of:
            0.42864335 = score(doc=4410,freq=6.0), product of:
              0.40473866 = queryWeight, product of:
                4.443085 = boost
                6.9177637 = idf(docFreq=118, maxDocs=44218)
                0.013168137 = queryNorm
              1.059062 = fieldWeight in 4410, product of:
                2.4494898 = tf(freq=6.0), with freq of:
                  6.0 = termFreq=6.0
                6.9177637 = idf(docFreq=118, maxDocs=44218)
                0.0625 = fieldNorm(doc=4410)
        0.2 = coord(5/25)
    
  5. Yu, M.; Sun, A.: Dataset versus reality : understanding model performance from the perspective of information need (2023) 0.14
    0.13904369 = sum of:
      0.13904369 = product of:
        0.69521844 = sum of:
          0.014721422 = weight(abstract_txt:available in 1073) [ClassicSimilarity], result of:
            0.014721422 = score(doc=1073,freq=1.0), product of:
              0.06258576 = queryWeight, product of:
                1.1050072 = boost
                4.3011656 = idf(docFreq=1628, maxDocs=44218)
                0.013168137 = queryNorm
              0.23521999 = fieldWeight in 1073, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.3011656 = idf(docFreq=1628, maxDocs=44218)
                0.0546875 = fieldNorm(doc=1073)
          0.009738216 = weight(abstract_txt:which in 1073) [ClassicSimilarity], result of:
            0.009738216 = score(doc=1073,freq=2.0), product of:
              0.043170035 = queryWeight, product of:
                1.1239944 = boost
                2.9167147 = idf(docFreq=6503, maxDocs=44218)
                0.013168137 = queryNorm
              0.22557814 = fieldWeight in 1073, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                2.9167147 = idf(docFreq=6503, maxDocs=44218)
                0.0546875 = fieldNorm(doc=1073)
          0.029030686 = weight(abstract_txt:researchers in 1073) [ClassicSimilarity], result of:
            0.029030686 = score(doc=1073,freq=2.0), product of:
              0.07811551 = queryWeight, product of:
                1.2345138 = boost
                4.805261 = idf(docFreq=983, maxDocs=44218)
                0.013168137 = queryNorm
              0.37163794 = fieldWeight in 1073, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                4.805261 = idf(docFreq=983, maxDocs=44218)
                0.0546875 = fieldNorm(doc=1073)
          0.19008693 = weight(abstract_txt:dataset in 1073) [ClassicSimilarity], result of:
            0.19008693 = score(doc=1073,freq=5.0), product of:
              0.23059557 = queryWeight, product of:
                2.5977564 = boost
                6.7410603 = idf(docFreq=141, maxDocs=44218)
                0.013168137 = queryNorm
              0.82433033 = fieldWeight in 1073, product of:
                2.236068 = tf(freq=5.0), with freq of:
                  5.0 = termFreq=5.0
                6.7410603 = idf(docFreq=141, maxDocs=44218)
                0.0546875 = fieldNorm(doc=1073)
          0.45164117 = weight(abstract_txt:datasets in 1073) [ClassicSimilarity], result of:
            0.45164117 = score(doc=1073,freq=6.0), product of:
              0.5124801 = queryWeight, product of:
                5.915614 = boost
                6.578893 = idf(docFreq=166, maxDocs=44218)
                0.013168137 = queryNorm
              0.8812853 = fieldWeight in 1073, product of:
                2.4494898 = tf(freq=6.0), with freq of:
                  6.0 = termFreq=6.0
                6.578893 = idf(docFreq=166, maxDocs=44218)
                0.0546875 = fieldNorm(doc=1073)
        0.2 = coord(5/25)