Document (#41369)

Author
Zubiaga, A.
Title
¬A longitudinal assessment of the persistence of twitter datasets
Source
Journal of the Association for Information Science and Technology. 69(2018) no.8, S.974-984
Year
2018
Abstract
Social media datasets are not always completely replicable. Having to adhere to requirements of platforms such as Twitter, researchers can only release a list of unique identifiers, which others can then use to recollect the data themselves. This leads to subsets of the data no longer being available, as content can be deleted or user accounts deactivated. To quantify the long-term impact of this in the replicability of datasets, we perform a longitudinal analysis of the persistence of 30 Twitter datasets, which include more than 147 million tweets. By recollecting Twitter datasets ranging from 0 to 4 years old by using the tweet IDs, we look at four different factors quantifying the extent to which recollected datasets resemble original ones: completeness, representativity, similarity, and changingness. Although the ratio of available tweets keeps decreasing as the dataset gets older, we find that the textual content of the recollected subset is still largely representative of the original dataset. The representativity of the metadata, however, keeps fading over time, both because the dataset shrinks and because certain metadata, such as the users' number of followers, keeps changing. Our study has important implications for researchers sharing and using publicly shared Twitter datasets in their research.
Content
Vgl.: https://onlinelibrary.wiley.com/doi/abs/10.1002/asi.24026.
Theme
Informetrie
Object
Twitter

Similar documents (content)

  1. Sedhai, S.; Sun, A.: ¬An analysis of 14 Million tweets on hashtag-oriented spamming* (2017) 0.25
    0.25051188 = sum of:
      0.25051188 = product of:
        0.8946853 = sum of:
          0.015502262 = weight(abstract_txt:content in 4683) [ClassicSimilarity], result of:
            0.015502262 = score(doc=4683,freq=1.0), product of:
              0.059344362 = queryWeight, product of:
                1.0731467 = boost
                4.1796083 = idf(docFreq=1847, maxDocs=44421)
                0.01323076 = queryNorm
              0.26122552 = fieldWeight in 4683, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.1796083 = idf(docFreq=1847, maxDocs=44421)
                0.0625 = fieldNorm(doc=4683)
          0.11385697 = weight(abstract_txt:tweet in 4683) [ClassicSimilarity], result of:
            0.11385697 = score(doc=4683,freq=3.0), product of:
              0.12339723 = queryWeight, product of:
                1.0942261 = boost
                8.523414 = idf(docFreq=23, maxDocs=44421)
                0.01323076 = queryNorm
              0.9226866 = fieldWeight in 4683, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                8.523414 = idf(docFreq=23, maxDocs=44421)
                0.0625 = fieldNorm(doc=4683)
          0.01114215 = weight(abstract_txt:which in 4683) [ClassicSimilarity], result of:
            0.01114215 = score(doc=4683,freq=2.0), product of:
              0.043262918 = queryWeight, product of:
                1.1222068 = boost
                2.9137893 = idf(docFreq=6552, maxDocs=44421)
                0.01323076 = queryNorm
              0.25754502 = fieldWeight in 4683, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                2.9137893 = idf(docFreq=6552, maxDocs=44421)
                0.0625 = fieldNorm(doc=4683)
          0.024663394 = weight(abstract_txt:metadata in 4683) [ClassicSimilarity], result of:
            0.024663394 = score(doc=4683,freq=1.0), product of:
              0.08087569 = queryWeight, product of:
                1.2527902 = boost
                4.87927 = idf(docFreq=917, maxDocs=44421)
                0.01323076 = queryNorm
              0.30495438 = fieldWeight in 4683, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.87927 = idf(docFreq=917, maxDocs=44421)
                0.0625 = fieldNorm(doc=4683)
          0.24410154 = weight(abstract_txt:tweets in 4683) [ClassicSimilarity], result of:
            0.24410154 = score(doc=4683,freq=7.0), product of:
              0.19489339 = queryWeight, product of:
                1.9447685 = boost
                7.574333 = idf(docFreq=61, maxDocs=44421)
                0.01323076 = queryNorm
              1.2524875 = fieldWeight in 4683, product of:
                2.6457512 = tf(freq=7.0), with freq of:
                  7.0 = termFreq=7.0
                7.574333 = idf(docFreq=61, maxDocs=44421)
                0.0625 = fieldNorm(doc=4683)
          0.09454927 = weight(abstract_txt:dataset in 4683) [ClassicSimilarity], result of:
            0.09454927 = score(doc=4683,freq=1.0), product of:
              0.22676983 = queryWeight, product of:
                2.5692575 = boost
                6.6710296 = idf(docFreq=152, maxDocs=44421)
                0.01323076 = queryNorm
              0.41693935 = fieldWeight in 4683, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.6710296 = idf(docFreq=152, maxDocs=44421)
                0.0625 = fieldNorm(doc=4683)
          0.3908698 = weight(abstract_txt:twitter in 4683) [ClassicSimilarity], result of:
            0.3908698 = score(doc=4683,freq=5.0), product of:
              0.40500543 = queryWeight, product of:
                4.432715 = boost
                6.905677 = idf(docFreq=120, maxDocs=44421)
                0.01323076 = queryNorm
              0.96509767 = fieldWeight in 4683, product of:
                2.236068 = tf(freq=5.0), with freq of:
                  5.0 = termFreq=5.0
                6.905677 = idf(docFreq=120, maxDocs=44421)
                0.0625 = fieldNorm(doc=4683)
        0.28 = coord(7/25)
    
  2. Saif, H.; He, Y.; Fernandez, M.; Alani, H.: Contextual semantics for sentiment analysis of Twitter (2016) 0.21
    0.21251878 = sum of:
      0.21251878 = product of:
        0.88549495 = sum of:
          0.09296383 = weight(abstract_txt:tweet in 3667) [ClassicSimilarity], result of:
            0.09296383 = score(doc=3667,freq=2.0), product of:
              0.12339723 = queryWeight, product of:
                1.0942261 = boost
                8.523414 = idf(docFreq=23, maxDocs=44421)
                0.01323076 = queryNorm
              0.75337046 = fieldWeight in 3667, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                8.523414 = idf(docFreq=23, maxDocs=44421)
                0.0625 = fieldNorm(doc=3667)
          0.007878689 = weight(abstract_txt:which in 3667) [ClassicSimilarity], result of:
            0.007878689 = score(doc=3667,freq=1.0), product of:
              0.043262918 = queryWeight, product of:
                1.1222068 = boost
                2.9137893 = idf(docFreq=6552, maxDocs=44421)
                0.01323076 = queryNorm
              0.18211183 = fieldWeight in 3667, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                2.9137893 = idf(docFreq=6552, maxDocs=44421)
                0.0625 = fieldNorm(doc=3667)
          0.09226172 = weight(abstract_txt:tweets in 3667) [ClassicSimilarity], result of:
            0.09226172 = score(doc=3667,freq=1.0), product of:
              0.19489339 = queryWeight, product of:
                1.9447685 = boost
                7.574333 = idf(docFreq=61, maxDocs=44421)
                0.01323076 = queryNorm
              0.47339582 = fieldWeight in 3667, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                7.574333 = idf(docFreq=61, maxDocs=44421)
                0.0625 = fieldNorm(doc=3667)
          0.09454927 = weight(abstract_txt:dataset in 3667) [ClassicSimilarity], result of:
            0.09454927 = score(doc=3667,freq=1.0), product of:
              0.22676983 = queryWeight, product of:
                2.5692575 = boost
                6.6710296 = idf(docFreq=152, maxDocs=44421)
                0.01323076 = queryNorm
              0.41693935 = fieldWeight in 3667, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.6710296 = idf(docFreq=152, maxDocs=44421)
                0.0625 = fieldNorm(doc=3667)
          0.30276644 = weight(abstract_txt:twitter in 3667) [ClassicSimilarity], result of:
            0.30276644 = score(doc=3667,freq=3.0), product of:
              0.40500543 = queryWeight, product of:
                4.432715 = boost
                6.905677 = idf(docFreq=120, maxDocs=44421)
                0.01323076 = queryNorm
              0.74756145 = fieldWeight in 3667, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                6.905677 = idf(docFreq=120, maxDocs=44421)
                0.0625 = fieldNorm(doc=3667)
          0.29507497 = weight(abstract_txt:datasets in 3667) [ClassicSimilarity], result of:
            0.29507497 = score(doc=3667,freq=2.0), product of:
              0.50982016 = queryWeight, product of:
                5.8845315 = boost
                6.548176 = idf(docFreq=172, maxDocs=44421)
                0.01323076 = queryNorm
              0.57878244 = fieldWeight in 3667, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                6.548176 = idf(docFreq=172, maxDocs=44421)
                0.0625 = fieldNorm(doc=3667)
        0.24 = coord(6/25)
    
  3. Fang, Z.; Dudek, J.; Costas, R.: ¬The stability of Twitter metrics : a study on unavailable Twitter mentions of scientific publications (2020) 0.16
    0.16084242 = sum of:
      0.16084242 = product of:
        0.8042121 = sum of:
          0.065735355 = weight(abstract_txt:tweet in 1036) [ClassicSimilarity], result of:
            0.065735355 = score(doc=1036,freq=1.0), product of:
              0.12339723 = queryWeight, product of:
                1.0942261 = boost
                8.523414 = idf(docFreq=23, maxDocs=44421)
                0.01323076 = queryNorm
              0.53271335 = fieldWeight in 1036, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                8.523414 = idf(docFreq=23, maxDocs=44421)
                0.0625 = fieldNorm(doc=1036)
          0.007878689 = weight(abstract_txt:which in 1036) [ClassicSimilarity], result of:
            0.007878689 = score(doc=1036,freq=1.0), product of:
              0.043262918 = queryWeight, product of:
                1.1222068 = boost
                2.9137893 = idf(docFreq=6552, maxDocs=44421)
                0.01323076 = queryNorm
              0.18211183 = fieldWeight in 1036, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                2.9137893 = idf(docFreq=6552, maxDocs=44421)
                0.0625 = fieldNorm(doc=1036)
          0.04734695 = weight(abstract_txt:original in 1036) [ClassicSimilarity], result of:
            0.04734695 = score(doc=1036,freq=2.0), product of:
              0.099151835 = queryWeight, product of:
                1.3871382 = boost
                5.4025183 = idf(docFreq=543, maxDocs=44421)
                0.01323076 = queryNorm
              0.47751966 = fieldWeight in 1036, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                5.4025183 = idf(docFreq=543, maxDocs=44421)
                0.0625 = fieldNorm(doc=1036)
          0.13047777 = weight(abstract_txt:tweets in 1036) [ClassicSimilarity], result of:
            0.13047777 = score(doc=1036,freq=2.0), product of:
              0.19489339 = queryWeight, product of:
                1.9447685 = boost
                7.574333 = idf(docFreq=61, maxDocs=44421)
                0.01323076 = queryNorm
              0.66948277 = fieldWeight in 1036, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                7.574333 = idf(docFreq=61, maxDocs=44421)
                0.0625 = fieldNorm(doc=1036)
          0.55277336 = weight(abstract_txt:twitter in 1036) [ClassicSimilarity], result of:
            0.55277336 = score(doc=1036,freq=10.0), product of:
              0.40500543 = queryWeight, product of:
                4.432715 = boost
                6.905677 = idf(docFreq=120, maxDocs=44421)
                0.01323076 = queryNorm
              1.3648542 = fieldWeight in 1036, product of:
                3.1622777 = tf(freq=10.0), with freq of:
                  10.0 = termFreq=10.0
                6.905677 = idf(docFreq=120, maxDocs=44421)
                0.0625 = fieldNorm(doc=1036)
        0.2 = coord(5/25)
    
  4. Ortega, J.L.: ¬The presence of academic journals on Twitter and its relationship with dissemination (tweets) and research impact (citations) (2017) 0.15
    0.1509544 = sum of:
      0.1509544 = product of:
        0.75477195 = sum of:
          0.015502262 = weight(abstract_txt:content in 410) [ClassicSimilarity], result of:
            0.015502262 = score(doc=410,freq=1.0), product of:
              0.059344362 = queryWeight, product of:
                1.0731467 = boost
                4.1796083 = idf(docFreq=1847, maxDocs=44421)
                0.01323076 = queryNorm
              0.26122552 = fieldWeight in 410, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.1796083 = idf(docFreq=1847, maxDocs=44421)
                0.0625 = fieldNorm(doc=410)
          0.007878689 = weight(abstract_txt:which in 410) [ClassicSimilarity], result of:
            0.007878689 = score(doc=410,freq=1.0), product of:
              0.043262918 = queryWeight, product of:
                1.1222068 = boost
                2.9137893 = idf(docFreq=6552, maxDocs=44421)
                0.01323076 = queryNorm
              0.18211183 = fieldWeight in 410, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                2.9137893 = idf(docFreq=6552, maxDocs=44421)
                0.0625 = fieldNorm(doc=410)
          0.0772205 = weight(abstract_txt:followers in 410) [ClassicSimilarity], result of:
            0.0772205 = score(doc=410,freq=1.0), product of:
              0.13738136 = queryWeight, product of:
                1.1545647 = boost
                8.993418 = idf(docFreq=14, maxDocs=44421)
                0.01323076 = queryNorm
              0.5620886 = fieldWeight in 410, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                8.993418 = idf(docFreq=14, maxDocs=44421)
                0.0625 = fieldNorm(doc=410)
          0.22599413 = weight(abstract_txt:tweets in 410) [ClassicSimilarity], result of:
            0.22599413 = score(doc=410,freq=6.0), product of:
              0.19489339 = queryWeight, product of:
                1.9447685 = boost
                7.574333 = idf(docFreq=61, maxDocs=44421)
                0.01323076 = queryNorm
              1.1595782 = fieldWeight in 410, product of:
                2.4494898 = tf(freq=6.0), with freq of:
                  6.0 = termFreq=6.0
                7.574333 = idf(docFreq=61, maxDocs=44421)
                0.0625 = fieldNorm(doc=410)
          0.4281764 = weight(abstract_txt:twitter in 410) [ClassicSimilarity], result of:
            0.4281764 = score(doc=410,freq=6.0), product of:
              0.40500543 = queryWeight, product of:
                4.432715 = boost
                6.905677 = idf(docFreq=120, maxDocs=44421)
                0.01323076 = queryNorm
              1.0572115 = fieldWeight in 410, product of:
                2.4494898 = tf(freq=6.0), with freq of:
                  6.0 = termFreq=6.0
                6.905677 = idf(docFreq=120, maxDocs=44421)
                0.0625 = fieldNorm(doc=410)
        0.2 = coord(5/25)
    
  5. Yu, M.; Sun, A.: Dataset versus reality : understanding model performance from the perspective of information need (2023) 0.14
    0.1371326 = sum of:
      0.1371326 = product of:
        0.685663 = sum of:
          0.014817342 = weight(abstract_txt:available in 2075) [ClassicSimilarity], result of:
            0.014817342 = score(doc=2075,freq=1.0), product of:
              0.06294447 = queryWeight, product of:
                1.1052185 = boost
                4.304519 = idf(docFreq=1630, maxDocs=44421)
                0.01323076 = queryNorm
              0.23540339 = fieldWeight in 2075, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.304519 = idf(docFreq=1630, maxDocs=44421)
                0.0546875 = fieldNorm(doc=2075)
          0.009749381 = weight(abstract_txt:which in 2075) [ClassicSimilarity], result of:
            0.009749381 = score(doc=2075,freq=2.0), product of:
              0.043262918 = queryWeight, product of:
                1.1222068 = boost
                2.9137893 = idf(docFreq=6552, maxDocs=44421)
                0.01323076 = queryNorm
              0.2253519 = fieldWeight in 2075, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                2.9137893 = idf(docFreq=6552, maxDocs=44421)
                0.0546875 = fieldNorm(doc=2075)
          0.028905736 = weight(abstract_txt:researchers in 2075) [ClassicSimilarity], result of:
            0.028905736 = score(doc=2075,freq=2.0), product of:
              0.07799919 = queryWeight, product of:
                1.2303096 = boost
                4.791714 = idf(docFreq=1001, maxDocs=44421)
                0.01323076 = queryNorm
              0.3705902 = fieldWeight in 2075, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                4.791714 = idf(docFreq=1001, maxDocs=44421)
                0.0546875 = fieldNorm(doc=2075)
          0.18499127 = weight(abstract_txt:dataset in 2075) [ClassicSimilarity], result of:
            0.18499127 = score(doc=2075,freq=5.0), product of:
              0.22676983 = queryWeight, product of:
                2.5692575 = boost
                6.6710296 = idf(docFreq=152, maxDocs=44421)
                0.01323076 = queryNorm
              0.81576663 = fieldWeight in 2075, product of:
                2.236068 = tf(freq=5.0), with freq of:
                  5.0 = termFreq=5.0
                6.6710296 = idf(docFreq=152, maxDocs=44421)
                0.0546875 = fieldNorm(doc=2075)
          0.44719923 = weight(abstract_txt:datasets in 2075) [ClassicSimilarity], result of:
            0.44719923 = score(doc=2075,freq=6.0), product of:
              0.50982016 = queryWeight, product of:
                5.8845315 = boost
                6.548176 = idf(docFreq=172, maxDocs=44421)
                0.01323076 = queryNorm
              0.87717056 = fieldWeight in 2075, product of:
                2.4494898 = tf(freq=6.0), with freq of:
                  6.0 = termFreq=6.0
                6.548176 = idf(docFreq=172, maxDocs=44421)
                0.0546875 = fieldNorm(doc=2075)
        0.2 = coord(5/25)