Document (#35472)

Author
Fu, T.
Abbasi, A.
Chen, H.
Title
¬A focused crawler for Dark Web forums
Source
Journal of the American Society for Information Science and Technology. 61(2010) no.6, S.1213-1231
Year
2010
Abstract
The unprecedented growth of the Internet has given rise to the Dark Web, the problematic facet of the Web associated with cybercrime, hate, and extremism. Despite the need for tools to collect and analyze Dark Web forums, the covert nature of this part of the Internet makes traditional Web crawling techniques insufficient for capturing such content. In this study, we propose a novel crawling system designed to collect Dark Web forum content. The system uses a human-assisted accessibility approach to gain access to Dark Web forums. Several URL ordering features and techniques enable efficient extraction of forum postings. The system also includes an incremental crawler coupled with a recall-improvement mechanism intended to facilitate enhanced retrieval and updating of collected content. Experiments conducted to evaluate the effectiveness of the human-assisted accessibility approach and the recall-improvement-based, incremental-update procedure yielded favorable results. The human-assisted approach significantly improved access to Dark Web forums while the incremental crawler with recall improvement also outperformed standard periodic- and incremental-update approaches. Using the system, we were able to collect over 100 Dark Web forums from three regions. A case study encompassing link and content analysis of collected forums was used to illustrate the value and importance of gathering and analyzing content from such online communities.
Theme
Internet
Suchmaschinen

Similar documents (author)

  1. Chen, Y.N.; Chen, S.J.: ¬A metadata practice of the OFLA FRBR model : a case study for the National Palace Museum in Taipai (2004) 4.34
    4.3394766 = sum of:
      4.3394766 = weight(author_txt:chen in 4384) [ClassicSimilarity], result of:
        4.3394766 = score(doc=4384,freq=2.0), product of:
          0.99999994 = queryWeight, product of:
            6.136947 = idf(docFreq=260, maxDocs=44421)
            0.16294746 = queryNorm
          4.339477 = fieldWeight in 4384, product of:
            1.4142135 = tf(freq=2.0), with freq of:
              2.0 = termFreq=2.0
            6.136947 = idf(docFreq=260, maxDocs=44421)
            0.5 = fieldNorm(doc=4384)
    
  2. Chen, C.C.; Chen, H.H.; Chen, K.H.: ¬The design of the XML/Metadata management system (2000) 3.99
    3.9860637 = sum of:
      3.9860637 = weight(author_txt:chen in 5633) [ClassicSimilarity], result of:
        3.9860637 = score(doc=5633,freq=3.0), product of:
          0.99999994 = queryWeight, product of:
            6.136947 = idf(docFreq=260, maxDocs=44421)
            0.16294746 = queryNorm
          3.986064 = fieldWeight in 5633, product of:
            1.7320508 = tf(freq=3.0), with freq of:
              3.0 = termFreq=3.0
            6.136947 = idf(docFreq=260, maxDocs=44421)
            0.375 = fieldNorm(doc=5633)
    
  3. Chen, W.Y.: Observations on cataloguing and classification (1991) 3.84
    3.8355918 = sum of:
      3.8355918 = weight(author_txt:chen in 4183) [ClassicSimilarity], result of:
        3.8355918 = score(doc=4183,freq=1.0), product of:
          0.99999994 = queryWeight, product of:
            6.136947 = idf(docFreq=260, maxDocs=44421)
            0.16294746 = queryNorm
          3.835592 = fieldWeight in 4183, product of:
            1.0 = tf(freq=1.0), with freq of:
              1.0 = termFreq=1.0
            6.136947 = idf(docFreq=260, maxDocs=44421)
            0.625 = fieldNorm(doc=4183)
    
  4. Chen, H.: Knowledge-based document retrieval : framework and design (1992) 3.84
    3.8355918 = sum of:
      3.8355918 = weight(author_txt:chen in 5282) [ClassicSimilarity], result of:
        3.8355918 = score(doc=5282,freq=1.0), product of:
          0.99999994 = queryWeight, product of:
            6.136947 = idf(docFreq=260, maxDocs=44421)
            0.16294746 = queryNorm
          3.835592 = fieldWeight in 5282, product of:
            1.0 = tf(freq=1.0), with freq of:
              1.0 = termFreq=1.0
            6.136947 = idf(docFreq=260, maxDocs=44421)
            0.625 = fieldNorm(doc=5282)
    
  5. Chen, P.S.: On inference rules of logic-based information retrieval systems (1994) 3.84
    3.8355918 = sum of:
      3.8355918 = weight(author_txt:chen in 6730) [ClassicSimilarity], result of:
        3.8355918 = score(doc=6730,freq=1.0), product of:
          0.99999994 = queryWeight, product of:
            6.136947 = idf(docFreq=260, maxDocs=44421)
            0.16294746 = queryNorm
          3.835592 = fieldWeight in 6730, product of:
            1.0 = tf(freq=1.0), with freq of:
              1.0 = termFreq=1.0
            6.136947 = idf(docFreq=260, maxDocs=44421)
            0.625 = fieldNorm(doc=6730)
    

Similar documents (content)

  1. Fu, T.; Abbasi, A.; Chen, H.: ¬A hybrid approach to Web forum interactional coherence analysis (2008) 0.14
    0.1437106 = sum of:
      0.1437106 = product of:
        0.51325214 = sum of:
          0.050542273 = weight(abstract_txt:outperformed in 2872) [ClassicSimilarity], result of:
            0.050542273 = score(doc=2872,freq=1.0), product of:
              0.078552954 = queryWeight, product of:
                1.0246117 = boost
                8.235732 = idf(docFreq=31, maxDocs=44421)
                0.009308956 = queryNorm
              0.6434166 = fieldWeight in 2872, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                8.235732 = idf(docFreq=31, maxDocs=44421)
                0.078125 = fieldNorm(doc=2872)
          0.016845409 = weight(abstract_txt:techniques in 2872) [ClassicSimilarity], result of:
            0.016845409 = score(doc=2872,freq=1.0), product of:
              0.047576267 = queryWeight, product of:
                1.1276861 = boost
                4.5321174 = idf(docFreq=1298, maxDocs=44421)
                0.009308956 = queryNorm
              0.35407168 = fieldWeight in 2872, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.5321174 = idf(docFreq=1298, maxDocs=44421)
                0.078125 = fieldNorm(doc=2872)
          0.013885778 = weight(abstract_txt:system in 2872) [ClassicSimilarity], result of:
            0.013885778 = score(doc=2872,freq=1.0), product of:
              0.052697837 = queryWeight, product of:
                1.6784347 = boost
                3.372775 = idf(docFreq=4140, maxDocs=44421)
                0.009308956 = queryNorm
              0.26349807 = fieldWeight in 2872, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                3.372775 = idf(docFreq=4140, maxDocs=44421)
                0.078125 = fieldNorm(doc=2872)
          0.08458175 = weight(abstract_txt:forum in 2872) [ClassicSimilarity], result of:
            0.08458175 = score(doc=2872,freq=2.0), product of:
              0.11072451 = queryWeight, product of:
                1.7203426 = boost
                6.9139757 = idf(docFreq=119, maxDocs=44421)
                0.009308956 = queryNorm
              0.7638936 = fieldWeight in 2872, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                6.9139757 = idf(docFreq=119, maxDocs=44421)
                0.078125 = fieldNorm(doc=2872)
          0.051624972 = weight(abstract_txt:recall in 2872) [ClassicSimilarity], result of:
            0.051624972 = score(doc=2872,freq=1.0), product of:
              0.11490519 = queryWeight, product of:
                2.1463895 = boost
                5.750825 = idf(docFreq=383, maxDocs=44421)
                0.009308956 = queryNorm
              0.44928318 = fieldWeight in 2872, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.750825 = idf(docFreq=383, maxDocs=44421)
                0.078125 = fieldNorm(doc=2872)
          0.033031236 = weight(abstract_txt:content in 2872) [ClassicSimilarity], result of:
            0.033031236 = score(doc=2872,freq=1.0), product of:
              0.10115776 = queryWeight, product of:
                2.599936 = boost
                4.1796083 = idf(docFreq=1847, maxDocs=44421)
                0.009308956 = queryNorm
              0.3265319 = fieldWeight in 2872, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.1796083 = idf(docFreq=1847, maxDocs=44421)
                0.078125 = fieldNorm(doc=2872)
          0.2627407 = weight(abstract_txt:forums in 2872) [ClassicSimilarity], result of:
            0.2627407 = score(doc=2872,freq=1.0), product of:
              0.42834592 = queryWeight, product of:
                5.86072 = boost
                7.85132 = idf(docFreq=46, maxDocs=44421)
                0.009308956 = queryNorm
              0.61338437 = fieldWeight in 2872, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                7.85132 = idf(docFreq=46, maxDocs=44421)
                0.078125 = fieldNorm(doc=2872)
        0.28 = coord(7/25)
    
  2. Yang, M.; Kiang, M.; Chen, H.; Li, Y.: Artificial immune system for illicit content identification in social media (2012) 0.14
    0.1431746 = sum of:
      0.1431746 = product of:
        0.5113379 = sum of:
          0.039128244 = weight(abstract_txt:postings in 980) [ClassicSimilarity], result of:
            0.039128244 = score(doc=980,freq=1.0), product of:
              0.07685278 = queryWeight, product of:
                1.0134629 = boost
                8.146119 = idf(docFreq=34, maxDocs=44421)
                0.009308956 = queryNorm
              0.50913244 = fieldWeight in 980, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                8.146119 = idf(docFreq=34, maxDocs=44421)
                0.0625 = fieldNorm(doc=980)
          0.013476327 = weight(abstract_txt:techniques in 980) [ClassicSimilarity], result of:
            0.013476327 = score(doc=980,freq=1.0), product of:
              0.047576267 = queryWeight, product of:
                1.1276861 = boost
                4.5321174 = idf(docFreq=1298, maxDocs=44421)
                0.009308956 = queryNorm
              0.28325734 = fieldWeight in 980, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.5321174 = idf(docFreq=1298, maxDocs=44421)
                0.0625 = fieldNorm(doc=980)
          0.064481884 = weight(abstract_txt:hate in 980) [ClassicSimilarity], result of:
            0.064481884 = score(doc=980,freq=1.0), product of:
              0.10722379 = queryWeight, product of:
                1.1970813 = boost
                9.622026 = idf(docFreq=7, maxDocs=44421)
                0.009308956 = queryNorm
              0.60137665 = fieldWeight in 980, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                9.622026 = idf(docFreq=7, maxDocs=44421)
                0.0625 = fieldNorm(doc=980)
          0.011370318 = weight(abstract_txt:approach in 980) [ClassicSimilarity], result of:
            0.011370318 = score(doc=980,freq=1.0), product of:
              0.048628196 = queryWeight, product of:
                1.396313 = boost
                3.741144 = idf(docFreq=2864, maxDocs=44421)
                0.009308956 = queryNorm
              0.2338215 = fieldWeight in 980, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                3.741144 = idf(docFreq=2864, maxDocs=44421)
                0.0625 = fieldNorm(doc=980)
          0.015709963 = weight(abstract_txt:system in 980) [ClassicSimilarity], result of:
            0.015709963 = score(doc=980,freq=2.0), product of:
              0.052697837 = queryWeight, product of:
                1.6784347 = boost
                3.372775 = idf(docFreq=4140, maxDocs=44421)
                0.009308956 = queryNorm
              0.298114 = fieldWeight in 980, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                3.372775 = idf(docFreq=4140, maxDocs=44421)
                0.0625 = fieldNorm(doc=980)
          0.069913946 = weight(abstract_txt:content in 980) [ClassicSimilarity], result of:
            0.069913946 = score(doc=980,freq=7.0), product of:
              0.10115776 = queryWeight, product of:
                2.599936 = boost
                4.1796083 = idf(docFreq=1847, maxDocs=44421)
                0.009308956 = queryNorm
              0.69113773 = fieldWeight in 980, product of:
                2.6457512 = tf(freq=7.0), with freq of:
                  7.0 = termFreq=7.0
                4.1796083 = idf(docFreq=1847, maxDocs=44421)
                0.0625 = fieldNorm(doc=980)
          0.29725716 = weight(abstract_txt:forums in 980) [ClassicSimilarity], result of:
            0.29725716 = score(doc=980,freq=2.0), product of:
              0.42834592 = queryWeight, product of:
                5.86072 = boost
                7.85132 = idf(docFreq=46, maxDocs=44421)
                0.009308956 = queryNorm
              0.6939652 = fieldWeight in 980, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                7.85132 = idf(docFreq=46, maxDocs=44421)
                0.0625 = fieldNorm(doc=980)
        0.28 = coord(7/25)
    
  3. Cohan, A.; Young, S.; Yates, A.; Goharian, N.: Triaging content severity in online mental health forums (2017) 0.11
    0.10656866 = sum of:
      0.10656866 = product of:
        0.5328433 = sum of:
          0.011370318 = weight(abstract_txt:approach in 4930) [ClassicSimilarity], result of:
            0.011370318 = score(doc=4930,freq=1.0), product of:
              0.048628196 = queryWeight, product of:
                1.396313 = boost
                3.741144 = idf(docFreq=2864, maxDocs=44421)
                0.009308956 = queryNorm
              0.2338215 = fieldWeight in 4930, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                3.741144 = idf(docFreq=2864, maxDocs=44421)
                0.0625 = fieldNorm(doc=4930)
          0.04784666 = weight(abstract_txt:forum in 4930) [ClassicSimilarity], result of:
            0.04784666 = score(doc=4930,freq=1.0), product of:
              0.11072451 = queryWeight, product of:
                1.7203426 = boost
                6.9139757 = idf(docFreq=119, maxDocs=44421)
                0.009308956 = queryNorm
              0.43212348 = fieldWeight in 4930, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.9139757 = idf(docFreq=119, maxDocs=44421)
                0.0625 = fieldNorm(doc=4930)
          0.0504741 = weight(abstract_txt:improvement in 4930) [ClassicSimilarity], result of:
            0.0504741 = score(doc=4930,freq=1.0), product of:
              0.1313466 = queryWeight, product of:
                2.2948172 = boost
                6.148508 = idf(docFreq=257, maxDocs=44421)
                0.009308956 = queryNorm
              0.38428175 = fieldWeight in 4930, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.148508 = idf(docFreq=257, maxDocs=44421)
                0.0625 = fieldNorm(doc=4930)
          0.05908807 = weight(abstract_txt:content in 4930) [ClassicSimilarity], result of:
            0.05908807 = score(doc=4930,freq=5.0), product of:
              0.10115776 = queryWeight, product of:
                2.599936 = boost
                4.1796083 = idf(docFreq=1847, maxDocs=44421)
                0.009308956 = queryNorm
              0.584118 = fieldWeight in 4930, product of:
                2.236068 = tf(freq=5.0), with freq of:
                  5.0 = termFreq=5.0
                4.1796083 = idf(docFreq=1847, maxDocs=44421)
                0.0625 = fieldNorm(doc=4930)
          0.36406416 = weight(abstract_txt:forums in 4930) [ClassicSimilarity], result of:
            0.36406416 = score(doc=4930,freq=3.0), product of:
              0.42834592 = queryWeight, product of:
                5.86072 = boost
                7.85132 = idf(docFreq=46, maxDocs=44421)
                0.009308956 = queryNorm
              0.8499303 = fieldWeight in 4930, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                7.85132 = idf(docFreq=46, maxDocs=44421)
                0.0625 = fieldNorm(doc=4930)
        0.2 = coord(5/25)
    
  4. Alqaraleh, S.; Ramadan, O.; Salamah, M.: Efficient watcher based web crawler design (2015) 0.08
    0.08480531 = sum of:
      0.08480531 = product of:
        0.42402655 = sum of:
          0.013476327 = weight(abstract_txt:techniques in 2627) [ClassicSimilarity], result of:
            0.013476327 = score(doc=2627,freq=1.0), product of:
              0.047576267 = queryWeight, product of:
                1.1276861 = boost
                4.5321174 = idf(docFreq=1298, maxDocs=44421)
                0.009308956 = queryNorm
              0.28325734 = fieldWeight in 2627, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.5321174 = idf(docFreq=1298, maxDocs=44421)
                0.0625 = fieldNorm(doc=2627)
          0.011370318 = weight(abstract_txt:approach in 2627) [ClassicSimilarity], result of:
            0.011370318 = score(doc=2627,freq=1.0), product of:
              0.048628196 = queryWeight, product of:
                1.396313 = boost
                3.741144 = idf(docFreq=2864, maxDocs=44421)
                0.009308956 = queryNorm
              0.2338215 = fieldWeight in 2627, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                3.741144 = idf(docFreq=2864, maxDocs=44421)
                0.0625 = fieldNorm(doc=2627)
          0.022276774 = weight(abstract_txt:human in 2627) [ClassicSimilarity], result of:
            0.022276774 = score(doc=2627,freq=1.0), product of:
              0.07613914 = queryWeight, product of:
                1.7472003 = boost
                4.681277 = idf(docFreq=1118, maxDocs=44421)
                0.009308956 = queryNorm
              0.2925798 = fieldWeight in 2627, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.681277 = idf(docFreq=1118, maxDocs=44421)
                0.0625 = fieldNorm(doc=2627)
          0.17427924 = weight(abstract_txt:crawling in 2627) [ClassicSimilarity], result of:
            0.17427924 = score(doc=2627,freq=4.0), product of:
              0.16512765 = queryWeight, product of:
                2.1008883 = boost
                8.443371 = idf(docFreq=25, maxDocs=44421)
                0.009308956 = queryNorm
              1.0554214 = fieldWeight in 2627, product of:
                2.0 = tf(freq=4.0), with freq of:
                  4.0 = termFreq=4.0
                8.443371 = idf(docFreq=25, maxDocs=44421)
                0.0625 = fieldNorm(doc=2627)
          0.2026239 = weight(abstract_txt:crawler in 2627) [ClassicSimilarity], result of:
            0.2026239 = score(doc=2627,freq=2.0), product of:
              0.26332387 = queryWeight, product of:
                3.2492552 = boost
                8.705735 = idf(docFreq=19, maxDocs=44421)
                0.009308956 = queryNorm
              0.76948553 = fieldWeight in 2627, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                8.705735 = idf(docFreq=19, maxDocs=44421)
                0.0625 = fieldNorm(doc=2627)
        0.2 = coord(5/25)
    
  5. Simeoni, F.; Yakici, M.; Neely, S.; Crestani, F.: Metadata harvesting for content-based distributed information retrieval (2008) 0.08
    0.084200226 = sum of:
      0.084200226 = product of:
        0.42100114 = sum of:
          0.049513992 = weight(abstract_txt:periodic in 2336) [ClassicSimilarity], result of:
            0.049513992 = score(doc=2336,freq=1.0), product of:
              0.089912064 = queryWeight, product of:
                1.0961931 = boost
                8.811096 = idf(docFreq=17, maxDocs=44421)
                0.009308956 = queryNorm
              0.5506935 = fieldWeight in 2336, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                8.811096 = idf(docFreq=17, maxDocs=44421)
                0.0625 = fieldNorm(doc=2336)
          0.022740636 = weight(abstract_txt:approach in 2336) [ClassicSimilarity], result of:
            0.022740636 = score(doc=2336,freq=4.0), product of:
              0.048628196 = queryWeight, product of:
                1.396313 = boost
                3.741144 = idf(docFreq=2864, maxDocs=44421)
                0.009308956 = queryNorm
              0.467643 = fieldWeight in 2336, product of:
                2.0 = tf(freq=4.0), with freq of:
                  4.0 = termFreq=4.0
                3.741144 = idf(docFreq=2864, maxDocs=44421)
                0.0625 = fieldNorm(doc=2336)
          0.12323403 = weight(abstract_txt:crawling in 2336) [ClassicSimilarity], result of:
            0.12323403 = score(doc=2336,freq=2.0), product of:
              0.16512765 = queryWeight, product of:
                2.1008883 = boost
                8.443371 = idf(docFreq=25, maxDocs=44421)
                0.009308956 = queryNorm
              0.7462956 = fieldWeight in 2336, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                8.443371 = idf(docFreq=25, maxDocs=44421)
                0.0625 = fieldNorm(doc=2336)
          0.07927497 = weight(abstract_txt:content in 2336) [ClassicSimilarity], result of:
            0.07927497 = score(doc=2336,freq=9.0), product of:
              0.10115776 = queryWeight, product of:
                2.599936 = boost
                4.1796083 = idf(docFreq=1847, maxDocs=44421)
                0.009308956 = queryNorm
              0.78367656 = fieldWeight in 2336, product of:
                3.0 = tf(freq=9.0), with freq of:
                  9.0 = termFreq=9.0
                4.1796083 = idf(docFreq=1847, maxDocs=44421)
                0.0625 = fieldNorm(doc=2336)
          0.14623751 = weight(abstract_txt:incremental in 2336) [ClassicSimilarity], result of:
            0.14623751 = score(doc=2336,freq=1.0), product of:
              0.29380456 = queryWeight, product of:
                3.963121 = boost
                7.963798 = idf(docFreq=41, maxDocs=44421)
                0.009308956 = queryNorm
              0.49773738 = fieldWeight in 2336, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                7.963798 = idf(docFreq=41, maxDocs=44421)
                0.0625 = fieldNorm(doc=2336)
        0.2 = coord(5/25)