Document (#30264)

Author
Hsu, C.-N.
Chang, C.-H.
Hsieh, C.-H.
Lu, J.-J.
Chang, C.-C.
Title
Reconfigurable Web wrapper agents for biological information integration
Source
Journal of the American Society for Information Science and Technology. 56(2005) no.5, S.505-517
Year
2005
Abstract
A variety of biological data is transferred and exchanged in overwhelming volumes on the World Wide Web. How to rapidly capture, utilize, and integrate the information on the Internet to discover valuable biological knowledge is one of the most critical issues in bioinformatics. Many information integration systems have been proposed for integrating biological data. These systems usually rely on an intermediate software layer called wrappers to access connected information sources. Wrapper construction for Web data sources is often specially hand coded to accommodate the differences between each Web site. However, programming a Web wrapper requires substantial programming skill, and is time-consuming and hard to maintain. In this article we provide a solution for rapidly building software agents that can serve as Web wrappers for biological information integration. We define an XML-based language called Web Navigation Description Language (WNDL), to model a Web-browsing session. A WNDL script describes how to locate the data, extract the data, and combine the data. By executing different WNDL scripts, we can automate virtually all types of Web-browsing sessions. We also describe IEPAD (Information Extraction Based on Pattern Discovery), a data extractor based on pattern discovery techniques. IEPAD allows our software agents to automatically discover the extraction rules to extract the contents of a structurally formatted Web page. With a programming-by-example authoring tool, a user can generate a complete Web wrapper agent by browsing the target Web sites. We built a variety of biological applications to demonstrate the feasibility of our approach.
Footnote
Beitrag in einem special issue on bioinformatics

Similar documents (author)

  1. Yang, T.-H.; Hsieh, Y.-L.; Liu, S.-H.; Chang, Y.-C.; Hsu, W.-L.: ¬A flexible template generation and matching method with applications for publication reference metadata extraction (2021) 2.75
    2.7496674 = sum of:
      2.7496674 = sum of:
        1.121949 = weight(author_txt:hsieh in 1064) [ClassicSimilarity], result of:
          1.121949 = score(doc=1064,freq=1.0), product of:
            0.5265256 = queryWeight, product of:
              8.523414 = idf(docFreq=23, maxDocs=44421)
              0.06177403 = queryNorm
            2.1308534 = fieldWeight in 1064, product of:
              1.0 = tf(freq=1.0), with freq of:
                1.0 = termFreq=1.0
              8.523414 = idf(docFreq=23, maxDocs=44421)
              0.25 = fieldNorm(doc=1064)
        1.6277184 = weight(author_txt:chang in 1064) [ClassicSimilarity], result of:
          1.6277184 = score(doc=1064,freq=1.0), product of:
            0.8501593 = queryWeight, product of:
              1.7970303 = boost
              7.6584163 = idf(docFreq=56, maxDocs=44421)
              0.06177403 = queryNorm
            1.9146041 = fieldWeight in 1064, product of:
              1.0 = tf(freq=1.0), with freq of:
                1.0 = termFreq=1.0
              7.6584163 = idf(docFreq=56, maxDocs=44421)
              0.25 = fieldNorm(doc=1064)
    
  2. Chang, R.: DBase, relational data models, and MARC records (1992) 2.03
    2.034648 = sum of:
      2.034648 = product of:
        4.069296 = sum of:
          4.069296 = weight(author_txt:chang in 5056) [ClassicSimilarity], result of:
            4.069296 = score(doc=5056,freq=1.0), product of:
              0.8501593 = queryWeight, product of:
                1.7970303 = boost
                7.6584163 = idf(docFreq=56, maxDocs=44421)
                0.06177403 = queryNorm
              4.78651 = fieldWeight in 5056, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                7.6584163 = idf(docFreq=56, maxDocs=44421)
                0.625 = fieldNorm(doc=5056)
        0.5 = coord(1/2)
    
  3. Chang, R.: ¬The development of indexing technology (1993) 2.03
    2.034648 = sum of:
      2.034648 = product of:
        4.069296 = sum of:
          4.069296 = weight(author_txt:chang in 7023) [ClassicSimilarity], result of:
            4.069296 = score(doc=7023,freq=1.0), product of:
              0.8501593 = queryWeight, product of:
                1.7970303 = boost
                7.6584163 = idf(docFreq=56, maxDocs=44421)
                0.06177403 = queryNorm
              4.78651 = fieldWeight in 7023, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                7.6584163 = idf(docFreq=56, maxDocs=44421)
                0.625 = fieldNorm(doc=7023)
        0.5 = coord(1/2)
    
  4. Chang, R.: Keyword searching and indexing (1993) 2.03
    2.034648 = sum of:
      2.034648 = product of:
        4.069296 = sum of:
          4.069296 = weight(author_txt:chang in 7222) [ClassicSimilarity], result of:
            4.069296 = score(doc=7222,freq=1.0), product of:
              0.8501593 = queryWeight, product of:
                1.7970303 = boost
                7.6584163 = idf(docFreq=56, maxDocs=44421)
                0.06177403 = queryNorm
              4.78651 = fieldWeight in 7222, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                7.6584163 = idf(docFreq=56, maxDocs=44421)
                0.625 = fieldNorm(doc=7222)
        0.5 = coord(1/2)
    
  5. Chang, R.H.: To classify or not to classify? : a new look at an old problem (1989) 2.03
    2.034648 = sum of:
      2.034648 = product of:
        4.069296 = sum of:
          4.069296 = weight(author_txt:chang in 2578) [ClassicSimilarity], result of:
            4.069296 = score(doc=2578,freq=1.0), product of:
              0.8501593 = queryWeight, product of:
                1.7970303 = boost
                7.6584163 = idf(docFreq=56, maxDocs=44421)
                0.06177403 = queryNorm
              4.78651 = fieldWeight in 2578, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                7.6584163 = idf(docFreq=56, maxDocs=44421)
                0.625 = fieldNorm(doc=2578)
        0.5 = coord(1/2)
    

Similar documents (content)

  1. Haslhofer, B.: ¬A Web-based mapping technique for establishing metadata interoperability (2008) 0.17
    0.1748036 = sum of:
      0.1748036 = product of:
        0.6242986 = sum of:
          0.026073378 = weight(abstract_txt:sources in 160) [ClassicSimilarity], result of:
            0.026073378 = score(doc=160,freq=4.0), product of:
              0.070364304 = queryWeight, product of:
                1.1129389 = boost
                4.743019 = idf(docFreq=1051, maxDocs=44421)
                0.013329878 = queryNorm
              0.37054837 = fieldWeight in 160, product of:
                2.0 = tf(freq=4.0), with freq of:
                  4.0 = termFreq=4.0
                4.743019 = idf(docFreq=1051, maxDocs=44421)
                0.0390625 = fieldNorm(doc=160)
          0.010237532 = weight(abstract_txt:based in 160) [ClassicSimilarity], result of:
            0.010237532 = score(doc=160,freq=3.0), product of:
              0.047536552 = queryWeight, product of:
                1.1203523 = boost
                3.1830752 = idf(docFreq=5005, maxDocs=44421)
                0.013329878 = queryNorm
              0.21536125 = fieldWeight in 160, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                3.1830752 = idf(docFreq=5005, maxDocs=44421)
                0.0390625 = fieldNorm(doc=160)
          0.030621292 = weight(abstract_txt:discovery in 160) [ClassicSimilarity], result of:
            0.030621292 = score(doc=160,freq=2.0), product of:
              0.098683946 = queryWeight, product of:
                1.318009 = boost
                5.616968 = idf(docFreq=438, maxDocs=44421)
                0.013329878 = queryNorm
              0.3102966 = fieldWeight in 160, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                5.616968 = idf(docFreq=438, maxDocs=44421)
                0.0390625 = fieldNorm(doc=160)
          0.0073365583 = weight(abstract_txt:information in 160) [ClassicSimilarity], result of:
            0.0073365583 = score(doc=160,freq=2.0), product of:
              0.054903436 = queryWeight, product of:
                1.7027681 = boost
                2.4188995 = idf(docFreq=10748, maxDocs=44421)
                0.013329878 = queryNorm
              0.13362658 = fieldWeight in 160, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                2.4188995 = idf(docFreq=10748, maxDocs=44421)
                0.0390625 = fieldNorm(doc=160)
          0.0470701 = weight(abstract_txt:integration in 160) [ClassicSimilarity], result of:
            0.0470701 = score(doc=160,freq=3.0), product of:
              0.13144001 = queryWeight, product of:
                1.8629644 = boost
                5.2929387 = idf(docFreq=606, maxDocs=44421)
                0.013329878 = queryNorm
              0.3581109 = fieldWeight in 160, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                5.2929387 = idf(docFreq=606, maxDocs=44421)
                0.0390625 = fieldNorm(doc=160)
          0.027355978 = weight(abstract_txt:data in 160) [ClassicSimilarity], result of:
            0.027355978 = score(doc=160,freq=3.0), product of:
              0.12141097 = queryWeight, product of:
                2.735005 = boost
                3.3302255 = idf(docFreq=4320, maxDocs=44421)
                0.013329878 = queryNorm
              0.22531718 = fieldWeight in 160, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                3.3302255 = idf(docFreq=4320, maxDocs=44421)
                0.0390625 = fieldNorm(doc=160)
          0.4756037 = weight(abstract_txt:wrapper in 160) [ClassicSimilarity], result of:
            0.4756037 = score(doc=160,freq=4.0), product of:
              0.61431956 = queryWeight, product of:
                4.650582 = boost
                9.909708 = idf(docFreq=5, maxDocs=44421)
                0.013329878 = queryNorm
              0.7741959 = fieldWeight in 160, product of:
                2.0 = tf(freq=4.0), with freq of:
                  4.0 = termFreq=4.0
                9.909708 = idf(docFreq=5, maxDocs=44421)
                0.0390625 = fieldNorm(doc=160)
        0.28 = coord(7/25)
    
  2. Rodríguez, A.; Carazo, J.M.; Trelles-Salazar, O.: Mining association rules from biological databases (2005) 0.14
    0.1374651 = sum of:
      0.1374651 = product of:
        0.5727713 = sum of:
          0.0868836 = weight(abstract_txt:bioinformatics in 261) [ClassicSimilarity], result of:
            0.0868836 = score(doc=261,freq=2.0), product of:
              0.11475353 = queryWeight, product of:
                1.0049933 = boost
                8.565973 = idf(docFreq=22, maxDocs=44421)
                0.013329878 = queryNorm
              0.75713223 = fieldWeight in 261, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                8.565973 = idf(docFreq=22, maxDocs=44421)
                0.0625 = fieldNorm(doc=261)
          0.034644037 = weight(abstract_txt:discovery in 261) [ClassicSimilarity], result of:
            0.034644037 = score(doc=261,freq=1.0), product of:
              0.098683946 = queryWeight, product of:
                1.318009 = boost
                5.616968 = idf(docFreq=438, maxDocs=44421)
                0.013329878 = queryNorm
              0.3510605 = fieldWeight in 261, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.616968 = idf(docFreq=438, maxDocs=44421)
                0.0625 = fieldNorm(doc=261)
          0.046412196 = weight(abstract_txt:extraction in 261) [ClassicSimilarity], result of:
            0.046412196 = score(doc=261,freq=1.0), product of:
              0.119926624 = queryWeight, product of:
                1.4529575 = boost
                6.192079 = idf(docFreq=246, maxDocs=44421)
                0.013329878 = queryNorm
              0.38700494 = fieldWeight in 261, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.192079 = idf(docFreq=246, maxDocs=44421)
                0.0625 = fieldNorm(doc=261)
          0.047444142 = weight(abstract_txt:pattern in 261) [ClassicSimilarity], result of:
            0.047444142 = score(doc=261,freq=1.0), product of:
              0.12169776 = queryWeight, product of:
                1.4636472 = boost
                6.2376356 = idf(docFreq=235, maxDocs=44421)
                0.013329878 = queryNorm
              0.38985223 = fieldWeight in 261, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.2376356 = idf(docFreq=235, maxDocs=44421)
                0.0625 = fieldNorm(doc=261)
          0.025270369 = weight(abstract_txt:data in 261) [ClassicSimilarity], result of:
            0.025270369 = score(doc=261,freq=1.0), product of:
              0.12141097 = queryWeight, product of:
                2.735005 = boost
                3.3302255 = idf(docFreq=4320, maxDocs=44421)
                0.013329878 = queryNorm
              0.20813909 = fieldWeight in 261, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                3.3302255 = idf(docFreq=4320, maxDocs=44421)
                0.0625 = fieldNorm(doc=261)
          0.33211696 = weight(abstract_txt:biological in 261) [ClassicSimilarity], result of:
            0.33211696 = score(doc=261,freq=2.0), product of:
              0.50978297 = queryWeight, product of:
                5.1885786 = boost
                7.370734 = idf(docFreq=75, maxDocs=44421)
                0.013329878 = queryNorm
              0.651487 = fieldWeight in 261, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                7.370734 = idf(docFreq=75, maxDocs=44421)
                0.0625 = fieldNorm(doc=261)
        0.24 = coord(6/25)
    
  3. Mahoui, M.; Miled, Z.B.; Godse, A.; Kulkarni, H.; Li, N.: BioFacets : faceted classification for biological information (2006) 0.12
    0.11688062 = sum of:
      0.11688062 = product of:
        0.73050386 = sum of:
          0.011821283 = weight(abstract_txt:based in 1779) [ClassicSimilarity], result of:
            0.011821283 = score(doc=1779,freq=1.0), product of:
              0.047536552 = queryWeight, product of:
                1.1203523 = boost
                3.1830752 = idf(docFreq=5005, maxDocs=44421)
                0.013329878 = queryNorm
              0.24867775 = fieldWeight in 1779, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                3.1830752 = idf(docFreq=5005, maxDocs=44421)
                0.078125 = fieldNorm(doc=1779)
          0.07686515 = weight(abstract_txt:integration in 1779) [ClassicSimilarity], result of:
            0.07686515 = score(doc=1779,freq=2.0), product of:
              0.13144001 = queryWeight, product of:
                1.8629644 = boost
                5.2929387 = idf(docFreq=606, maxDocs=44421)
                0.013329878 = queryNorm
              0.5847926 = fieldWeight in 1779, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                5.2929387 = idf(docFreq=606, maxDocs=44421)
                0.078125 = fieldNorm(doc=1779)
          0.054711957 = weight(abstract_txt:data in 1779) [ClassicSimilarity], result of:
            0.054711957 = score(doc=1779,freq=3.0), product of:
              0.12141097 = queryWeight, product of:
                2.735005 = boost
                3.3302255 = idf(docFreq=4320, maxDocs=44421)
                0.013329878 = queryNorm
              0.45063436 = fieldWeight in 1779, product of:
                1.7320508 = tf(freq=3.0), with freq of:
                  3.0 = termFreq=3.0
                3.3302255 = idf(docFreq=4320, maxDocs=44421)
                0.078125 = fieldNorm(doc=1779)
          0.58710545 = weight(abstract_txt:biological in 1779) [ClassicSimilarity], result of:
            0.58710545 = score(doc=1779,freq=4.0), product of:
              0.50978297 = queryWeight, product of:
                5.1885786 = boost
                7.370734 = idf(docFreq=75, maxDocs=44421)
                0.013329878 = queryNorm
              1.1516773 = fieldWeight in 1779, product of:
                2.0 = tf(freq=4.0), with freq of:
                  4.0 = termFreq=4.0
                7.370734 = idf(docFreq=75, maxDocs=44421)
                0.078125 = fieldNorm(doc=1779)
        0.16 = coord(4/25)
    
  4. Handbook of metadata, semantics and ontologies (2014) 0.09
    0.08860867 = sum of:
      0.08860867 = product of:
        0.3692028 = sum of:
          0.013374254 = weight(abstract_txt:based in 134) [ClassicSimilarity], result of:
            0.013374254 = score(doc=134,freq=2.0), product of:
              0.047536552 = queryWeight, product of:
                1.1203523 = boost
                3.1830752 = idf(docFreq=5005, maxDocs=44421)
                0.013329878 = queryNorm
              0.28134674 = fieldWeight in 134, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                3.1830752 = idf(docFreq=5005, maxDocs=44421)
                0.0625 = fieldNorm(doc=134)
          0.0266637 = weight(abstract_txt:variety in 134) [ClassicSimilarity], result of:
            0.0266637 = score(doc=134,freq=1.0), product of:
              0.08287836 = queryWeight, product of:
                1.2078575 = boost
                5.1475344 = idf(docFreq=701, maxDocs=44421)
                0.013329878 = queryNorm
              0.3217209 = fieldWeight in 134, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.1475344 = idf(docFreq=701, maxDocs=44421)
                0.0625 = fieldNorm(doc=134)
          0.029096432 = weight(abstract_txt:called in 134) [ClassicSimilarity], result of:
            0.029096432 = score(doc=134,freq=1.0), product of:
              0.087845735 = queryWeight, product of:
                1.2435277 = boost
                5.2995505 = idf(docFreq=602, maxDocs=44421)
                0.013329878 = queryNorm
              0.3312219 = fieldWeight in 134, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                5.2995505 = idf(docFreq=602, maxDocs=44421)
                0.0625 = fieldNorm(doc=134)
          0.053487763 = weight(abstract_txt:rapidly in 134) [ClassicSimilarity], result of:
            0.053487763 = score(doc=134,freq=1.0), product of:
              0.1318248 = queryWeight, product of:
                1.523329 = boost
                6.4919815 = idf(docFreq=182, maxDocs=44421)
                0.013329878 = queryNorm
              0.40574884 = fieldWeight in 134, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.4919815 = idf(docFreq=182, maxDocs=44421)
                0.0625 = fieldNorm(doc=134)
          0.011738494 = weight(abstract_txt:information in 134) [ClassicSimilarity], result of:
            0.011738494 = score(doc=134,freq=2.0), product of:
              0.054903436 = queryWeight, product of:
                1.7027681 = boost
                2.4188995 = idf(docFreq=10748, maxDocs=44421)
                0.013329878 = queryNorm
              0.21380253 = fieldWeight in 134, product of:
                1.4142135 = tf(freq=2.0), with freq of:
                  2.0 = termFreq=2.0
                2.4188995 = idf(docFreq=10748, maxDocs=44421)
                0.0625 = fieldNorm(doc=134)
          0.23484217 = weight(abstract_txt:biological in 134) [ClassicSimilarity], result of:
            0.23484217 = score(doc=134,freq=1.0), product of:
              0.50978297 = queryWeight, product of:
                5.1885786 = boost
                7.370734 = idf(docFreq=75, maxDocs=44421)
                0.013329878 = queryNorm
              0.4606709 = fieldWeight in 134, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                7.370734 = idf(docFreq=75, maxDocs=44421)
                0.0625 = fieldNorm(doc=134)
        0.24 = coord(6/25)
    
  5. Park, H.; You, S.; Wolfram, D.: Informal data citation for data sharing and reuse is more common than formal data citation in biomedical fields (2018) 0.08
    0.08485116 = sum of:
      0.08485116 = product of:
        0.4242558 = sum of:
          0.020858703 = weight(abstract_txt:sources in 544) [ClassicSimilarity], result of:
            0.020858703 = score(doc=544,freq=1.0), product of:
              0.070364304 = queryWeight, product of:
                1.1129389 = boost
                4.743019 = idf(docFreq=1051, maxDocs=44421)
                0.013329878 = queryNorm
              0.2964387 = fieldWeight in 544, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.743019 = idf(docFreq=1051, maxDocs=44421)
                0.0625 = fieldNorm(doc=544)
          0.046412196 = weight(abstract_txt:extraction in 544) [ClassicSimilarity], result of:
            0.046412196 = score(doc=544,freq=1.0), product of:
              0.119926624 = queryWeight, product of:
                1.4529575 = boost
                6.192079 = idf(docFreq=246, maxDocs=44421)
                0.013329878 = queryNorm
              0.38700494 = fieldWeight in 544, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                6.192079 = idf(docFreq=246, maxDocs=44421)
                0.0625 = fieldNorm(doc=544)
          0.024271013 = weight(abstract_txt:software in 544) [ClassicSimilarity], result of:
            0.024271013 = score(doc=544,freq=1.0), product of:
              0.08910797 = queryWeight, product of:
                1.5339069 = boost
                4.3580413 = idf(docFreq=1545, maxDocs=44421)
                0.013329878 = queryNorm
              0.27237758 = fieldWeight in 544, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                4.3580413 = idf(docFreq=1545, maxDocs=44421)
                0.0625 = fieldNorm(doc=544)
          0.09787172 = weight(abstract_txt:data in 544) [ClassicSimilarity], result of:
            0.09787172 = score(doc=544,freq=15.0), product of:
              0.12141097 = queryWeight, product of:
                2.735005 = boost
                3.3302255 = idf(docFreq=4320, maxDocs=44421)
                0.013329878 = queryNorm
              0.80611926 = fieldWeight in 544, product of:
                3.8729835 = tf(freq=15.0), with freq of:
                  15.0 = termFreq=15.0
                3.3302255 = idf(docFreq=4320, maxDocs=44421)
                0.0625 = fieldNorm(doc=544)
          0.23484217 = weight(abstract_txt:biological in 544) [ClassicSimilarity], result of:
            0.23484217 = score(doc=544,freq=1.0), product of:
              0.50978297 = queryWeight, product of:
                5.1885786 = boost
                7.370734 = idf(docFreq=75, maxDocs=44421)
                0.013329878 = queryNorm
              0.4606709 = fieldWeight in 544, product of:
                1.0 = tf(freq=1.0), with freq of:
                  1.0 = termFreq=1.0
                7.370734 = idf(docFreq=75, maxDocs=44421)
                0.0625 = fieldNorm(doc=544)
        0.2 = coord(5/25)