• DocumentCode
    3739162
  • Title

    Context-Aware Approximate String Matching for Large-Scale Real-Time Entity Resolution

  • Author

    Peter Christen;Ross W Gayler

  • Author_Institution
    Res. Sch. of Comput. Sci., Australian Nat. Univ., Acton, ACT, Australia
  • fYear
    2015
  • Firstpage
    211
  • Lastpage
    217
  • Abstract
    Techniques for approximate string matching have been widely studied over several decades. They are required in many applications, including entity resolution, spell checking, similarity joins, and biological sequence comparison. Most existing techniques for approximate string matching used in entity resolution only consider the two strings that are compared. They neglect contextual information such as the frequency of how often strings occur in a database, the likelihood of the character edits between strings, or how many other similar strings there are in a database. In this paper we investigate if incorporating such contextual information into edit distance based approximate string matching can improve matching quality for real-time entity resolution. In this application, query records have to be matched in sub-second time to records in a large database that refer to the same entity. We evaluate our approach on two large real data sets and compare it to several baseline approaches. Our results show that considering edit frequency and the neighborhood size of a string can improve matching results, while taking string frequencies into account can actually make results worse.
  • Keywords
    "Real-time systems","Indexes","Positron emission tomography","Conferences","Australia","Electronic mail"
  • Publisher
    ieee
  • Conference_Titel
    Data Mining Workshop (ICDMW), 2015 IEEE International Conference on
  • Electronic_ISBN
    2375-9259
  • Type

    conf

  • DOI
    10.1109/ICDMW.2015.152
  • Filename
    7395673