• DocumentCode
    140904
  • Title

    Combining information extraction and human computing for crowdsourced knowledge acquisition

  • Author

    Kondreddi, Sarath Kumar ; Triantafillou, P. ; Weikum, G.

  • Author_Institution
    Database & Inf. Syst. Group, Max Planck Inst. for Inf., Saarbrucken, Germany
  • fYear
    2014
  • fDate
    March 31 2014-April 4 2014
  • Firstpage
    988
  • Lastpage
    999
  • Abstract
    Automatic information extraction (IE) enables the construction of very large knowledge bases (KBs), with relational facts on millions of entities from text corpora and Web sources. However, such KBs contain errors and they are far from being complete. This motivates the need for exploiting human intelligence and knowledge using crowd-based human computing (HC) for assessing the validity of facts and for gathering additional knowledge. This paper presents a novel system architecture, called Higgins, which shows how to effectively integrate an IE engine and a HC engine. Higgins generates game questions where players choose or fill in missing relations for subject-relation-object triples. For generating multiple-choice answer candidates, we have constructed a large dictionary of entity names and relational phrases, and have developed specifically designed statistical language models for phrase relatedness. To this end, we combine semantic resources like WordNet, ConceptNet, and others with statistics derived from a large Web corpus. We demonstrate the effectiveness of Higgins for knowledge acquisition by crowdsourced gathering of relationships between characters in narrative descriptions of movies and books.
  • Keywords
    computational linguistics; human factors; information retrieval; knowledge acquisition; semantic networks; very large databases; HC; Higgins system architecture; IE; KBs; automatic information extraction; crowd-based human computing; crowdsourced knowledge acquisition; entity names; multiple-choice answer candidate generation; phrase relatedness; relational phrases; semantic resources; statistical language models; subject-relation-object triples; very large knowledge bases; Context; Dictionaries; Encyclopedias; Engines; Games; Motion pictures; Semantics;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Data Engineering (ICDE), 2014 IEEE 30th International Conference on
  • Conference_Location
    Chicago, IL
  • Type

    conf

  • DOI
    10.1109/ICDE.2014.6816717
  • Filename
    6816717