Title :
Scholarly big data information extraction and integration in the CiteSeerχ digital library
Author :
Williams, Kresimir ; Jian Wu ; Choudhury, Sagnik Ray ; Khabsa, Madian ; Giles, C. Lee
Author_Institution :
Inf. Sci. & Technol., Pennsylvania State Univ., University Park, PA, USA
fDate :
March 31 2014-April 4 2014
Abstract :
CiteSeerχ is a digital library that contains approximately 3.5 million scholarly documents and receives between 2 and 4 million requests per day. In addition to making documents available via a public Website, the data is also used to facilitate research in areas like citation analysis, co-author network analysis, scalability evaluation and information extraction. The papers in CiteSeerχ are gathered from the Web by means of continuous automatic focused crawling and go through a series of automatic processing steps as part of the ingestion process. Given the size of the collection, the fact that it is constantly expanding, and the multiple ways in which it is used both by the public to access scholarly documents and for research, there are several big data challenges. In this paper, we provide a case study description of how we address these challenges when it comes to information extraction, data integration and entity linking in CiteSeerχ. We describe how we: aggregate data from multiple sources on the Web; store and manage data; process data as part of an automatic ingestion pipeline that includes automatic metadata and information extraction; perform document and citation clustering; perform entity linking and name disambiguation; and make our data and source code available to enable research and collaboration.
Keywords :
Big Data; Web sites; citation analysis; data integration; digital libraries; information retrieval; meta data; pattern clustering; CiteSeerχ digital library; automatic ingestion pipeline; automatic metadata; automatic processing steps; big data information extraction; big data information integration; citation analysis; information extraction; public Website; Data handling; Data mining; Data storage systems; Feature extraction; Information management; Information retrieval; Joining processes;
Conference_Titel :
Data Engineering Workshops (ICDEW), 2014 IEEE 30th International Conference on
Conference_Location :
Chicago, IL
DOI :
10.1109/ICDEW.2014.6818305