HCDA/cikm-paper Changeset - 0ddc6c26cba0 · Centrum Wiskunde & Informatica (CWI)

Changeset - 0ddc6c26cba0

Parent rev.

Child rev.

[Not reviewed]

Merge

0 1 0

Gebrekirstos Gebremeskel - 11 years ago 2014-06-12 02:35:20
destinycome@gmail.com

Merge branch 'master' of https://scm.cwi.nl/IA/cikm-paper
auto merge

1 file changed with 8 insertions and 7 deletions:

mypaper-final.tex

0 comments (0 inline, 0 general)

mypaper-final.tex

➞

Show inline comments

@@ @@ -90,59 +90,60 @@ @@
 % \alignauthor
 % G.K.M. Tobin\titlenote{The secretary disavows
 % any knowledge of this author's actions.}\\
 %        \affaddr{Institute for Clarity in Documentation}\\
 %        \affaddr{P.O. Box 1212}\\
 %        \affaddr{Dublin, Ohio 43017-6221}\\
 %        \email{webmaster@marysville-ohio.com}
 % }
 % There's nothing stopping you putting the seventh, eighth, etc.
 % author on the opening page (as the 'third row') but we ask,
 % for aesthetic reasons that you place these 'additional authors'
 % in the \additional authors block, viz.
 % Just remember to make sure that the TOTAL number of authors
 % is the number that will appear on the first page PLUS the
 % number that will appear in the \additionalauthors section.
 \maketitle
 \begin{abstract}
 Cumulative citation recommendation refers to the problem faced by
 knowledge base curators, who need to continuously screen the media for
 updates regarding the knowledge base entries they manage. Automatic
 system support for this entity-centric information processing problem
 requires complex pipe\-lines involving both natural language
 processing and information retrieval components. The default pipeline
 processing and information retrieval components. The pipeline
 encountered in a variety of systems that approach this problem
 involves four stages: filtering, classification, ranking (or scoring),
 and evaluation. Filtering is an initial step, that reduces the
+and evaluation. Filtering is only an initial step, that reduces the
 web-scale corpus of news and other relevant information sources that
 may contain entity mentions into a working set of documents that should
 be more manageable for the subsequent stages.
 This step has a large impact on the recall that can be achieved.
 Keeping the subsequent steps constant, we therefore zoom in into the
 filtering stage, and conduct an in-depth analysis of the main design
 decisions here:
 cleansing noisy web data, the methods to create entity profiles, the
 Nevertheless, this step has a large impact on the recall that can be
 maximally attained! Therefore, in this study, we have focused on just
 this filtering stage and conduct an in-depth analysis of the main design
 decisions here: how to cleans the noisy text obtained online,
 the methods to create entity profiles, the
 types of entities of interest, document type, and the grade of
 relevance of the document-entity pair under consideration.
 We analyze how these factors (and the design choices made in their
 corresponding system components) affect filtering performance.
 We identify and characterize the relevant documents that do not pass
 the filtering stage by examining their contents. This way, we give
 estimate of a practical upper-bound of recall for entity-centric stream
 filtering.
 \end{abstract}
 % A category with the (minimum) three required fields
 \category{H.4}{Information Filtering}{Miscellaneous}
 %A category including the fourth, optional field follows...
 %\category{D.2.8}{Software Engineering}{Metrics}[complexity measures, performance measures]
 \terms{Theory}
 \keywords{Information Filtering; Cumulative Citation Recommendation; knowledge maintenance; Stream Filtering;  emerging entities} % NOT required for Proceedings
 \section{Introduction}
   In 2012, the Text REtrieval Conferences (TREC) introduced the Knowledge Base Acceleration (KBA) track  to help Knowledge Bases(KBs) curators. The track is crucial to address a critical need of KB curators: given KB (Wikipedia or Twitter) entities, filter  a stream  for relevant documents, rank the retrieved documents and recommend them to the KB curators. The track is crucial and timely because  the number of entities in a KB on one hand, and the huge amount of new information content on the Web on the other hand make the task of manual KB maintenance challenging.   TREC KBA's main task, Cumulative Citation Recommendation (CCR), aims at filtering a stream to identify   citation-worthy  documents, rank them,  and recommend them to KB curators.

0 comments (0 inline, 0 general)