HCDA/cikm-paper Changeset - 222076072307 · Centrum Wiskunde & Informatica (CWI)

Changeset - 222076072307

Parent rev.

Child rev.

[Not reviewed]

0 1 0

Arjen de Vries (arjen) - 11 years ago 2014-06-12 06:54:00
arjen.de.vries@cwi.nl

latex glitch

1 file changed with 1 insertions and 1 deletions:

mypaper-final.tex

0 comments (0 inline, 0 general)

mypaper-final.tex

➞

Show inline comments

 \subsection{Relevance judgments}
 TREC-KBA provided relevance judgments for training and
 testing. Relevance judgments are given as a document-entity
 pairs. Documents with citation-worthy content to a given entity are
 annotated  as \emph{vital},  while documents with tangentially
 relevant content, or documents that lack freshliness o  with content
 that can be useful for initial KB-dossier are annotated as
 \emph{relevant}. Documents with no relevant content are labeled
 \emph{neutral} and spam is labeled as \emph{garbage}.
 %The inter-annotator agreement on vital in 2012 was 70\% while in 2013 it
 %is 76\%. This is due to the more refined definition of vital and the
 %distinction made between vital and relevant.
 \subsection{Breakdown of results by document source category}
 %The results of the different entity profiles on the raw corpus are
 %broken down by source categories and relevance rank% (vital, or
 %relevant).
 In total, the dataset contains 24162 unique entity-document
 pairs, vital or relevant; 9521 of these have been labelled as vital,
 and the remaining 17424 as relevant.
 All documents are categorized into 8 source categories: 0.98\%
 arxiv(a), 0.034\% classified(c), 0.34\% forum(f), 5.65\% linking(l),
-.53\% mainstream-news(m-n), 18.40\% news(n), 12.93\% social(s) and
+.53\% main\-stream-news(m-n), 18.40\% news(n), 12.93\% social(s) and
 .2\% weblog(w). We have regrouped these source categories into three
 groups ``news'', ``social'', and ``other'', for two reasons: 1) some groups
 are very similar to each other. Mainstream-news and news are
 similar. The reason they exist separately, in the first place,  is
 because they were collected from two different sources, by different
 groups and at different times. we call them news from now on.  The
 same is true with weblog and social, and we call them social from now
 on.   2) some groups have so small number of annotations that treating
 them independently does not make much sense. Majority of vital or
 relevant annotations are social (social and weblog) (63.13\%). News
 (mainstream +news) make up 30\%. Thus, news and social make up about
 \% of all annotations.  The rest make up about 7\% and are all
 grouped as others.
  \section{Stream Filtering}\label{sec:fil}
  The TREC Filtering track defines filtering as a ``system that sifts
  through stream of incoming information to find documents that are
  relevant to a set of user needs represented by profiles''
  \cite{robertson2002trec}. Its information needs are long-term and are
  represented by persistent profiles, unlike the traditional search system
  whose adhoc information need is represented by a search
  query. Adaptive Filtering, one task of the filtering track,  starts
  with  a persistent user profile and a very small number of positive

0 comments (0 inline, 0 general)