-
Notifications
You must be signed in to change notification settings - Fork 0
Expand file tree
/
Copy pathch2_judgment.tex
More file actions
291 lines (190 loc) · 40.1 KB
/
Copy pathch2_judgment.tex
File metadata and controls
291 lines (190 loc) · 40.1 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
\chapter{Human Judgments}
\label{c-human-judgment}
The goal of collecting human judgments is to estimate the satisfaction of actual users of a search system, by asking explicit questions to judges (or assessors) who simulate the actual users.
A canonical example is collecting a binary relevance judgment for a document given a search topic. The form of human judgments can be quite varied, however, depending on the type of search task and judging target.
We will start with an example to make the discussion more concrete. Figure~\ref{fig:human_judgment_overview} shows a judging interface for evaluating the quality of a search engine results page (SERP) given the query 'crowdsourcing'. This example presents basic ingredients in collecting human judgments -- search tasks and judging targets. From this example one can imagine a myriad of possibilities in designing a human judgment task.
%shows a list of possible search tasks about the topic of \textit{crowdsourcing} on the left side, and a few samples from existing web search results for query `crowdsourcing' on the right side.
%\emine{Would it be better to show a more standard judging UI here? Like a query and a web page?}
%\emine{I think this example is confusing as it is not clear what the task is; there are three tasks and it is not clear what the judge is judging. Why not use an interface where the task is clear and feedback mechanisms are also clear? This may confuse the reader since it doesnt show how the judgments are collected. }
%\jin{@emine we do show judging I/F later.}%
%\mdr{I would suggest a different example, one that is not directly related to the subject matter of the survey.}
\begin{figure}
\begin{center}
\includegraphics[scale=0.5]{images/judging_interface}
\caption{An example UI for human judgment collection.}
\label{fig:human_judgment_overview}
\end{center}
\end{figure}
%%\paul{In Figure~\ref{fig:human_judgment_overview}: ``what is crowdsourcing?'' (no ``the''), ``to learn \emph{about} crowdsourcing research''}
While this is a simple example, it presents numerous trade-offs one can make in collecting human judgments. You can use either a (potentially ambiguous) keyword query, a well-defined topic description, or a description of a larger information-seeking task. You can collect judgments for a web document or any SERP element, including instant answers or a list of news articles. Queries, topics, or tasks could be created in many ways. And so on.
The rest of this chapter is to give you guidance in collecting human judgments, in the light of recent literature on this topic. We will look over how to collect search tasks and how to determine a judging target. Various considerations in designing a judging interface will be examined, as well as methods for finding and managing human judges.
%The first step in offline evaluation is collecting labels from human judges. In this chapter, we describe various considerations in collecting high-quality labels from human judges at scale. We first discuss the method for collecting search tasks, followed by the design of a judging method. We then discuss the collection of actual judgments, which is an non-trivial task to perform at scale. We also cover the trade-off and in using different types of judging resources -- in-house vs. crowd judges. (20-25 pages)
\section{Collecting Search Tasks}
Before considering judgment design, one needs to collect or construct search tasks against which search results will be evaluated. Search tasks are users' information needs that are typically represented as a description or as a query. In a setting where the search engine is used by actual users, the job of collecting search tasks can be as simple as sampling from queries users issue, whereas without access to such resources one needs to create tasks based on assumptions of target users and information needs.
\subsection{Creating Search Tasks}
In many cases one needs to perform offline evaluation without a working system -- e.g.\ in building a new product, or in an academic setting. In such cases it is essential to collect hypothetical search tasks, often called simulated search or work tasks (where work includes search and other things). \cite{Borlund:2003} summarizes the role of simulated work tasks as follows:
\mdr{I don't learn from 2.1.1 how I should go about creating a search task.}
% consider non-simulated search tasks?
\begin{quote}
``A simulated work task situation, which is a short `cover story', serves two main functions: 1) it triggers and develops a simulated information need by allowing for user interpretations of the situation, leading to cognitively individual information need interpretations as in real life; and 2) it is the platform against which situational relevance is judged. Further, by being the same for all test persons experimental control is provided. Hence, the concept of a simulated work task situation ensures the experiment both realism and control.''
\end{quote}
`Task' can mean different things for different people, and the IR literature has seen long debate over the definition of search task (see \cite{kelly2009methods} for a summary). For our purpose, it is sufficient to understand it as a information need which can be represented in a way that a human judge can use to judge the quality of given result.
%\emine{Is that the definition of task? Maybe we should use a more proper definition?}
%\paul{task $\neq$ need, but I think this use is blessed by so much past use}
The design of search tasks takes a few considerations which can critically affect evaluation results. First, there is the question of where the task comes from and how much the judge is interested in or knowledgable about the task, or the corresponding domain. \cite{Edwards:2016} show that judges' interests in the task has effects on how they perceive and perform the tasks. Judges in general had more knowledge on the tasks they were interested in, perceived the tasks as easier, and had higher engagement in terms of time spent. It is also known that judges' knowledge of the task can affect the quality of the outcome, with small but measureable differences between experts and non-experts \citep{Bailey:2008}.
\mdr{What are the implications for setting up an assessment exercise yourself? How to take these findings into account?}
\jin{I would recommend have a separate group of people designing evaluation, but this would be feasible mostly for industry settings.}
Another dimension of task creation is complexity, which again has many aspects. \cite{Kelly:2015} looked at this problem using a cognitive complexity framework. They found that participants spent more effort (queries, clicks and time to completion) in performing tasks with higher cognitive complexity (create, evaluate and analyze) than tasks with lower cognitive complexity (apply, understand, remember). In order to ensure the representativeness of the evaluation outcome, it would be sensible to balance the tasks of varying complexity in a way that matches actual users' workloads.
%\paul{how is this relevant to collecting/creating tasks for offline evaluation?}
In summary, these results show that the characteristics of search task can affect the quality of corresponding human judgments, and therefore is an important dimension in designing an offline evaluation. Unless you have good understanding of your target users, it is a good idea to interview potential users to find out what the target distribution of tasks should be. It is also recommended to collect information about task characteristics and design experiments accordingly so that one can control the effect of these factors in evaluation.
\mdr{Can you make this more precise? Give more details? What does "important" mean? This is all fairly abstract?}
\jin{some details added}
\subsection{Sampling Query Logs}
Assuming you have a working search engine with real users, it is natural to collect search tasks from query log data. While this is a seemingly straightforward task, there are a few considerations. We outline some below, along with recommendations based on recent studies.
\mdr{I would expect the outcome of 2.1.2 to be a clear recipe for sampling. But that 's not the case.}
\paragraph{Evaluation Goals} The appropriate sampling strategy depends on evaluation goals. In a typical scenario, it is reasonable to start with a \textit{representative} sample of the traffic. Measurements based on this sampling strategy would lead to the characterization of \textit{average} performance, but there are scenarios where average performance is not informative.
For example, \cite{Zaragoza:2010} suggested techniques to identify segments useful for measurement. They introduce the notion of `disruptive sets', which are a set of queries with high quality results in one engine, but not in another. Using a disruptive set, one can focus on the set of queries with a goal to gain competitive advantage.
Other goals can also dictate the choice of sample. For instance, in industry one often targets a specific query segment (e.g., queries with fresh or local search intent); or perhaps on \textit{hard} queries where there is more room for improvement and metrics are more sensitive. In these cases sampling from the particular segment maximizes the evaluation efficiency.
\paragraph{Characteristics of Search Traffic} The characteristics of search traffic also needs to be considered. \cite{Baeza-Yates:2015} shows that web search query logs follow a power distribution, with longer tails. He suggests a sampling technique to generate a sample that follows this distribution. The main idea is to bin the queries based on the frequency, which allows the sampled queries to match the distribution of original query set.
Generally speaking, this is a form of stratified sampling using the query frequency as criteria, which can be extended to another query characteristics such as location, time and user cohorts. For instance, one can imagine binning the traffic by location and then sample equally across the bins to get the sample of traffic with balanced geographic representation.
%\paul{so, stratified and re-balanced?}
%\jin{more details?}
\paragraph{Query vs. Task Description} While it is possible to ask judges to imagine a search task given a query, it is open to question whether using a query to represent an information need is worthwhile.\paul{I've tried to rephrase this} Unlike search tasks, which should contain sufficient details of user context and information need, queries in a typical search engine are often in an abbreviated form, ambiguous and/or with typographical errors. \jin{any recommended citation? i.e., \% of queries with errors}
These characteristics of user queries can be a significant source of noise because 1) there can be many query forms for the same information need~\citep{Bailey:2015:UVI}, and 2) inferring true information needs from queries can be hard. On the other hand, \cite{Yilmaz:2014:EID} argued that the choice of intent descriptions can also cause large variability in evaluation results and therefore the judging should be done based on queries.
All in all, despite limitations, user queries are still the most readily available sources of task information, and therefore are widely used for judging search results. One can mitigate the noise and ambiguity of the search query by training judges and presenting possible meanings of the query -- i.e., a SERP from a commercial search engine. Alternatively, generating and using a set of task descriptions corresponding to each query would be an more expensive yet potentially better way to mitigate this ambiguity issue. % We discuss this in detail in Section~\ref{s:judging-context}.
%\mdr{I did not find a discussion of noise and ambiguity in 2.2.1}
\subsection{Summary}
In summary, here is our list of recommendations for collecting tasks for human judgments.
\begin{enumerate}
\item Decide whether to create search task or sample from query logs.
\begin{enumerate}
\item If query logs are available and queries are easy to understand, using query as the search task would be fine.
\item If query logs are available yet queries are not easy to understand, consider using seed queries to generate simulated search tasks.
\item If query logs are not available, consider reaching out potential users to collect search tasks.
\end{enumerate}
\item In sampling queries from logs, use an appropriate sampling strategy.
\begin{enumerate}
\item If the goal is to collect a representative sample of traffic, use random sampling.
\item If the goal is to collect a biased sample of traffic according to certain criteria, use stratified sampling based on that criteria.
\end{enumerate}
\item Always collect metadata (task type, user con text, etc.) along with search tasks to facilitate further analysis. User context can also be presented as a part of the search task.
\end{enumerate}
\section{Designing a Judging Interface}
Once the search tasks are collected, we are ready to design a system to gather judgments. There are several main considerations in designing a judging interface: we cover these in what follows.
\begin{enumerate}
\item How do we describe the context of a search task? \\(user location, preferences, previous queries in the session, etc.)
\item What should be the target of each judgment? \\(webpage, SERP elements or whole SERP)
\item What should be the scale of judgment? \\(absolute vs. relative, numeric scales vs. Likert-type scales vs. magnitude estimation)
\item What are the quality dimensions we want to measure? \\(relevance, usefulness, novelty, trustworthiness, etc.)
\end{enumerate}
\subsection{Judging Context}
\label{s:judging-context}
There are many contextual variables that affect user satisfaction with any given search result: users' knowledge and preference, language, timing and location of the search, just to name a few. Even with well-defined search tasks, it is hard to specify all these factors, let alone with terse keyword queries. Providing some of this contextual information to judges can potentially reduce the user-judge gap, thereby increasing the judgment quality. \paul{refs?} \mdr{Measured how?}
The choice of what context to provide depends again on the evaluation goal -- what do you want judges to know about the search task? For instance, if you think user location is crucial in judging the relevance of results (which is the case in many tasks), you should present the user's location alongside the query text. Note that, if possible, the location information should be collected along with user queries to get a realistic sample of actual user locations.
Relevance judgments are also affected by what user already did during the session, so it is reasonable to present some part of user session as judging context. Several authors have examined this. \cite{Chandar2013} used a document as context, with the goal of collecting judgments when the context document has already been read. They proposed an evaluation framework for novelty and diversity evaluation which captures subtopics implicitly and at finer-grained levels. \cite{Golbus:2014:CDR} also experimented with using a document as a context, and found that the metrics based on conditional judgments correlate better with user preference at SERP-level.
While one may assume that adding more and more context can only increase the quality of judgments by reducing the user-judge gap further, it should be noted that more context means more effort for judges in digesting and applying the information. Moreover, more context can increase judging cost by adding a further source of variability. That is, instead of collecting judgment for every search task, these judgments should now be collected for every query and context pairs, which can potentially make the evaluation prohibitively expensive.
%\paul{but you just suggested sampling e.g. location at the same time; so there'll be a 1:1 mapping query:context. But it's true that if you want to examine the effect of one more variable (market, location, time, \dots\ then you'll need more data)}\jin{Yes, judging based on query+context will add variance, which necessitates more data}
Therefore, one should carefully consider the cost/value trade-off in adding the context to a judging task. As an extreme example, \cite{Mao:2016} used the entire session as a judging context for collecting judgments on usefulness (as opposed to relevance) and found that usefulness metrics show higher inter-assessor agreement and better correlation with task-level satisfaction elicited from actual users. However, since adding the whole session as judging context increase both the effort needed for individual judgment and the number of judgments required, they recommend using usefulness evaluation only for post-hoc analysis of the experiments.
\subsection{Judging Target}
Judging target defines the basic form of judgment (i.e., what to present and how many), and it is the most critical decision as the details of judging interface depends on it. While creating a judging interface, we should also decide the granularity of judgment, and whether the judgment should be given for a single item, or a set of items.
\subsubsection{Judging Unit}
The judging unit is the unit at which judgments should be collected, i.e., at what granularity do we want to collect judgments? In web search, for example, the judging unit can be a webpage, SERP elements or a whole SERP, as shown in Figure \ref{fig:judging_units}.
\begin{figure}
\begin{center}
\includegraphics[scale=0.5]{images/judging_units}
\caption{Various judging units for web search results.}
\label{fig:judging_units}
\end{center}
\end{figure}
The judging unit should be determined by the goal of evaluation: if you care about the quality of a ranked list, collecting judgments for each individual result seems like a natural choice. If the presentation of the whole SERP is a primary concern, the entire SERP might be the right unit to collect judgments at.
On the other hand, if the judging target is reasonably complex with multiple sub-components, it is also possible to collect judgments at smaller units (i.e., SERP elements) and then calculate scores for large unit (i.e., the whole SERP) by combining unit scores in a sensible way. This is how most IR evaluation metrics (i.e., MAP or NDCG) work.
Now, if we want to collect judgments for SERPs, should we collect element-wise judgments and then combine, or collect single SERP-level judgments? This question can be generalized into the decision of judging unit when the judging target is complex. There is no hard and fast rule to determine the right judging unit, but here we describe a few trade-offs.
A smaller judging unit means a simpler judging task, which can be faster and more reliable. However, the number of judgments to evaluate a larger unit (i.e., a SERP) can be quite high if the judging unit is small, making overall judging cost higher than collecting a single judgment for the whole larger unit. Studying this trade-off would be an interesting venue for research.
\paul{reference? or other evidence?}\jin{I don't know of any.}
A smaller judging unit also means better reusability of individual labels, because you can combine labels for each SERP element to evaluate arbitrary configurations (e.g., arbitrary rankings of URLs on a SERP). This means that the cost of collecting judgments can be amortized over multiple experiments. In fact, query-URL relevance judgments have been so widely used in TREC and other settings because it allows the creation of test collection which can be used to evaluate any ranked list.
On the other hand, using a smaller judging unit makes an assumption that each label can be collected independent of other elements -- for example, that the quality of an item at rank~2 on a SERP can be assessed without knowing anything about ranks 1 or~3. This is hardly true in a typical search scenario where the concept and criteria of relevance can evolve over time. In this regards, larger judging units have the benefit of providing rich context for judges.
More importantly, larger judging units can capture various set-level properties -- including the comprehensiveness, redundancy between elements. For instance, SERP-level judging can reveal whether the SERP captures all the reasonable intents for a given query. Also, the redundancy among documents in a ranked list can be captured only at the list-level.
In literature, as briefly mentioned above, document-level judgment has been most prevalent. However, there has been some works that focus work SERP-level evaluation. \cite{Bailey2010} introduce a judgment scheme which can capture the interaction among SERP elements as well as element-level quality.
SERP-level judgments were introduced by \cite{Thomas2006}, who propose a pairwise judging interface in order to minimize the complexity of defining judging criteria (more about this in the following section). Several other works including \cite{Kim:2013} refined this idea to include dimensional relevance judgments as well as overall SERP-level comparison.
% \cite{Al-Maskari2007} and
\paul{I will add work by Falk et al.\ on judging snippets}
\subsubsection{Absolute vs. Relative Judgments}
Another consideration in determining a judging target is the type of judgment, which can be either absolute or relative. In absolute judging, judgments are collected for a single judging target, whereas relative judgment asks for a pairwise preference between two targets. Figure~\ref{fig:judgment_types} shows the two types of judgments in evaluating web search results.
\begin{figure}
\begin{center}
\includegraphics[scale=0.5]{images/judgment_types}
\caption{Absolute vs. relative judgments.}
\label{fig:judgment_types}
\end{center}
\end{figure}
%\paul{in Figure~\ref{fig:judgment_types}, ``compare THE two SETS OF results''?}
%\emine{In Figure 2.3 here we show a ranked list of results and ask the user how they rate the search result, which is confusing. I think this should either be individual document or ask a different question for the whole page}
Now, how should one choose between absolute and relative judgments? In general, absolute judging requires objective criteria to distinguish amongst different levels, whereas relative judgments can avoid the issue. \cite{CarteretteBCD08} have also suggested that relative judgments tend to be more accurate for document-level judging, while \cite{Kazai:2013} found that a pairwise judging interface improves crowdsourcing quality so that it can be on par with that of trained judges.
Relative judgments have been used in various evaluation settings. \cite{Chandar2013} employed document-level pairwise judging using another document as a context, to evaluate novelty and diversity. \cite{Arguello:2011} proposed an evaluation scheme for aggregated search based on pairwise preference judgment at element level, and \cite{Zhou:2012} used SERP-level pairwise preference judgments as part of the evaluation framework for aggregated search.
On the other hand, the number of relative judgments grows with the square of the number of items. Since preferences may be weak, and may also be nontransitive, in principle each possible pair needs to be labeled. \cite{CarteretteBCD08} On the other hand, absolute judgments are reusable in that you can compare among any items for which you have item-level labels. Therefore, if you want to reuse judgments in an environment where multiple generations of ranking techniques should be compared against each other, absolute judgments may save cost in the long run. This is also the reason that TREC has employed absolute judgment since its inception.\cite{}
%\subsection{Scales}
\paul{I'll add notes on: different types of scales, e.g. Likert-type vs numeric, Falk's work on magnitude estimation, IIiX paper on semantic differentials?}
\paul{I'll add notes on: Diane et al.\ on the effect of question mode? Can't remember if this is relevant}
\subsection{Judging Criteria}
The central assumption of offline evaluation is that human judges can represent real users, and we often want judges to tell us if the judging target would be relevant to the potential user. \mdr{Should relevance be the core criterion here? Why not "utility" (see eg Belkin).} However, this is not a trivial task for judges given the contextual and multi-faceted nature of relevance \citep{Borlund:2003}, and for example \cite{Chouldechova:2013} report increased judging quality when done by query owners (users who did the search themselves) compared to query non-owners.
Also, while the concept of relevance is broad, it typically specifies the relationship between an information need and an object, and is not sufficient to capture the true value of the item in the context of a search session. Therefore, it has been argued that IR as a field should move beyond relevance to evaluate usefulness in the context of search tasks \citep{Belkin:2015:SAL}. The TREC Session track \citep{carterette2014overview} and TREC Task Track \cite{yilmaz2015overview} is another movement in the same spirit.
%\emine{Is there a reason why we used session track here but not tasks track? Tasks track used usefulness based judgments and focuses on tasks}
%\jin{Relavance seems to subsume usefulness according to Borlund:2003. But Belkin:2015 seems to use a narrow definition of relevance.}
Recent work has tried to address this problem from multiple angles. The role of user effort and effort-based judging has been proposed \citep{Yilmaz:2014,VermaYC16}, where it is shown that effort should be incorporated as an additional factor in human judgment to build retrieval systems that optimize user satisfaction. \cite{Carterette:2011:SEU} also analyzed existing evaluation metrics from the view of expected utility and expected efforts. \cite{Golbus:2014:CDR} and \cite{Kim:2013} also experimented with multi-dimensional judgment collection, which is useful in finding the relationship between different aspects of relevance.
Another thread of work looked at relevance judgments in the context of other items, or even the whole session. \cite{Chandar2013} proposed judging methods for novelty and diversity, where they employed preference-based judgment between document A and B in the context of a third document (C). The resulting method has the benefit of allowing the evaluation of novelty and diversity without requiring the collection of sub-topical judgments.
\cite{Mao:2016} proposed collecting usefulness judgment in the context of whole session. They showed that high relevance by assessors is a necessary but not sufficient condition for high usefulness for users, and that usefulness judgments better correlate with behavioral signals such as click cumulative gains. But since usefulness judgments are costly to collect, they advised only collecting them post-hoc.
Overall, the current literature suggests many ways to set judging criteria for relevance, with different methods having different emphases. If the goal is to focus on query-document relevance, a simple interface as seen at the top of Figure~\ref{fig:judgment_types} will do. However, one can add another document or even whole session history as a context if the goal is to capture the value of the item in the context of a broader search task. \mdr{In terms of practical hands on advice on setting up label collection efforts, this is a bit vague.}
\subsection{Summary}
In summary, here is our list of recommendations for designing an interface for human judgments collection.
\begin{enumerate}
\item Consider presenting each search task with context to reduce variability of the results.
\item Use the smallest judging unit at the beginning to collect fine-grained information with least amount of noise, yet consider collecting more coarse-grained (set-level) judgments as well to capture interactions among items.
\item Use absolute rating scale when it is possible to define clear criteria for each rating. Use relative rating scale otherwise, especially when employing crowd judges.
\item Judging interface design is an iterative process. Test multiple versions with small group of judges before scaling up.
\end{enumerate}
\section{Collecting Judgments}
Once the judging interface is designed, the next task is to find judges to work with. Here we discuss considerations in choosing judge groups. Recently crowdsourcing has become a standard way to collect judgments at large-scale. Since quality control is more challenging when working with crowd judges, we also discuss considerations in crowdsourcing human judgments.
\subsection{Choosing Judge Groups}
There are quite a few options from which you can find judges, but you can roughly put them into four categories: 1) team members who work on the project, 2) expert judges who typically sit in-house with the team, 3) crowd judges who work remotely and can be reached via platforms like Amazon Mechanical Turk, 4) people who actually use the system. %\paul{you're ruling out users themselves? they're discussed a couple of paragraphs below}
How should we decide on which option to choose? First, it is recommended to start some judging exercise with the team (Group 1) before outsourcing the judging task, because you need to make sure you provide high-quality interfaces and descriptions to get judgments of reasonable quality. But this approach soon hits scalability issues\mdr{Please explain. How many judges are needed? You don't say this?}, so we focus on expert judges (Group 2) and crowd judges (Group 3) in this paper.
%\emine{Why do we have to start with the team? Why does it have scalability issues? This part is not clear to me} \paul{I'd always pilot internally first, I think that's all Jin's saying here}\jin{Yes, exactly}
There has been some recent work comparing human judges of different characteristics. \cite{Bailey:2008} is a classic work where they found that judges' level of expertise on the domain can result in small yet consistent difference on system scores and rankings. Similarly, \cite{Chouldechova:2013} looked at judgments done by query owners (users who did the search themselves) vs. query non-owners, where they concluded that query owners are can distinguish a higher quality set of search results from a lower quality set in a blind comparison.
However, neither finding domain experts nor using search tasks from judges themselves are feasible if you need judgments at large scale, or the goal is to collect judgments from representative sample of user traffic. Typically the options available are either in-house judges with some form of training or crowd judges.
Among these groups, \cite{Kazai:2013} found that trained judges are significantly more likely to agree with each other and with users than crowd workers. But when they compared third-party judgments with clicks from real users, they found that the judgments from trained judges does not necessarily show higher agreement with metrics based on user clicks.%\paul{I've paraphrased slightly here, is it still correct?}\jin{Yes}
\mdr{Implications for setting up your own labeling effort?}
\jin{I would say avoid this as much as possible.}
\subsection{Crowdsourcing Relevance Judgments}
\label{s-crowdsourcing}
Collecting labels from humans used to require finding and managing a group of people one by one, which is often an expensive and time-consuming process. Compared to this, crowdsourcing -- hiring subjects from remotely using services such as Amazon Mechanical Turk -- has a clear benefit in cost and scalability, and therefore it has gathered a lot of attention from research community, including a large body of work produced in IR community as well. \cite{Alonso2012} provides a comprehensive survey of research and best practice in this area.
Along with the availability of cheap workforce from across the globe, the challenge in managing the quality of outcome has emerged. A standard approach in reducing errors has been aggregating redundant judgments from a group of independent assessors, and several works has focused on collecting and aggregating redundant labels.
\cite{Venanzi:2014} proposed a community-based Bayesian label aggregation model which is based on finding latent groups among crowd workers and aggregating labels based on them. \cite{Davtyan2015} proposed using textual similarity to aggregate crowd judgments, where the relevance labels from similar documents are propagated. Companies such as Crowdflower\footnote{https://www.crowdflower.com/} provide a service by which high quality labels are automatically calculated based on redundant judgments. They also provide resources on how to design a crowdsourcing task for search relevance judgments.\footnote{How to: Run a Search Relevance Job https://goo.gl/gfsEYi} \mdr{What's the point? How are readers of this survey going to benefit from this comment?} \jin{This can be an example where ideas from research are put into practice}
Another approach to improving the quality of crowdsourced judgments is by improving the judging interface design. This section already dealt with design decisions on judging interface design, and \cite{Kazai2012} provide further guidance in deciding the complexity of judging tasks and the amount of payment per judgment. They recommend 1) pricing the hit according to expected efforts from judges, because both paying too little or much has downsides, 2) reducing the complexity of task such that cheating take approximately the same effort as faithfully completing the tasks, and 3) having multiple ways to detect the quality of the work such as questions with known answers or judges' behavior.
Recently, there has been several proposals regarding how to embed quality control as a part of natural judging workflow. \cite{Alonso:2015} proposed adding a simple tasks which can prepare judges for actual tasks and allow easy validation of the crowd judges' faithfulness at the same time. For instance, in collecting relevance judgments for social media posts, one can ask whether the post contain a person's name, which is an easy task that can be answered only by reading it. Similarly, \cite{mcdonnell2016relevant} suggested simple annotation scheme which can also be useful in results validation. The idea is to ask for annotation of relevant part of the judging target, which forces reading by judges, and then allows simple validation without extra efforts.
As large-scale crowdsourcing has become commonplace in a industry setting, several authors have recently investigated workflow design for crowdsourcing. At microscopic level, \cite{Scholer:2013} and \cite{Shokouhi:2015} looked at the effect of previous assessments on the quality of a judgment, and showed that the human annotators are likely to assign different relevance labels to a document depending on the quality of the last document they had judged for the same query. At a macroscopic level, \cite{Megorskaya2015} explored various parameters in designing workflow, and argue for having a communication channel between judges and 3--5-way overlap in a production environment.
\subsection{Summary}
In summary, here is our list of recommendations for collecting human judgments.
\begin{enumerate}
\item Always start the judging task internally and with small number of judges before scaling up to avoid wasting judging efforts.
\item For simpler judging tasks, try crowdsourcing first, along with the following considerations:
\begin{enumerate}
\item Vary the amount of overlap to find the right trade-off between the judging cost and the precision of the outcome.
\item Use a simple interface that takes minimum instructions to use. Set the price of the task to match the expected efforts.
\item Make sure the UI has more than one built-in quality control mechanisms, such as trap questions and malicious behavior detection.
\end{enumerate}
\item For more involved judging tasks, consider hiring in-house judges. Try to hire people with domain expertise to further improve the quality.
\end{enumerate}
\section{Open Issues}
So far in this section, we looked at issues in collecting human judgments, and provided guidance based on latest research. However, search is rapidly evolving and as such new research areas are emerging. Before moving on to the next topic, here we discuss several open issues.
\paragraph{New Judging Targets} Most existing research considers document-level judging. But modern SERPs contain rich results beyond documents, such as instant answers and multimedia results. Extending document-based judging model into these new judging targets would be an interesting problem. This includes judging methods for snippets, instant answers and rich SERPs with all these elements.
\cite{al2010evaluating} investigated several methods for summarizing the contents of a webpage, and found that adding visual summary of a webpage such as thumbnail, sailent image, tag clouds does reduce the time for relevant judgments to be made, yet does not improve the accuracy in doing so. This double-sided effect of visual elements reassure that adding an flashy UI element is not always a good idea.
\paragraph{New Endpoints for Search} Smart phones are becoming standard devices for accessing the internet;\footnote{http://a16z.com/2014/10/28/mobile-is-eating-the-world/} and recently conversational agents have become a major focus for many tech companies. We are yet to learn how these new environments can affect judgment collection, yet changes in device size (smaller) and interaction modality (from click to touch and voice) is sure to change how to evaluate the quality of search results. Recent work such provide some hints at what needs to change for these new environments.
\cite{VermaY16} investigated the difference in relevance judgment collection between mobile and desktop interfaces, and found that judging time in mobile documents is higher than in desktop, which is somewhat counterintuitive given the small screen size in mobile environment. They also found that viewport features are useful in predicting relevance judgments in mobile environment, where smaller screen size makes it easier to pinpoint where the user is reading. \cite{Kiseleva:2016} focused on building a predictive model of success in conversational setting with new featues such as voice and touch interaction features, and found that dialogue-style interaction necessitates task-level modeling as opposed to query-level modeling.
\mdr{This is too short/abstract to be meaningful.} \jin{Details added}
\paragraph{New Judging Methods for Personalized Search} Standard judging methods collect labels given a search task and a single, or a pair of, search results. However, this model may not work in environments where search is highly contextual and personal, such as in searching with conversational agents. Several recent works such as those by \cite{Xu:2009} and \cite{Moraveji:2011} explored task-based judgment collection, where judges perform search given a (possibly personalised) search engine to make their judgments.
While allowing judges to perform searches themselves certainly allows more degrees of freedom for judges in evaluating given search engine, this in turn adds variability in outcome and careful experiment design is required. In comparing two search engines using the task-based judging method, \cite{Xu:2009} proposes a cross-over design that balances the assignment of search engine and tasks across judges. This reduces the variance of estimated delta between two engines by an order of magnitude.
\mdr{Make this a useful bit of information for your readers.}\jin{Details added}
%\paul{do we want to also mention synthetic collections, e.g.~Jin's PhD work, Leif et al.'s synthetic queries?} \jin{Feel free to add, although we're focusing on web search evaluation so not sure if they're relevant}
\paragraph{Closing Remarks}
In this chapter we discussed human judgments collection: how to collect search tasks, design a judging interface, and hiring and managing judges to actually collect judgments. Despite recent changes in search user interfaces, basics learned in this chapter would still be useful as guiding various decisions in designing and collecting human judgments. In subsequent chapters, the judgments collected will be used to calculate metrics and draw conclusions at the level of experiments.
\mdr{Missing: a look ahead, i.e., a statement on how the choices made in this chapter (setting up the label collection) affects the next two stages in the offline evaluation pipeline, and vice versa.}
\jin{Added}