Saraschandra KaranamHuman - AI Interactions Researcher · Enterprise AI
Back to Research
RESEARCH PROGRAM • 2012-14

Human Computation / Crowdsourcing

Benchmarking of crowdworker's performance mapping task quality and efficiency on crowdsourcing platforms leading to development of a recommendation engine - CrowdUtility

Task <Description ; QoS ; $ Budget>Digitization?Translation?Image Labeling?CROWD UTILITYPlatform StatisticsPlatform ModelsRecommendation Engine1234Cost/UnitResponse TimeAccuracyAmazon MTurkCloud FactoryCrowd FlowerMobileWorksOther Platforms
Representative Visualization

CrowdUtility Recommendation Engine

Architectural diagram of CrowdUtility - a recommendation engine for crowdsourcing platforms.

1 PoC
2 Platforms(Crowdflower, MobileWorks)
5 PatentsGranted
7 Publications(ICML, HCOMP, CHI, ACM WebScience)
02 · WHY THIS RESEARCH MATTERED

Efficiency & Quality Variance

Crowd workers (unlike in a typical organization), exhibit varying work patterns, expertise, and performance - with little or no control that can be imposed on them. Requesters (e.g. enterprises) also exhibit diverse requirements in terms of the size, complexity and timings of the tasks, as well as SLAs (performance expectations). Clearly, the heterogeneity makes the choice of a platform suited for a given task difficult for the user.

Platform 1Platform 203060901201500h4h8h12h16h20h24hCompletion Time (s)Time of Day (24-Hour Cycle)
Mean Task Completion Time across multiple hours for Platform 1 and Platform 2
RESEARCH QUESTION

"How can business owners decide which crowdsourcing platform to choose to meet enterprise SLAs?"

Focusing only on Externally Observable Characteristics (EOCs), without any internal knowledge of the platforms, this research built an end-to-end recommendation engine starting from requirements specification, to platform recommendation to execution of the tasks.

Stream 01

Benchmarking

Collect performance data of crowdworkers

Stream 02

Build models

Create statistical models that characterize each platform over time

Stream 03

Recommendation

Compute recommendations based on input task requirements and existing behavior models

Stream 04

Dispatch Tasks

Send tasks to the recommended platform for execution.

Stream 05

Update models

On completion of tasks, new performance data is fed back to update the models in real-time.

03 · EMPIRICAL DEVELOPMENT

Research Program

Connected studies conducted progressively answered different aspects of the central research question.

Study 01

Benchmarking Crowdworker Performance

2012 - 2013
Question

"How does crowdworkers performance vary across time of the day, day of the week, geography, cost, number of tasks, task complexity?"

Methods
Benchmarking
Outcome

Baseline crowdworker performance

Study 02

Development of CrowdUtility recommendation engine

2013 - 2014
Question

"How to build a recommendation engine that dynamically assigns tasks to platforms based on cost, accuracy, and deadline constraints?"

Methods
Machine LearningSimulationTask Scheduling
Outcome

CrowdUtility Recommendation Engine, Task Scheduler

04 · REPRESENTATIVE CASE STUDY

Understanding Crowd Workers Dynamic Performance Variability across Crowdsourcing Platforms

Background

Crowd workers exhibit varying work patterns, expertise, and quality - which, in turn, lead to wide variability in the performance of platforms. Very little work, however, has been done in terms of understanding these variations and leveraging the information for better selection of platforms.

Research Questions

  1. Identify parameters along which two crowdsourcing platform's performance differ?
  2. Understand the patterns / interactions across these parameters.

Research Approach

Benchmarking
(a)Mean Task-Completion TimeMean Task Accuracy200.82220.84240.86260.89280.91300.93320.95340.98361.001234Mean Task-Completion Time (minutes)Mean Task AccuracyCost (number of Cents)
Mean Task Completion Time and Mean Task Accuracy as incentive increases

Key Findings

01

Geography

Crowdsourcing platforms varied in the availability of workers across geographies

02

Day of the Week

One platform was faster on weekdays while the other was faster on weekends

03

Incentive

As payment increases, we observed faster task completion times - at the cost of accuracy.

04

Task Complexity

As task complexity increased, task accuracy dropped on both crowdsourcing platforms

Research Contribution

Baseline studies on worker performance benchmarking directly informed the recommendation engine parameters developed in the second phase.

05 · PROGRAM SYNTHESIS

What We Learned

Learnings from this program

Externally Observable Characteristics

EOCs were sufficient to model crowdsourcing platform performance characteristics

06 · RESEARCH ASSETS

Research Artifacts

Empirical assets and frameworks generated to guide future enterprise-wide design and engineering direction.

Benchmarking data

Performance data of two crowdsourcing platforms

Classification models

Characterizing the performance of the two crowdsourcing platforms

Task Schedulers

That decide the batch size and schedule tasks

Proof of Concept

A working prototype of the end to end CrowdUtility system

07 · PROGRAM PORTFOLIO

Additional Studies in this Research Program

This representative deep dive is one part of a much larger program of research.

Study 01

CrowdUtility: A Recommendation System for Crowdsourcing Platforms

Study 02

Task Scheduling algorithms

08 · PROGRAM OUTCOMES

Research Impact

The structural, organizational, and methodological contributions generated by this program.

Product Contributions

  • A working prototype of the end to end CrowdUtility system

Organizational Alignment

  • PoC built and tested with a client
  • 4 patents
  • 7 publications

Methodological Value

  • Using externally observable characteristics to model
  • Task schedulers