Skip to content

Archive · 2012–2014

Crowdsourcing and Scalable Active Learning

Using active learning to make large-scale crowdsourced data acquisition more efficient.

How can active learning focus human effort when crowdsourced data acquisition scales to millions of tasks?

Research on using active learning to reduce the cost of large-scale crowdsourced data acquisition.

The problem

Crowdsourcing can become prohibitively expensive when every item receives the same amount of human attention.

The approach

Use active learning to identify which labels or tasks are most informative and allocate crowd effort selectively.

Research on using active learning to reduce the cost of large-scale crowdsourced data acquisition.

Main contributions

  • Research framing and system design
  • Methods, implementation, and empirical evaluation
  • Open research artifacts and scholarly dissemination

Publications

Active Learning for Crowd-Sourced Databases

Barzan Mozafari, Purnamrita Sarkar, Michael J. Franklin, Michael I. Jordan, Samuel Madden

CoRR 2012 · Computing Research Repository

Project ↗
Cite
@article{0cebd645-dd1c-4dc8-9f14-a20d779e8067,
  title = {Active Learning for Crowd-Sourced Databases},
  author = {Barzan Mozafari and Purnamrita Sarkar and Michael J. Franklin and Michael I. Jordan and Samuel Madden},
  journal = {Computing Research Repository},
  year = {2012}
}

Scaling Up Crowd-Sourcing to Very Large Datasets: A Case for Active Learning

Barzan Mozafari, Purnamrita Sarkar, Michael J. Franklin, Michael I. Jordan, Samuel Madden

PVLDB 2014 · Proceedings of the VLDB Endowment

Paper ↗Data ↗Project ↗
Cite
@article{d44b8133-c57f-42fc-8bcf-5b45a5abcdfc,
  title = {Scaling Up Crowd-Sourcing to Very Large Datasets: A Case for Active Learning},
  author = {Barzan Mozafari and Purnamrita Sarkar and Michael J. Franklin and Michael I. Jordan and Samuel Madden},
  journal = {Proceedings of the VLDB Endowment},
  year = {2014},
  doi = {10.14778/2735471.2735474}
}