Improving big citizen science data: Moving beyond haphazard sampling.

Corey T CallaghanJodi J L RowleyWilliam K Cornwell Alistair G B Poore Richard E Major

Published in: PLoS biology (2019)

Citizen science is mainstream: millions of people contribute data to a growing array of citizen science projects annually, forming massive datasets that will drive research for years to come. Many citizen science projects implement a "leaderboard" framework, ranking the contributions based on number of records or species, encouraging further participation. But is every data point equally "valuable?" Citizen scientists collect data with distinct spatial and temporal biases, leading to unfortunate gaps and redundancies, which create statistical and informational problems for downstream analyses. Up to this point, the haphazard structure of the data has been seen as an unfortunate but unchangeable aspect of citizen science data. However, we argue here that this issue can actually be addressed: we provide a very simple, tractable framework that could be adapted by broadscale citizen science projects to allow citizen scientists to optimize the marginal value of their efforts, increasing the overall collective knowledge.

Keyphrases