NICAR 2015: Data from scratch — How to crowdsource data

We know data tells us a lot. We write programs to automate data scraping. We spend hours creating data visualizations that help readers see what they need to see. We use data to make claims and generate stories that are reliable and have impact.

Data is important and we seem to be surrounded by it. But that's not quite true. Sometimes, there is no data?

A session at NICAR that really resonated with me was Data from Scratch: When data doesn’t exist, led by Griff Palmer, Ricardo Brom and Lisa Pickoff-White. Pickoff-White shared her experience building PriceCheck, a crowdsourced project that KQED launched last year to answer the question “How much does health care cost?”

The team wanted to compare and contrast the costs of certain procedures or services with and without insurance. The biggest problem was that contracts between patients and their insurance providers were confidential, so no one could get to the information. Except the patient, that is.

So they crowdsourced for data. KQED was able to get hundreds of users to enter their personal information about insurance benefits (see above) onto the site and then create a database to search for procedures and their respective costs. For example, a particular chest x-ray within 50 miles of my own hometown in the Bay Area costs $107, but one patient paid $36 because her insurance was able to cover the rest.

Strategies

For this project, Pickoff-White mentioned specific strategies they used to make sure they got enough accurate and useful data to “make apples-to-apples comparisons.” Here are seven strategies and tactics she used to get clean data that you can apply to your own projects:

Get the ball rolling. This is pretty simple. The team used social media and their on-air broadcast presence to take advantage of the trust that people already had in KQED.
Report during the process. Here is a list of stories the team wrote when they found particularly large disparities in costs with and without insurance. This kept their information relevant and allowed them to push updates to users as they continued to ask for data.
Ask for one thing at a time. For instance, they would push a request for information on chest x-ray costs and then later another on mammograms. By splitting up the procedures they were asking for, they could target specific people in a wide range of patients and get all the information on one thing at a time.

Do the hard work for the users. The team found that users were making mistakes when filling out the survey on their health benefits, so they implemented autocompleting for procedures they looked for and used Google places to standardize the input of the medical care providers.
Explain benefits to users. A lot of times, they found that patients didn’t know how to properly read their benefits. This was an opportunity for them to help their target audience learn and contribute their information at the same time.
Use common sense. The data reporters had a general understanding about health benefits and costs, so they were able to pick out careless mistakes. For example, Pickoff-White noticed that $3592.50 was unusually steep for a certain procedure. She contacted the user, who corrected the mistake to $359.25. This led to the next strategy in which they would...
Ask for contact information. They included a space for the user to enter in an email address, to correct errors exactly like the one above.

“Some data is better than no data”

Pickoff-White admits it’s hard to determine how much data is enough data to mean something. But because of the specificity of each patient’s experience, the database displays all cases separately which reflects the transparency they aim for. It’s main function is so that someone can search a database and see information on and cost disparities for someone else in a similar situation. Their goal is to get as much data as possible but not necessarily to generalize every patient’s experience into one whopping conclusion.

Here is the powerpoint from the rest of the Data from Scratch session at NICAR.

About the author

Ashley Wu

Undergraduate Fellow

Designing, developing and studying journalism at Northwestern. Also constantly scouting the campus for free food.

Tagged

crowdsource Data Ricardo Brom Griff Palmer Lisa Pickoff-White KQED

Latest Posts

Lab , projects | Oct 6, 2023

A Big Change That Will Probably Affect Your Storymaps

by Joe Germuska | joegermuska

A big change is coming to StoryMapJS, and it will affect many, if not most existing storymaps. When making a storymap, one way to set a style and tone for your project is to set the "map type," also known as the "basemap." When we launched StoryMapJS, it included options for a few basemaps created by Stamen Design. These included the "watercolor" style, as well as the default style for new storymaps, "Toner Lite." Stamen...

Continue Reading
People | Jan 31, 2023

Introducing AmyJo Brown, Knight Lab Professional Fellow

AmyJo Brown, a veteran journalist passionate about supporting and reshaping local political journalism and who it engages, has joined the Knight Lab as a 2022-2023 professional fellow. Her focus is on building The Public Ledger, a data tool structured from local campaign finance data that is designed to track connections and make local political relationships – and their influence – more visible. “Campaign finance data has more stories to tell – if we follow the...

Continue Reading
Ideas | May 31, 2022

Interactive Entertainment: How UX Design Shapes Streaming Platforms

by Max Johnson

As streaming develops into the latest age of entertainment, how are interfaces and layouts being designed to prioritize user experience and accessibility? The Covid-19 pandemic accelerated streaming services becoming the dominant form of entertainment. There are a handful of new platforms, each with thousands of hours of content, but not much change or differentiation in the user journeys. For the most part, everywhere from Netflix to illegal streaming platforms use similar video streaming UX standards, and...

Continue Reading
Lab projects | Dec 13, 2021

Innovation with collaborationExperimenting with AI and investigative journalism in the Americas.

by Mago Torres | magiccia

Lee este artículo en español. How might we use AI technologies to innovate newsgathering and investigative reporting techniques? This was the question we posed to a group of seven newsrooms in Latin America and the US as part of the Americas Cohort during the 2021 JournalismAI Collab Challenges. The Collab is an initiative that brings together media organizations to experiment with AI technologies and journalism. This year, JournalismAI, a project of Polis, the journalism think-tank at...

Continue Reading
Lab projects , En Español | Dec 13, 2021

Innovación con colaboraciónCuando el periodismo de investigación experimenta con inteligencia artificial.

by Mago Torres | magiccia

Read this article in English. ¿Cómo podemos usar la inteligencia artificial para innovar las técnicas de reporteo y de periodismo de investigación? Esta es la pregunta que convocó a un grupo de siete organizaciones periodísticas en América Latina y Estados Unidos, el grupo de las Américas del 2021 JournalismAI Collab Challenges. Esta iniciativa de colaboración reúne a medios para experimentar con inteligencia artificial y periodismo. Este año, JournalismAI, un proyecto de Polis, la think-tank de periodismo...

Continue Reading
Studio | Journalism AI Readiness Scorecard

AI, Automation, and Newsrooms: Finding Fitting Tools for Your Organization

by Hannah Barton , Helen Bradshaw , Joshua Hoeflich , Grace Lee , Sammie Pyo

If you’d like to use technology to make your newsroom more efficient, you’ve come to the right place. Tools exist that can help you find news, manage your work in progress, and distribute your content more effectively than ever before, and we’re here to help you find the ones that are right for you. As part of the Knight Foundation’s AI for Local News program, we worked with the Associated Press to interview dozens of......

Continue Reading