How to be 'Choosy': Wrangling Big Datasets
“How to be Choosy” provides a framework for reducing the size of very large datasets (with too many rows of columns) according to pedagogical goals. It describes technical methods for reducing dataset size using CODAP or computational notebooks (Jupyter or CoLab), provides interactive templates and practice activites, and offers a DIY guide to help educators and curriculum designers make large datasets manageable for classroom instruction.
Intended Audience & Use Case: These resources are designed for K-12 and undergraduate statistics and data science instructors, as well as curriculum designers. Though some of the ideas and activities might be useful for students as well, we recommend providing them with additional scaffolding and support than is available here.
Available Resources
- Billboard Hot 100 Wrangling Demo (Jupyter / CODAP)
- EPA Toxic Release Inventory Wrangling Demo (Jupyter / CODAP)
- How to be ‘Choosy’ Quick Reference Guide (PDF)
- See everything on Github
Associated Publications The collection accompanies the Teaching Statistics publication by Wilkerson et al. (2025) called How to be Choosy: Wrangling Big Datasets for the Classroom.