Episode
Jack Cushman: Inside Harvard’s Data.gov Archive (EP 3)
- Podcast
- PodQueue
- Published
- Nov 20, 2025
- Duration seconds
- 4833
- Processing state
not_requested- Canonical source
- https://podqueue.fm/playlist_items/h8ystefekd132c0i9ij7gl58
Actions
POST https://stenobird.com/v1/public/podcasts/podqueue-6583043/episodes/jack-cushman-inside-harvard-s-data-gov-archive-ep-3/transcription-requests
Idempotently request low-priority transcript generation for this episode.GET https://stenobird.com/podcast/podqueue-6583043/jack-cushman-inside-harvard-s-data-gov-archive-ep-3.md
Read the agent-friendly Markdown representation of this episode resource.
Summary
Join us for a conversation with Jack Cushman from the Harvard Law School Library Innovation Lab about their new archive of Data.gov—more than 311,000 datasets harvested in 2024–2025, updated daily, and published on Source Cooperative. We’ll dig into two threads:- BagIt for durability: How Library of Congress–standard packaging, checksums, and signatures support authenticity, provenance, and long-term citation.- Discovery without a server: how browser-based querying over static data makes 17.9 TB of datasets findable and fast to explore. We’ll also talk about practical choices that matter when you’re archiving government data: what to bag, what metadata to preserve, how to track change over time, and how to make it usable for researchers, journalists, and agencies.