Episode

Jack Cushman: Inside Harvard’s Data.gov Archive (EP 3)

Podcast
PodQueue
Published
Nov 20, 2025
Duration seconds
4833
Processing state
not_requested
Canonical source
https://podqueue.fm/playlist_items/h8ystefekd132c0i9ij7gl58
Audio
https://podqueue.fm/proxy/lyJtf2d_LXAmc1NfimVDsg.mp3
JSON
/v1/public/podcasts/podqueue-6583043/episodes/jack-cushman-inside-harvard-s-data-gov-archive-ep-3
Markdown
/podcast/podqueue-6583043/jack-cushman-inside-harvard-s-data-gov-archive-ep-3.md

Actions

  • POST https://stenobird.com/v1/public/podcasts/podqueue-6583043/episodes/jack-cushman-inside-harvard-s-data-gov-archive-ep-3/transcription-requests
    Idempotently request low-priority transcript generation for this episode.
  • GET https://stenobird.com/podcast/podqueue-6583043/jack-cushman-inside-harvard-s-data-gov-archive-ep-3.md
    Read the agent-friendly Markdown representation of this episode resource.

Summary

Join us for a conversation with Jack Cushman from the Harvard Law School Library Innovation Lab about their new archive of Data.gov—more than 311,000 datasets harvested in 2024–2025, updated daily, and published on Source Cooperative. We’ll dig into two threads:- BagIt for durability: How Library of Congress–standard packaging, checksums, and signatures support authenticity, provenance, and long-term citation.- Discovery without a server: how browser-based querying over static data makes 17.9 TB of datasets findable and fast to explore. We’ll also talk about practical choices that matter when you’re archiving government data: what to bag, what metadata to preserve, how to track change over time, and how to make it usable for researchers, journalists, and agencies.