Episode

"How good are slop-vestigators?" by Hasan Baig, OscarGilg, Hamzah

Podcast
LessWrong (Curated & Popular)
Published
Sep 9, 2026
Duration seconds
824
Processing state
not_requested
Canonical source
https://www.buzzsprout.com/2037297/episodes/19781284-how-good-are-slop-vestigators-by-hasan-baig-oscargilg-hamzah.mp3
Audio
https://www.buzzsprout.com/2037297/episodes/19781284-how-good-are-slop-vestigators-by-hasan-baig-oscargilg-hamzah.mp3
JSON
/v1/public/podcasts/lesswrong-curated-popular-5643401/episodes/how-good-are-slop-vestigators-by-hasan-baig-oscargilg-hamzah
Markdown
/podcast/lesswrong-curated-popular-5643401/how-good-are-slop-vestigators-by-hasan-baig-oscargilg-hamzah.md

Actions

  • POST https://stenobird.com/v1/public/podcasts/lesswrong-curated-popular-5643401/episodes/how-good-are-slop-vestigators-by-hasan-baig-oscargilg-hamzah/transcription-requests
    Idempotently request low-priority transcript generation for this episode.
  • GET https://stenobird.com/podcast/lesswrong-curated-popular-5643401/how-good-are-slop-vestigators-by-hasan-baig-oscargilg-hamzah.md
    Read the agent-friendly Markdown representation of this episode resource.

Summary

TLDR: We release MessageBoardAuditBench: a benchmark to measure how well agents can replicate the recent investigation into a swarm of OpenAI agents colluding via a message board on an online wiki. We open-source the benchmark as an Inspect eval.We find that top models cover up to 51% of findings under our rubric and that model performance improves with time budget and general capability.We observe OpenAI models are less likely than other models to suggest the incident came from an interna...