Episode

Beyond Black Box Scores: How Musubi Trains Custom AI for Trust and Safety Teams

Podcast
Just Now Possible
Published
Jun 11, 2026
Duration seconds
4360
Processing state
not_requested
Canonical source
https://justnowpossible.podigee.io/27-trust-and-safety-at-musubi
Audio
https://audio.podigee-cdn.net/2513034-m-bbef551be86993145534d497cf0f2ddb.mp3?source=feed
JSON
/v1/public/podcasts/just-now-possible-7481489/episodes/beyond-black-box-scores-how-musubi-trains-custom-ai-for-trust-and-safety-teams
Markdown
/podcast/just-now-possible-7481489/beyond-black-box-scores-how-musubi-trains-custom-ai-for-trust-and-safety-teams.md

Actions

  • POST https://stenobird.com/v1/public/podcasts/just-now-possible-7481489/episodes/beyond-black-box-scores-how-musubi-trains-custom-ai-for-trust-and-safety-teams/transcription-requests
    Idempotently request low-priority transcript generation for this episode.
  • GET https://stenobird.com/podcast/just-now-possible-7481489/beyond-black-box-scores-how-musubi-trains-custom-ai-for-trust-and-safety-teams.md
    Read the agent-friendly Markdown representation of this episode resource.

Summary

What do you do when off-the-shelf moderation scores aren't good enough—and the alternative is paying human contractors to spend their days reviewing traumatizing content at scale? In this episode of Just Now Possible, Teresa Torres talks with Nikki Marinsek (Data Scientist), Brian McCaffrey (Software Engineer), and Dan Means (Machine Learning Engineer) from Musubi, an AI-native trust and safety toolkit for content platforms. Musubi builds custom-trained ML models and LLM-powered moderation tools that adapt to each platform's unique policies—from dating apps to social networks to AI inference endpoints. They walk through the full journey: training the first prototype on tabular data, discovering their AI was sometimes catching things human moderators missed, and building a policy optimizer that uses agentic flows to help teams iterate on their moderation policies without needing a data scientist in the room. You'll hear how they balance latency, accuracy, and cost for clients handling hundreds of millions of actions per month, why pushing eval tools directly to customers is their core product strategy, and what's next as they build flexible agentic orchestration for non-technical trust and safety teams.