# Using the Smartest AI to Rate Other AI Page: https://stenobird.com/podcast/unsupervised-learning/using-the-smartest-ai-to-rate-other-ai Text version: https://stenobird.com/podcast/unsupervised-learning/using-the-smartest-ai-to-rate-other-ai.md Podcast: [Unsupervised Learning](https://stenobird.com/podcast/unsupervised-learning) Published: 2025-04-19T06:00:00+00:00 Episode link: https://omny.fm/shows/unsupervised-learning/using-the-smartest-ai-to-rate-other-ai Audio file: https://mgln.ai/e/p5837/pscrb.fm/rss/p/traffic.omny.fm/d/clips/070af456-729b-4a0f-9c09-a6c100397b59/3b159371-276d-429e-ae86-a6c1003b01c4/8a5bf7e7-84d2-4b9b-9080-b2c1013e0fd5/audio.mp3?utm_source=Podcast&in_playlist=7b61d4e1-bd3d-4d3f-97c2-a6c1003b01c9 Processing state: failed JSON: https://stenobird.com/v1/public/podcasts/unsupervised-learning/episodes/using-the-smartest-ai-to-rate-other-ai Duration seconds: 575 ## Resource In this episode, I walk through a Fabric Pattern that assesses how well a given model does on a task relative to humans. This system uses your smartest AI model to evaluate the performance of other AIs—by scoring them across a range of tasks and comparing them to human intelligence levels. I talk about: 1. Using One AI to Evaluate Another The core idea is simple: use your most capable model (like Claude 3 Opus or GPT-4) to judge the outputs of another model (like GPT-3.5 or Haiku) against a task and input. This gives you a way to benchmark quality without manual review. 2. A Human-Centric Grading System Models are scored on a human scale—from “uneducated” and “high school” up to “PhD” and “world-class human.” Stronger models consistently rate higher, while weaker ones rank lower—just as expected. 3. Custom Prompts That Push for Deeper Evaluation The rating prompt includes instructions to emulate a 16,000+ dimensional scoring system, using expert-level heuristics and attention to nuance. The system also asks the evaluator to describe what would have been required to score higher, making this a meta-feedback loop for improving future performance. Note: This episode was recorded a few months ago, so the AI models mentioned may not be the latest—but the framework and methodology still work perfectly with current models. Subscribe to the newsletter at: https://danielmiessler.com/subscribe Join the UL community at: https://danielmiessler.com/upgrade Follow on X: https://x.com/danielmiessler Follow on LinkedIn: https://www.linkedin.com/in/danielmiessler See you in the next one! Become a Member: https://danielmiessler.com/upgrade See omnystudio.com/listener for privacy information. ## Actions - request_transcript: `POST https://stenobird.com/v1/public/podcasts/unsupervised-learning/episodes/using-the-smartest-ai-to-rate-other-ai/transcription-requests` — Idempotently request low-priority transcript generation for this episode. - read_markdown: `GET https://stenobird.com/podcast/unsupervised-learning/using-the-smartest-ai-to-rate-other-ai.md` — Read the agent-friendly Markdown representation of this episode resource. A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed. ## Transcript Full transcripts are not published on public pages unless there is a clear rights basis.