# #325 Phelim Brady: Why AI's Future Depends on Human Judgement Page: https://stenobird.com/podcast/eye-on-ai/325-phelim-brady-why-ai-s-future-depends-on-human-judgement Text version: https://stenobird.com/podcast/eye-on-ai/325-phelim-brady-why-ai-s-future-depends-on-human-judgement.md Podcast: [Eye On A.I.](https://stenobird.com/podcast/eye-on-ai) Published: 2026-03-09T20:00:00+00:00 Episode link: https://aneyeonai.libsyn.com/325-phelim-brady-why-ais-future-depends-on-human-judgement Audio file: https://dts.podtrac.com/redirect.mp3/pscrb.fm/rss/p/pdst.fm/e/arttrk.com/p/VI4CS/prfx.byspotify.com/e/mgln.ai/e/p58841/pscrb.fm/rss/p/traffic.libsyn.com/secure/aneyeonai/Sequence_21_1.mp3?dest-id=727317 Processing state: not_requested JSON: https://stenobird.com/v1/public/podcasts/eye-on-ai/episodes/325-phelim-brady-why-ai-s-future-depends-on-human-judgement Duration seconds: 2836 ## Resource AI often looks fully automated. But behind the scenes, a huge amount of human judgment is shaping how these systems actually work. In this episode, Craig Smith speaks with Phelim Bradley, co-founder and CEO of Prolific, a platform that connects millions of real people with researchers and AI labs to evaluate and improve AI systems. They explore the hidden human layer behind modern AI, why traditional benchmarks are becoming less reliable, and why AI companies increasingly rely on real human feedback to measure model performance in the real world. Phelim also explains how demographic differences influence how models are evaluated, why human judgment remains critical even as AI improves, and how the collaboration between humans and AI will shape the next phase of development. This conversation reveals the human backbone behind today's AI systems. Stay Updated: Craig Smith on X: https://x.com/craigss Eye on A.I. on X: https://x.com/EyeOn_AI (00:00) Preview and Intro (02:45) Founding Prolific And Early Pain Points (06:30) From Mechanical Turk To Representativeness (09:55) Academic Research And AI Use Cases Split (13:40) Vetting Real Participants And Fighting Fraud (17:45) Scale, Community Growth, And Talent Mix (22:00) High-Complexity Projects Over Commoditised Labeling (26:40) Measuring Model Persuasion With Live Conversations (30:20) Demographic-Aware Model Preference Benchmarks (34:10) The Rise Of Human Evaluation Over Benchmarks (38:00) Enterprise Model Choice And Continuous Evaluation (42:00) Why Humans Won't Disappear From The Loop ## Actions - request_transcript: `POST https://stenobird.com/v1/public/podcasts/eye-on-ai/episodes/325-phelim-brady-why-ai-s-future-depends-on-human-judgement/transcription-requests` — Idempotently request low-priority transcript generation for this episode. - read_markdown: `GET https://stenobird.com/podcast/eye-on-ai/325-phelim-brady-why-ai-s-future-depends-on-human-judgement.md` — Read the agent-friendly Markdown representation of this episode resource. A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed. ## Transcript Full transcripts are not published on public pages unless there is a clear rights basis.