Episode
Grok 4.5's Split Verdict: Best Agentic Tool Use on the Board, a Doubled Hallucination Rate, and the New Math of Model Trust - July 11, 2026
- Published
- Jul 11, 2026
- Duration seconds
- 905
- Processing state
not_requested
Actions
POST https://stenobird.com/v1/public/podcasts/dx-today-no-hype-podcast-news-about-ai-dx-6434212/episodes/grok-4-5-s-split-verdict-best-agentic-tool-use-on-the-board-a-doubled-hallucination-rate-and-the-new-math-of-model-trust-july-11-2026/transcription-requests
Idempotently request low-priority transcript generation for this episode.GET https://stenobird.com/podcast/dx-today-no-hype-podcast-news-about-ai-dx-6434212/grok-4-5-s-split-verdict-best-agentic-tool-use-on-the-board-a-doubled-hallucination-rate-and-the-new-math-of-model-trust-july-11-2026.md
Read the agent-friendly Markdown representation of this episode resource.
Summary
Grok 4.5's Split Verdict: Best Agentic Tool Use on the Board, a Doubled Hallucination Rate, and the New Math of Model Trust Grok 4.5's first independent report card is in, and it is split down the middle: the best agentic tool use score on the Artificial Analysis board and a number one claim on the SWE marathon, alongside a hallucination rate that doubled from 25 percent to 54 percent. Chris and Laura unpack how a model can know more yet be wrong more often, why grounded professional work a...