Episode

Grok 4.5's Split Verdict: Best Agentic Tool Use on the Board, a Doubled Hallucination Rate, and the New Math of Model Trust - July 11, 2026

Podcast
DX Today | No-Hype Podcast & News About AI & DX
Published
Jul 11, 2026
Duration seconds
905
Processing state
not_requested
Canonical source
https://www.buzzsprout.com/2207817/episodes/19478522-grok-4-5-s-split-verdict-best-agentic-tool-use-on-the-board-a-doubled-hallucination-rate-and-the-new-math-of-model-trust-july-11-2026.mp3
Audio
https://www.buzzsprout.com/2207817/episodes/19478522-grok-4-5-s-split-verdict-best-agentic-tool-use-on-the-board-a-doubled-hallucination-rate-and-the-new-math-of-model-trust-july-11-2026.mp3
JSON
/v1/public/podcasts/dx-today-no-hype-podcast-news-about-ai-dx-6434212/episodes/grok-4-5-s-split-verdict-best-agentic-tool-use-on-the-board-a-doubled-hallucination-rate-and-the-new-math-of-model-trust-july-11-2026
Markdown
/podcast/dx-today-no-hype-podcast-news-about-ai-dx-6434212/grok-4-5-s-split-verdict-best-agentic-tool-use-on-the-board-a-doubled-hallucination-rate-and-the-new-math-of-model-trust-july-11-2026.md

Actions

  • POST https://stenobird.com/v1/public/podcasts/dx-today-no-hype-podcast-news-about-ai-dx-6434212/episodes/grok-4-5-s-split-verdict-best-agentic-tool-use-on-the-board-a-doubled-hallucination-rate-and-the-new-math-of-model-trust-july-11-2026/transcription-requests
    Idempotently request low-priority transcript generation for this episode.
  • GET https://stenobird.com/podcast/dx-today-no-hype-podcast-news-about-ai-dx-6434212/grok-4-5-s-split-verdict-best-agentic-tool-use-on-the-board-a-doubled-hallucination-rate-and-the-new-math-of-model-trust-july-11-2026.md
    Read the agent-friendly Markdown representation of this episode resource.

Summary

Grok 4.5's Split Verdict: Best Agentic Tool Use on the Board, a Doubled Hallucination Rate, and the New Math of Model Trust Grok 4.5's first independent report card is in, and it is split down the middle: the best agentic tool use score on the Artificial Analysis board and a number one claim on the SWE marathon, alongside a hallucination rate that doubled from 25 percent to 54 percent. Chris and Laura unpack how a model can know more yet be wrong more often, why grounded professional work a...