# An LLM Evaluation Framework for High-Stakes AI Page: https://stenobird.com/podcast/software-engineering-institute-sei-podcast-series-389154/an-llm-evaluation-framework-for-high-stakes-ai Text version: https://stenobird.com/podcast/software-engineering-institute-sei-podcast-series-389154/an-llm-evaluation-framework-for-high-stakes-ai.md Podcast: [Software Engineering Institute (SEI) Podcast Series](https://stenobird.com/podcast/software-engineering-institute-sei-podcast-series-389154) Published: 2026-06-11T18:29:00+00:00 Episode link: https://cmu-sei-podcasts.libsyn.com/an-llm-evaluation-framework-for-high-stakes-ai Audio file: https://traffic.libsyn.com/clean/secure/cmu-sei-podcasts/SEIP_05_01_TURRI_v2.mp3?dest-id=762491 Processing state: not_requested JSON: https://stenobird.com/v1/public/podcasts/software-engineering-institute-sei-podcast-series-389154/episodes/an-llm-evaluation-framework-for-high-stakes-ai Duration seconds: 993 ## Resource Experimentation and validation of LLM performance is critical when building LLM-driven systems that must reliably deliver a service, from customer service chat bots to intelligence analysis tools. To help teams meet the need for rigorous evaluation methods, a research team in the SEI's AI Division led by Violet Turri has developed the Evaluating Large Language Models (ELM) library, which is built on best practices for LLM evaluation and benchmarking. In the latest episode from the Carnegie Mellon University Software Engineering Institute, Turri sits down with Katie Robinson, a design researcher also in the SEI's AI division, to discuss the ELM library, which turns evaluation from an ad-hoc process into a repeatable, extensible framework. ## Actions - request_transcript: `POST https://stenobird.com/v1/public/podcasts/software-engineering-institute-sei-podcast-series-389154/episodes/an-llm-evaluation-framework-for-high-stakes-ai/transcription-requests` — Idempotently request low-priority transcript generation for this episode. - read_markdown: `GET https://stenobird.com/podcast/software-engineering-institute-sei-podcast-series-389154/an-llm-evaluation-framework-for-high-stakes-ai.md` — Read the agent-friendly Markdown representation of this episode resource. A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed. ## Transcript Full transcripts are not published on public pages unless there is a clear rights basis.