# OLMo: Everything You Need to Train an Open Source LLM with Akshita Bhagia - #674 Page: https://stenobird.com/podcast/twiml-ai-podcast/olmo-everything-you-need-to-train-an-open-source-llm-with-akshita-bhagia-674 Text version: https://stenobird.com/podcast/twiml-ai-podcast/olmo-everything-you-need-to-train-an-open-source-llm-with-akshita-bhagia-674.md Podcast: [The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)](https://stenobird.com/podcast/twiml-ai-podcast) Published: 2024-03-04T20:10:00+00:00 Episode link: https://twimlai.com/podcast/twimlai/olmo-everything-you-need-to-train-an-open-source-llm/ Audio file: https://pscrb.fm/rss/p/traffic.megaphone.fm/MLN6116602529.mp3?updated=1709582779 Processing state: failed JSON: https://stenobird.com/v1/public/podcasts/twiml-ai-podcast/episodes/olmo-everything-you-need-to-train-an-open-source-llm-with-akshita-bhagia-674 Duration seconds: 1932 ## Resource Today we’re joined by Akshita Bhagia, a senior research engineer at the Allen Institute for AI. Akshita joins us to discuss OLMo, a new open source language model with 7 billion and 1 billion variants, but with a key difference compared to similar models offered by Meta, Mistral, and others. Namely, the fact that AI2 has also published the dataset and key tools used to train the model. In our chat with Akshita, we dig into the OLMo models and the various projects falling under the OLMo umbrella, including Dolma, an open three-trillion-token corpus for language model pretraining, and Paloma, a benchmark and tooling for evaluating language model performance across a variety of domains. The complete show notes for this episode can be found at twimlai.com/go/674. ## Actions - request_transcript: `POST https://stenobird.com/v1/public/podcasts/twiml-ai-podcast/episodes/olmo-everything-you-need-to-train-an-open-source-llm-with-akshita-bhagia-674/transcription-requests` — Idempotently request low-priority transcript generation for this episode. - read_markdown: `GET https://stenobird.com/podcast/twiml-ai-podcast/olmo-everything-you-need-to-train-an-open-source-llm-with-akshita-bhagia-674.md` — Read the agent-friendly Markdown representation of this episode resource. A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed. ## Transcript Full transcripts are not published on public pages unless there is a clear rights basis.