# OpenAI’s custom inference chip & MoE fine-tuning gets faster - AI News (Jun 26, 2026) Page: https://stenobird.com/podcast/the-automated-daily-ai-news-edition-6657064/openai-s-custom-inference-chip-moe-fine-tuning-gets-faster-ai-news-jun-26-2026 Text version: https://stenobird.com/podcast/the-automated-daily-ai-news-edition-6657064/openai-s-custom-inference-chip-moe-fine-tuning-gets-faster-ai-news-jun-26-2026.md Podcast: [The Automated Daily - AI News Edition](https://stenobird.com/podcast/the-automated-daily-ai-news-edition-6657064) Published: 2026-06-26T12:31:25+00:00 Episode link: https://theautomateddaily.com/episodes/2026-06-26-openai-s-custom-inference-chip-moe-fine-tuning-gets-faster Audio file: https://dts.podtrac.com/redirect.mp3/cdn.theautomateddaily.com/audio/hn-ai/2026-06-26/en/episode.mp3 Processing state: not_requested JSON: https://stenobird.com/v1/public/podcasts/the-automated-daily-ai-news-edition-6657064/episodes/openai-s-custom-inference-chip-moe-fine-tuning-gets-faster-ai-news-jun-26-2026 Duration seconds: 451 ## Resource Please support this podcast by checking out our sponsors: - KrispCall: Agentic Cloud Telephony - https://try.krispcall.com/tad - Invest Like the Pros with StockMVP - https://www.stock-mvp.com/?via=ron - Discover the Future of AI Audio with ElevenLabs - https://try.elevenlabs.io/tad Support The Automated Daily directly: Buy me a coffee: https://buymeacoffee.com/theautomateddaily Today's topics: OpenAI’s custom inference chip - OpenAI and Broadcom revealed “Jalapeño,” a purpose-built inference accelerator aimed at lowering LLM serving cost and boosting performance-per-watt at data-center scale. MoE fine-tuning gets faster - NVIDIA and Hugging Face pushed MoE training forward with NeMo AutoModel, highlighting throughput and memory gains that could reduce the GPU barrier for fine-tuning large MoE LLMs. Apple’s AI-first Mac roadmap - A report says Apple may skip higher-end M6 variants and jump pro Macs to AI-heavier M7 Pro/Max/Ultra chips, signaling on-device AI as a core silicon priority. Gemini adds built-in computer use - Google folded “computer use” into Gemini 3.5 Flash, making it easier to build agents that can see interfaces and take actions while adding safeguards against prompt injection. Amazon versus Perplexity’s agent browser - Amazon sued Perplexity over its Comet agentic browser, raising big questions about bot disclosure, user-agent spoofing, and who’s accountable when AI acts inside logged-in sessions. Anthropic alleges mass distillation attack - Anthropic told the U.S. Senate it believes Alibaba-linked operators ran a large-scale model distillation campaign using fraudulent accounts, escalating the policy fight over AI capability “theft.” Diffusion to engine-ready 3D geometry - Google Research introduced FLAT, a way to decode diffusion video latents directl… ## Actions - request_transcript: `POST https://stenobird.com/v1/public/podcasts/the-automated-daily-ai-news-edition-6657064/episodes/openai-s-custom-inference-chip-moe-fine-tuning-gets-faster-ai-news-jun-26-2026/transcription-requests` — Idempotently request low-priority transcript generation for this episode. - read_markdown: `GET https://stenobird.com/podcast/the-automated-daily-ai-news-edition-6657064/openai-s-custom-inference-chip-moe-fine-tuning-gets-faster-ai-news-jun-26-2026.md` — Read the agent-friendly Markdown representation of this episode resource. A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed. ## Transcript Full transcripts are not published on public pages unless there is a clear rights basis.