# Memory-Efficient Looped Transformer: Decoupling Compute from Memory in Looped Language Models

Page: https://stenobird.com/podcast/daily-paper-cast-7079649/memory-efficient-looped-transformer-decoupling-compute-from-memory-in-looped-language-models
Text version: https://stenobird.com/podcast/daily-paper-cast-7079649/memory-efficient-looped-transformer-decoupling-compute-from-memory-in-looped-language-models.md
Podcast: [Daily Paper Cast](https://stenobird.com/podcast/daily-paper-cast-7079649)
Published: 2026-05-13T04:31:39+00:00
Episode link: https://share.transistor.fm/s/51524a66
Audio file: https://media.transistor.fm/51524a66/8bb6ac2a.mp3
Processing state: not_requested
JSON: https://stenobird.com/v1/public/podcasts/daily-paper-cast-7079649/episodes/memory-efficient-looped-transformer-decoupling-compute-from-memory-in-looped-language-models
Duration seconds: 1349

## Resource

🤗 Upvotes: 21 | cs.CL, cs.AI, cs.LG Authors: Victor Conchello Vendrell, Arnau Padres Masdemont, Niccolò Grillo, Jordi Ros-Giralt, Arash Behboodi, Fabio Valerio Massoli Title: Memory-Efficient Looped Transformer: Decoupling Compute from Memory in Looped Language Models Arxiv: http://arxiv.org/abs/2605.07721v1 Abstract: Recurrent LLM architectures have emerged as a promising approach for improving reasoning, as they enable multi-step computation in the embedding space without generating intermediate tokens. Models such as Ouro perform reasoning by iteratively updating internal representations while retaining a standard Key-Value (KV) cache across iterations, causing memory consumption to grow linearly with reasoning depth. Consequently, increasing the number of reasoning iterations can lead to prohibitive memory usage, limiting the practical scalability of such architectures. In this work, we propose Memory-Efficient Looped Transformer (MELT), a novel architecture that decouples reasoning depth from memory consumption. Instead of using a standard KV cache per layer and loop, MELT maintains a single KV cache per layer that is shared across reasoning loops. This cache is updated over time via a learnable gating mechanism. To enable stable and efficient training under this architecture, we propose to train MELT using chunk-wise training in a two phase procedure: interpolated transition, followed by attention-aligned distillation, both from the LoopLM starting model to MELT. Empirically, we show that MELT models fine-tuned from pretrained Ouro parameters outperform standard LLMs of comparable size, while maintaining a memory footprint comparable to those models and dramatically smaller than Ouro's. Overall, MELT achieves constant-memory iterative reasoning without sacrificin…

## Actions

- request_transcript: `POST https://stenobird.com/v1/public/podcasts/daily-paper-cast-7079649/episodes/memory-efficient-looped-transformer-decoupling-compute-from-memory-in-looped-language-models/transcription-requests` — Idempotently request low-priority transcript generation for this episode.
- read_markdown: `GET https://stenobird.com/podcast/daily-paper-cast-7079649/memory-efficient-looped-transformer-decoupling-compute-from-memory-in-looped-language-models.md` — Read the agent-friendly Markdown representation of this episode resource.

A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed.

## Transcript

Full transcripts are not published on public pages unless there is a clear rights basis.