# LatentPress: Context Compression Beyond Text and Vision Page: https://stenobird.com/podcast/daily-paper-cast-7079649/latentpress-context-compression-beyond-text-and-vision Text version: https://stenobird.com/podcast/daily-paper-cast-7079649/latentpress-context-compression-beyond-text-and-vision.md Podcast: [Daily Paper Cast](https://stenobird.com/podcast/daily-paper-cast-7079649) Published: 2026-09-04T08:21:45+00:00 Episode link: https://share.transistor.fm/s/d0235eed Audio file: https://media.transistor.fm/d0235eed/ad827368.mp3 Processing state: not_requested JSON: https://stenobird.com/v1/public/podcasts/daily-paper-cast-7079649/episodes/latentpress-context-compression-beyond-text-and-vision Duration seconds: 1222 ## Resource 🤗 Upvotes: 52 | cs.LG, cs.AI Authors: Zhengze Zhou, Hejian Sang Title: LatentPress: Context Compression Beyond Text and Vision Arxiv: http://arxiv.org/abs/2609.01507v2 Abstract: Compressed context is usually carried as human-readable text or as rendered images that must be decoded, even when its consumer is a language model. We introduce LatentPress, which writes conversational histories and long documents into a third representation: continuous memory tokens that a frozen decoder reads directly through its input-embedding interface, with no text reconstruction at inference. A small reader-matched writer compresses $4$-$16\times$ while training only an adapter (4.2M-26.2M parameters, $\sim\!0.1\%$ of the decoder). On LongMemEval, LatentPress reaches $0.504$ accuracy at $7.70\times$ compression versus $0.490$ for uncompressed evidence, outperforming text summaries (0.184) and OCR-based compression (0.426 to 0.312). On LongBench-QA, in-domain writers match or exceed raw-context reading at $4$-$8\times$ compression, while $16\times$ trails raw. Writing takes 43ms per conversation, roughly an order of magnitude faster than text summarization or OCR reconstruction, and reading is $5$-$9\times$ faster than raw context or cached OCR. We validate the interface under two transfer settings, zero-shot from UltraChat to LongMemEval memory QA and from LongMemEval-derived QA to unseen LongBench document domains, establishing direct soft tokens as a practical machine-facing context interface beyond text and vision. The implementation of the experiments could be found at: https://github.com/HJSang/LatentPress . ## Actions - request_transcript: `POST https://stenobird.com/v1/public/podcasts/daily-paper-cast-7079649/episodes/latentpress-context-compression-beyond-text-and-vision/transcription-requests` — Idempotently request low-priority transcript generation for this episode. - read_markdown: `GET https://stenobird.com/podcast/daily-paper-cast-7079649/latentpress-context-compression-beyond-text-and-vision.md` — Read the agent-friendly Markdown representation of this episode resource. A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed. ## Transcript Full transcripts are not published on public pages unless there is a clear rights basis.