Episode

Solving the Memory Wall: A Deep Dive into AI Inference with Sandra Rivera

Podcast
Amelia's Weekly Fish Fry
Published
Apr 10, 2026
Duration seconds
1000
Processing state
not_requested
Canonical source
https://techfocus.podbean.com/e/inside-visora-how-a-french-startup-is-rewiring-ai-inference/
Audio
https://mcdn.podbean.com/mf/web/fy9wdqmrhqdmm8dy/FF_-_April_10_take_16gun5.mp3
JSON
/v1/public/podcasts/amelia-s-weekly-fish-fry-166630/episodes/solving-the-memory-wall-a-deep-dive-into-ai-inference-with-sandra-rivera
Markdown
/podcast/amelia-s-weekly-fish-fry-166630/solving-the-memory-wall-a-deep-dive-into-ai-inference-with-sandra-rivera.md

Actions

  • POST https://stenobird.com/v1/public/podcasts/amelia-s-weekly-fish-fry-166630/episodes/solving-the-memory-wall-a-deep-dive-into-ai-inference-with-sandra-rivera/transcription-requests
    Idempotently request low-priority transcript generation for this episode.
  • GET https://stenobird.com/podcast/amelia-s-weekly-fish-fry-166630/solving-the-memory-wall-a-deep-dive-into-ai-inference-with-sandra-rivera.md
    Read the agent-friendly Markdown representation of this episode resource.

Summary

This week, I'm excited to welcome Sandra Rivera from VSORA! We dive into a discussion on why AI inference is essential for deployment at scale, specifically focusing on how VSORA’s patented software architecture addresses the "memory wall" by collapsing memory layers. We explore their recent tape-out, which promises approximately 3X the performance at half the power of leading GPUs. We also chat about deployment use cases, the need for low latency and high determinism, future plans for OEM modules and MLPerf benchmarking, and even get a brief look into Sandra’s family llama farm.