Episode
SE Radio 696: Flavia Saldanha on Data Engineering for AI
- Published
- Nov 25, 2025
- Duration seconds
- 4465
- Processing state
failed- Canonical source
- https://se-radio.net/2025/11/se-radio-696-flavia-saldanha-on-data-engineering-for-ai/
Actions
POST https://stenobird.com/v1/public/podcasts/software-engineering-radio/episodes/se-radio-696-flavia-saldanha-on-data-engineering-for-ai/transcription-requests
Idempotently request low-priority transcript generation for this episode.GET https://stenobird.com/podcast/software-engineering-radio/se-radio-696-flavia-saldanha-on-data-engineering-for-ai.md
Read the agent-friendly Markdown representation of this episode resource.
Summary
Flavia Saldanha, a consulting data engineer, joins host Kanchan Shringi to discuss the evolution of data engineering from ETL (extract, transform, load) and data lakes to modern lakehouse architectures enriched with vector databases and embeddings. Flavia explains the industry's shift from treating data as a service to treating it as a product, emphasizing ownership, trust, and business context as critical for AI-readiness. She describes how unified pipelines now serve both business intelligence and AI use cases, combining structured and unstructured data while ensuring semantic enrichment and a single source of truth. She outlines key components of a modern data stack, including data marketplaces, observability tools, data quality checks, orchestration, and embedded governance with lineage tracking. This episode highlights strategies for abstracting tooling, future-proofing architectures, enforcing data privacy, and controlling AI-serving layers to prevent hallucinations. Saldanha concludes that data engineers must move beyond pure ETL thinking, embrace product and NLP skills, and work closely with MLOps, using AI as a co-pilot rather than a replacement. Brought to you by IEEE Computer Society and IEEE Software magazine.