# How Data Scientists Use Nearest Neighbors for Anomaly Detection Page: https://stenobird.com/podcast/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/how-data-scientists-use-nearest-neighbors-for-anomaly-detection Text version: https://stenobird.com/podcast/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/how-data-scientists-use-nearest-neighbors-for-anomaly-detection.md Podcast: [The Data Science Podcast with Fexingo: Analytics, Machine Learning, and Data-Driven Conversations](https://stenobird.com/podcast/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831) Published: 2026-07-07T08:49:58+00:00 Episode link: https://audio.fexingo.com/business/the-data-science-podcast/episode-0095.mp3 Audio file: https://audio.fexingo.com/business/the-data-science-podcast/episode-0095.mp3 Processing state: not_requested JSON: https://stenobird.com/v1/public/podcasts/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/episodes/how-data-scientists-use-nearest-neighbors-for-anomaly-detection Duration seconds: 540 ## Resource In Episode 95 of The Data Science Podcast with Fexingo, Lucas and Luna dive into a practical yet underappreciated technique: using k-nearest neighbors for anomaly detection. They kick off with a real-world story from a major credit card processor that flagged a series of fraudulent transactions by measuring distance to the nearest legitimate patterns. Lucas explains why distance-based methods can outperform deep learning in low-signal, high-stakes settings, especially when you need interpretable reasons for each flag. Luna challenges him on scalability and the curse of dimensionality, and they discuss how companies like Stripe and PayPal have used variants of k-NN in production fraud pipelines. They also touch on the trade-offs between global and local outlier factors, and how to choose k when the definition of 'normal' shifts over time. A concrete segment on choosing distance metrics — Euclidean vs. Manhattan vs. cosine — gives listeners an actionable guideline. Mid-episode, they weave in a natural request for listener support, tying it back to the value of open-source tools. If you've ever wondered when to reach for a simple nearest-neighbor approach instead of a neural network, this episode gives you the framework. #DataScience #MachineLearning #AnomalyDetection #KNearestNeighbors #FraudDetection #OutlierDetection #DistanceMetrics #LocalOutlierFactor #Stripe #PayPal #CurseOfDimensionality #Interpretability #Technology #FexingoBusiness #BusinessPodcast #TechPodcast #DataSciencePodcast #ProductionML Keep every episode free: buymeacoffee.com/fexingo ## Actions - request_transcript: `POST https://stenobird.com/v1/public/podcasts/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/episodes/how-data-scientists-use-nearest-neighbors-for-anomaly-detection/transcription-requests` — Idempotently request low-priority transcript generation for this episode. - read_markdown: `GET https://stenobird.com/podcast/the-data-science-podcast-with-fexingo-analytics-machine-learning-and-data-driven-conversations-7871831/how-data-scientists-use-nearest-neighbors-for-anomaly-detection.md` — Read the agent-friendly Markdown representation of this episode resource. A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed. ## Transcript Full transcripts are not published on public pages unless there is a clear rights basis.