Back

Explore every episode of the podcast Data Skeptic

Dive into the complete episode list for Data Skeptic. Each episode is cataloged with detailed descriptions, making it easy to find and explore specific topics. Keep track of all episodes from your favorite podcast and never miss a moment of insightful content.

Rows per page:

1–50 of 599

TitlePub. DateDuration
Ant Encounters26 Aug 202400:31:26

In this interview with author Deborah Gordon, Kyle asks questions about the mechanisms at work in an ant colony and what ants might teach us about how to build artificial intelligence. Ants are surprisingly adaptive creatures whose behavior emerges from their complex interactions. Aspects of network theory and the statistical nature of ant behavior are just some of the interesting details you'll get in this episode.

 
Computing Toolbox19 Aug 202400:38:44

This season it's become clear that computing skills are vital for working in the natural sciences. In this episode, we were fortunate to speak with Madlen Wilmes, co-author of the book "Computing Skills for Biologists: A Toolbox". We discussed the book and why it's a great resource for students and teachers. In addition to the book, Madlen shared her experience and advice on transitioning from academia to an industry career and how data analytic skills transfer to jobs that your professionals might not always consider. Join us and learn more about the book and careers using transferable skills.

Learn to Code18 Jun 202400:49:31

Do you code or are you interested in learning to code? Join us today and hear from three individuals that are at very different stages of their coding journeys. Becky Hansis-O'Neill (also our co-host this season) shares her experiences as a newbie who wants to learn more. Dr. Malia Gehan, a self-taught developer interested in studying plant phenotypes, explains why and how she and her colleagues learned to code and developed PlantCV. Finally, Dr. John Wilmes discusses his work as a professional mathematician and Machine Learning Research Engineer. Whether you are thinking about learning to code or an expert, we're sure you will see a bit of yourself in this episode. 

Fairness in e-Commerce Search05 Sep 202200:40:46

When we search for products in e-commerce stores, we do not care what goes on under the hood to generate the results. However, there may be an intentional algorithmic effort to gravitate us toward a particular product. On the show, today, Abhisek Dash and Saptarshi Ghosh discuss their research on fairness in the search result of Amazon smart speakers.

Fraudulent Amazon Reviewers29 Aug 202200:41:18

Chances are that you have bought a product online majorly because of the reviews you saw. Unfortunately, not all reviews are genuine. Today, Rajvardhan Oak shares some insight from his research on fraudulent Amazon reviews. He explained the inner workings of fraudulent reviews and revealed key insights from his qualitative and quantitative study.

Ad Targeting in Amazon Smart Speakers22 Aug 202200:32:40

While we give attention to textual data on the web, many do not know the unique power of echo interactions with smart devices for ad targeting. Today, our guest, Umar Iqbal joins us to discuss his study on using Amazon Smart Speakers for ad targeting. He gave interesting revelations about how voice data is captured and analysed for ad purposes. Listen to find out more.

Adwords with Unknown Budgets15 Aug 202200:34:09

Rajan Udwani, an Assistant Professor at the University of California Berkeley joins us to discuss his work on AdWords with unknown budgets. He discussed the previous approaches to ad allocation, as well as his maiden approach that introduced randomization for better results. Listen for more.

ML Ops Best Practices12 Aug 202200:30:12

Today, we are joined by Piotr Niedźwiedź, Founder and CEO of Neptune.ai. Piotr discusses common MLOps activities by data science teams and how they can take advantage of Neptune.ai for better experiment tracking and efficiency. Listen for more!

Affiliate Marketing Rabbithole08 Aug 202200:52:20

Affiliate marketing creates an opportunity for marketers to gain a commission by promoting a product or service.  Cookies are typically used for tracking and the advertiser whose product or service is being featured pays the marketing only on transactions.

Today's episode covers those approaches and is also a story of conflict between two large companies and how one affiliate marketer got caught in the middle.

Monetization of Youtube Conspiracy Theorists01 Aug 202200:54:06

Cameron Ballard joins us today to discuss his work around YouTube conspiracy theories. He revealed interesting observations about conspiracy theories on YouTube including how predatory ads are most common in conspiracy theory videos and how YouTube's algorithm subtly works for predatory ads. 

User Perceptions of Problematic Ads25 Jul 202200:37:53

Eric Zeng joins us to discuss his study around understanding bad ads and efforts that can be taken to limit bad ads online. He discussed how he and his co authors scrapped a large amount of ad data, applied a machine learning algorithm, and commensurate statistical results.

Political Digital Advertising Analysis21 Jul 202200:35:35

NaLette Brodnax, a political scientist and an Assistant Professor in the McCourt School of Public Policy at Georgetown University joins us to discuss her work on analyzing digital advertisements for political campaigns. She used data for electoral campaigns on Facebook to answer questions that help us better understand how digital ads affect the outcome of elections.

 

Click here for additional show notes!

Thanks to our sponsor!
https://neptune.ai/ Log, store, query, display, organize and compare all your model metadata in a single place

Fraud Detection in Crowdfunding Campaigns18 Jul 202200:35:47
Animal Computer Interaction10 Jun 202400:42:49

You've heard of Human Computer Interaction (HCI), now get ready for Animal Computer Interaction (ACI). Ilyena has made a career developing computer interfaces for non-human animals. She has worked with dogs, parrots, primates, and even giraffes. This is challenging because animals have a wide range of abilities and preferences. Parrots, for example, use their tongues to make selections on touchscreens. Listen in on our conversation and learn about interface development and testing with animals and how technology may improve animal welfare. 

Artificial Intelligence and Auction Design11 Jul 202200:43:13
Privacy Preference Signals04 Jul 202200:33:28

Have you ever wondered what goes on under the hood when you accept a website's cookies? Today, Maximilian Hils, a PhD student in Computer Science, at the University of Innsbruck, Austria, dissects the ad tech industry and the standards put in place to protect users' data. He also shares his thoughts on the use of VPNs as well as other tools that help shield your data from prying eyes on the internet.

Click here for additional show notes

Thanks to our sponsor:
https://clear.ml/ ClearML is an open-source MLOps solution users love to customize, helping you easily Track, Orchestrate, and Automate ML workflows at scale.

Neural Architecture Search for CTR Prediction27 Jun 202200:28:02

Ravi Krishna joins us today to talk about his recent work on a differentiable NAS framework for ads CTR prediction. He discussed what CTR prediction is about and why his NAS framework helps in building neural networks for better ads recommendation. Listen to learn about methodology, related literature and his results.

Click for additional show notes

Thanks to our sponsor:
https://astrato.io Astrato is a modern BI and analytics platform built for the Snowflake Data Cloud. A next-generation live query data visualization and analytics solution, empowering everyone to make live data decisions.

Algorithmic PPC Management21 Jun 202200:43:56

Effectively managing a large budget of pay per click advertising demands software solutions. When spending multi-million dollar budgets on hundreds of thousands of keywords, an effective algorithmic strategy is required to optimize marketing objectives.

In this episode, Nathan Janos joins us to share insights from his work in the ad tech industry.

Click for additional show notes

Thanks to our sponsor!
https://wandb.com/ The developer-first MLOps platform. Build better models faster with experiment tracking, dataset versioning, and model management.

Data Skeptic: Ad Tech18 Jun 202200:42:24

Increasingly, people get most if not all of the information they consume online. Alongside the web sites, videos, apps, and other destinations, we're consistently served advertisements alongside the organic content we search for or discover. Targetted ads make it possible for you to discover relevant new products you might otherwise not have heard about. Targetting can also open a pandora's box of ethical considerations. Online advertising is a complex network of automated systems. Algorithms controlling algorithms controlling what we see.

This season of Data Skeptic will focus on the applications of data science to digital advertising technology. In this first episode in particular, Kyle shares some of his own personal experiences and insights working in pay-per-click marketing.

Click for additional show notes

 

 

The Reliability of Mobile Phone Data13 Jun 202200:49:32

Our mobile phones generate an incredible amount of data inbound and outbound. In today's episode, Nishant Kishore, a PhD graduate of Harvard University in Infectious Disease Epidemiology, explains how mobility data from mobile phones can be captured and analysed to understand the spread of infectious diseases.

Click here for additional show notes

Thanks to our sponsor!
https://neptune.ai/ Log, store, query, display, organize, and compare all your model metadata in a single place

Haywire Algorithms06 Jun 202200:33:33

The pandemic changed how we lived. And this had a ripple effect on the performance of machine learning models. Ravi Parikh joins us today to discuss how the pandemic has affected the performance of machine learning models in clinical care and some actionable steps to fix it.

Click here for additional show notes

Thanks to our sponsor:
Astera Centerprise is a no-code data integration platform that allows users to build ETL/ELT pipelines for modern data warehousing and analytics.

School Reopening Analysis30 May 202200:33:17

Carly Lupton-Smith joins us today to speak about her research which investigated the consistency between household and county measures of school reopening. Carly is a doctoral researcher in Biostatistics at Johns Hopkins Bloomberg School of Public Health. Listen to know about her findings.

Click here for additional show notes on our website!

Thanks to our sponsor!
ClearML is an open-source MLOps solution users love to customize, helping you easily Track, Orchestrate, and Automate ML workflows at scale.

Astera Centerprise is a no-code data integration platform that allows users to build ETL/ELT pipelines for modern data warehousing and analytics.

 

Modern Data Stacks26 May 202200:34:33

Today, we are joined by Alexander Thor, a Product Manager at Vizlib, makers of Astrato. Astrato is a data analytics and business intelligence tool built on the cloud and for the cloud. Alexander discusses the features and capabilities of Astrato for data professionals.

Visit our website for additional show notes!

 

Emoji as a Predictor23 May 202200:21:25

Emojis are arguably one of the most effective ways to express emotions when texting. In today's episode, Xuan Lu shares her research on the use of emojis by developers. She explains how the study of emojis can track the emotions of remote workers and predict future behavior. Listen to find out more!

Ape Gestures03 Jun 202400:49:26

Cat observes great apes in the wild and in the lab to crack the code of their gestural communication. We discussed the challenges and benefits of studying apes in the wild vs in the lab. Cat also shared how her lab identifies and studies ape gestures. It turns out that humans are pretty good at guessing what apes are trying to communicate with one another. Join us in this episode to learn more about the evolution of communication in great apes, and what we can learn from our closest relatives. 

Polarizing Trends in the Gig Economy16 May 202200:46:17
On the show today, Fabian Braesemann, a research fellow at the University of Oxford, joins us to discuss his study analyzing the gig economy. He revealed the trends he discovered since remote work became mainstream, the factors causing spatial polarization and some downsides of the gig economy. Listen to learn what he found. 
Remote Learning in Applied Engineering12 May 202200:25:15

On the show today, we interview Mouhamed Abdulla, a professor of Electrical Engineering at Sheridan Institute of Technology. Mouhamed joins us to discuss his study on remote teaching and learning in applied engineering. He discusses how he embraced the new approach after the pandemic, the challenges he faced and how he tackled them. Listen to find out more.

Click here for additional show notes on our website!

Thanks to our sponsor!
https://neptune.ai/

Log, store, query, display, organize, and compare all your model metadata in a single place

 

Remote Productivity09 May 202200:29:48
It is difficult to estimate the effect on remote working across the board. Darja Šmite, who speaks with us today, is a professor of Software Engineering at the Blekinge Institute of Technology. In her recently published paper, she analyzed data on several companies' activities before and after remote working became prevalent. She discussed the results found, why they were and some subtle drawbacks of remote working. Check it out!

 

Click here for additional show notes on our website!

Does Remote Learning Work?01 May 202200:48:10
We explore this complex question in two interviews today.  First, Kasey Wagoner describes 3 approaches to remote lab sessions and an analysis of which was the most instrumental to students.  Second, Tahiya Chowdhury shares insights about the specific features of video-conferencing platforms that are lacking in comparison to in-person learning.

Click here for additional show notes on our website!

Thanks to our sponsor!
ClearML is an open-source MLOps solution users love to customize, helping you easily Track, Orchestrate, and Automate ML workflows at scale.

 

Covid-19 Impact on Bicycle Usage25 Apr 202200:31:12
In this episode, we speak with Abdullah Kurkcu, a Lead Traffic Modeler. Abdullah joins us to discuss his recent study on the effect of COVID-19 on bicycle usage in the US. He walks us through the data gathering process, data preprocessing, feature engineering, and model building. Abdullah also disclosed his results and key takeaways from the study. Listen to find out more. 

Click here for additional show notes on our website.

Thanks to our sponsor!
Astrato is a modern BI and analytics platform built for the Snowflake  Data Cloud. A next-generation live query data visualization and analytics solution, empowering everyone to make live data decisions.

 

 

Learning Digital Fabrication Remotely22 Apr 202200:33:31

Today, we are joined by Jennifer Jacobs and Nadya Peek, who discuss their experience in teaching remote classes for a course that is largely hands-on. The discussion was focused on digital fabrication, why it is important, the prospect for the future, the challenges with remote lectures, and everything in between.

Click here for additional show notes on our website!

Thanks to our sponsor!
https://neptune.ai/

Log, store, query, display, organize, and compare all your model metadata in a single place

Remote Software Development18 Apr 202200:37:33

Today, we are joined by Denae Ford, a Senior Researcher at Microsoft Research and an Affiliate Assistant Professor at the University of Washington. Denae discusses her work around remote work and its culminating impact on workers. She narrowed down her research to how COVID-19 has affected the working system of software engineers and the emerging challenges it brings.

 

 

Click here to access additional show notes on our website!

 

Thanks to our sponsor! 

Weights & Biases : The developer-first MLOps platform. Build better models faster with experiment tracking, dataset versioning, and model management.

 

Quantum K-Means11 Apr 202200:39:52

In this episode, we interview Jonas Landman, a Postdoc candidate at the University of Edinburg. Jonas discusses his study around quantum learning where he attempted to recreate the conventional k-means clustering algorithm and spectral clustering algorithm using quantum computing. 

Click here to access additional show notes on our website!

K-Means in Practice04 Apr 202200:30:41
K-means is widely used in real-life business problems. In this episode, Mujtaba Anwer, a researcher and Data Scientist walks us through some use cases of k-means. He also spoke extensively on how to prepare your data for clustering, find the best number of clusters to use, and turn the 'abstract' result into real business value. Listen to learn.  Click here to access additional show notes on our website! Thanks to our sponsor!
ClearML is an open-source MLOps solution users love to customize, helping you easily Track, Orchestrate, and Automate ML workflows at scale.
Fair Hierarchical Clustering28 Mar 202200:34:26
Building a fair machine learning model has become a critical consideration in today's world. In this episode, we speak with Anshuman Chabra, a Ph.D. candidate in Computer Networks. Chhabra joins us to discuss his research on building fair machine learning models and why it is important. Find out how he modeled the problem and the result found.

Click here to access additional show notes on our webiste!

Thanks to our sponsor!
https://astrato.io

Astrato is a modern BI and analytics platform built for the Snowflake Data Cloud. A next-generation live query data visualization and analytics solution, empowering everyone to make live data decisions.

Evaluating AI Abilities27 May 202400:49:40

In this episode, Kozzy discusses his endeavors to compare the cognitive abilities of humans, animals, and AI programs. Specifically, we discussed object permanence, the ability to understand an object still exists in space even when you can't see it. Our conversation traverses both philosophical and practical questions surrounding AI evaluation. We also learned about Animal AI 3, a gaming environment developed in Unity where AI programs and humans can go head-to-head to solve different problems in a gaming environment.

Matrix Factorization For k-Means21 Mar 202200:30:07

Many people know K-means clustering as a powerful clustering technique but not all listeners will be as familiar with spectral clustering. In today's episode, Sibylle Hess from the Data Mining group at TU Eindhoven joins us to discuss her work around spectral clustering and how its result could potentially cause a massive shift from the conventional neural networks. Listen to learn about her findings.

Visit our website for additional show notes

Thanks to our sponsor, Weights & Biases

Breathing K-Means14 Mar 202200:42:55

In this episode, we speak with Bernd Fritzke, a proficient financial expert and a Data Science researcher on his recent research - the breathing K-means algorithm. Bernd discussed the perks of the algorithms and what makes it stand out from other K-means variations. He extensively discussed the working principle of the algorithm and the subtle but impactful features that enables it produce top-notch results with low computational resources. Listen to learn about this algorithm.

Power K-Means07 Mar 202200:32:38

In today's episode, Jason, an Assistant Professor of Statistical Science at Duke University talks about his research on K power means. K power means is a newly-developed algorithm by Jason and his team, that aims to solve the problem of local minima in classical K-means, without demanding heavy computational resources. Listen to find out the outcome of Jason's study.

Click here to access additional show notes on our website!

Thanks to our Sponsors:
ClearML is an open-source MLOps solution users love to customize, helping you easily Track, Orchestrate, and Automate ML workflows at scale. https://clear.ml

Springboard
Springboard offers end-to-end online data career programs that encompass data science, data analytics, data engineering, and machine learning engineering.

Explainable K-Means03 Mar 202200:25:53

In this episode, Kyle interviews Lucas Murtinho about the paper "Shallow decision treees for explainable k-means clustering" about the use of decision trees to help explain the clustering partitions. 

Check out our website for extended show notes!

Thanks to our Sponsors:
ClearML is an open-source MLOps solution users love to customize, helping you easily Track, Orchestrate, and Automate ML workflows at scale.

Customer Clustering28 Feb 202200:22:03

Have you ever wondered how you can use clustering to extract meaningful insight from a time-series single-feature data? In today's episode, Ehsan speaks about his recent research on actionable feature extraction using clustering techniques. Want to find out more? Listen to discover the methodologies he used for his research and the commensurate results.

Visit our website for extended show notes!

https://clear.ml/

ClearML is an open-source MLOps solution users love to customize, helping you easily Track, Orchestrate, and Automate ML workflows at scale.

k-means Image Segmentation22 Feb 202200:23:01

Linh Da joins us to explore how image segmentation can be done using k-means clustering.  Image segmentation involves dividing an image into a distinct set of segments.  One such approach is to do this purely on color, in which case, k-means clustering is a good option. 

Check out our website for extended show notes and images!

Thanks to our Sponsors: Visit Weights and Biases mention Data Skeptic when you request a demo! &
Nomad Data 

In the image below, you can see the k-means clustering segmentation results for the same image with the values of 2, 4, 6, and 8 for k.

 
Tracking Elephant Clusters18 Feb 202200:26:27

In today's episode, Gregory Glatzer explained his machine learning project that involved the prediction of elephant movement and settlement, in a bid to limit the activities of poachers. He used two machine learning algorithms, DBSCAN and K-Means clustering at different stages of the project. Listen to learn about why these two techniques were useful and what conclusions could be drawn.

Click here to see additional show notes on our website!

Thanks to our sponsor, Astrato

k-means clustering14 Feb 202200:24:22
Welcome to our new season, Data Skeptic: k-means clustering.  Each week will feature an interview or discussion related to this classic algorithm, it's use cases, and analysis.

This episode is an overview of the topic presented in several segments.

Snowflake Essentials07 Feb 202200:46:43
Frank Bell, Snowflake Data Superhero, and SnowPro, joins us today to talk about his book "Snowflake Essentials: Getting Started with Big Data in the Cloud." 

Thanks to our Sponsors:

  • Find Better Data Faster with Nomad Data. Visit nomad-data.com
  • Visit Springboard and use promo code DATASKEPTIC to receive a $750 discount
Explainable Climate Science31 Jan 202200:34:50

Zack Labe, a Post-Doctoral Researcher at Colorado State University, joins us today to discuss his work "Detecting Climate Signals using Explainable AI with Single Forcing Large Ensembles."
Works Mentioned
"Detecting Climate Signals using Explainable AI with Single Forcing Large Ensembles"
by Zachary M. Labe, Elizabeth A. Barnes

Sponsored by:
Astrato
and
BBEdit by Bare Bones Software

HMMs for Behavior20 May 202400:45:11

Théo Michelot has made a career out of tackling tough ecological questions using time-series data. How do scientists turn a series of GPS location observations over time into useful behavioral data? GPS tech has improved to the point that modern data sets are large and complex. In this episode, Théo takes us through his research and the application of Hidden Markov Models to complex time series data. If you have ever wondered what biologists do with data from those GPS collars you have seen on TV, this is the episode for you! 

Energy Forecasting Pipelines24 Jan 202200:43:21

Erin Boyle, the Head of Data Science at Myst AI, joins us today to talk about her work with Myst AI, a time series forecasting platform and service with the objective for positively impacting sustainability.

https://docs.myst.ai/docs Visit Weights and Biases at wandb.me/dataskeptic Find Better Data Faster with Nomad Data. Visit nomad-data.com

Matrix Profiles in Stumpy17 Jan 202200:39:09

Sean Law, Principle Data Scientist, R&D at a Fortune 500 Company, comes on to talk about his creation of the STUMPY Python Library.

Sponsored by Hello Fresh and mParticle:

Go to Hellofresh.com/dataskeptic16 for up to 16 free meals AND 3 free gifts!

Visit mparticle.com to learn how teams at Postmates, NBCUniversal, Spotify, and Airbnb use mParticle's customer data infrastructure to accelerate their customer data strategies.

The Great Australian Prediction Project14 Jan 202200:25:19

Data scientists and psychics have at least one major thing in common. Both professions attempt to predict the future. In the case of a data scientist, this is done using algorithms, data, and often comes with some measure of quality such as a confidence interval or estimated accuracy. In contrast, psychics rely on their intuition or an appeal to the supernatural as the source for their predictions. Still, in the interest of empirical evidence, the quality of predictions made by psychics can be put to the test.

The Great Australian Psychic Prediction Project seeks to do exactly that. It's the longest known project tracking annual predictions made by psychics, and the accuracy of those predictions in hindsight. Richard Saunders, host of The Skeptic Zone Podcast, joins us to share the results of this decadal study.

Read the full report: https://www.skeptics.com.au/2021/12/09/psychic-project-full-results-released/

And follow the Skeptics Zone: https://www.skepticzone.tv/

 

© My Podcast Data