Explorez tous les épisodes du podcast Data Brew by Databricks
| Titre | Date | Durée | |
|---|---|---|---|
| Reinforcement Fine-Tuning and the Future of Specialized AI Models | 05 août 2025 | 00:40:24 | |
What if building a custom AI model for your business was as simple as giving feedback—no massive labeled datasets required? In this episode, we sit down with Travis Addair, CTO and Co-Founder of Predibase, creators of the first reinforcement fine-tuning platform, to explore the future of specialized AI. Discover how reinforcement fine-tuning is revolutionizing model customization, enabling you to start fast, adapt to your unique data, and keep improving through human feedback. Whether you’re an AI enthusiast or a business leader, you’ll learn how this breakthrough is making advanced AI accessible to everyone. Highlights:
| |||
| Benchmarking Domain Intelligence | Data Brew | Episode 45 | 24 avr. 2025 | 00:31:41 | |
In this episode, Pallavi Koppol, Research Scientist at Databricks, explores the importance of domain-specific intelligence in large language models (LLMs). She discusses how enterprises need models tailored to their unique jargon, data, and tasks rather than relying solely on general benchmarks. | |||
| SWE-bench & SWE-agent | Data Brew | Episode 44 | 17 avr. 2025 | 00:36:22 | |
In this episode, Kilian Lieret, Research Software Engineer, and Carlos Jimenez, Computer Science PhD Candidate at Princeton University, discuss SWE-bench and SWE-agent, two groundbreaking tools for evaluating and enhancing AI in software engineering. | |||
| Enterprise AI: Research to Product | Data Brew | Episode 43 | 10 avr. 2025 | 00:38:03 | |
In this episode, Dipendra Kumar, Staff Research Scientist, and Alnur Ali, Staff Software Engineer at Databricks, discuss the challenges of applying AI in enterprise environments and the tools being developed to bridge the gap between research and real-world deployment. | |||
| Multimodal AI | Data Brew | Episode 42 | 07 avr. 2025 | 00:42:14 | |
In this episode, Chang She, CEO and Co-founder of LanceDB, discusses the challenges of handling multimodal data and how LanceDB provides a cutting-edge solution. He shares his journey from contributing to Pandas to building a database optimized for images, video, vectors, and subtitles. | |||
| Age of Agents | Data Brew | Episode 41 | 27 mars 2025 | 00:40:47 | |
In this episode, Michele Catasta, President of Replit, explores how AI-driven agents are transforming software development by making coding more accessible and automating application creation. | |||
| Reward Models | Data Brew | Episode 40 | 20 mars 2025 | 00:39:58 | |
In this episode, Brandon Cui, Research Scientist at MosaicML and Databricks, dives into cutting-edge advancements in AI model optimization, focusing on Reward Models and Reinforcement Learning from Human Feedback (RLHF). | |||
| Retrieval, rerankers, and RAG tips and tricks | Data Brew | Episode 39 | 20 févr. 2025 | 00:45:22 | |
In this episode, Andrew Drozdov, Research Scientist at Databricks, explores how Retrieval Augmented Generation (RAG) enhances AI models by integrating retrieval capabilities for improved response accuracy and relevance. | |||
| The Power of Synthetic Data | Data Brew | Episode 38 | 04 févr. 2025 | 00:42:28 | |
In this episode, Yev Meyer, Chief Scientist at Gretel AI, explores how synthetic data transforms AI and ML by improving data access, quality, privacy, and model training. | |||
| Secret to Production AI: Tools & Infrastructure | Data Brew | Episode 37 | 22 janv. 2025 | 00:37:14 | |
In this episode, Julia Neagu, CEO & co-founder of Quotient AI, explores the challenges of deploying Generative AI and LLMs, focusing on model evaluation, human-in-the-loop systems, and iterative development. | |||
| Mixture of Memory Experts (MoME) | Data Brew | Episode 36 | 10 janv. 2025 | 00:41:24 | |
In this episode, Sharon Zhou, Co-Founder and CEO of Lamini AI, shares her expertise in the world of AI, focusing on fine-tuning models for improved performance and reliability. | |||
| Mixed Attention & LLM Context | Data Brew | Episode 35 | 21 nov. 2024 | 00:39:11 | |
In this episode, Shashank Rajput, Research Scientist at Mosaic and Databricks, explores innovative approaches in large language models (LLMs), with a focus on Retrieval Augmented Generation (RAG) and its impact on improving efficiency and reducing operational costs. | |||
| Kumo AI & Relational Deep Learning | Data Brew | Episode 34 | 14 oct. 2024 | 00:43:27 | |
In this episode, Jure Leskovec, Co-founder of Kumo AI and Professor of Computer Science at Stanford University, discusses Relational Deep Learning (RDL) and its role in automating feature engineering. | |||
| LLMs: Internals, Hallucinations, and Applications | Data Brew | Episode 33 | 21 juil. 2023 | 00:38:50 | |
Our fifth season dives into large language models (LLMs), from understanding the internals to the risks of using them and everything in between. While we're at it, we'll be enjoying our morning brew. | |||
| Demonstrate–Search–Predict Framework | Data Brew | Episode 32 | 29 juin 2023 | 00:33:14 | |
We will dive into LLMs for our fifth season, from understanding the internals to the risks of using them and everything in between. While we’re at it, we’ll be enjoying our morning brew. In this session, we interviewed Omar Khattab - Computer Science Ph.D. Student at Stanford, creator of DSP (Demonstrate–Search–Predict Framework), to discuss DSP, common applications, and the future of NLP. | |||
| Generative AI Risks | Data Brew | Episode 31 | 08 juin 2023 | 00:34:38 | |
We will dive into LLMs for our fifth season, from understanding the internals to the risks of using them and everything in between. While we’re at it, we’ll be enjoying our morning brew. | |||
| John Snow Labs & SparkNLP | Data Brew | Episode 30 | 01 juin 2023 | 00:43:17 | |
We are back and we will dive into LLMs from understanding the internals to the risks of using them and everything in between. While we’re at it, we’ll be enjoying our morning brew. | |||
| Data Brew Season 4 Episode 6: Professional Athletes | 09 juin 2022 | 00:35:49 | |
For our fourth season, we focus on connected health and how data & AI augment and improve our daily health. While we’re at it, we’ll be enjoying our morning brew. | |||
| Data Brew Season 4 Episode 5: Public Health: Education, Access, and Policy | 05 mai 2022 | 00:34:39 | |
For our fourth season, we focus on connected health and how data & AI augment and improve our daily health. While we’re at it, we’ll be enjoying our morning brew. | |||
| Data Brew Season 4 Episode 4: 1283 Days of Running (and Counting) | 14 avr. 2022 | 00:35:54 | |
For our fourth season, we focus on connected health and how data & AI augment and improve our daily health. While we’re at it, we’ll be enjoying our morning brew. | |||
| Data Brew Season 4 Episode 3: Last Man Standing | 31 mars 2022 | 00:41:20 | |
For our fourth season, we focus on connected health and how data & AI augment and improve our daily health. While we’re at it, we’ll be enjoying our morning brew. | |||
| Data Brew Season 4 Episode 2: NBA Analytics | 10 mars 2022 | 00:30:16 | |
For our fourth season, we focus on connected health and how data & AI augment and improve our daily health. While we’re at it, we’ll be enjoying our morning brew. | |||
| Data Brew Season 4 Episode 1: Reducing Injury & Increasing Retention of Industrial Athletes | 24 févr. 2022 | 00:33:58 | |
For our fourth season, we focus on connected health and how data & AI augment and improve our daily health. While we’re at it, we’ll be enjoying our morning brew. | |||
| Data Brew Season 3 Episode 6: Open Source | 28 oct. 2021 | 00:33:49 | |
For our third season, we focus on how leaders use data for change. Whether it’s building data teams or using data as a constructive catalyst, we interview subject matter experts from industry to dive deeper into these topics. See more at databricks.com/data-brew | |||
| Data Brew Season 3 Episode 5: Sustainability & Sake | 14 oct. 2021 | 00:32:26 | |
For our third season, we focus on how leaders use data for change. Whether it’s building data teams or using data as a constructive catalyst, we interview subject matter experts from industry to dive deeper into these topics. See more at databricks.com/data-brew | |||
| Data Brew Season 3 Episode 4: Executive Education | 07 oct. 2021 | 00:38:46 | |
For our third season, we focus on how leaders use data for change. Whether it’s building data teams or using data as a constructive catalyst, we interview subject matter experts from industry to dive deeper into these topics. See more at databricks.com/data-brew | |||
| Data Brew Season 3 Episode 3: 3 T’s to Securing AI Systems: Tests, tests, and more tests | 30 sept. 2021 | 00:35:01 | |
For our third season, we focus on how leaders use data for change. Whether it’s building data teams or using data as a constructive catalyst, we interview subject matter experts from industry to dive deeper into these topics. See more at databricks.com/data-brew | |||
| Data Brew Season 3 Episode 2: Data Culture Outside ‘The Valley’ | 23 sept. 2021 | 00:35:39 | |
For our third season, we focus on how leaders use data for change. Whether it’s building data teams or using data as a constructive catalyst, we interview subject matter experts from industry to dive deeper into these topics. | |||
| Data Brew Season 3 Episode 1: Disrupt: Challenge your Business Assumptions | 16 sept. 2021 | 00:29:45 | |
For our third season, we focus on how leaders use data for change. Whether it’s building data teams or using data as a constructive catalyst, we interview subject matter experts from industry to dive deeper into these topics. | |||
| Data Brew Season 2 Episode 9: Data Driven Software | 21 juil. 2021 | 00:31:12 | |
For our second season of Data Brew, we will be focusing on machine learning, from research to production. We will interview folks in academia and industry to discuss topics such as data ethics, production-grade infrastructure for ML, hyperparameter tuning, AutoML, and many more. | |||
| Data Brew Season 2 Episode 8: Feature Engineering | 09 juil. 2021 | 00:31:17 | |
For our second season of Data Brew, we will be focusing on machine learning, from research to production. We will interview folks in academia and industry to discuss topics such as data ethics, production-grade infrastructure for ML, hyperparameter tuning, AutoML, and many more. | |||
| Data Brew Season 2 Episode 7: Interpretable Machine Learning | 01 juil. 2021 | 00:37:07 | |
For our second season of Data Brew, we will be focusing on machine learning, from research to production. We will interview folks in academia and industry to discuss topics such as data ethics, production-grade infrastructure for ML, hyperparameter tuning, AutoML, and many more. | |||
| Data Brew Season 2 Episode 6: AutoML | 17 juin 2021 | 00:35:55 | |
For our second season of Data Brew, we will be focusing on machine learning, from research to production. We will interview folks in academia and industry to discuss topics such as data ethics, production-grade infrastructure for ML, hyperparameter tuning, AutoML, and many more. | |||
| Data Brew Season 2 Episode 5: ML Applications | 10 juin 2021 | 00:32:40 | |
For our second season of Data Brew, we will be focusing on machine learning, from research to production. We will interview folks in academia and industry to discuss topics such as data ethics, production-grade infrastructure for ML, hyperparameter tuning, AutoML, and many more. | |||
| Data Brew Season 2 Episode 4: Hyperparameter and Neural Architecture Search | 13 mai 2021 | 00:33:25 | |
For our second season of Data Brew, we will be focusing on machine learning, from research to production. We will interview folks in academia and industry to discuss topics such as data ethics, production-grade infrastructure for ML, hyperparameter tuning, AutoML, and many more. | |||
| Data Brew Season 2 Episode 3: Infrastructure for ML | 05 mai 2021 | 00:30:34 | |
For our second season of Data Brew, we will be focusing on machine learning, from research to production. We will interview folks in academia and industry to discuss topics such as data ethics, production-grade infrastructure for ML, hyperparameter tuning, AutoML, and many more. | |||
| Data Brew Season 2 Episode 2: Data Ethics | 28 avr. 2021 | 00:25:47 | |
For our second season of Data Brew, we will be focusing on machine learning, from research to production. We will interview folks in academia and industry to discuss topics such as data ethics, production-grade infrastructure for ML, hyperparameter tuning, AutoML, and many more. | |||
| Data Brew Season 2 Episode 1: ML in Production | 22 avr. 2021 | 00:30:49 | |
For our second season, we will be focusing on machine learning, from research to production. We will interview folks in academia and industry to discuss topics such as data ethics, production-grade infrastructure for ML, hyperparameter tuning, AutoML, and many more. | |||
| Data Brew Season 1 Episode 6: Journey of Big Data | 18 févr. 2021 | 00:40:16 | |
Jules Damji and Tathagata Das guide us through their journey in big data and the evolution of data architecture in the past 30 years. They discuss some of the biggest changes in industry they’ve seen, as well as trends to look forward to in the coming years. This is a fun episode connecting all four authors of the Learning Spark, 2nd Edition book. | |||
| Data Brew Season 1 Episode 5: Combining Machine Learning and MLflow with your Lakehouse | 06 janv. 2021 | 00:36:00 | |
Ellissa Verseput, ML Engineer at Quby, joins Denny and Brooke to discuss how Quby leverages ML to extract additional value from their data lake and how they manage this process. | |||
| Data Brew Season 1 Episode 4: BI on Data Lakes - Making it Real for Retail | 22 déc. 2020 | 00:29:05 | |
In this session, we discuss the lessons learned with Lara Minor, Senior Enterprise Data Manager at Columbia Sportswear, on how her team achieved a 70% reduction in pipeline creation time. This had reduced ETL workload times from four hours with previous data warehouses to minutes enabling near real-time analytics. Her team migrated from multiple legacy data warehouses, run by individual lines of business, to a single scalable, reliable, performant data lake. | |||
| Data Brew Season 1 Episode 3: Demystifying Delta Lake | 06 déc. 2020 | 00:25:51 | |
Delta Lake is an open source storage layer that brings reliability to data lakes. Delta Lake offers ACID transactions, scalable metadata handling, and unifies streaming and batch data processing. It runs on top of your existing data lake and is fully compatible with Apache Spark APIs. For our “Demystifying Delta Lake” session, we will interview Michael Armbrust - committer and PMC member of Apache Spark™ and the original creator of Spark SQL. He currently leads the team at Databricks that designed and built Structured Streaming and Delta Lake. See more at databricks.com/data-brew | |||
| Data Brew Season 1 Episode 2: Welcome to Lakehouse | 12 nov. 2020 | 00:26:10 | |
Legacy approaches have failed to deliver on the promise of a single data architecture that can support every downstream use case from BI to AI. Lakehouse aspires to address this by combining the best of data warehouses and data lakes. Ali Ghodsi, Co-Founder and CEO of Databricks, and David Meyer, SVP of Product at Databricks, explain how. | |||
| Data Brew Season 1 Episode 1: From data warehousing to data lakes in 40 minutes | 28 oct. 2020 | 00:44:48 | |
In our inaugural episode, we’d like to welcome data warehouse luminaries Barry Devlin, Susan O’Connell, and Donald Farmer to discuss the evolution of data warehouses, data lakes, and lakehouses. | |||