Explorez tous les épisodes du podcast TalkRL: The Reinforcement Learning Podcast
| Titre | Date | Durée | |
|---|---|---|---|
| NeurIPS 2024 - Posters and Hallways 3 | 09 Mar 2025 | 00:10:01 | |
Posters and Hallway episodes are short interviews and poster summaries. Recorded at NeurIPS 2024 in Vancouver BC Canada. Featuring
| |||
| NeurIPS 2024 - Posters and Hallways 2 | 05 Mar 2025 | 00:08:48 | |
Posters and Hallway episodes are short interviews and poster summaries. Recorded at NeurIPS 2024 in Vancouver BC Canada. Featuring
| |||
| NeurIPS 2024 - Posters and Hallways 1 | 03 Mar 2025 | 00:09:32 | |
Posters and Hallway episodes are short interviews and poster summaries. Recorded at NeurIPS 2024 in Vancouver BC Canada. Featuring
| |||
| Abhishek Naik on Continuing RL & Average Reward | 10 Feb 2025 | 01:21:40 | |
Abhishek Naik was a student at University of Alberta and Alberta Machine Intelligence Institute, and he just finished his PhD in reinforcement learning, working with Rich Sutton. Now he is a postdoc fellow at the National Research Council of Canada, where he does AI research on Space applications. Featured References Reinforcement Learning for Continuing Problems Using Average Reward Reward Centering Learning and Planning in Average-Reward Markov Decision Processes Discounted Reinforcement Learning Is Not an Optimization Problem
| |||
| Neurips 2024 RL meetup Hot takes: What sucks about RL? | 23 Dec 2024 | 00:17:45 | |
What do RL researchers complain about after hours at the bar? In this "Hot takes" episode, we find out! Recorded at The Pearl in downtown Vancouver, during the RL meetup after a day of Neurips 2024. Special thanks to "David Beckham" for the inspiration :) | |||
| RLC 2024 - Posters and Hallways 5 | 20 Sep 2024 | 00:13:17 | |
Posters and Hallway episodes are short interviews and poster summaries. Recorded at RLC 2024 in Amherst MA. Featuring:
| |||
| RLC 2024 - Posters and Hallways 4 | 19 Sep 2024 | 00:04:52 | |
Posters and Hallway episodes are short interviews and poster summaries. Recorded at RLC 2024 in Amherst MA. Featuring:
| |||
| RLC 2024 - Posters and Hallways 3 | 18 Sep 2024 | 00:06:43 | |
Posters and Hallway episodes are short interviews and poster summaries. Recorded at RLC 2024 in Amherst MA. Featuring:
| |||
| RLC 2024 - Posters and Hallways 2 | 16 Sep 2024 | 00:15:52 | |
Posters and Hallway episodes are short interviews and poster summaries. Recorded at RLC 2024 in Amherst MA. Featuring:
| |||
| RLC 2024 - Posters and Hallways 1 | 10 Sep 2024 | 00:05:46 | |
Posters and Hallway episodes are short interviews and poster summaries. Recorded at RLC 2024 in Amherst MA. Featuring:
| |||
| Finale Doshi-Velez on RL for Healthcare @ RCL 2024 | 02 Sep 2024 | 00:07:35 | |
Finale Doshi-Velez is a Professor at the Harvard Paulson School of Engineering and Applied Sciences. This off-the-cuff interview was recorded at UMass Amherst during the workshop day of RL Conference on August 9th 2024. Host notes: I've been a fan of some of Prof Doshi-Velez' past work on clinical RL and hoped to feature her for some time now, so I jumped at the chance to get a few minutes of her thoughts -- even though you can tell I was not prepared and a bit flustered tbh. Thanks to Prof Doshi-Velez for taking a moment for this, and I hope to cross paths in future for a more in depth interview. References
| |||
| David Silver 2 - Discussion after Keynote @ RCL 2024 | 28 Aug 2024 | 00:16:17 | |
Thanks to Professor Silver for permission to record this discussion after his RLC 2024 keynote lecture. Recorded at UMass Amherst during RCL 2024. Due to the live recording environment, audio quality varies. We publish this audio in its raw form to preserve the authenticity and immediacy of the discussion.
| |||
| David Silver @ RCL 2024 | 26 Aug 2024 | 00:11:27 | |
David Silver is a principal research scientist at DeepMind and a professor at University College London. This interview was recorded at UMass Amherst during RLC 2024. References
| |||
| Vincent Moens on TorchRL | 08 Apr 2024 | 00:40:14 | |
Dr. Vincent Moens is an Applied Machine Learning Research Scientist at Meta, and an author of TorchRL and TensorDict in pytorch. Featured References TorchRL: A data-driven decision-making library for PyTorch
| |||
| Arash Ahmadian on Rethinking RLHF | 25 Mar 2024 | 00:33:30 | |
Arash Ahmadian is a Researcher at Cohere and Cohere For AI focussed on Preference Training of large language models. He’s also a researcher at the Vector Institute of AI. Featured Reference Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs Arash Ahmadian, Chris Cremer, Matthias Gallé, Marzieh Fadaee, Julia Kreutzer, Olivier Pietquin, Ahmet Üstün, Sara Hooker Additional References
| |||
| Glen Berseth on RL Conference | 11 Mar 2024 | 00:21:38 | |
Glen Berseth is an assistant professor at the Université de Montréal, a core academic member of the Mila - Quebec AI Institute, a Canada CIFAR AI chair, member l'Institute Courtios, and co-director of the Robotics and Embodied AI Lab (REAL). Featured Links Reinforcement Learning Conference Closing the Gap between TD Learning and Supervised Learning--A Generalisation Point of View | |||
| Ian Osband | 07 Mar 2024 | 01:08:26 | |
Ian Osband is a Research scientist at OpenAI (ex DeepMind, Stanford) working on decision making under uncertainty. We spoke about: - Information theory and RL - Exploration, epistemic uncertainty and joint predictions - Epistemic Neural Networks and scaling to LLMs
Reinforcement Learning, Bit by Bit From Predictions to Decisions: The Importance of Joint Predictive Distributions Zheng Wen, Ian Osband, Chao Qin, Xiuyuan Lu, Morteza Ibrahimi, Vikranth Dwaracherla, Mohammad Asghari, Benjamin Van Roy
Ian Osband, Zheng Wen, Seyed Mohammad Asghari, Vikranth Dwaracherla, Morteza Ibrahimi, Xiuyuan Lu, Benjamin Van Roy
Ian Osband, Zheng Wen, Seyed Mohammad Asghari, Vikranth Dwaracherla, Morteza Ibrahimi, Xiuyuan Lu, Benjamin Van Roy
| |||
| Sharath Chandra Raparthy | 12 Feb 2024 | 00:40:41 | |
Sharath Chandra Raparthy on In-Context Learning for Sequential Decision Tasks, GFlowNets, and more! Sharath Chandra Raparthy is an AI Resident at FAIR at Meta, and did his Master's at Mila.
| |||
| Pierluca D'Oro and Martin Klissarov | 13 Nov 2023 | 00:57:24 | |
Pierluca D'Oro and Martin Klissarov on Motif and RLAIF, Noisy Neighborhoods and Return Landscapes, and more! Pierluca D'Oro is PhD student at Mila and visiting researcher at Meta.
Motif: Intrinsic Motivation from Artificial Intelligence Feedback To keep doing RL research, stop calling yourself an RL researcher | |||
| Martin Riedmiller | 22 Aug 2023 | 01:13:56 | |
Martin Riedmiller of Google DeepMind on controlling nuclear fusion plasma in a tokamak with RL, the original Deep Q-Network, Neural Fitted Q-Iteration, Collect and Infer, AGI for control systems, and tons more!
Neural fitted Q iteration–first experiences with a data efficient neural reinforcement learning method | |||
| Max Schwarzer | 08 Aug 2023 | 01:10:18 | |
Max Schwarzer is a PhD student at Mila, with Aaron Courville and Marc Bellemare, interested in RL scaling, representation learning for RL, and RL for science. Max spent the last 1.5 years at Google Brain/DeepMind, and is now at Apple Machine Learning Research. Featured References
| |||
| Julian Togelius | 25 Jul 2023 | 00:40:04 | |
Julian Togelius is an Associate Professor of Computer Science and Engineering at NYU, and Cofounder and research director at modl.ai
Featured References Julian Togelius, Georgios N. Yannakakis Learning Controllable 3D Level Generators Zehua Jiang, Sam Earle, Michael Cerny Green, Julian Togelius PCGRL: Procedural Content Generation via Reinforcement Learning Ahmed Khalifa, Philip Bontrager, Sam Earle, Julian Togelius Illuminating Generalization in Deep Reinforcement Learning through Procedural Level Generation Niels Justesen, Ruben Rodriguez Torrado, Philip Bontrager, Ahmed Khalifa, Julian Togelius, Sebastian Risi | |||
| Jakob Foerster | 08 May 2023 | 01:03:45 | |
Jakob Foerster on Multi-Agent learning, Cooperation vs Competition, Emergent Communication, Zero-shot coordination, Opponent Shaping, agents for Hanabi and Prisoner's Dilemma, and more. Jakob Foerster is an Associate Professor at University of Oxford. Featured References Learning with Opponent-Learning Awareness Model-Free Opponent Shaping Off-Belief Learning Adversarial Cheap Talk
| |||
| Danijar Hafner 2 | 12 Apr 2023 | 00:45:21 | |
Danijar Hafner on the DreamerV3 agent and world models, the Director agent and heirarchical RL, realtime RL on robots with DayDreamer, and his framework for unsupervised agent design! Danijar Hafner is a PhD candidate at the University of Toronto with Jimmy Ba, a visiting student at UC Berkeley with Pieter Abbeel, and an intern at DeepMind. He has been our guest before back on episode 11.
Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, Timothy Lillicrap
Deep Hierarchical Planning from Pixels [ blog ] Action and Perception as Divergence Minimization [ blog ]
| |||
| Jeff Clune | 27 Mar 2023 | 01:11:11 | |
AI Generating Algos, Learning to play Minecraft with Video PreTraining (VPT), Go-Explore for hard exploration, POET and Open Endedness, AI-GAs and ChatGPT, AGI predictions, and lots more! Professor Jeff Clune is Associate Professor of Computer Science at University of British Columbia, a Canada CIFAR AI Chair and Faculty Member at Vector Institute, and Senior Research Advisor at DeepMind.
Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videos [ Blog Post ] Robots that can adapt like animals Illuminating search spaces by mapping elites Paired Open-Ended Trailblazer (POET): Endlessly Generating Increasingly Complex and Diverse Learning Environments and Their Solutions First return, then explore | |||
| Natasha Jaques 2 | 14 Mar 2023 | 00:46:02 | |
Hear about why OpenAI cites her work in RLHF and dialog models, approaches to rewards in RLHF, ChatGPT, Industry vs Academia, PsiPhi-Learning, AGI and more! Dr Natasha Jaques is a Senior Research Scientist at Google Brain. Featured References Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog Basis for Intentions: Efficient Inverse Reinforcement Learning using Past Experience
| |||
| Jacob Beck and Risto Vuorio | 07 Mar 2023 | 01:07:05 | |
Jacob Beck and Risto Vuorio on their recent Survey of Meta-Reinforcement Learning. Jacob and Risto are Ph.D. students at Whiteson Research Lab at University of Oxford.
| |||
| John Schulman | 18 Oct 2022 | 00:44:21 | |
John Schulman is a cofounder of OpenAI, and currently a researcher and engineer at OpenAI.
WebGPT: Browser-assisted question-answering with human feedback Training language models to follow instructions with human feedback Additional References
| |||
| Sven Mika | 19 Aug 2022 | 00:34:56 | |
Sven Mika is the Reinforcement Learning Team Lead at Anyscale, and lead committer of RLlib. He holds a PhD in biomathematics, bioinformatics, and computational biology from Witten/Herdecke University.
RLlib Documentation: RLlib: Industry-Grade Reinforcement Learning RLlib: Abstractions for Distributed Reinforcement Learning
Register at raysummit.org and use code RAYSUMMIT22RL for a further 25% off the already reduced prices. | |||
| Karol Hausman and Fei Xia | 16 Aug 2022 | 01:03:09 | |
Karol Hausman is a Senior Research Scientist at Google Brain and an Adjunct Professor at Stanford working on robotics and machine learning. Karol is interested in enabling robots to acquire general-purpose skills with minimal supervision in real-world environments. Fei Xia is a Research Scientist with Google Research. Fei Xia is mostly interested in robot learning in complex and unstructured environments. Previously he has been approaching this problem by learning in realistic and scalable simulation environments (GibsonEnv, iGibson). Most recently, he has been exploring using foundation models for those challenges. Featured References Inner Monologue: Embodied Reasoning through Planning with Language Models Additional References
Register at raysummit.org and use code RAYSUMMIT22RL for a further 25% off the already reduced prices. | |||
| Sai Krishna Gottipati | 01 Aug 2022 | 01:08:11 | |
Saikrishna Gottipati is an RL Researcher at AI Redefined, working on RL, MARL, human in the loop learning. Featured References Cogment: Open Source Framework For Distributed Multi-actor Training, Deployment & Operations Do As You Teach: A Multi-Teacher Approach to Self-Play in Deep Reinforcement Learning Learning to navigate the synthetically accessible chemical space using reinforcement learning Additional References
Episode sponsor: Anyscale Register at raysummit.org and use code RAYSUMMIT22RL for a further 25% off the already reduced prices. | |||
| Aravind Srinivas 2 | 09 May 2022 | 00:58:33 | |
Aravind Srinivas is back! He is now a research Scientist at OpenAI. Featured References Decision Transformer: Reinforcement Learning via Sequence Modeling VideoGPT: Video Generation using VQ-VAE and Transformers | |||
| Rohin Shah | 12 Apr 2022 | 01:37:04 | |
Dr. Rohin Shah is a Research Scientist at DeepMind, and the editor and main contributor of the Alignment Newsletter. Featured References The MineRL BASALT Competition on Learning from Human Feedback Preferences Implicit in the State of the World Benefits of Assistance over Reward Learning On the Utility of Learning about Humans for Human-AI Coordination Evaluating the Robustness of Collaborative Agents
| |||
| Jordan Terry | 22 Feb 2022 | 01:03:48 | |
Jordan Terry is a PhD candidate at University of Maryland, the maintainer of Gym, the maintainer and creator of PettingZoo and the founder of Swarm Labs.
PettingZoo: Gym for Multi-Agent Reinforcement Learning
| |||
| Robert Lange | 20 Dec 2021 | 01:10:57 | |
Robert Tjarko Lange is a PhD student working at the Technical University Berlin. Featured References Learning not to learn: Nature versus nurture in silico On Lottery Tickets and Minimal Task Representations in Deep Reinforcement Learning Semantic RL with Action Grammars: Data-Efficient Learning of Hierarchical Task Abstractions MLE-Infrastructure on Github
| |||
| NeurIPS 2021 Political Economy of Reinforcement Learning Systems (PERLS) Workshop | 18 Nov 2021 | 00:24:07 | |
We hear about the idea of PERLS and why its important to talk about.
| |||
| Amy Zhang | 27 Sep 2021 | 01:09:35 | |
Amy Zhang is a postdoctoral scholar at UC Berkeley and a research scientist at Facebook AI Research. She will be starting as an assistant professor at UT Austin in Spring 2023. Featured References Invariant Causal Prediction for Block MDPs Multi-Task Reinforcement Learning with Context-based Representations MBRL-Lib: A Modular Library for Model-based Reinforcement Learning
| |||
| Xianyuan Zhan | 30 Aug 2021 | 00:41:30 | |
Xianyuan Zhan is currently a research assistant professor at the Institute for AI Industry Research (AIR), Tsinghua University. He received his Ph.D. degree at Purdue University. Before joining Tsinghua University, Dr. Zhan worked as a researcher at Microsoft Research Asia (MSRA) and a data scientist at JD Technology. At JD Technology, he led the research that uses offline RL to optimize real-world industrial systems. Featured References | |||
| Eugene Vinitsky | 18 Aug 2021 | 01:06:02 | |
Eugene Vinitsky is a PhD student at UC Berkeley advised by Alexandre Bayen. He has interned at Tesla and Deepmind.
A learning agent that acquires social norms from public sanctions in decentralized multi-agent settings Optimizing Mixed Autonomy Traffic Flow With Decentralized Autonomous Vehicles and Multi-Agent RL Lagrangian Control through Deep-RL: Applications to Bottleneck Decongestion The Surprising Effectiveness of PPO in Cooperative Multi-Agent Games
| |||
| Jess Whittlestone | 20 Jul 2021 | 01:31:36 | |
Dr. Jess Whittlestone is a Senior Research Fellow at the Centre for the Study of Existential Risk and the Leverhulme Centre for the Future of Intelligence, both at the University of Cambridge.
The Societal Implications of Deep Reinforcement Learning Artificial Canaries: Early Warning Signs for Anticipatory and Democratic Governance of AI
| |||
| Aleksandra Faust | 06 Jul 2021 | 00:54:30 | |
Dr Aleksandra Faust is a Staff Research Scientist and Reinforcement Learning research team co-founder at Google Brain Research. Featured References Learning Navigation Behaviors End-to-End with AutoRL Evolving Rewards to Automate Reinforcement Learning Evolving Reinforcement Learning Algorithms John D Co-Reyes, Yingjie Miao, Daiyi Peng, Esteban Real, Quoc V Le, Sergey Levine, Honglak Lee, Aleksandra Faust
Additional References
| |||
| Sam Ritter | 21 Jun 2021 | 01:40:35 | |
Sam Ritter is a Research Scientist on the neuroscience team at DeepMind. Featured References Unsupervised Predictive Memory in a Goal-Directed Agent (MERLIN) Meta-RL without forgetting: Been There, Done That: Meta-Learning with Episodic Recall Meta-Reinforcement Learning with Episodic Recall: An Integrative Theory of Reward-Driven Learning Meta-RL exploration and planning: Rapid Task-Solving in Novel Environments Synthetic Returns for Long-Term Credit Assignment Additional References
| |||
| Thomas Krendl Gilbert | 17 May 2021 | 01:12:14 | |
Thomas Krendl Gilbert is a PhD student at UC Berkeley’s Center for Human-Compatible AI, specializing in Machine Ethics and Epistemology. Featured References Mapping the Political Economy of Reinforcement Learning Systems: The Case of Autonomous Vehicles AI Development for the Public Interest: From Abstraction Traps to Sociotechnical Risks
| |||
| Marc G. Bellemare | 13 May 2021 | 00:57:40 | |
Professor Marc G. Bellemare is a Research Scientist at Google Research (Brain team), An Adjunct Professor at McGill University, and a Canada CIFAR AI Chair. Featured References The Arcade Learning Environment: An Evaluation Platform for General Agents
| |||
| Robert Osazuwa Ness | 08 May 2021 | 01:18:43 | |
Robert Osazuwa Ness is an adjunct professor of computer science at Northeastern University, an ML Research Engineer at Gamalon, and the founder of AltDeep School of AI. He holds a PhD in statistics. He studied at Johns Hopkins SAIS and then Purdue University.
| |||
| Marlos C. Machado | 12 Apr 2021 | 01:31:31 | |
Dr. Marlos C. Machado is a research scientist at DeepMind and an adjunct professor at the University of Alberta. He holds a PhD from the University of Alberta and a MSc and BSc from UFMG, in Brazil.
Revisiting the Arcade Learning Environment: Evaluation Protocols and Open Problems for General Agents Contrastive Behavioral Similarity Embeddings for Generalization in Reinforcement Learning [ video ] Efficient Exploration in Reinforcement Learning through Time-Based Representations A Laplacian Framework for Option Discovery in Reinforcement Learning [ video ] Eigenoption Discovery through the Deep Successor Representation Exploration in Reinforcement Learning with Deep Covering Options Autonomous navigation of stratospheric balloons using reinforcement learning Generalization and Regularization in DQN
| |||
| Nathan Lambert | 22 Mar 2021 | 00:50:35 | |
Nathan Lambert is a PhD Candidate at UC Berkeley. Featured References Learning Accurate Long-term Dynamics for Model-based Reinforcement Learning Objective Mismatch in Model-based Reinforcement Learning Low Level Control of a Quadrotor with Deep Model-Based Reinforcement Learning On the Importance of Hyperparameter Optimization for Model-based Reinforcement Learning
| |||
| Kai Arulkumaran | 16 Mar 2021 | 00:46:26 | |
Kai Arulkumaran is a researcher at Araya in Tokyo. Featured References AlphaStar: An Evolutionary Computation Perspective Analysing Deep Reinforcement Learning Agents Trained with Domain Randomisation Training Agents using Upside-Down Reinforcement Learning
| |||
| Michael Dennis | 26 Jan 2021 | 01:00:50 | |
Michael Dennis is a PhD student at the Center for Human-Compatible AI at UC Berkeley, supervised by Professor Stuart Russell. I'm interested in robustness in RL and multi-agent RL, specifically as it applies to making the interaction between AI systems and society at large to be more beneficial.--Michael Dennis
Emergent Complexity and Zero-shot Transfer via Unsupervised Environment Design [PAIRED] Adam Gleave, Michael Dennis, Cody Wild, Neel Kant, Sergey Levine, Stuart Russell Accumulating Risk Capital Through Investing in Cooperation
| |||
| Roman Ring | 11 Jan 2021 | 00:42:23 | |
Roman Ring is a Research Engineer at DeepMind. Featured References Grandmaster level in StarCraft II using multi-agent reinforcement learning Replicating DeepMind StarCraft II Reinforcement Learning Benchmark with Actor-Critic Methods
| |||