Back

Explore every episode of the podcast How I AI

Dive into the complete episode list for How I AI. Each episode is cataloged with detailed descriptions, making it easy to find and explore specific topics. Keep track of all episodes from your favorite podcast and never miss a moment of insightful content.

Rows per page:

1–50 of 116

TitlePub. DateDuration
How OpenAI uses ChatGPT Sites (live at DevDay!) | Kath Korevec (Product Lead)05 Oct 202600:35:40

Kath Korevec is a member of the Product staff at OpenAI working on Codex, and she spent over a year building and using ChatGPT Sites internally before its public launch. She’s been on the front lines of shipping Plugin Insights, MCP plugin hosting, and the connector ecosystem, which now includes around 60 integrations.


What you’ll learn:

  1. What Plugin Insights actually does, and why it changes who can use a site
  2. The incident command site Kath built for her OpenAI team, and how it uses live Slack and Notion connectors
  3. The one phrase that tells Codex to wire up connectors for you
  4. The infrastructure layer inside Sites that most people haven’t touched yet
  5. How Kath fixed her Spotify after her kids wrecked it, using Reddit and computer use
  6. The skill distribution model behind her community dungeon crawler, and why it’s a new way to think about collaboration
  7. Why model speed is what actually determines how creative you get
  8. Where Kath draws a hard line on AI acting in her name

—

Brought to you by:

Merge—Connective infrastructure for production AI

Vanta—Automate compliance and simplify security

—

In this episode, we cover:

(00:00) Welcome and intro

(01:00) Sites: internal testing and use cases

(05:53) Sites infrastructure

(09:00) Kath’s favorite connectors

(11:10) Unique ways to use Sites

(12:28) Curating a custom Spotify playlist

(15:50) Game design and development

(25:42) Bringing inference into Sites: the widget experiment

(30:00) Awesome Sites gallery

(31:57) Kath’s prompting strategy

(34:20) Wrap-up and how to find Kath

—

Tools referenced:

• ChatGPT Sites: https://chatgpt.com/sites

• OpenAI Codex: https://openai.com/codex

—

Other references:

• OpenAI DevDay: https://openai.com/devday

• Awesome Sites: https://awesomesites.ai

—

Where to find Kath Korevec:

LinkedIn: https://www.linkedin.com/in/kathleensimpson/?isSelfProfile=false

X: https://x.com/simpsoka

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

Jev: 8 real use cases for the fastest, cheapest model I’ve ever used | John Lindquist30 Sep 202600:46:06

John Lindquist created egghead.io, a developer education platform used by hundreds of thousands of working engineers. These days he’s building mega.dev, a hands-on program specifically for developers who want to do real work with AI agents, not just prototype them.


What you’ll learn:

  1. Why Jev is a decision engine, not a chatbot, and what that distinction actually changes about how you build
  2. How John built a real-time voice to-do app that classifies and executes commands with no visible pause
  3. The data deduplication pattern that merges messy records in milliseconds using confidence scores
  4. Why Jev works best as a router, and how a single text input can navigate users deep into an app
  5. What a chess match between Jev and a low-reasoning LLM reveals about speed, cost, and when to use which
  6. The multi-step classification pattern John reaches for when one Jev pass isn’t enough
  7. Where Jev falls short, and when you should still reach for a full generative model

—

Brought to you by:

Vanta—Automate compliance and simplify security

—

In this episode, we cover:

(00:00) John Lindquist returns for Jev week

(04:32) What Jev actually outputs

(06:15) Demo: real-time voice to-do app

(08:17) How sequential Jev calls chain together

(10:38) Demo: plain English to function name (grocery cart)

(11:50) Demo: data deduplication and record merging

(13:45) Confidence scores and multi-model validation

(15:06) Demo: Jev as a multi-level app router

(18:23) Architecting around Jev

(19:35) Demo: Jev vs. traditional LLM at chess (speed and cost benchmarks)

(24:29) DOM interactions as a decision set, not an infinite canvas

(28:21) Demo: Wikipedia “path to philosophy” route mapper

(30:28) Demo: multi-agent coordination and collision avoidance

(33:36) Demo: real-time presentation coach

(36:56) Quick recap

(39:54) Lightning round and final thoughts

—

Tools referenced:

• Jev (TypeSafe AI decision model): https://typesafe.ai/blog/introducing-system-one-models-and-jev

• Vercel AI Gateway: https://vercel.com/docs/ai-gateway

• OpenRouter: https://openrouter.ai

• Opus 5.5 (mentioned in context of iterative demo building): https://www.anthropic.com/claude-opus-5-5

—

Where to find John Lindquist:

LinkedIn: linkedin.com/in/john-lindquist-84230766

X: https://x.com/johnlindquist

Mega.dev: https://mega.dev/

Egghead.io: https://egghead.io/

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

OpenAI Dev Day 2026: The releases that actually matter30 Sep 202600:24:20

I spent the day at OpenAI’s DevDay in San Francisco, and I have good news and bad news: OpenAI released a lot of stuff.


In this episode, I break down the announcements worth paying attention to - and show you what happened when I tested some of them early. We’ll meet my Dot, explore why Spaces and Sites could matter for how teams work, and get into the model and API updates I’m most excited about as a developer.


I use the Decisions API to find podcast thumbnails where nobody looks awkward, build a collaborative sketchpad with Astra ultrafast, and let my kids redesign a 3D world in real time. That last experiment cost about $97. My wallet has thoughts.


These are my early impressions: what’s promising, what still feels rough, and what I think you should try first.


What you’ll learn:

  1. What OpenAI’s Dots can do, how I’ve been using mine, and why I’m waiting to give a full verdict
  2. Why Spaces might be one of the most underhyped announcements for collaboration between humans and agents
  3. How Sites with connectors and plugins could help teams share internal tools with the right data permissions
  4. Where GPT-6.1 Sol fits in my model stack—and why speed and cost matter
  5. What vision adds to the Decisions API, including my thumbnail-selection and hot dog demos
  6. What Astra ultrafast makes possible for interactive AI apps, from collaborative drawing to a changing 3D game
  7. Where the speed feels magical, where the experience still needs work, and what it costs

—

In this episode, we cover:

(00:00) OpenAI DevDay recap—and pressing the Codex reset button

(00:58) Dots: early impressions and rough edges

(06:57) Spaces: working with humans and agents

(10:37) Sites, connectors, and sharing internal tools

(13:06) Models and platform: GPT-6.1 Sol

(14:36) Decisions API: fast decisions with vision

(15:27) Finding better podcast thumbnails with AI

(16:29) Hot dog or not hot dog?

(17:17) Astra ultrafast: speed, pricing, and possibilities

(18:50) The Other Pencil: drawing alongside AI

(19:45) Little Starship: a 3D world you can change with a prompt

(21:24) The $97 AI game—and what it makes possible

(22:25) Agents API, computer use, plugins, and plan updates

(23:03) What I’d try first

—

Tools referenced:

• ChatGPT: Dots, Spaces, and Sites: https://chatgpt.com/

• Codex: https://openai.com/codex/

• OpenAI API — GPT-6.1 Sol, Decisions API, and Astra ultrafast: https://platform.openai.com/

• Jev: https://typesafe.ai/

—

Other references:

• OpenAI DevDay 2026: https://devday.openai.com/

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

Jev for beginners: how to use it and what to build28 Sep 202600:26:24

Jev is TypeSafe AI’s new decision model. It returns type-safe structured values (a choice, a score, a probability) instead of generated text, at 4 cents per million input tokens with no output charge. This week I ran it on five real projects: PR categorization, a meta-analysis of my own Claude and Codex sessions, Gmail triage, the ChatPRD product insights graph, and a live audience dashboard built from 4,500 YouTube comments.


What you’ll learn:

  1. What makes Jev fundamentally different from every other model I’ve used
  2. How I analyzed 1,700 PRs for 9 cents and what I found out about where my engineering effort actually went
  3. The personal meta-analysis you can run on your own Claude and Codex sessions right now
  4. Why I stopped using Jev alone, and what I pair it with now
  5. How I turned 4,500 YouTube comments into a searchable audience dashboard for almost nothing
  6. The real-time app I built in an afternoon that shows something surprising about Jev’s speed
  7. Why Jev’s pricing model is different from any LLM I’ve used, and what it makes practical to build
  8. The ChatPRD product insights project: 1,100 signals, 200,000 classifications, and what it cost me

—

Brought to you by:

OpenArt—An all-in-one AI creation platform for images, videos, music, audio, and more

—

In this episode, we cover:

(00:00) Jev launch and what makes it different from every other model

(02:49) Type-safe values explained

(05:28) Understanding Jev outputs

(07:39) Use case 1: PR categorization and pairwise clustering

(11:12) Use case 2: analyzing your own local Claude Code and Codex sessions

(13:00) Use case 3: Gmail triage with Jev scoring and LLM follow-up

(14:30) Use case 4: ChatPRD’s product insights graph

(18:17) Demo: How I AI audience signal dashboard

(22:14) Demo: voice-to-color emotion-mapping app

(25:16) Jev week recap and what’s coming in episode 2

—

Tools referenced:

• Jev (TypeSafe AI): https://typesafe.ai

• Vercel: https://vercel.com/ai

• GitHub API: https://docs.github.com/en/rest

• YouTube Data API v3: https://developers.google.com/youtube/v3

• OpenAI Realtime Voice API: https://platform.openai.com/docs/guides/realtime

• Gemini 3.5 Flash-Lite: https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash-lite

• API Ninjas Quotes API: https://api-ninjas.com/api/quotes

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

Opus 5.5 vs. GPT-6 Sol: which model won my blind taste test?22 Sep 202600:38:51

I got up early to record an Opus 5.5 review. Then Anthropic and OpenAI dropped new models on the same morning, and I decided to do something I’d never done before: take the How I AI bench live. I put GPT-6 Astra, GPT-6 Sol, Claude Opus 5.5, and more through the work I actually care about: emails, PRDs, frontend prototypes, backend work, long-running agents, SVGs, and video editing. I scored the outputs without knowing which model made them, so you get to watch me make predictions, change my mind, and reveal my own very inconsistent taste. Astra won my heart. Opus 5.5 won my week. Sol still has me split. There’s a creative result I got completely wrong, an LLM judge that disagreed with me, and a return to Barbie Bench: the 3D fashion game that keeps reminding me how far we have to go. The hands are tragic. AGI has not arrived.


What you’ll learn:

  1. How I run the How I AI bench blind, and what gets an output a bad score before I even know which model made it
  2. Why Astra won my heart while Opus 5.5 might be overall strongest, especially for long-running agents and B2B frontend
  3. Where Sol still wins me over on clear writing, readable PRDs, and price
  4. The character SVG results that completely overturned my prediction about Anthropic
  5. What happened when I asked these models to edit video, and why I think skills explain part of the disappointment
  6. Why an LLM judge disagreed with my rankings, and what it was rewarding that I wasn’t

—

In this episode, we cover:

(00:00) LIVE setup and new model launches

(01:30) What’s new in Opus 5.5, Sol, and Luna

(04:11) Guardrails, personality, and speed

(09:00) The How I AI bench and blind evaluation process

(11:31) Email and personal-productivity results

(13:50) Frontend prototype vibe checks

(24:10) Backend, agent personality, and long-running tasks

(28:25) SVG illustration test

(29:48) AI video-editing results

(30:43) Predictions before the reveal

(31:20) Barbie Bench: the 3D fashion-game test

(34:17) Results: Astra, Sol, and Opus 5.5

(35:04) Writing clarity and creative surprises

(36:51) Why the LLM judge disagreed with me

(37:24) What each model is actually best for

—

Tools referenced:

• Claude Opus 5.5: https://www.anthropic.com/claude-opus-5-5

• GPT-6 Sol and Luna: https://openai.com/index/introducing-gpt-6-sol-and-luna/

• Codex (OpenAI): https://openai.com/codex

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

I left Claude for months. Opus 5.5 is why I'm back22 Sep 202600:24:52

I’ve been off Claude for months. Not because it got dumb, but because it got annoying. The rambling, the hedging, the preachy little disclaimers on tasks that didn’t need them. I moved most of my daily work to Codex and I didn’t miss it. Then Anthropic shipped Opus 5.5: 40% cheaper than Opus 5, faster, and with what they’re calling a fundamentally different alignment approach. I ran it for a week across real work, including four long-running agentic tasks, a full ChatPRD homepage redesign, an SVG benchmark, and one very firm refusal, and I’m ready to give you the honest verdict. There’s a lot to like. There are still two things that drive me a little crazy. And there’s one capability I genuinely wasn’t expecting.


What you’ll learn:

  1. Why I walked away from Claude entirely, and what it took for me to come back
  2. The real cost math on Opus 5.5 and why pricing matters more for agentic work than single prompts
  3. What happened when I ran four long-running agentic tasks, including one that tried to manipulate Claude mid-run
  4. Why Opus 5.5 is now my go-to for frontend prototyping, and where it still lets me down
  5. The one capability I genuinely didn’t see coming, and no other model in my stack can match it
  6. The moment Opus 5.5 told me flat-out no, and what that says about where Anthropic’s safety posture actually lands in practice
  7. Where Codex still wins, and how I’m splitting my model stack after a full week of testing

—

In this episode:

(00:00) Why I stopped using Claude

(01:02) What Anthropic says Opus 5.5 is

(01:54) Cost, speed, and benchmark overview

(03:20) Safety, alignment, and the cybersecurity limits

(05:02) How I AI bench

(05:39) Voice test: is it actually not annoying?

(07:54) Long-running agentic task results

(10:50) Frontend prototyping

(17:23) Writing voice and email

(19:41) SVG illustrations

(20:46) Video editing

(21:42) My verdict: what it’s good at, what it still isn’t

—

Tools referenced:

• Claude Opus 5.5: https://www.anthropic.com/claude-opus-5-5

• ElevenLabs MCP connector: https://elevenlabs.io/mcp

• Codex (OpenAI): https://openai.com/codex

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

How Warp ships 2,000 PRs a month with AI factories | Zach Lloyd (CEO, Warp)21 Sep 202600:46:51

Zach Lloyd is the co-founder and CEO of Warp, an AI-powered terminal and software factory platform used by tens of thousands of engineers. Before Warp, he spent nearly a decade at Google, including time as a principal engineer on Google Sheets. He built Warp from the ground up as a modern, AI-native alternative to legacy terminals, and the team has since expanded into software factories: a full cloud-based system that takes an idea in Slack all the way through to a merged PR.


In this episode:

  1. Why a software factory is more than a coding agent
  2. The public Slack → Linear → GitHub → QA workflow
  3. Human interactions per PR as a signal of automation and throughput
  4. Why human review is still the bottleneck
  5. Scoring agent runs, finding failure modes, and self-improving agent workflows
  6. Replaying real tasks to choose model cost and quality tradeoffs
  7. CEO workflows with Figma MCP, Granola, and research agents

—

Brought to you by:

DX—Engineering intelligence for the AI era

OpenArt—An all-in-one AI creation platform for images, videos, music, audio, and more

—

In this episode, we cover:

(00:00) Intro

(02:35) Warp’s AI software factory, Wilson

(09:23) Automatic factory triggers

(11:12) The engineering leader dashboard Zach wishes he’d had

(15:18) How code review is changing in an AI factory

(17:08) Tracking cost per PR across model configs

(18:47) Using LLM-as-a-judge to score every agent run

(20:02) Catching redundant tests

(22:19) How the factory self-improves from failed runs

(26:03) Quick recap

(28:33) Building a cost-quality Pareto chart for model selection

(31:35) How Zach uses AI for non-technical CEO work

(32:10) Figma MCP demo

(35:43) Granola MCP demo

(36:41) GOG CLI demo

(38:20) Thinking in parallel tasks instead of sequential ones

(40:42) Zach’s prompting strategy for factory tasks

(44:48) Where to find Zach

—

Tools referenced:

• Warp (AI terminal and software factories): https://warp.dev

• Warp Factories: https://warp.dev/factories

• Linear (project and issue tracking): https://linear.app

• GitHub (version control and PR management): https://github.com

• Slack (team communication and factory input layer): https://slack.com

• Sentry (crash reporting and automated issue triggers): https://sentry.io

• Figma (design, used via Figma MCP): https://figma.com

• Granola (AI meeting notes and MCP integration): https://granola.so

• Grok Bot (fast inference, cost/quality trade-off): https://x.ai/bot/guides/grok-bot-101

—

Where to find Zach:

X: https://x.com/ZachLloydTweets

—

Where to find Claire:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

Muse review: The personal AI agent that gets consumer UX right16 Sep 202600:37:11

I spent a few hours putting Meta’s Muse, its new personal AI agent, through a real first-pass test: onboarding, calendar management, goal setting, a one-shot family morning newsletter, browser-based shopping, and the animated avatar that honestly surprised me.


What you’ll learn:

  1. Why Muse is the best-designed personal agent I’ve tested, and what specifically made it feel that way
  2. The one-shot family PDF Muse produced that Claude and Codex never quite nailed
  3. How Muse’s permission model works, and why it’s different from every other agent I’ve used
  4. Why I set up a sleep training goal in Muse, and what it revealed about agent tone
  5. The activity feed feature I immediately wished Codex and Claude Code had
  6. Where Muse failed, and what it says about the limits of this category right now
  7. The animated avatar decision that showed me what top-of-craft AI product design actually looks like

—

Brought to you by:

Optimizely—Your AI agent orchestration platform for marketing and digital teams

OpenArt—An all-in-one AI creation platform for images, videos, music, audio, and more

—

In this episode, we cover:

(00:00) What Muse is and who it’s actually built for

(04:41) Signing in and the onboarding flow

(07:16) The activity feed and its task lineage

(08:24) First real task: managing the family calendar and deleting soccer practice

(09:48) Requesting a morning newsletter PDF

(14:29) The personalized news feed and how I set it up

(16:05) The “Ideas” feature as an out-of-the-box prompt library

(17:10) Setting up personal goals (water, shoes, and sleep training)

(21:40) Library: documents, websites, images, videos, and podcasts

(23:11) Quick recap and what I love

(23:56) Activity feed design deep dive: tool calls and step-by-step lineage

(25:18) How Muse handles permissions

(26:12) The animated avatar: Polly becomes Slime, the teal dragon

(29:34) Browser use test: shopping for New Balance 9060s (not great)

(31:15) Browser use test 2: buying IMAX tickets for The Odyssey (much better)

(33:34) TL;DR and what I’ll actually use Muse for going forward

—

Tools referenced:

• Muse: https://muse.ai/

• Stripe Link (payment method featured in Muse): https://link.com

• 1Password (future Muse integration mentioned): https://1password.com

• OpenClaw (Claire’s previous personal agent setup): https://openclaw.ai/

• Grok Bot (Grok-based agent from prior stack): https://x.ai/news/introducing-grok-bot

• Codex (OpenAI coding agent, comparison point): https://openai.com/codex

• NotebookLM (Google, comparison to Muse’s podcast generation): https://notebooklm.google.com

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

How Grok Bot designers use AI agents to build personal sites and product prototypes | John Bai & Peng Zheng14 Sep 202600:41:34

John Bai and Peng Zheng are designers on the Grok Bot team at SpaceXAI, where they’re building one of the most talked-about AI products right now. John writes publicly about his design process (his piece “Designing Grok Bot with Grok Bot” has already made the rounds) and shares bot templates with the design community. Peng brings a product-design sensibility to personal tools, and his website doubles as a live demo of what he builds.


What you’ll learn:

  1. How Peng built a self-updating personal website using Grok Bot as the entire backend pipeline, with no CMS and no Figma file
  2. The exact check-in bot setup that lets Peng send a photo or a place name and have his portfolio update itself automatically
  3. How John’s Figma Bro bot handles production design tasks while he’s at the gym
  4. How John uses voice memos to direct Figma work through an MCP connection without opening his laptop
  5. The “shower thought to prototype” workflow John uses with DevBot to test interaction ideas without first going through a product manager or engineer
  6. The “trash can method” of software development
  7. How both designers organize their personal bot ecosystems
  8. What John and Peng actually think AI means for the future of design as a craft

—

Brought to you by:

WorkOS—Make your app enterprise-ready, with SSO, SCIM, RBAC, and more

Vanta—Automate compliance and simplify security

—

In this episode, we cover:

(00:00) Introducing John and Peng

(02:53) The Grok Bot hype train

(04:35) Peng’s self-updating personal website built with Grok Bot

(15:12) How AI makes design more accessible

(19:35) Website update result

(20:13) John’s Figma Bro bot

(23:35) Creating marketing materials for the bot marketplace

(26:00) DevBot: from shower thoughts to working prototypes

(28:48) The trash can method of software development

(31:19) Other bots John and Peng are using

(39:05) Practical tips for when bots don’t do what you want

—

Tools referenced:

• Grok Bot (xAI): https://x.ai/bot

• Figma: https://www.figma.com

• Figma MCP server: https://www.figma.com/mcp-catalog/

• Google Places API: https://developers.google.com/maps/documentation/places/web-service

• Notion: https://www.notion.so

• Swarm (Foursquare): https://www.swarmapp.com

—

Other references:

• Designing Grok Bot with Grok Bot: https://x.ai/bot/guides/designing-grok-bot-with-grok-bot

• Figma Bro bot template (shared by John Bai): https://x.ai/bot/marketplace/bots/figma-bro

• From zero coding background to hardware hacker: How Cursor + a Raspberry Pi makes AI fun: https://www.lennysnewsletter.com/p/from-zero-coding-background-to-hardware?utm_source=publication-search

—

Where to find John and Peng:

John Bai on X: https://x.com/johnbai

Peng Zheng on X: https://x.com/pengzheng_

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

Build your own company brain: the enterprise AI playbook from Stripe’s engineering team | Sharadh Krishnamurthy07 Sep 202600:50:24

Sharadh Krishnamurthy is an engineering manager at Stripe, where he helped build Kai, the company’s internal AI agent used by more than 10,000 employees every week. He’s worked across several of Stripe’s core infrastructure teams, including data and developer experience, which gives him a grounded, systems-level perspective on what it actually takes to make AI work at enterprise scale. He’s currently focused on the governance, skills, and infrastructure layers that let every Stripe employee use AI safely and effectively, regardless of their technical background.


What you’ll learn:

  1. Why Stripe built Kai from scratch instead of buying, and what tipped the decision
  2. What Kai knows about you by default and what you actually control
  3. Why “projects” at Stripe are a governance mechanism, not just a folder
  4. How Stripe structured its data layer so agents can query safely at scale
  5. Why the infrastructure Stripe built for human developers turned out to be exactly what agents needed
  6. How Kai’s skills platform lets any employee package a workflow, and what happens when you have 2,000 of them
  7. What Sharadh learned the hard way when agents nearly took down production systems

—

Brought to you by:

DX—Engineering intelligence for the AI era

Hyperagent—Deploy fleets of agents that handle real work

—

In this episode, we cover:

(00:00) Introducing Sharadh

(02:46) Why Stripe built an AI agent (Kai) instead of buying tools

(05:18) What Kai knows about you (and what you can turn off)

(06:51) Projects as a governance layer

(10:04) Live demo: Kai builds a dashboard

(12:18) Tools, skills, and the secure sandbox

(17:22) Why Stripe has benefited so much from AI

(19:20) Agentic identity, load shedding, and rogue agents

(20:41) Iterating on the dashboard

(25:01) How they rolled out Kai across the team

(29:07) How projects work

(34:18) Bespoke agents for bespoke use cases

(35:58) The skill builder workflow

(40:40) Skill quality, evals, and telemetry

(43:01) Recap

(45:13) Lightning round

—

Tools referenced:

• Trino: https://trino.io/

• Anthropic: https://www.anthropic.com/

• Gemini: https://gemini.google.com/

• Cursor: https://www.cursor.com/

—

Where to find Sharadh Krishnamurthy:

LinkedIn: https://www.linkedin.com/in/sharadhk

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

GPT-6 Astra is a banger - here’s everything I’ve built03 Sep 202600:32:19

I got early access to GPT-6 Astra: when I say this model broke through tasks I couldn’t crack with 5.6 Sol or Fable, I mean it specifically: the ChatPRD product intelligence feature, building 3D games, the hardware hack, and a handful of one-shot coding projects I’d tried and failed on repeatedly.


What you’ll learn:

  1. Why Astra’s computer use feels different, and which production tools I’m trusting it with
  2. The one feature I’d thrown every model at for six months, and what finally got it to 90%
  3. How I’m using browser use for QA, not building, and what it found that I would’ve missed
  4. Why I think UI is genuinely back, and what that means for SaaS and MCPs
  5. The hardware hack I’d been chasing since GPT-5.5, and how Astra finally cracked it
  6. What Astra built me in Blender in one shot, and why 3D is my new capability benchmark
  7. The AIM-style Mac app Astra made in one shot, and what it signals about desktop development now
  8. An honest take on speed, cost, and whether Astra is worth making your daily driver

—

In this episode, we cover:

(00:00) GPT-6 Astra overview

(03:44) Browser/computer use test on my CRM

(09:08) Flora thumbnail generation

(13:00) Browser use for QA

(15:20) Coding: ChatPRD product intelligence feature, finally one-shotted

(18:36) Hardware hack: Divoom MiniToo CLI and live streaming display

(22:23) Building an AIM-style Mac app

(24:24) Blender and 3D assets: Barbie Bench and the kids’ family app

(28:52) Summary: what Astra is great at and what to try first

—

Tools referenced:

• GPT-6 Astra: https://openai.com/index/gpt-6-astra/

• Codex: https://openai.com/codex

• Flora (node-based AI image/video editing): https://flora.ai/

• Figma: https://www.figma.com

• Blender: https://www.blender.org

• GPT Image 2: https://developers.openai.com/api/docs/models/gpt-image-2

• Divoom MiniToo: https://divoom.com/products/minitoo

• cxo.dev: https://www.cxo.dev/

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

Grok Bot vs. OpenClaw: How I replaced my entire agent stack02 Sep 202600:36:17

I’m running about 30 active agents at any given moment, and in this episode I break down my full Grok Bot setup: what it is, how it compares to OpenClaw, and the nine bots I’ve built for work and my personal life. We go deep on Chief (my chief-of-staff bot sweeping six inboxes and multiple Slack workspaces), TradBot (the family agent that prints a kitchen-table newspaper for my kids), two engineering bots handling my PR queue and SOC 2 compliance monitoring, Holly Helpdesk, and a handful of personal bots I didn’t expect to actually love. I also walk through how I migrated everything from OpenClaw, including the script I used to export and transplant each agent’s identity and schedule.


What you’ll learn:

  1. The three primitives Grok Bot is built on, and why one of them changes what agents can actually do
  2. How Chief, my general-purpose chief of staff, handles a scope I didn’t think a single bot could manage
  3. The writing quirk I noticed immediately with the Grok model, and what I did about it before letting it near my inbox
  4. Why I created a family agent, what it produces every morning, and the design principle I used that has nothing to do with a screen
  5. The two engineering bots doing work I used to do myself, and how one of them handles compliance in a way that surprised me
  6. How Holly Helpdesk started getting five-star reviews from customers who had no idea they were talking to a bot
  7. The personal bots I built mostly on a whim, and the one I now look forward to every Monday morning

—

Brought to you by:

WorkOS—Make your app enterprise-ready, with SSO, SCIM, RBAC, and more

Hyperagent—Deploy fleets of agents that handle real work

—

In this episode, we cover:

(00:00) Why I migrated from OpenClaw to Grok Bot

(02:10) Grok Bot overview: the three core primitives

(07:41) Chief: my chief-of-staff bot

(11:51) OpenClaw vs. Grok Bot

(12:41) TradBot: my family agent

(19:51) LGTM the PR Closer

(22:03) Lockdown: SOC 2 control monitoring bot

(24:00) Holly Helpdesk: customer support

(26:58) Penny Pincher: subscription audit, insurance negotiation, Rolex shopping

(29:37) ShopZilla and Sylvie Style: personal shopping and wardrobe bots

(32:54) How to migrate your OpenClaws

(34:26) Final take

—

Tools referenced:

• Grok Bot (SpaceXAI multi-agent platform): https://x.ai/news/introducing-grok-bot

• OpenClaw (previous agent platform): https://openclaw.ai/

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

How I turned Claude into a self-improving PM assistant | Daniel Blum (PM, Melio)31 Aug 202600:46:07

Daniel Blum is a product manager at Melio, a B2B payments company, and one of the most systematic thinkers I’ve had on the show when it comes to personal AI infrastructure. He’s spent the past year building a Claude- and Cowork-based productivity system that manages his Notion board, processes his Slack and email, and runs self-improvement loops every week without needing to be prompted. Beyond his own workflow, Daniel built and scaled a “Workstation” onboarding plugin that gets any Melio employee up and running with a personalized Claude setup in about 15 minutes.


What you’ll learn:

  1. Why Daniel says the two rules that make any AI system powerful aren’t about the tool you pick
  2. How his weekly prep automation fills an entire Notion board from scratch every Sunday, without his touching it
  3. The morning brief feature that teaches Claude new internal terms on its own, so company jargon never slows it down
  4. Why he describes Notion as “read-only” now, and what that says about how PM workflows are changing
  5. The self-improvement loop that watches Daniel’s edits, spots recurring friction, and suggests new skills to build
  6. How he uses a skill called “Improve” to filter the endless flood of AI tips without drowning in them
  7. What he built to scale his personal system to every PM at Melio, and the UX lesson he learned the hard way
  8. The capability gap that’s still keeping him from running 100% of his work through Claude

—

Brought to you by:

Optimizely—Your AI agent orchestration platform for marketing and digital teams

Jira AI SDLC—Get your tokens’ worth with Jira

—

In this episode, we cover:

(00:00) Daniel’s background and the PM overhead problem he needed to solve

(03:30) His AI stack at Melio

(05:00) The two rules that make any AI system genuinely powerful

(06:00) The Notion board Cowork built for him (and manages on his behalf)

(07:30) How he contextualizes Claude with voice memos, links, and recurring updates

(09:00) His weekly prep automation

(11:00) His morning brief

(15:00) How Claude flags unknown internal terms and saves them to context

(17:30) Running 70% to 80% of his workday through Cowork

(19:00) Chrome connector vs. MCPs for tools without integrations

(20:00) The real ROI question: why the early weeks feel slow, and why you push through anyway

(25:00) Scaling the system to the team with the Workstation plugin

(26:30) The self-improvement loop

(31:00) How the Improve skill separates actually useful AI tips from the hype

(32:00) The Workstation onboarding flow, and the UX lesson from distributing “Spectacular”

(38:00) The 20% Claude still can’t do, and what changes when it can

(41:00) What Daniel spends his reclaimed time on

(42:30) Claude rage

—

Tools referenced:

• Claude: https://claude.ai

• Notion: https://notion.so

—

Other references:

• From a $6.90 newsletter to $3M API: How a non-coder built Memelord | Jason Levin: https://www.lennysnewsletter.com/p/from-a-690-newsletter-to-3m-api-how?utm_source=publication-search

• How the founder of Morning Brew built a Claude content machine that never runs out of ideas and never sounds like slop | Alex Lieberman: https://www.lennysnewsletter.com/p/how-the-founder-of-morning-brew-built?utm_source=publication-search

—

Where to find Daniel Blum:

LinkedIn: https://www.linkedin.com/in/blumd/

Website: https://www.imdanielblum.com

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

I spent $20,000 on Devin in a month. Here’s what I learned | Ryan Carson (solo founder)24 Aug 202600:44:13

Ryan Carson is a five-time founder and the current solo founder of Untangle, a B2B SaaS platform for family law firms. Before Untangle, he co-founded Treehouse, an online coding education platform, and has spent the better part of two decades building and leading tech companies. He’s active on X, where he shares his solo founder journey in real time, including what he actually spends on AI tools each month.


What you’ll learn:

  1. Why Ryan manages 15 concurrent Devin agents with a folder system and a piece of paper, not a dashboard
  2. The Watchdog playbook: what he built to replace a customer success team across every law firm account
  3. How his LAN PR skill closes the loop on 40 daily PRs without a QA team reviewing a single one
  4. Why he moved off local agents almost entirely, and the one situation where he still reaches for Codex
  5. The design workflow we’re both using: Claude Design into a Markdown spec, then Codex to build the real thing
  6. What he found when he got away from his computer and met a real customer, and why it changed his entire product direction
  7. How he’s hiring his first engineer without a single phone screen or interview
  8. Why we both think more AI output is actually the wrong goal, and what to optimize for instead

—

Brought to you by:

WorkOS—Make your app enterprise-ready, with SSO, SCIM, RBAC, and more

Jira AI SDLC—Get your tokens’ worth with Jira

—

In this episode, we cover:

(00:00) Introduction to Ryan Carson

(02:55) Ryan’s update: Untangle, the divorce PMF pivot, and B2B growth

(07:35) Ryan’s current Devin stack: folders, P0 threads, and the paper list

(16:29) Watchdog playbook: account monitoring across every firm

(18:09) Managing agent decision fatigue: Ryan’s method vs. Claire’s

(20:32) Producing more output does not make a better product

(22:47) Using cloud agents for ops beyond just code

(25:14) When to use Codex vs. Devin vs. Claude Code

(27:15) Merge Mommy recap

(28:15) LAN PR skill: review loops, video walkthrough, and auto-merge

(30:20) Slack vs. Devin threads for async team communication

(35:58) Claude Design plus Codex for building a technical design system

(39:00) EA tools: Polly the OpenClaw vs. Claude Code on a Mac Mini

(42:09) Quick recap and final thoughts

—

Tools referenced:

• Devin (Cognition): https://www.cognition.ai/

• Codex (OpenAI): https://openai.com/codex

• Claude Code/Claude Design (Anthropic): https://www.anthropic.com/claude

• OpenClaw (Claude-based desktop client): https://openclaw.ai

• Cursor: https://www.cursor.com/

• BugBot (Devin’s built-in PR review): https://cursor.com/bugbot

• Sentry (error monitoring referenced in Watchdog): https://sentry.io/

• Ugmonk (analog to-do system): https://ugmonk.com/

—

Other references:

• Devin playbooks/skills documentation: https://docs.cognition.ai/

• Jack Dorsey/Buzz (async-first communication referenced): https://buzz.new/

• Merge Mommy (Claire’s Eve agent for PR risk scoring, deployed on Vercel): https://www.lennysnewsletter.com/p/build-an-ai-code-review-bot-in-30

—

Where to find Ryan Carson:

X: https://x.com/ryancarson

Untangle: https://untangle.us

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

I tested Grok Bot, Grok 4.6, and Cursor Origin - here’s my honest take18 Aug 202600:27:14

This week I’m doing a solo breakdown of everything xAI and Cursor have shipped recently, including Grok Bot, Cursor Origin, and the Grok 4.6 model. I set up five Grok Bots, ran Grok 4.6 through my Claire Weighted Index against GPT-5.6 Sol, Claude Sonnet 5, and Opus 5, and spent time actually using Origin as a GitHub replacement. Here’s what’s worth your attention, what’s overhyped, and where I’m personally putting my time.


What you’ll learn:

  1. The one Grok Bot feature no other agent platform has shipped yet, and why it made me actually use the product
  2. What a week of real Grok Bot use revealed, and why I still reach for my OpenClaws
  3. Whether Cursor Origin is a GitHub replacement or just a pretty redesign
  4. Where Grok 4.6 landed on the Claire Index, and the one category where it genuinely surprised me

—

Brought to you by:

Bolt.new—Turn your idea into a real product

Jira AI SDLC—Get your tokens’ worth with Jira

—

In this episode, we cover:

(00:00) Why everyone’s quietly switching to Grok

(01:52) Grok Bot overview and setup

(03:22) My 5 Grok Bots

(04:30) The killer feature: multi-account connectors

(06:07) Grok Bot’s virtual machine and how it actually works

(06:41) Experience overview

(07:35) What I don’t love about Grok Bot

(10:08) Grok Bot use cases and my honest verdict

(12:20) Cursor Origin: the agent-native GitHub replacement

(13:47) What Origin actually looks like in practice

(14:59) Why I’m not switching from GitHub yet

(17:42) What would get me to move over

(18:52) Grok 4.6 and the How I AI Vibe bench

(20:41) Claire Index results: where Grok 4.6 ranked

(23:03) Design evals: where Grok surprised me

(25:00) My conclusion and how I’m splitting my time now

—

Tools referenced:

• Grok Bot: https://x.ai/bot

• Cursor: https://cursor.com/home

• Cursor Origin: https://cursor.com/origin

• OpenClaw: https://openclaw.ai/

• GitHub: https://github.com

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

How a solo founder used Codex and ChatGPT to launch a fashion brand without engineers | Yana Welinder 17 Aug 202600:32:52

Yana Welinder is the solo founder of Yana Bana, an AI-native fashion brand built with AI as her technical co-founder, starting from hand-drawn sketches and ending with runway photos, CAD files for 3D printing, and a live Stripe-connected pre-order site—no engineers required. A former product leader, she brings an operator’s rigor to her creative process: her “fashion prompt” is a detailed spec covering silhouette, volume, fabric behavior, movement, and sound, and watching her use Codex plus computer use to navigate 3D design software that’s entirely new to her is a clarifying demo of what today’s toolset actually makes possible.


What you’ll learn:

  1. How Yana uses a custom fashion prompt as a technical spec to get consistent, realistic, on-design outputs
  2. Why ChatGPT Images 2.0 outperforms other models for fashion design
  3. How she uses Codex plus computer use to operate CAD and fashion software she’s never personally learned
  4. The workflow for taking a garment from hand-drawn sketch to product photo, runway photo, and influencer shot in a single session
  5. How she ran vendor outreach end to end using deep research and browser use
  6. How she built a full e-commerce site with voting, databases, and Stripe integration
  7. Why she’s testing human patternmakers and Codex in parallel

—

Brought to you by:

Merge—Connective infrastructure for production AI

Jira AI SDLC—Get your tokens’ worth with Jira

—

In this episode, we cover:

(00:00) Introducing Yana Welinder and Yana Bana

(02:38) Tour of the Yana Bana site

(05:20) The fashion prompt stack

(07:39) Live demo: generating a jacket from a prompt in ChatGPT

(10:01) Why Image Gen 2.0 beats other models

(11:51) The “prompt as spec” principle

(14:02) Iterating the design

(17:12) Using Codex and computer use to build CAD files in 3D software

(20:50) Vendor research, outreach emails, and Superhuman browser use

(23:34) Building the full e-commerce site

(27:40) Quick recap and what’s still hard

(30:05) How Yana prompts when AI pushes back

(31:15) Where to find Yana and how to vote on her garments

—

Tools referenced:

• ChatGPT (Images 2.0): https://chat.openai.com

• Codex (OpenAI): https://openai.com/codex

• CLO 3D (fashion pattern software): https://www.clo3d.com

• Vercel: https://vercel.com

• GitHub: https://github.com

• Stripe: https://stripe.com

• Superhuman: https://superhuman.com

—

Other references:

• Ruth Asawa: https://ruthasawa.com

• SFMOMA (Ruth Asawa): https://www.sfmoma.org/artist/Ruth_Asawa/

—

Where to find Yana Welinder:

LinkedIn: https://www.linkedin.com/in/ywelinder/

X: https://x.com/yanabana

Website: https://www.yanabana.com

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

Claude Code for normal people: skills, voice mode, and how to collaborate with AI10 Aug 202600:43:08

Grace Clarke is an AI educator and former marketing consultant who taught herself Claude Code earlier this year and built a curriculum out of the process. She now runs her entire service business on tools she’s built with Claude, including a pipeline operator, a proposal maker, and a Gmail replacement she created in under 30 minutes, and teaches individuals and teams to do the same.


What you’ll learn:

  1. How to build an hourly pipeline in Claude that moves clients through your process automatically
  2. Why Grace ditched traditional proposals for password-protected, interactive HTML documents built in Claude
  3. How she uses a “voice guide” skill file so every Claude output sounds like her, not like AI slop
  4. The two-step forcing function she teaches non-technical clients to build the muscle of opening Claude
  5. Why she started building in Claude Code, then handed the work off to Cowork via a Markdown session file
  6. How she replaced Gmail entirely with a custom inbox
  7. Why she teaches “intent engineering” instead of prompt engineering, and what that looks like in practice
  8. How she uses Claude on her phone, on walks, to track workouts and manage plants alongside client work

—

Brought to you by:

Bolt.new—Turn your idea into a real product

Hyperagent—Deploy fleets of agents that handle real work

—

In this episode, we cover:

(00:00) Grace’s background and why she started building with Claude

(04:48) The pipeline operator: what it is and how it runs her business every hour

(08:48) Building the muscle memory to use AI

(12:02) What goes into building a skill file (voice guide, proposal rules, versioning)

(13:50) How she built her proposal maker

(16:15) The voice guide: teaching Claude how she thinks, not just how she writes

(21:22) Live demo of the custom Gmail replacement built in Cowork

(30:44) Workout tracking, plant photos, and tiny daily Claude habits

(34:51) The biggest misconception holding people back from adopting AI

(38:36) What Grace does when Claude is not giving her what she wants

(40:38) Claude builds a proposal for Claire in real time

—

Tools referenced:

• Claude: https://claude.ai

• Claude Code: https://claude.ai/code

• Netlify: https://www.netlify.com

• Google Forms: https://forms.google.com

• Google Sheets: https://sheets.google.com

• Google Cloud (for service accounts and custom connectors): https://cloud.google.com

—

Other reference:

• Stratechery by Ben Thompson: https://stratechery.com

—

Where to find Grace Clarke:

LinkedIn: https://www.linkedin.com/in/gracegclarke/

X: https://x.com/graceclarke

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

Build an AI code review bot in 30 minutes with Vercel Eve05 Aug 202600:24:13

AI writes most of my code now, and that created a new problem: a PR queue I couldn’t keep up with. In this episode, I walk through how I built Merge Mommy, a Vercel Eve agent that reads every PR after checks pass, scores it across six risk dimensions, auto-approves the low-risk ones, and pings me in Slack for anything that needs a human. I built the whole thing in one Codex session, it’s SOC 2 compatible, and it’s already cleared my backlog.


What you’ll learn:

  1. Why AI-generated PRs create a review bottleneck and why the answer isn’t reviewing all of them
  2. How Intercom 5x’d PR approval speed and reduced revert rates by putting AI in the review loop
  3. Why Vercel Eve is the simplest framework I’ve found for deploying AI agents in Slack and GitHub
  4. How I built a full PR review agent in Codex with one prompt and a few steering turns
  5. The six components I use to score PR risk (blast radius, reversibility, data security, ops impact, verification gap, and change surface)
  6. How I used Chrome browser use to handle Slack bot and GitHub app configuration so I never had to click through setup screens manually
  7. Why auto-approved PRs can be SOC 2 compliant as long as the process is auditable, queryable, and in your risk policy
  8. How to set up Slack escalation so low-risk PRs become a two-click merge with no manual review

—

Brought to you by:

WorkOS—Make your app Enterprise Ready today

—

In this episode, we cover:

(00:00) The PR review backlog problem nobody’s talking about

(02:35) Why you don’t have to review every AI-generated PR

(05:14) How Intercom built AI-approved PRs (and proved they’re safer)

(06:10) How the Eve framework works (directory, skills, channels, connectors)

(09:16) The Codex prompt I used to build the entire bot

(11:36) What the agent actually does: read, score, approve, or escalate

(13:07) Setting up your Eve agent

(15:47) The six-component risk scoring model

(17:23) Merge Mommy in action: three live PR examples

(21:10) Recap and how to build your own version

—

Tools referenced:

• Vercel Eve: https://vercel.com/eve

• Vercel AI SDK: https://sdk.vercel.ai/

• Vercel Chat SDK: https://chat-sdk.dev/

• Codex (OpenAI): https://openai.com/codex

—

Other references:

• AI is approving our pull requests: Here’s how we made it safe: https://www.intercom.com/blog/ai-is-approving-our-pull-requests-heres-how-we-made-it-safe/

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

ChatGPT Codex Voice + browser + Sites: an expert’s AI workflow | Nick Baumann (OpenAI)03 Aug 202600:41:34

Nick Baumann is on the Developer Experience team at OpenAI, where he spends his days building with, testing, and communicating the capabilities of ChatGPT Codex and ChatGPT Work. In this episode, Nick walks me through several features that have launched or evolved recently: the new voice interface with its screen-reading orb, the Heartbeats automation system in ChatGPT Work on mobile, the live ChatGPT Sites deployment feature, and his personal use case for AI-assisted UGC video editing.


What you’ll learn:

  1. How two-person voice chat works
  2. How Heartbeats work
  3. How to build and deploy a live website with ChatGPT Sites
  4. How to delegate a flight search, hotel booking, and expense report to Codex in a single voice conversation without opening a single app manually
  5. Why ChatGPT Work on mobile is the most underutilized AI workflow for people already using the ChatGPT app
  6. How to use a custom UGC Video plugin to feed 50 raw clips into ChatGPT, let it pull transcripts, pick the best takes, and assemble a finished vertical video overnight

—

Brought to you by:

Bolt.new—Turn your idea into a real product

Hyperagent—Deploy fleets of agents that handle real work

—

In this episode, we cover:

(00:00) Introduction to Nick Baumann

(02:56) What’s new in Codex

(05:40) ChatGPT Work and Heartbeats

(06:40) Live Codex voice demo

(13:25) Latency vs. intelligence

(14:36) Quick recap

(15:04) Voice on mobile and the ChatGPT Sites workflow

(21:24) Live UGC video demo

(32:30) How I AI website results

(34:04) Lightning round and final thoughts

—

Tools referenced:

• ChatGPT Codex: https://chatgpt.com/codex

• ChatGPT Sites: https://chatgpt.site

—

Where to find Nick Baumann:

LinkedIn: linkedin.com/in/nick--baumann

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

From zero coding background to hardware hacker: How Cursor + a Raspberry Pi makes AI fun27 Jul 202600:28:05

Maddie Reese is a vibe coder, hardware tinkerer, and builder. She builds things at the intersection of software and hardware, including a thermal receipt printer that people around the world can message directly, a fully functional Twitter pager running on a Raspberry Pi, and a personal API that tells you her coffee order so you don’t have to ask. Maddie approaches hardware the same way she approaches software: dump the idea into Cursor, let it interview her, get a shopping list, triple-check the parts before buying, and build. She got her start after her dad introduced her to Lovable, and she locked herself in her room and didn’t come up for air.


What you’ll learn:

  1. How Maddie built a thermal receipt printer that accepts messages from anywhere in the world using a Raspberry Pi and Bluetooth
  2. How to use Cursor’s agent view to brainstorm a hardware project
  3. What belongs in a personal API and why agents, not just humans, will be the ones using it
  4. How to read just enough code to do some damage, without needing to understand all of it
  5. Why building for fun, not practicality, is the fastest path to actually shipping physical projects

—

Brought to you by:

Firecrawl—Power AI agents with clean web data

Customer.io—Build customer engagement campaigns from a single prompt

—

In this episode, we cover:

(00:00) Intro

(02:00) Maddie’s AI pill moment

(03:53) The thermal receipt printer: live demo and how it works

(11:10) The pager project

(17:23) Why she uses Cursor’s clean agent view instead of terminals and browsers

(19:05) The personal API: coffee order, pets, favorite snacks, and more

(22:57) Lightning round and final thoughts

—

Tools referenced:

• Cursor: https://www.cursor.com/

• Lovable: https://lovable.dev/

• Raspberry Pi: https://www.raspberrypi.com/

• Resend: https://resend.com/

• Cloudflare Workers: https://workers.cloudflare.com/

• Supabase (Conduct database referenced): https://supabase.com/

• Twitter/X API: https://developer.x.com/

• Spoke pager network: https://www.spoke.com/

• OpenClaw: https://openclaw.ai/

—

Where to find Maddie Reese:

Website: https://maddiedreese.com

Message her directly: https://maddiedreese.com/message

X: https://x.com/maddiedreese

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

Claude Opus 5 review: this model is brilliant (but annoying)24 Jul 202600:24:51

I’m tired of new models. Every week there’s a new benchmark, a new frontier intelligence claim, a new thing to test. But here we are, because Opus 5 just dropped and I’ve had real hands-on time with it, so you’re getting the honest version.


This is my full Opus 5 review: personality analysis, live benchmark results from my 7-model How I AI eval, and an actual verdict on whether I’m swapping it in. Spoiler: the answer surprised me.


What you’ll learn:

  1. Why I think we’ve hit an intelligence overhang and what that means for which model variables actually matter now
  2. How Opus 5’s “neurotic” personality showed up in real coding sessions, including a merge conflict it refused to touch
  3. What I learned from asking both Opus 5 and GPT‑5.6 Sol “who’s smarter, you or me?”
  4. Where Opus 5, GPT‑5.6 Sol, Sonnet 5, and Gemini 3.1 Pro actually landed on the HIA benchmark leaderboard
  5. The one use case where Opus 5 earned straight 5s from me
  6. My actual plan for using Opus 5 going forward

—

In this episode, I cover:

(00:00) Opus 5 is here

(03:15) First impressions

(06:12) Opus 5 vs. GPT‑5.6 Sol personality comparison

(14:39) Claude Slop: the verbosity problem and why it makes my blood boil

(16:55) How the How I AI benchmark works (7 models, 6 tasks, blind scoring)

(18:30) Live benchmark results: the leaderboard reveal

(23:25) My verdict and how I’ll actually use Opus 5

—

Tools referenced:

• Claude Opus 5:

• Anthropic blog: https://www.anthropic.com/news

• GPT‑5.6 Sol: https://openai.com/index/previewing-gpt-5-6-sol/

• Sonnet 5: https://www.anthropic.com/news/claude-sonnet-5

• Gemini 3.1 Pro: https://deepmind.google/models/gemini/pro/

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

Computer & browser use in Codex (5 real examples)22 Jul 202600:27:40

Today I’m walking you through one of my absolute favorite AI features right now: browser and computer use via Codex (the ChatGPT desktop app). I use this every single day, personally and professionally, and I wanted to share the specific workflows I’ve built, the moments that surprised me, and the mental model that makes it actually click.


What you’ll learn:

  1. How browser use and computer use work, and why the Codex desktop app plus Chrome extension is the combo I rely on
  2. How I use Codex to QA my onboarding flow, including exhaustive mobile testing I would never do manually
  3. Why under-prompting frontier models gets better results than detailed step-by-step instructions
  4. How my husband EJ Lawless’s persona-impersonation trick surfaces friction points I can’t see as the builder
  5. How I use browser use to get through my LinkedIn inbox without touching it myself
  6. How I had Codex shop Free People’s sale and add 10 medium items to my cart (breastfeeding-friendly and Hawaii-ready)
  7. How computer use can control iPhone mirroring so your Mac can technically operate your phone
  8. Three more computer-use shortcuts: filling annoying forms, creating Google Sheets mid-workflow, and managing router

—

Brought to you by:

Runway—The creative AI platform for images, video and more

Hyperagent—Deploy fleets of agents that handle real work

—

In this episode, we cover:

(00:00) Intro

(01:46) What browser use and computer use actually are

(03:08) Why I use Codex specifically and how the desktop app plus Chrome extension works

(04:15) Use case 1: QA testing my onboarding flow

(10:41) Results: 11 issues, one high-severity blocker, one Google Sheet with screenshots

(12:10) Use case 2: persona testing

(18:20) Use case 3: LinkedIn inbox, hands-free

(20:37) Use case 4: AI personal shopper

(23:47) Rapid-fire uses: forms, iPhone mirroring, router access from out of state, Google Docs

(26:50) Wrap-up

—

Tools referenced:

• Codex (ChatGPT desktop app): https://openai.com/codex

• Claude desktop app: https://claude.ai/download

• Monologue (voice dictation for AI): https://monologue.app

• iPhone mirroring (Apple): https://support.apple.com/en-us/111775

• Google Sheets: https://sheets.google.com

—

Other references:

• Jesse Genet episode (How I AI): https://www.lennysnewsletter.com/p/5-openclaw-agents-run-my-home-finances?utm_source=publication-search

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

How the founder of Morning Brew built a Claude content machine that never runs out of ideas and never sounds like slop | Alex Lieberman20 Jul 202600:42:58

Alex Lieberman co-founded Morning Brew in college and grew it into one of the most-read business newsletters in the world before selling it to Business Insider. Now he’s the co-founder and co-managing partner of Tenex. In this episode, Alex explains why distribution is becoming a durable moat, why founders and teams need to “climb Cringe Mountain,” and how he rebuilt his content process around AI without letting it produce generic slop. He walks us through every step of his Content Machine live: an Oracle that scans internal systems and the internet for content spikes, an interview panel that pulls out his real ideas, voice and style files that keep drafts sounding like him, an editorial council that scores and revises posts, and a lessons loop that learns from his feedback.


What you’ll learn:

  1. Why the blank page is the biggest friction point in content creation, and how an AI Oracle eliminates it
  2. How to map your current workflow before you add any AI
  3. How Alex built a six-step Content Machine in Claude that goes from idea spike to publishable post
  4. Why the interview step (not the drafting step) is where AI slop actually comes from
  5. How to codify your voice in a Markdown file so an AI drafts in your register, not the internet’s average
  6. Why your employees are your most underleveraged marketing channel right now
  7. How the Tenex Creator Cup turned content creation into a team sport with a $5,000 prize pool

—

Brought to you by:

Firecrawl—Power AI agents with clean web data

Customer.io—Build customer engagement campaigns from a single prompt

—

In this episode, we cover:

(00:00) Introduction to Alex Lieberman

(02:35) Why Alex built a content machine

(06:56) Alex’s thoughts on AI slop

(09:00) Mapping the workflow from scratch

(13:24) The six-step Content Machine setup

(23:11) Live demo: Oracle, Interview Panel, and Writer’s Council in action

(30:38) Employee advocacy: the Tenex Creator Cup and $5K prize pool

(36:45) Lightning round: great engineers, AI use cases, slop fixes

—

Tools referenced:

• Claude / Claude Code (Anthropic): https://claude.ai

• Wispr Flow (voice-to-text transcription): https://wisprflow.ai

• Notion: https://notion.so

• Linear: https://linear.app

• Slack: https://slack.com

—

Other references:

• Morgan Housel: https://www.morganhousel.com

• David Perell: https://perell.com

• Shaan Puri / My First Million podcast: https://www.mfmpod.com

• Gary Vaynerchuk: https://garyvaynerchuk.com

—

Where to find Alex Lieberman:

X: https://x.com/businessbarista

LinkedIn: https://www.linkedin.com/in/alex-lieberman/

Tenex: https://www.tenex.co/

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

This solo builder runs 24/7 local AI on his own hardware | Alex Finn13 Jul 202600:35:50

Alex Finn is an AI builder, YouTuber, and the creator of Vibe Code Academy, a community for people learning to build with AI tools. He runs one of the most ambitious local AI setups I’ve come across: three Mac Studio 512 GB machines, a DGX Spark, and a custom RTX 5090 build, all coordinated through a fleet dashboard he built himself. He’s spent five months figuring out which local models belong on which machines, how to wire them to Claude Code loops, and how to get a software factory running without babysitting it.


What you’ll learn:

  1. How Alex chose between a Mac Studio (512 GB unified memory), DGX Spark, and RTX 5090, and what each is actually good for
  2. Why Tailscale is worth installing even on a single machine, and how it lets one agent manage your entire hardware fleet
  3. How the build loop and review loop in Claude Code work
  4. How to allocate tasks by machine and model
  5. Why unlimited local inference changes the use-case math in a way a $20 cloud subscription never can
  6. What OpenClaw and Hermes are each best suited for, and why Alex runs five agents total with failover baked in

—

Brought to you by:

Runway—The creative AI platform for images, video, and more

Jira Product Discovery—Prioritize with insights, build with confidence

—

In this episode, we cover:

(00:00) Intro

(02:58) Alex's hardware stack

(03:48) What "ambient AI" means

(04:15) Alex's red-pill moment with OpenClaw

(07:04) Mac Studio vs. DGX Spark vs. RTX 5090

(13:24) How to set up local models with no technical knowledge (Tailscale + OpenClaw/Hermes)

(17:16) Fleet control dashboard: assigning 24/7 tasks across machines

(20:42) Local models as security scanners feeding Claude Code

(22:25) How Alex allocates GLM 5.2, Qwen 3.6, and Ornith 1.0 by task

(24:28) OpenClaw vs. Hermes: the honest comparison

(26:55) The software factory: build loop, review loop, rocket emoji

(31:55) Lightning round: favorite hardware, favorite model, prompting style

(34:46) Where to find Alex

—

Tools referenced:

• Claude Code: https://claude.ai/code

• OpenClaw: https://openclaw.ai/

• Hermes: https://hermes-agent.nousresearch.com/

• Tailscale: https://tailscale.com/

• Codex (OpenAI): https://openai.com/codex

• GLM 5.2 (z.ai): https://huggingface.co/zai-org/GLM-5.2

• Qwen 3.6 (Alibaba): https://huggingface.co/Qwen/Qwen3.6-35B-A3B

• Ornith 1.0: https://github.com/deepreinforce-ai/Ornith-1

• Gemma 4: https://huggingface.co/collections/google/gemma-4

• Playwright (browser testing): https://playwright.dev/

• Vercel (preview deploys): https://vercel.com/

—

Other references:

• DGX Spark (Nvidia): https://www.nvidia.com/en-us/products/workstations/dgx-spark/

• Mac Studio (Apple): https://www.apple.com/mac-studio/

• How to design AI agent loops: schedules, goals, and subagents in Claude Code and Codex: https://www.lennysnewsletter.com/p/how-to-design-ai-agent-loops-schedules

—

Where to find Alex Finn:

LinkedIn: https://www.linkedin.com/in/alex-finn-1848684a

YouTube: https://www.youtube.com/@AlexFinnOfficial

X: https://x.com/AlexFinn

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

GPT-5.6 Sol vs. Claude Fable: Why OpenAI’s new model crushes my benchmark09 Jul 202600:36:40

GPT-5.6 Sol is back, and I ran it through my full How I AI vibe benchmark against GPT-5.6 Terra, Luna, Claude Fable 5, and Sonnet 5 across five categories: PRDs, prototypes, wireframes, debugging, and agentic voice. Sol won by a meaningful margin on my Claire Weighted Index (70% my taste, 30% Terminal Bench 2.1), and I also tested two use cases I can't stop thinking about: building a gamified homework tracking app for my kids in one shot with Codex, and browser automation with Chrome that burned through 500 LinkedIn replies while I did literally nothing.


What you’ll learn:

  1. How I scored five AI models (including GPT 5.6 Sol, Fable 5, and Sonnet 5) using my “Claire Weighted Index” benchmark across PRDs, prototypes, code, and agentic voice
  2. The difference between GPT-5.6 Sol (Terra) and Sol for PRD writing
  3. How Fable’s precision and pedantry made it harder to collaborate with, and the exact moment Sol broke through where Fable got stuck
  4. Why Sonnet 5 is still my go-to for agentic voice in OpenClaw, even after this whole benchmark
  5. How I used GPT-5.6 Sol in Codex to build a fully gamified homework tracking app for my kids in one shot
  6. The video editing use case that saved me hours clipping a talk I gave at Cursor’s event
  7. How to use Codex plus GPT-5.6 and Chrome for browser automation, and why this is my single most-loved use case right now

—

In this episode, I cover:

(00:00) Intro

(01:10) The three GPT-5.6 models: Sol, Terra, Luna

(02:17) Pricing: Sol vs. Fable API costs

(03:24) The How I AI benchmark

(05:03) Claire-weighted Index results

(07:00) Per-task winners: prototypes, PRDs, agentic voice

(11:59) What Claire actually rewards

(13:20) Full-fidelity prototype side-by-sides (Sol vs. Fable)

(17:45) Wireframes

(18:19) Agentic voice

(19:15) Where Sol is better than other models

(23:56) Gamified kids’ homework app, built in one shot

(28:02) Fable’s pedantry problem and how Sol broke through it

(31:49) Two bonus use cases: video editing and browser use

(35:08) Final summary and model recommendations

—

Tools referenced:

• GPT 5.6 (Sol, Terra, Luna): https://help.openai.com/en/articles/20001325-a-preview-of-gpt-56-sol-terra-and-luna

• Codex: https://openai.com/codex

• ChatPRD: https://www.chatprd.ai/

• CapCut: https://www.capcut.com/

• Math Academy: https://www.mathacademy.com/

—

Other references:

• Cursor event where Claire spoke on the future of PM: https://www.youtube.com/watch?v=4CAFK-rc26A

• ChatPRD blog (where benchmark outputs will be published): https://www.chatprd.ai/

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

What a harness is and how to build one with Claude Agent SDK08 Jul 202600:24:35

Everybody is saying, “It’s not the model, it’s the harness,” but almost nobody stops to explain what a harness actually is. So I did. I built one live on the show: a Sentry bug-debugging harness for my company ChatPRD, using the Claude Agent SDK, a custom terminal UI built with the Ink library, and opinionated adapters for Sentry, Linear, GitHub, and Vercel. The harness handles evidence gathering, root-cause analysis, and follow-up artifact creation, all without me needing to type “dear agent, please fix this bug” ever again. I also walk through the architecture, share the code structure, and give you the exact process I used so you can build your own harness for any repetitive, structured workflow in your business.


What you’ll learn:

  1. What a harness actually is
  2. When to build a harness versus when to stick with a general-purpose tool like Claude Code or Codex
  3. How to encode specific permissions into a harness
  4. The three components every harness needs
  5. How I used GPT-5.5 and Claude Opus to build the harness code itself (and where they both initially resisted)
  6. How to structure the artifacts your harness produces so the whole team can use the output

—

Brought to you by:

Bolt.new—Turn your idea into a real product

Customer.io—Build customer engagement campaigns from a single prompt

—

In this episode, we cover:

(00:00) What is an AI harness?

(03:19) When to build a harness

(04:33) Why Claire picked bug triage

(06:00) Why not just use Claude Code?

(07:48) Demo: The custom harness interface

(11:04) Architecture: runs, tasks, tools, and artifacts

(13:44) Building it with Codex and Claude

(15:08) Code map and file layout

(16:51) A look at the code

(19:18) The live investigation result

(21:01) How to build your own harness

—

Tools referenced:

• Claude Agent SDK (Anthropic): https://code.claude.com/docs/en/agent-sdk/overview

• Claude Sonnet 4.6 (model used inside the harness): https://www.anthropic.com/news/claude-sonnet-4-6

• Claude Opus (used to build the harness): https://www.anthropic.com/claude/opus

• GPT-5.5 (Codex, used to build the harness): https://openai.com/index/introducing-gpt-5-5/

• Ink (terminal UI library for Node.js): https://github.com/vadimdemedes/ink

• Sentry (error monitoring): https://sentry.io/

• Linear (project management): https://linear.app/

• GitHub: https://github.com/

• Vercel: https://vercel.com/

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

How I run autonomous coding agents from my phone with OpenAI Symphony + Linear | Alessio Fanelli (Kernel Labs)06 Jul 202600:35:54

Alessio Fanelli, founder of Kernel Labs and co-host of Latent Space podcast, walks us through two very different AI workflows: (1) a fully autonomous coding setup using OpenAI Symphony + Linear, where Linear acts as a state machine and Symphony manages agents through the whole dev lifecycle with zero babysitting; (2) Codex with browser access searching eBay for underpriced Pokémon cards—autonomously browsing, extracting PSA certificate numbers, and flagging deals on $10K–$20K cards for his San Carlos card shop, Merlin Games.


What you’ll learn:

  1. Why “agent manager” is a better mental model than “agent prompter”
  2. Why local Mac Minis don’t scale, and what a cloud VPS unlocks
  3. How to wire Symphony and Linear together as an agent state machine
  4. How to track token costs per task (and what 221 million tokens buys you)
  5. What Glimpse does, and why better agent senses extend autonomous runs
  6. Why your CLAUDE.md probably needs a full purge, not more instructions
  7. How Codex scouts underpriced $10K Pokémon cards on eBay at scale
  8. The new category of small business that AI just made possible

—

Brought to you by:

Firecrawl—Power AI agents with clean web data

Jira Product Discovery—Prioritize with insights, build with confidence

—

In this episode, we cover:

(00:00) Intro

(02:24) Prompter vs. agent manager

(04:31) Live demo: Symphony + Linear

(09:31) Setting up Symphony

(14:15) Purging your skills files

(18:06) The benefits of this system

(19:10) Demo: Using Codex to hunt for Pokémon cards

(24:17) The benefit of AI for small businesses

(28:23) Lightning round

—

Tools referenced:

• OpenAI Codex: https://openai.com/codex

• OpenAI Symphony (open-source framework): https://github.com/openai/symphony

• Linear (project management/agent state machine): https://linear.app

• PSA (Professional Sports Authenticator) grading: https://www.psacard.com

• TCGplayer (card pricing): https://www.tcgplayer.com

• eBay (used for card price scouting): https://www.ebay.com

—

Other references:

• Meta Ray-Ban glasses: https://www.ray-ban.com/usa/ray-ban-meta-smart-glasses

• The Monk and the Riddle by Randy Komisar: https://www.amazon.com/Monk-Riddle-Creating-Making-Living/dp/1578516447/ref=sr_1_1

• The Divine Comedy by Dante Alighieri: https://www.amazon.com/dp/0451208633

• AS Roma (football club Alessio and Claire are both fans of): https://www.asroma.com/en

—

Where to find Alessio Fanelli:

X: https://x.com/FanaHOVA

Latent Space podcast: https://www.latent.space/

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

Sonnet 5 review: I ran 64 generations to find out if it's worth it30 Jun 202600:25:56

I’ve been testing every major frontier model release since the start of the year, and when Anthropic dropped Sonnet 5, I wanted more than a vibe check. I got tired of one-off tests I couldn’t repeat or compare over time, so I built something better: the How I AI Bench, a repeatable eval harness I constructed live using Claude Code while recording this episode. I ran Sonnet 5 blind against four other frontier models (Sonnet 4.6, Opus 4.8, GPT-5.5, and Gemini 3 Pro) across PRD quality, prototype generation, agentic task completion, and agent personality. The results were not what I expected.


What you’ll learn:

  1. What Anthropic claims Sonnet 5 improves over Sonnet 4.6, and where the benchmark data actually backs that up
  2. How I built the How I AI Bench in under 45 minutes using Claude Code, starting from my own stored session history
  3. Why I combined human vibe scoring (70%) with LLM as judge scoring (30%) instead of trusting either alone
  4. How to set up a local HTML scoring page so you can rate AI outputs on gut feel and export those scores as JSON
  5. Which model I recommend for PRDs, which for complex prototypes, and which for chatting with an agent daily

—

Brought to you by:

Runway—The creative AI platform for images, video and more

Hyperagent—Deploy fleets of agents that handle real work

—

In this episode, we cover:

(00:00) Sonnet 5 is out

(01:55) What Anthropic claims

(04:02) Why I’m done with one-off vibe checks

(05:05) Building the How I AI Bench live with Claude Code

(07:42) The scoring system

(10:43) Agent voice eval

(11:57) Quick recap

(13:58) Results: The How I AI index leaderboard

(21:21) What I’m improving for the next run

(22:16) Generating a Claire-weighted index

(23:53) Model-by-task recommendations

—

Tools referenced:

• Claude Sonnet 5: https://www.anthropic.com/news/claude-sonnet-5

• Claude Opus 4.8: https://www.anthropic.com/news/claude-opus-4-8

• GPT-5.5 (OpenAI): https://openai.com/index/introducing-gpt-5-5/

• Gemini 3 Pro (Google DeepMind): https://deepmind.google/models/gemini/pro/

• Cursor: https://www.cursor.com/

—

Other references:

• SWE-bench Pro (agentic coding benchmark referenced): https://www.swebench.com/

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

No Figma. No Jira. No docs. How Gusto built a new product line with Claude Code | Eddie Kim (CTO)29 Jun 202600:51:51

Eddie Kim is the co-founder and CTO of the payroll and HR platform Gusto, which just crossed $1 billion in revenue and serves more than 500,000 small businesses. Recently he did something most CTOs don’t: he went back to writing code. With three other engineers and one designer, Eddie built Gusto Cofounder, a net-new AI product, from zero code to a tier-one launch in 10 weeks. He walks through how that team actually worked, why they threw out nearly every process, and how anyone can copy the approach.


What you’ll learn:

  1. The trash-can method: how to write, review, and delete a full PR as a product decision instead of a planning doc
  2. The two-tool agent stack behind Gusto Cofounder
  3. The exact “perma-Zoom” setup that replaced standups, retros, and Slack threads for 10 weeks
  4. How a designer with no engineering background hit the 94th percentile for shipping code
  5. The eval-first workflow Eddie uses to fix real customer bugs with Claude Code
  6. How a non-technical leader can prototype an idea to win buy-in, then carry it all the way to production-quality code

—

Brought to you by:

Magic Patterns—Prototypes that look like your product

Jira Product Discovery—Prioritize with insights, build with confidence

—

In this episode, we cover:

(00:00) Intro: five people, 10 weeks

(02:38) The origins of Cofounder

(08:32) Inside the 10-week build process

(12:50) Building with no PMs

(14:38) The “trash can” method

(17:15) The stack architecture

(19:10) Shipping to production from day one

(22:03) How a designer became a top engineer

(29:05) Demo: Cofounder over text and Slack

(31:45) Demo: running a real payroll

(36:26) Live coding with evals in Claude Code

(39:39) Recap: prototype, small team, permission

(43:17) Lightning round

(48:44) Where to find Eddie and Cofounder

—

Tools referenced:

• Gusto Cofounder (early access/waitlist): https://gusto.com/cofounder

• Claude Code (Anthropic): https://claude.ai/code

• Cloudflare Workers: https://workers.cloudflare.com/

• Vercel AI SDK: https://sdk.vercel.ai/

• DX (engineering analytics): https://getdx.com/

• Wispr Flow (voice-to-text): https://wisprflow.ai

• OpenClaw: https://openclaw.ai/

—

Other references:

• Gusto (the main product, “Gusto Classic”): https://gusto.com

• Mindbody (referenced as customer data source): https://www.mindbodyonline.com/

—

Where to find Eddie Kim:

LinkedIn: https://www.linkedin.com/in/edawerd/

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

GLM 5.2: why I’m replacing Opus in Claude Code with this new model24 Jun 202600:27:13

I put GLM 5.2, the open-weight coding model from Z.AI, through four real tasks inside my actual codebase: a codebase architecture audit, a UI redesign, and a 45-minute autonomous bug-hunting session pulling from Sentry and Vercel logs. Total cost: $3.36 for roughly 6 million tokens, a prioritized bug-fix dashboard I’m actually shipping from, and a landing page redesign that matched Chat PRD’s design system on the first try.


What you’ll learn:

  1. What “open-weight” actually means and why it matters for cost and vendor independence
  2. How to connect GLM 5.2 to Cursor and Claude Code
  3. How it performs on codebase exploration and autonomous architecture summarization in a real production Next.js app
  4. Whether GLM 5.2 can match an existing design system
  5. How the model handles a 45-minute long-running autonomous task
  6. Where GLM 5.2 stumbled 
  7. The actual cost breakdown

—

Brought to you by:

Mercury—Radically different banking loved by over 300K entrepreneurs

—

In this episode, we cover:

(00:00) What open-weight models are and why GLM 5.2 is worth testing

(01:38) GLM 5.2 model overview

(04:02) Capabilities and benchmark results

(06:02) How to set up GLM 5.2 in Cursor

(08:37) How to set up GLM 5.2 in Claude Code

(11:04) Live test 1: codebase exploration and architecture audit on ChatPRD

(12:43) Live test 2: generating an HTML architecture and roadmap page

(16:37) Live test 3: redesigning the How I AI landing page in Cursor

(20:57) Live test 4: 45-minute autonomous task, pulling Sentry errors and Vercel logs

(22:35) Where it struggled

(23:49) My verdict on the output

(25:23) Cost breakdown

—

Tools referenced:

—

Other references:

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

How Claude Mythos found a 15-year-old bug in Mozilla Firefox | Brian Grinstead22 Jun 202600:48:28

Brian Grinstead is a distinguished engineer at Mozilla, where he’s worked on Firefox and the web platform since 2013 (he joined to help launch Firefox DevTools). Recently he and his team pointed an agentic bug-finding pipeline at Firefox—a codebase with tens of thousands of files and tens of millions of lines of code—and shipped a record month of security fixes. The viral chart everyone saw gave the credit to Anthropic’s new Mythos model. Brian’s take is that the harness and pipeline did just as much of the work, and he walks through exactly how it runs and how anyone can build a starter version.


What you’ll learn:

  1. How to build a basic bug-finding harness by running Claude Code or Codex with one prompt and the -p flag, no SDK required
  2. Why pointing an agent at a whole codebase fails, and how an LLM judge can score and rank files before you spend any compute
  3. How a verifier subagent kills false positives by catching the agent when it cheats
  4. The goal-loop pattern: give an agent a tightly scoped problem, a clear pass/fail signal, and let it retry far past the point a human would quit
  5. Why teams that already invested in fuzzing, CI, and dev tooling are so far ahead
  6. How to weigh model versus harness, and why Brian splits the credit close to 50-50
  7. How a non-engineer can reuse the same score, verify, and fix the loop for design quality, conversion rate, or tech debt
  8. Why AI-generated patches still can’t ship on their own, and where humans stay in the loop

—

Brought to you by:

WorkOS—Make your app enterprise-ready today

Metaview—The agentic recruiting platform for winning teams

—

In this episode, we cover:

(00:00) Introduction to Brian Grinstead

(02:43) The viral chart: Firefox Security Bug Fixes by Month

(05:32) How the custom harness works

(10:22) Goal loops and guardrails

(14:45) How they built it

(16:55) Real bugs, including a 15-year-old one

(23:00) Open-sourcing it

(26:26) Why humans still review every fix

(32:30) Live demo and prioritizing files

(40:18) Mobilizing the team and recap

(42:33) Lightning round

—

Tools referenced:

• Claude Code: https://claude.ai/code

• Claude Agent SDK: https://code.claude.com/docs/en/agent-sdk/overview

• Codex: https://openai.com/index/openai-codex/

• OpenAI Agent SDK: https://developers.openai.com/api/docs/guides/agents

• VS Code: https://code.visualstudio.com/

• Docker: https://www.docker.com/

• Firefox: https://www.mozilla.org/firefox/

• Address Sanitizer: https://github.com/google/sanitizers

• RLBox: https://rlbox.dev/

—

Other references:

• Mozilla Bug Bounty Program: https://www.mozilla.org/security/bug-bounty/

• Mozilla GitHub: https://github.com/mozilla

—

Where to find Brian Grinstead:

LinkedIn: https://www.linkedin.com/in/bgrins/

GitHub: https://github.com/bgrins

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

How to design AI agent loops: schedules, goals, and subagents in Claude Code and Codex17 Jun 202600:29:06

I break down every loop type from scratch—what a heartbeat, cron, hook, and goal loop actually are, when each one fits, and the five things any effective loop needs before it touches production. Then I build two live loops: a daily aging-PR reviewer in Claude Code that schedules itself at 10:15 a.m. and spins off its own subagents, and a weekly skills-identification loop in Codex that spawns goal-based subagents to validate its own output in real time.


What you’ll learn:

  1. The plain-English definition of a loop—and why it’s just an automated prompt, not a scary new paradigm
  2. The four loop types (heartbeat, cron, hook, and goal) and when each one actually fits your workflow
  3. How to think about loop design using the “onboarding an employee” mental model
  4. The five things every effective loop needs: work trees, skills, plugins/connectors, subagents, and state tracking
  5. How to build a scheduled PR-review routine in Claude Code that babysits aging PRs and alerts your team
  6. How to set up a weekly skills-identification automation in Codex that spawns its own validating subagents
  7. Why goal-based loops are the hardest to write well—and where most people burn tokens for nothing
  8. The two warning signs that your loop is going to get expensive before it gets useful

—

Brought to you by:

WorkOS—Make your app enterprise-ready today

Runway—The creative AI platform for images, video, and more

—

In this episode, we cover:

(00:00) Prompts are out and loops are in

(02:30) Defining a loop

(03:03) The four ways to automate a prompt: heartbeat, cron, hooks, and goals

(06:03) Five things every effective loop needs

(09:26) The “onboarding an employee” framework for designing loops

(11:58) Live build #1: Daily aging PR loop in Claude Code

(17:08) Subagents inside loops

(19:00) Live build #2: Weekly skills identification loop in Codex

(22:57) Watching subagents spin up in real time

(25:28) Warning signals around loops

(27:31) What listeners are doing with loops

—

Tools referenced:

• Claude Code: https://claude.ai/code

• Codex: https://chatgpt.com/codex

• OpenClaw: https://openclaw.ai/

—

Other references:

• Claire’s article “Why OpenClaw Feels Alive Even Though It’s Not”: https://x.com/clairevo/article/2017741569521271175

• Addy Osmani’s article on loop engineering: https://addyosmani.com/blog/loop-engineering/

• Using Goals in Codex: https://developers.openai.com/cookbook/examples/codex/using_goals_in_codex

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

How Braintrust uses AI agents, evals, and CI to ship better software | Ankur Goyal15 Jun 202600:40:11

In this episode, I sit down with Ankur Goyal, founder and CEO of Braintrust, the AI evals and observability platform used by teams like Notion, Stripe, Vercel, and Zapier. This one is for the senior engineers, staff engineers, VPs of engineering, and CTOs in my audience. We get into how coding agents can take on deeply technical architecture and infrastructure work that no single human engineer could tackle before, and then we demystify evals so you can use them to make your AI products better without touching the implementation.


What you’ll learn:

  1. How Ankur uses Codex to run week-long benchmark experiments across database indexes, column store formats, and execution engines to speed up slow queries
  2. Why he argues there’s no excuse to skip rigorous benchmarking now that agents can run them tirelessly
  3. The “agent line” framework: how to decide which decisions, directions, and interactions you can hand off to an agent
  4. How I think about the practical vs. theoretical quality of AI on hard technical problems, and why human attention decays on tedious work
  5. Why evals are the modern version of a PRD, and how to encode “what good looks like” so a model can figure out the “how”
  6. How to build a scoring function live and let an agent improve your prompt inside a safe playground
  7. How Ankur turned his designer David’s taste into a repeatable eval so quality scales beyond one person
  8. Why fixing your CI is the highest-leverage way to speed up engineering velocity

—

Brought to you by:

Guru—The AI layer of truth

Persona—Trusted identity verification for any use case

—

In this episode, we cover:

(00:00) Introduction to Ankur Goyal

(03:00) Using AI agents for database optimization

(06:10) Running exhaustive benchmarks with coding agents

(09:03) Why staff engineers are wrong about AI limitations

(11:30) The “agent line” framework for delegation

(14:00) Ankur’s workflow: running 4 to 6 concurrent agents

(17:16) Technical setup: foreground agents, background agents, and cloud environments

(20:32) Spending time with AI tools

(23:06) Demystifying evals

(26:02) Live demo: Building an eval for documentation answers

(30:20) The alternative to evals: vibe checks and whack-a-mole

(32:09) Capturing designer taste in scoring functions

(33:13) Quick recap

(33:44) Managing velocity and throughput

(35:40) Why CI/CD investment is critical for AI-accelerated teams

(37:30) Ankur’s prompting strategy when agents fail

(39:10) Closing thoughts and how to connect

—

Tools referenced:

• Braintrust: https://www.braintrust.dev/

• Codex: https://openai.com/codex/

• GPT 5.4: https://developers.openai.com/api/docs/models/gpt-5.4

• Claude: https://claude.ai/

—

Other references:

• GPT 5.5 just did what no other model could: https://www.lennysnewsletter.com/p/gpt-55-just-did-what-no-other-model

• Paul Graham’s Maker vs. Manager Schedule: http://www.paulgraham.com/makersschedule.html

• tmux: https://github.com/tmux/tmux

• Chris Tate at Vercel: https://www.linkedin.com/in/ctatedev/

—

Where to find Ankur Goyal:

LinkedIn: https://www.linkedin.com/in/ankrgyl/

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

Claude Fable 5 review: what the new Mythos model gets right (and very wrong)09 Jun 202600:17:24

Claude Fable 5 is the first Mythos-class intelligence model to be generally available, and I got early access to test it before launch. In this episode, I walk through what Anthropic is promising, what actually stood out when I used it on real work, and where I think it fits in your AI stack.

—

In this episode, we cover:

(00:00) Introduction: Fable 5 is finally here

(00:31) What Anthropic says about the model

(05:14) Token-intensive by design

(06:28) Safety classifiers and the new fallback concept

(07:46) Is this or is this not Mythos?

(08:30) New product launches: Managed Agents and more

(09:20) Crushing benchmarks

(09:55) What it’s actually like to use (the good and the bad)

(11:40) Test 1: product graph spec

(12:56) Test 2: designing a skills registry

(14:04) Conservative on execution

(14:43) Test 3: multi-agent orchestration

(15:39) My takeaways

—

Tools referenced:

• Claude Fable 5: https://www.anthropic.com/news/claude-fable-5-mythos-5

• Claude Managed Agents: https://platform.claude.com/docs/en/managed-agents/overview

—

Other reference:

• SWBench Pro benchmark: https://www.swebench.com/

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

Shopping with Claude: How to find quality brands, automate returns, and buy things that last 100 years | Nicole Ruiz08 Jun 202600:36:56

Nicole Ruiz is a writer and parent who has built a comprehensive AI-powered shopping system to help her family buy high-quality, long-lasting items while avoiding the noise of drop-shipping brands, paid ads, and poorly made products. She writes an interview series on Substack about how technology is changing the household.


What you’ll learn:

  1. How to build a Claude Project with custom instructions for vetting brands based on heritage, craftsmanship, and return policies
  2. The shopping criteria that help surface century-old manufacturers over trendy direct-to-consumer brands
  3. How to use Claude to search through trusted vendor websites that have terrible UX
  4. Why AI actually helps small artisans and heritage brands compete against Amazon’s infrastructure
  5. How to use Claude Cowork to automate returns by finding receipts in your email and drafting refund requests
  6. The technique for getting Claude to analyze whether a brand is legitimate or just a drop-shipping operation
  7. How to shop within a specific budget or with gift cards using AI assistance

—

Brought to you by:

Orkes—The enterprise platform for reliable applications and agentic workflows

Metaview—The agentic recruiting platform for winning teams

—

In this episode, we cover:

(00:00) Introduction to Nicole and AI-powered shopping

(02:29) The problem

(04:55) Building a Claude Project for household purchasing

(07:44) The “anti-to-do list” concept for reducing mental overhead

(10:30) Shopping for a can opener: the system in action

(15:53) How AI helps century-old brands with terrible websites

(18:45) Processing returns with Claude Cowork

(25:06) Using gift cards strategically

(26:33) Vetting brands

(29:40) Recap, lightning round, and final thoughts

—

Tools referenced:

• Claude: https://claude.ai/

• Claude Cowork: https://www.anthropic.com/product/claude-cowork

—

Other references:

• Boston General Store: https://bostongeneralstore.com/

• L.L.Bean: https://www.llbean.com/

• Manufactum: https://www.manufactum.com/

• 5 OpenClaw agents run my home, finances, and code | Jesse Genet: https://www.lennysnewsletter.com/p/5-openclaw-agents-run-my-home-finances

• From a $6.90 newsletter to $3M API: How a non-coder built Memelord | Jason Levin: https://www.lennysnewsletter.com/p/from-a-690-newsletter-to-3m-api-how

—

Where to find Nicole Ruiz:

X: https://x.com/nwilliams030

Substack (The Third Oikos): https://www.thirdoikos.com/

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

Gemini Omni: Clone yourself with AI in under 15 minutes03 Jun 202600:20:35

In this experimental episode, I document my real-time attempt to create an AI avatar of myself using Google Flow and the new Gemini Omni video generation model. I walk through the entire process—from scanning my face with my phone to generating a complete one-minute hype video for the podcast, all in about 15 minutes.


What you’ll learn:

  1. How to create an AI avatar using Google Flow in under five minutes
  2. Why video AI tools unlock creative possibilities for people with zero video production skills
  3. The step-by-step process of generating a full storyboard using AI as your creative producer
  4. How to use character consistency features to generate multiple video scenes with the same avatar
  5. The uncanny-valley moments you’ll encounter when your AI clone doesn’t quite nail emotions or physics
  6. How to stitch together AI-generated scenes into a complete video using built-in editing tools

—

Brought to you by:

Merge—Connective infrastructure for production AI

Jira Product Discovery—Prioritize with insights, build with confidence

—

In this episode, we cover:

(00:00) Getting started with Google Flow and Gemini Omni

(01:38) The avatar creation process: scanning and photo capture

(02:55) Using Flow to brainstorm a hype video storyboard

(06:59) Generating the first video scene with the avatar

(08:41) Troubleshooting: accidentally generating images instead of videos

(09:32) Generating all seven scenes for the complete video

(11:37) Reviewing the avatar videos

(13:13) Stitching the videos together in the browser-based editor

(14:32) The complete How I AI hype video

(15:32) What worked and what didn’t

(19:04) Final thoughts

—

Tools referenced:

• Google Flow: https://labs.google/fx/tools/flow

• Gemini Omni: https://gemini.google/overview/video-generation/

• Veo 3: https://deepmind.google/technologies/veo/

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

Building an iPhone app with zero technical skills | Bryce Rattner Keithley 01 Jun 202600:46:33

Bryce Rattner Keithley has spent her career in talent and recruiting, working with technical leaders but never writing a line of code herself. Yet she managed to build Daily Hundred—a fitness app featuring custom AI-generated videos of anthropomorphic animals demonstrating exercises—and ship it to the App Store before her software engineer friends. Using Replit, Claude, Gemini, and a relentless beginner’s mindset, Bryce proves that in the AI era, execution is no longer the constraint on good ideas.


What you’ll learn:

  1. How to build and ship an iPhone app using Replit without any coding knowledge
  2. The step-by-step process for creating custom AI-generated workout videos by combining Gemini images with real exercise footage
  3. How to use Claude as your technical architect and Claude Code as your software engineer
  4. How to navigate App Store submission requirements (including fixing rejection feedback)
  5. Why being hyper-literal in your prompts unlocks better AI results
  6. Why a beginner’s mind is actually an advantage when building with AI tools

—

Brought to you by:

WorkOS—Make your app enterprise-ready today

Metaview—The agentic recruiting platform for winning teams

—

In this episode, we cover:

(00:00) Introduction to Bryce and Daily Hundred

(04:48) Building with Replit

(06:16) The beginner’s mindset advantage

(11:17) Creating anthropomorphic animals

(22:55) Moving from static image to video

(27:15) The floating genie and other anthropomorphic animal generations

(30:46) Shifting from web app to App Store submission

(36:24) User feedback

(37:41) Lightning round and final thoughts

—

Tools referenced:

• Replit: https://replit.com/

• Lovable: https://lovable.dev/

• Claude: https://claude.ai/

• Claude Code: https://claude.ai/code

• Gemini: https://gemini.google.com/

• Higgsfield: https://higgsfield.ai/

• Kling: https://kling.ai/

• Railway: https://railway.app/

• TestFlight: https://developer.apple.com/testflight/

—

Other references:

• How a 91-year-old vibe coded a complex event management system using Claude and Replit | John Blackman: https://www.lennysnewsletter.com/p/how-a-91-year-old-vibe-coded-a-complex

• What Got You Here Won’t Get You There: https://www.amazon.com/What-Got-Here-Wont-There/dp/1401301304

• How Women Rise: https://www.amazon.com/How-Women-Rise-Holding-Careers/dp/0316440124

• A Whole New Mind: https://www.amazon.com/Whole-New-Mind-Right-Brainers-Future/dp/1594481717

• How to Win Friends and Influence People: https://www.amazon.com/How-Win-Friends-Influence-People/dp/0671027034

—

Where to find Bryce Rattner Keithley:

LinkedIn: https://www.linkedin.com/in/brycerattner/

GitHub: https://github.com/brk-bot/

Daily Hundred on the App Store: https://apps.apple.com/us/app/daily100-fitness-challenge/id6762108062

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

Claude Opus 4.8 is here. Is it as good as they say?28 May 202600:13:39

I got a few hours of early-access testing with Anthropic’s newly released model Opus 4.8. I walk through real coding, design, and strategy tasks across Claude Code and Claude Cowork, and give you my unfiltered view on what impressed me and what didn’t.

—

What you’ll learn:

  1. Where Opus 4.8 excels: greenfield prototypes, one-shot features, and fast execution
  2. Where it struggles: the last 10%, edge cases in existing codebases, and hallucinations
  3. How Opus 4.8 compares to Opus 4.7 on business strategy work
  4. Why I’m still reaching for Opus 4.7 on data-heavy strategy and roadmap work
  5. The new features shipping alongside the model: dynamic workflows with parallel subagents and effort control in Claude.ai and Cowork
  6. The prompting and harness strategy I’d use to get the most out of it

—

In this episode, we cover:

(00:00) Introduction to Opus 4.8 

(00:44) Benchmark performance and pricing

(01:53) First coding test: Building a prototyping tool

(03:00) Where it failed: The last 10% problem

(03:27) The hallucination problem

(04:23) Testing Opus 4.8 on existing codebases

(05:24) The ambition test: Building games for a 9-year-old

(07:03) Business strategy test: 4.7 vs 4.8

(08:23) The roadmap test

(09:17) Final verdict

—

References:

• System Card: Claude Opus 4.8: https://cdn.sanity.io/files/4zrzovbb/website/c886650a2e96fc0925c805a1a7ca77314ccbf4a6.pdf

• Introducing Claude Opus 4.8 on X: https://x.com/claudeai/status/2060042702150930686?s=20

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

The Codex feature that works while you sleep27 May 202600:30:20

In this 30-minute episode, I walk through my favorite feature in Codex: the /goal command. I show how Goals transform AI from a turn-based assistant that needs constant ‘what’s next?’ prompting into an autonomous agent that can work for hours on complex, multi-step tasks. I share three real examples: eliminating thousands of Sentry errors, cleaning 3,900 emails down to 68, and organizing hundreds of Linear tasks.


What you’ll learn:

  1. What Goals are and how they differ from standard prompts
  2. How I used /goal to eliminate hundreds of error logs in my codebase over a five-hour autonomous run
  3. The non-technical use cases that make Goals incredibly powerful: cleaning up 3,900 emails in under four hours and organizing hundreds of project management tasks in Linear
  4. How to write effective /goal prompts with measurable outcomes, verification methods, and constraints
  5. When not to use Goals and what makes a strong versus weak Goal
  6. Why Goals represent a fundamental shift in how we work with AI, from babysitting the model to managing it

—

Brought to you by:

Mercury—Radically different banking loved by over 300K entrepreneurs

—

In this episode, we cover:

(00:00) Introduction

(01:50) What is /goal and when should you use it?

(02:45) The difference between prompts and Goal-based loops

(04:06) Claire’s first five-hour 45-minute autonomous coding task

(05:05) How to manage a Goal lifecycle: view, pause, resume, and clear

(06:06) How to write strong goals: outcomes vs. outputs

(07:34) The six components of effective Goals

(08:57) Example: Reducing P95 checkout latency with /goal

(09:36) Demo: Using /goal to eliminate Sentry errors in ChatPRD

(13:18) Demo: Burning down Vercel API errors

(17:28) Non-technical use case: Cleaning 3,900 emails with /goal

(21:24) Demo: Using /goal to clean up Linear project tasks

(24:41) When not to use /goal

(26:10) Why /goal changes everything

—

Tools referenced:

• Codex: https://openai.com/codex/

• Sentry: https://sentry.io/

• Vercel: https://vercel.com/

• Linear: https://linear.app/

—

Other reference:

• OpenAI blog post “Using Goals in Codex”: https://developers.openai.com/cookbook/examples/codex/using_goals_in_codex

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

How the engineer behind Claude Cowork actually uses Claude | Felix Rieseberg (Anthropic)25 May 202600:59:25

Felix Rieseberg is the engineering lead for Claude Cowork and Claude Code Desktop at Anthropic. He previously spent five years at Slack building developer tools. In this episode, Felix demonstrates how he uses Claude to solve real-life problems: analyzing floor plans to build interactive 3D house walkthroughs, automatically tracking promises he makes on Twitter, and building a $20 hardware device that physically approves Claude actions with a button press.


What you’ll learn:

  1. How to use Claude Cowork to turn a 2D floor plan into an interactive 3D walkthrough where you can move furniture around
  2. The “go one abstraction layer up” philosophy: why you should never manually enter data Claude can find itself
  3. How to use your email as an inventory database for furniture, clothing, and personal purchases
  4. When to use Opus vs. Sonnet 4.6 (hint: it’s about how well you can scope the problem, not technical complexity)
  5. How live artifacts work and why they’re powerful for dashboards that refresh with real-time data from your connectors
  6. The product philosophy behind making latency delightful
  7. How to build your own $20 hardware device using Claude Code (no hardware experience required)
  8. Why Felix never reads the code Claude writes and judges it purely on output

—

Brought to you by:

Magic Patterns—Prototypes that look like your product

Guru—The AI layer of truth

—

In this episode, we cover:

(00:00) Introduction to Felix Rieseberg

(02:40) Felix’s role at Anthropic

(03:25) The multiple tabs in Claude and why they exist

(05:55) Using Claude Cowork to design a new house using floor plans

(09:52) When to use Opus versus Sonnet 4.6

(12:37) Building an interactive 3D furniture planner

(14:30) Using your email as a source of truth for personal inventory

(15:58) The anti-to-do list: going one abstraction layer up

(23:14) Introduction to live artifacts

(26:02) Building a personal dashboard with live data

(28:37) Being polite to Claude (and why it matters for your humanity)

(30:28) Claude interaction tips

(32:33) Looking at the daily dashboard

(33:55) How live artifacts work with connectors

(35:02) Redesigning the dashboard

(37:55) The biggest gap: people don’t know what problems AI can solve

(41:52) The reverse interview

(42:30) Making latency delightful through asynchronous design

(44:05) The redesigned dashboard

(45:28) AI should free up your creative energy

(46:44) Building a $20 hardware Claude buddy

(52:33) Why kids are magical AI users

(54:30) Recap and final thoughts

—

Tools referenced:

• Claude Cowork: https://www.anthropic.com/product/claude-cowork

• Claude Code: https://claude.ai/code

• Claude for Chrome: https://code.claude.com/docs/en/chrome

• Claude Desktop: https://claude.ai/download

• Live Artifacts: https://support.claude.com/en/articles/14729249-use-live-artifacts-in-claude-cowork

• Connectors (Spotify, Gmail, Calendar, Notion): https://claude.ai/settings/connectors

• Slack: https://slack.com/

—

Where to find Felix Rieseberg:

Website: https://felixrieseberg.com/

LinkedIn: https://www.linkedin.com/in/felixrieseberg/

X: https://x.com/felixrieseberg

GitHub: https://github.com/felixrieseberg

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

What launched at Google I/O 2026 (30-minute day 1 recap)20 May 202600:33:52

Today is day one of Google I/O 2026, and I walk through every major announcement live—from the new Gemini 3.5 model family to Anti-Gravity 2.0, Google AI Studio, Gemini’s consumer redesign, the Omni video model, Flow, Stitch, and Pomelli. I test them in real time and tell you exactly which ones delivered.


What you’ll learn:

  1. How Gemini 3.5 Flash benchmarks against Claude and GPT models on speed and agentic coding tasks
  2. How Anti-Gravity 2.0’s new features (projects, scheduled tasks, subagents, slash commands) compare to Codex and Claude Code
  3. Why the /grill-me slash command could be a more aggressive alternative to Claude Code’s clarification flow—and how to use it
  4. How Google AI Studio’s new Workspace integration is designed to own the internal productivity app use case
  5. How Google’s new creative tools work in practice: Omni (video generation), Flow (cinematic video editing and character consistency), Stitch (streaming UI design with inline edits), and Pomelli (brand identity and asset generation)
  6. Why Google’s launch-to-availability gap is still a problem—and what to do when a featured product doesn’t actually work yet

—

Brought to you by:

Magic Patterns—Prototypes that look like your product

Thoughtspot—Build AI-powered analytics into your product

—

In this episode, we cover:

(00:00) Google I/O 2026 day 1 overview

(01:47) Gemini 3.5 flash

(04:19) Antigravity updates

(06:32) CLI test and agent features

(07:59) Core agent features released today—May 19th, 2026

(09:43) New slash commands

(11:20) Antigravity test results and takeaways

(12:25) AI Studio updates

(13:52) Access issues

(15:20) Gemini redesign

(17:24) Gemini image gen test

(19:16) Omni (video generation)

(22:56) Flow (cinematic editing)

(24:31) Avatar creation test

(26:45) Pomelli and Stitch

(31:13) Recap and final thoughts

—

Tools referenced:

• Gemini 3.5 Flash: https://deepmind.google/technologies/gemini/

• Antigravity: https://antigravity.google/

• Google AI Studio: https://aistudio.google.com/

• Google Gemini: https://gemini.google.com/

• Omni (video generation): https://gemini.google/overview/video-generation/

• Google Flow: https://flow.google/

• Stitch: https://stitch.withgoogle.com/

• Pomelli (Google brand tool): https://labs.google.com/pomelli/about/

—

Other references:

• Google I/O 2026 announcements: https://blog.google/innovation-and-ai/sundar-pichai-io-2026/

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

HTML is the new Markdown: How Anthropic engineers are building with Claude Code | Thariq Shihipar18 May 202600:35:58

Thariq Shihipar is an engineer at Anthropic working on the Claude Code team. He’s spent the past several months experimenting with HTML as a replacement for Markdown in planning and implementation workflows, discovering that richer visual formats lead to better human engagement—and, ultimately, better products. In this episode, filmed at Anthropic’s Code with Claude event in San Francisco, Thariq demonstrates how to use HTML artifacts to create interactive plans, build throwaway UIs for specific problems, and maintain living design systems that travel with your codebase.


What you’ll learn:

  1. Why HTML has replaced Markdown as the ideal format for AI agent communication and planning
  2. How to brainstorm in HTML to get visual mockups and interactive demos instead of text lists
  3. The technique for building throwaway micro-UIs to edit specific parts of your plan
  4. How to create a living design system in HTML that lives in your repo and travels with every project
  5. Why “complexity has to earn its keep” and how HTML helps you stay in the loop without over-constraining Claude
  6. The prompting technique that gives Claude flexibility while ensuring that you get what you need
  7. Why 99% of your AI-generated tokens should go to planning, interfaces, and communication—not production code

—

Brought to you by:

Celigo—Intelligent automation built for AI

Persona—Trusted identity verification for any use case

—

In this episode, we cover:

(00:00) Introduction

(02:39) HTML as the new Markdown

(04:30) The compute allocator mindset

(05:51) How HTML makes specs more engaging

(06:48) Demo: Brainstorming in HTML with Claude Code

(09:24) From brainstorm to full implementation plan

(11:20) Prompting philosophy: Trust Claude but give it constraints

(13:50) The future of PRDs and tech specs

(18:16) Making HTML specs editable

(20:23) The abundance mindset

(24:17) Just-in-time documentation and throwaway software

(25:39) Using plans as artifacts for implementation

(26:39) Demo: Living design systems in HTML

(30:16) Adding comments and annotations to HTML plans

(31:42) Recap: The HTML workflow

(32:21) Lightning round and final thoughts

—

Tools referenced:

• Claude Code: https://claude.ai/code

• Claude Design: https://claude.ai/design

• AWS: https://aws.amazon.com/

• Figma: https://www.figma.com/

• GitHub: https://github.com/

—

Other references:

• Anthropic Code with Claude event: https://claude.com/code-with-claude

• SpaceX partnership announcement: https://www.anthropic.com/news/higher-limits-spacex

• Jevons paradox: https://en.wikipedia.org/wiki/Jevons_paradox

—

Where to find Thariq Shihipar:

Website: https://www.thariq.io/

LinkedIn: https://www.linkedin.com/in/thariqshihipar/

X: https://x.com/trq212

GitHub: https://github.com/ThariqS

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

Spec-driven development: The AI engineering workflow at Notion | Ryan Nystrom11 May 202600:47:53

Ryan Nystrom is a software engineer at Notion. He joined in December 2024 after Notion acquired Campsite, the team communication platform he co-founded with Brian Lovin. At Notion, he’s been a core builder of Notion AI and the Custom Agents feature launched in February 2026. He manages a team of six to seven engineers while still writing code himself, currently running Project Afterburner, a push to cut Notion’s CI time to a quarter of its current duration.


What you’ll learn:

  1. How to build a Notion AI custom agent that auto-generates your daily standup pre-read by pulling from Slack, GitHub, Honeycomb metrics, and yesterday’s meeting transcript
  2. How to configure subagents and MCP integrations within Notion AI
  3. How Notion’s internal “Boxy” system lets engineers @mention Codex from within Notion comments and get a full pull request with screenshots in 20 minutes
  4. The spec-first development workflow: dictate an idea into Whisper, have Codex format it as a proper spec, commit it to the repo, and let the agent implement and verify it autonomously
  5. Why fast CI is absolutely critical in the age of AI coding agents
  6. How to prompt AI coding agents to defend their reasoning under pushback
  7. Why engineering managers and even senior executives should keep writing code

—

Brought to you by:

WorkOS—Make your app enterprise-ready today

Orkes—The enterprise platform for reliable applications and agentic workflows

—

In this episode, we cover:

(00:00) Introduction to Ryan Nystrom

(02:48) How AI has upended 12+ years of the same working routine

(04:30) Project Afterburner: Notion’s push to cut CI time to a quarter

(09:00) Why high-frequency, high-quality meetings beat lower-frequency standups

(11:10) How automated context surfaces every engineer’s work equally

(12:15) Why cutting meeting prep is a burnout protection mechanism

(14:26) The case for engineering managers writing code

(16:13) Inside “Boxy”: Notion’s internal VM-based background agent system

(20:30) Old World vs. New World code review

(24:51) Prompting Codex from Notion comments

(29:20) The emotions around code review

(31:01) Quick recap

(32:00) Spec-first development: writing and checking agent specs into the repo

(35:10) The spec as changelog: version control for how a feature actually works

(37:53) How engineers’ roles are evolving

(39:00) Lightning round

(45:21) Where to find Ryan

—

Tools referenced:

• Notion AI: https://www.notion.com/product/ai

• Notion Custom Agents: https://www.notion.com/blog/introducing-custom-agents

• Codex (OpenAI): https://openai.com/codex

• Claude Code (Anthropic): https://claude.ai/code

• Honeycomb (observability + MCP): https://www.honeycomb.io

• Whisper (OpenAI voice transcription): https://openai.com/research/whisper

• Slack: https://slack.com

• GitHub: https://github.com

—

Other references:

• How Stripe built “minions”—AI coding agents that ship 1,300 PRs weekly from Slack reactions | Steve Kaliski (Stripe): https://www.chatprd.ai/how-i-ai/stripes-ai-minions-ship-1300-prs-weekly-from-a-slack-emoji

• Notion 3.3 Custom Agents launch (February 24, 2026): https://www.notion.com/releases/2026-02-24

—

Where to find Ryan Nystrom:

X: https://x.com/ryannystrom

LinkedIn: https://www.linkedin.com/in/ryannystrom/

GitHub: https://github.com/rnystrom

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

Code with Claude: The 5 biggest updates explained07 May 202600:11:50

Claire breaks down the biggest announcements from Anthropic’s “Code with Claude” event and what they actually mean for builders shipping AI products today. From scheduled AI routines to outcome-based agents, multi-agent orchestration, and new memory systems, Claire walks through the features she’s most excited to use immediately—and how they could reshape the future of agentic software.


What you’ll learn:

  1. How Claude Code routines let you automate recurring workflows on schedules or webhooks
  2. What “Outcomes” are and how rubric-based agent grading works
  3. How multi-agent orchestration enables specialized AI teams with different roles and tools
  4. Why Anthropic’s new “Dreams” memory system matters for long-term agent behavior
  5. Why increased Claude Code usage limits are a bigger deal than they sound
  6. How Claire thinks about building practical agentic products today

—

Resources:

• Code with Claude: https://claude.com/code-with-claude

• Claude Code Routines Docs: https://code.claude.com/docs/en/routines

• Define Outcomes Docs: https://platform.claude.com/docs/en/managed-agents/define-outcomes

• Dreams Docs: https://platform.claude.com/docs/en/managed-agents/dreams

• Multi-Agent Docs: https://platform.claude.com/docs/en/managed-agents/multi-agent

• Managed Agent Webhooks Docs: https://platform.claude.com/docs/en/managed-agents/webhooks#supported-event-types

• Codex (OpenAI): https://openai.com/codex

• GitHub: https://github.com

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

Quests, token leaderboards, and a skills marketplace: The elite AI adoption playbook | John Kim (Sendbird)06 May 202600:42:19

John Kim is the co-founder and CEO of Delight.ai, a customer experience platform that’s transforming how companies deploy AI. But what makes John’s story fascinating isn’t just his product; it’s how he’s turned his entire company into an AI-native organization. His marketing team built a fully functional e-commerce swag store with Stripe integration in days. His sales team built their own CRM tools. His recruiting team automated their entire workflow. And it’s all tracked, measured, and celebrated through an internal platform called Automators.


What you’ll learn:

  1. How Sendbird’s marketing team built a fully functional swag store with Stripe integration in a day (with no engineering support)
  2. How the Automators platform works—an internal marketplace where anyone can request AI tools and engineers (or AI agents) can build them
  3. How to create secure, compliant templates so non-technical teams can ship to production safely
  4. How Sendbird built a token usage dashboard with five tiers (beginner through AI God) and why tracking the smoothness of the curve matters more than the total
  5. Why visible leadership usage is the most powerful adoption signal
  6. Why Sendbird rewrote job descriptions to prioritize curiosity, agency, and energy over years of experience
  7. How John uses AI for his own learning

—

Brought to you by:

WorkOS—Make your app enterprise-ready today

ThoughtSpot—Build AI-powered analytics into your product

—

In this episode, we cover:

(00:00) Introduction to John Kim

(02:45) The Delight.ai swag store built by marketing in two days

(05:51) The before times: when fun had to earn its place on the roadmap

(07:55) Demo: The Automators platform and quest system

(13:47) The AI Engineer for Internal Operations role

(16:06) Demo: The company-wide skills marketplace

(17:19) Treating AI adoption as a product

(18:43) Real wins: team-level and campaign examples

(21:51) Why SaaS isn’t dead—it’s being rebuilt internally

(23:46) Demo: The token tracking dashboard

(26:32) Measuring without fear: setting expectations, not punishments

(28:54) Quick recap

(30:51) Personal AI use cases: endless knowledge at your fingertips

(36:15) Lightning round and final thoughts

—

Tools referenced:

• Claude Code: https://claude.ai/code

• Codex (OpenAI): https://openai.com/codex

• Obsidian: https://obsidian.md

• GitHub: https://github.com

• Stripe: https://stripe.com

—

Other references:

• Jason Levin (CEO of Memelord) on How I AI: https://www.lennysnewsletter.com/p/from-a-690-newsletter-to-3m-api-how

• Konami Code: https://en.wikipedia.org/wiki/Konami_Code

• Andrew Huberman’s podcast: https://hubermanlab.com/

• Y Combinator: https://www.ycombinator.com/

—

Where to find John Kim:

X: https://x.com/doshkim

Instagram: https://instagram.com/dosh

LinkedIn: https://www.linkedin.com/in/doshkim/

Company: https://delight.ai

Delight.ai Spark Conference (May 7, SF): https://delight.ai/spark

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

The internal AI tool that’s transforming how Stripe designs products | Owen Williams04 May 202600:54:44

Owen Williams is a design manager at Stripe who built Protodash, an internal AI-powered prototyping platform that lets designers and PMs create high-quality Stripe dashboard prototypes without writing code. What started as a bundle of Cursor rules and React components evolved into a full web-based prototyping studio that runs in dev boxes, complete with design review modes, variant testing, and AI-powered iteration. Surprisingly, PMs now use Protodash just as much as designers, fundamentally changing how Stripe approaches prototyping, design reviews, and engineering handoffs.


What you’ll learn:

  1. How Stripe built an internal AI prototyping tool using Cursor rules, MCPs, and their design system
  2. Why “blurple slop” happens when designers use generic AI tools—and how to fix it
  3. The architecture behind Protodash: React router, design system components, and MCP integrations
  4. How Stripe prototypes in dev boxes so designers never have to worry about local setup
  5. Why “demos, not memos” transformed Stripe’s design review culture
  6. How Stripe built design review modes, variant testing, and AI annotation directly into your prototyping tool
  7. Why internal tools don’t need to be production-grade to be transformative

—

Brought to you by:

Celigo—Intelligent automation built for AI

Cursor—The best way to code with AI

—

In this episode, we cover:

(00:00) Welcome and intro to Owen Williams

(02:19) The “blurple slop” problem with AI design tools

(03:50) Protodash: an internal vibe-coding tool for Stripe prototypes

(05:26) Why an engineering background helped Owen lower the bar for designers

(07:55) The Cursor rules that taught the Stripe design system

(09:04) Running prototypes on dev boxes vs. locally

(10:30) “Demos, not memos” and rewiring design reviews at Stripe

(14:50) Building Protodash Studio: a browser-based wrapper for prototyping

(19:04) Live demo: variants, line charts, and remixing prototypes in browser

(21:02) Self-testing prototypes that take screenshots and check their work

(23:20) Multiple variant features

(26:08) The annotate-for-AI button for in-canvas feedback

(27:21) Design review mode: comments, summaries, and AI follow-up

(29:39) Why building internal tools beats buying off-the-shelf

(32:50) PMs as the surprise power users of Protodash

(35:20) Live demo: a Black Friday/Cyber Monday pet store dashboard

(42:03) Lo-fi modes, monospace fonts, and “Comic Sans for WIP” at Shopify

(44:45) Quick recap

(45:35) The Radar prototype that changed engineering handoff

(49:08) Lightning round and final thoughts

—

Blog & detailed workflow walkthroughs from this episode:

Stripe’s Owen Williams on Killing ‘Blurple Slop’ with an Internal Prototyping Studio: http://chatprd.ai/how-i-ai/stripe-owen-williams-on-buildling-internal-prototyping-studio

↳ How To Connect a Design System to an AI Code Editor for High Fidelity Prototypes: https://www.chatprd.ai/how-i-ai/workflows/how-to-connect-a-design-system-to-an-ai-code-editor-for-high-fidelity-prototypes

↳ Streamline Design Reviews with an AI-Powered Prototyping Studio: https://www.chatprd.ai/how-i-ai/workflows/streamline-design-reviews-with-an-ai-powered-prototyping-studio

↳ Build a Personal AI App to Track Purchases and User Manuals: https://www.chatprd.ai/how-i-ai/workflows/build-a-personal-ai-app-to-track-purchases-and-user-manuals

—

Tools referenced:

• v0: https://v0.app/

• Cursor: https://cursor.com/

• Claude Code: https://www.claude.com/product/claude-code

• Claude Design: https://www.anthropic.com/news/claude-design-anthropic-labs

• Figma: https://www.figma.com/

• Stripe Radar: https://stripe.com/radar

• Balsamiq: https://balsamiq.com/

—

Where to find Owen Williams:

X: https://x.com/ow

Website: https://owenwillia.ms/

LinkedIn: https://www.linkedin.com/in/owenpwilliams

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

From a $6.90 newsletter to $3M API: How a non-coder built Memelord | Jason Levin27 Apr 202600:51:53

Jason Levin is the CEO and founder of Memelord, an AI-powered meme creation platform that helps brands and individuals create contextual, trending memes. He started Memelord as a $6.90-per-month newsletter sending subscribers to a Google Slides deck, grew it to $100K ARR on Bubble without hiring engineers, then raised $3M to build it into an API-first product.


What you’ll learn:

  1. How Jason grew Memelord from a $6.90/month newsletter to $100K ARR without writing a single line of code
  2. Why “no UX is the best UX” and how agents are becoming Memelord’s primary users
  3. The mandatory vibe-coding rule for his marketing team and how it unlocks unprecedented creativity
  4. Why free tools are the new PDF downloads and how they’ve generated hundreds of thousands of emails
  5. Jason’s hardware hacking projects, including a bedside keyboard that creates Linear tickets without waking his wife
  6. Why AI can be funny (but humans are still funnier) and which model is the funniest
  7. The philosophy of building hyper-personalized software just for yourself

—

Brought to you by:

WorkOS—Make your app enterprise-ready today

Persona—Trusted identity verification for any use case

—

In this episode, we cover:

(00:00) Introduction to Jason Levin and Memelord

(04:28) Demo: Agentic meme creation with OpenClaw

(06:55) “No UX is the best UX”—building for an agent-first future

(08:35) How Memelord started as a $6.90 newsletter with Google Slides

(12:35) Building to $100K ARR on Bubble with 395 workflows

(15:20) Demo: Free tools section that generates hundreds of thousands of emails

(17:59) Why Cursor is perfect for non-technical founders

(20:20) Let your marketers cook—or watch them leave

(24:19) Commit graph that shows the vibe-coding inflection point

(25:25) Tools: Claude, Gemini, Linear, PostHog

(28:19) Build weird stuff in the real world

(33:24) Creative AI use cases

(39:56) Using OpenClaw for calendar analysis

(43:37) Can AI be funny? Which model is funniest?

(45:26) Memes are not slop

(46:45) What Jason doesn’t use AI for

(48:12) Final thoughts

—

Blog & detailed workflow walkthroughs from this episode:

How I AI: Jason Levin’s Workflows for Agentic Memes, Vibe Coding, and Hardware Hacking: https://www.chatprd.ai/how-i-ai/jason-levins-workflows-for-agentic-memes-vibe-coding-and-hardware-hacking

↳ Build a Custom Bedside Keyboard for Idea Capture with Raspberry Pi and ChatGPT: https://www.chatprd.ai/how-i-ai/workflows/build-a-custom-bedside-keyboard-for-idea-capture-with-raspberry-pi-and-chatgpt

↳ Build Free Marketing Tools as Lead Magnets Using AI Code Assistants: https://www.chatprd.ai/how-i-ai/workflows/build-free-marketing-tools-as-lead-magnets-using-ai-code-assistants

↳ Automate Meme Marketing with an AI Agent and OpenClaw: https://www.chatprd.ai/how-i-ai/workflows/automate-meme-marketing-with-an-ai-agent-and-openclaw

—

Tools referenced:

• Memelord API: https://memelord.com/api

• Cursor: https://cursor.com/

• Bubble: https://bubble.io/

• OpenClaw: https://openclaw.ai

• Claude: https://claude.ai/

• ChatGPT: https://chat.openai.com/

• Gemini: https://gemini.google.com/

• Grok: https://grok.x.ai/

• Linear: https://linear.app/

• PostHog: https://posthog.com/

• Zapier: https://zapier.com/

—

Other references:

• Diego Zaks—“The best UX is no UX”: https://x.com/diegozaks/status/1966526522136649980

• Sam Lessin: https://wlessin.com/

• “Stop giving me advice”: https://stopgivingmeadvice.com

• Memelord free tools: https://memelord.com/tools

—

Where to find Jason Levin:

Twitter: https://twitter.com/iamjasonlevin

Instagram: https://instagram.com/iamjasonlevin

LinkedIn: https://www.linkedin.com/in/iamjasonlevin/

Memelord: https://memelord.com

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

GPT 5.5 just did what no other model could23 Apr 202600:23:36

In this mini episode, I break down OpenAI’s new GPT 5.5 and GPT 5.5 Pro after weeks of early testing. I walk through three real jobs I threw at the model:  building an app for me to teach my second grader more advanced subtraction concepts, tackling a tech debt problem in the ChatPRD codebase, and hacking into a proprietary Bluetooth pixel display that every other model had failed me on. My verdict: higher intelligence, better efficiency, and genuinely autonomous long-running loops that change what I think is worth tackling.


What you’ll learn:

  1. How I think about GPT 5.5 Pro’s pricing vs engineering time, and when I believe the “intelligence tax” is worth paying
  2. Why I treat GPT 5.5 as a developer model first, and why I couldn’t find a consumer use case that justified its intelligence
  3. The exact prompt pattern I use to unlock a long-running autonomous subagent loop
  4. How I got a near-six-hour autonomous run to one-shot 98% of edge cases in a migration over millions of chat threads and drop my Sentry error rate to the floor
  5. Why I’m now throwing GPT 5.5 at tech debt, flaky tests, and security backlogs first
  6. How I combined a Bluetooth packet sniffer and GPT 5.5 to reverse-engineer a proprietary pixel speaker after Claude Code and GPT 5.4 both gave up
  7. How I use the /personality command inside Codex to swap the default “baked potato” tone for something I actually enjoy working with

—

In this episode, I cover:

(00:00) Introduction to GPT 5.5 testing

(00:40) What is GPT 5.5 and how much does it cost?

(03:23) Testing GPT 5.5 in ChatGPT: the intelligence overhang problem

(07:12) Moving to Codex: where GPT 5.5 really shines

(16:01) Hacking a Chinese Bluetooth speaker

(21:47) Final thoughts on GPT 5.5’s intelligence and efficiency

—

Tools referenced:

• GPT 5.5 and GPT 5.5 Pro: https://openai.com/index/introducing-gpt-5-5/

• Codex: https://openai.com/codex/

• ChatGPT: https://chat.openai.com/

• Claude Code: https://claude.ai/code

• Sentry: https://sentry.io/

• Divoom MiniToo: https://divoom.com/products/minitoo

—

Other references:

• OpenAI Codex Security: https://openai.com/index/codex-security-now-in-research-preview/

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

What Claude Design is actually good for (and why Figma isn’t dead, yet)22 Apr 202600:27:33

In this mini episode, I do a full walkthrough of the AI design tools that dropped in April 2026: Anthropic’s new Claude Design, OpenAI’s GPT Images 2.0, and Google Labs’ open-source DESIGN.md format. I import a full design system from Lenny’s Newsletter, build a landing page, turn my own article into a polished deck, generate a brand kit for ChatPRD, and run a personal color analysis from a photo.


What you’ll learn:

  1. How Claude Design handles design system imports and whether it can actually replace Figma
  2. The three best use cases for Claude Design: marketing landing pages, slide decks, and creative redesigns
  3. Why ChatGPT Images 2.0 is a breakthrough for brand kits and layout work
  4. Google’s new DESIGN.md standard
  5. The practical limits of AI design tools (spoiler: you’ll hit credit limits fast)

—

Brought to you by:

WorkOS—Make your app enterprise-ready today

Rippling—Stop wasting time on admin tasks, build your startup faster

—

In this episode, we cover:

(00:00) Welcome and what’s in the spring 2026 AI design drop

(01:45) Claude Design overview

(03:05) Importing Lenny’s Newsletter design system into Claude Design

(04:06) How Claude Design structures a design system

(05:42) Google Labs’ DESIGN.md standard

(06:41) Building Lenny Doc, a PRD generator landing page using the Lenny design system

(09:44) Why the three-variation output is Claude Design’s smartest UX choice

(10:20) Hitting the Claude Design limit and paying $200 to keep going

(11:05) Where Figma still wins

(13:20) Reviewing Lenny Doc

(16:19) Turning an Open Claude article into a branded slide deck

(17:57) The ’90s GeoCities “Lenny’s Product Zone” redesign

(19:44) Claude Design recap

(20:15) ChatGPT Images 2.0 and what makes it the first “thinking” image model

(21:25) Generating a multi-page brand kit for ChatPRD and iterating with reference images

(23:43) Personal color analysis demo

(26:02) Recap

—

Detailed workflow walkthroughs from this episode:

• How I Put Claude Design and GPT Images 2.0 to the Test: Building Landing Pages, Slides, and Brand Kits: https://www.chatprd.ai/how-i-ai/claude-design-and-gpt-images-2-building-landing-pages-slides-and-brand-kits

• How to Generate a Professional Brand Kit with GPT Images 2.0: https://www.chatprd.ai/how-i-ai/workflows/how-to-generate-a-professional-brand-kit-with-gpt-images-2-0

• How to Convert an Article into a Polished Slide Deck with AI: https://www.chatprd.ai/how-i-ai/workflows/how-to-convert-an-article-into-a-polished-slide-deck-with-ai

• How to Build a High-Fidelity Landing Page with Claude Design: https://www.chatprd.ai/how-i-ai/workflows/how-to-build-a-high-fidelity-landing-page-with-claude-design

—

Tools referenced:

• Claude Design: https://claude.ai/design

• ChatGPT Images 2.0: https://openai.com/index/introducing-chatgpt-images-2-0/

• Midjourney: https://www.midjourney.com/

—

Other references:

• Google’s DESIGN.md: https://stitch.withgoogle.com/docs/design-md/overview

• Lenny’s Newsletter: https://www.lennysnewsletter.com/

• Jamie Gannon “How I AI” episode on reference styles: https://www.lennysnewsletter.com/p/mastering-midjourney-how-to-create

• Brand prompt inspiration: https://x.com/riomadeit/status/2046682442791071787

• Figma team “How I AI” episode on design systems: https://www.lennysnewsletter.com/p/from-figma-to-claude-code-and-back

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

How Intercom 2x’d their engineering velocity in 9 months with Claude Code | Brian Scanlan20 Apr 202601:18:45

Brian Scanlan is a senior principal engineer at Intercom, where he’s led the company’s transformation to AI-first engineering. In just nine months, Intercom doubled their R&D throughput while maintaining code quality, with 100% of engineers—plus designers, PMs, and TPMs—now shipping code via Claude Code.


What you’ll learn:

  1. How Intercom doubled their merged PRs per R&D employee in just nine months using Claude Code
  2. The telemetry infrastructure they built to measure AI adoption and quality across hundreds of engineers
  3. Why they built a skills repository with hooks that enforce engineering standards automatically
  4. How they’re preparing their product for an agent-first world with CLIs, MCPs, and ephemeral APIs
  5. The permission and accountability framework that enabled rapid AI adoption
  6. Why backlog zero is now achievable and what that means for engineering culture

—

Brought to you by:

Celigo—Intelligent automation built for AI

Cursor—The best way to code with AI

—

In this episode, we cover:

(00:00) Introduction to Brian Scanlan

(02:40) Why Intercom went all-in on AI for both product and engineering

(05:01) The breakthrough moment with Opus 4.6 and Christmas break 2025

(07:02) Demo: Intercom’s merged PRs per R&D head

(12:50) Agent-first work as a fundamental reimagining of technical workflows

(14:27) The cost tradeoff: treating AI spend as an investment

(16:47) Measuring quality

(21:22) Demo: Shipping a redirect in the Rails monolith with Claude Code

(24:03) Creating a custom PR skill

(26:33) Building a software factory with predictable quality standards

(30:15) Telemetry infrastructure: Honeycomb for skill usage tracking

(32:10) Session data collection and personalized usage insights

(36:08) Quick overview

(39:20) Walking through Intercom’s skills repository

(42:16) Deep dive: The flaky spec skill and how it reached 100x capability

(46:44) The “and then” workflow for building comprehensive skills

(52:31) The live website and overview of workflows

(53:32) How internal AI experience informs customer product decisions

(56:18) Making SaaS products agent-friendly with CLIs and helpful hints

(01:03:49) Why conversion drop-off is invisible in agent-driven workflows

(01:05:28) Lightning round and final thoughts

—

Detailed workflow walkthroughs from this episode:

• How Intercom Doubled Engineering Output: Brian Scanlan's 4 AI Workflows for Claude Code: https://www.chatprd.ai/how-i-ai/how-intercom-doubled-engineering-output-brian-scanlan-ai-workflows-for-claude-code

• Design an Agent-Friendly CLI to Automate SaaS Product Onboarding: https://www.chatprd.ai/how-i-ai/workflows/design-an-agent-friendly-cli-to-automate-saas-product-onboarding

• Build a Self-Improving AI Agent to Automatically Fix Flaky Tests: https://www.chatprd.ai/how-i-ai/workflows/build-a-self-improving-ai-agent-to-automatically-fix-flaky-tests

• Automate High-Quality Pull Request Descriptions with a Custom AI Skill: https://www.chatprd.ai/how-i-ai/workflows/automate-high-quality-pull-request-descriptions-with-a-custom-ai-skill

—

Tools referenced:

• Claude Code: https://claude.ai/code

• Cursor: https://cursor.com/

• Honeycomb: https://www.honeycomb.io/

• Snowflake: https://www.snowflake.com/

• Fin AI: https://www.intercom.com/fin

• Vercel: https://vercel.com/

—

Other references:

• Intercom GitHub Repo: https://github.com/intercom

• Google API Go Client Repo: https://github.com/googleapis/google-api-go-client

—

Where to find Brian Scanlan:

X: https://x.com/brian_scanlan

LinkedIn: https://www.linkedin.com/in/scanlanb/

Company: https://www.intercom.com

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

© My Podcast Data · Independent project · Data from Apple & Spotify