How I AI, hosted by Claire Vo, is for anyone wondering how to actually use these magical new tools to improve the quality and efficiency of their work. In each episode, guests will share a specific, practical, and impactful way they’ve learned to use AI in their work or life. Expect 30-minute episodes, live screen sharing, and tips/tricks/workflows you can copy immediately. If you want to demystify AI and learn the skills you need to thrive in this new world, this podcast is for you.
Site
RSS
Spotify
Apple
Data updated on 06/10/2026
Recent rankings
Latest chart positions across Apple Podcasts and Spotify rankings.
Shared links between episodes and podcasts
Links found in episode descriptions and other podcasts that share them.
How OpenAI uses ChatGPT Sites (live at DevDay!) | Kath Korevec (Product Lead)
Monday, October 5, 2026 • Duration 35:40
Kath Korevec is a member of the Product staff at OpenAI working on Codex, and she spent over a year building and using ChatGPT Sites internally before its public launch. She’s been on the front lines of shipping Plugin Insights, MCP plugin hosting, and the connector ecosystem, which now includes around 60 integrations.
What you’ll learn:
What Plugin Insights actually does, and why it changes who can use a site
The incident command site Kath built for her OpenAI team, and how it uses live Slack and Notion connectors
The one phrase that tells Codex to wire up connectors for you
The infrastructure layer inside Sites that most people haven’t touched yet
How Kath fixed her Spotify after her kids wrecked it, using Reddit and computer use
The skill distribution model behind her community dungeon crawler, and why it’s a new way to think about collaboration
Why model speed is what actually determines how creative you get
Where Kath draws a hard line on AI acting in her name
Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.
Jev: 8 real use cases for the fastest, cheapest model I’ve ever used | John Lindquist
Wednesday, September 30, 2026 • Duration 46:06
John Lindquist created egghead.io, a developer education platform used by hundreds of thousands of working engineers. These days he’s building mega.dev, a hands-on program specifically for developers who want to do real work with AI agents, not just prototype them.
What you’ll learn:
Why Jev is a decision engine, not a chatbot, and what that distinction actually changes about how you build
How John built a real-time voice to-do app that classifies and executes commands with no visible pause
The data deduplication pattern that merges messy records in milliseconds using confidence scores
Why Jev works best as a router, and how a single text input can navigate users deep into an app
What a chess match between Jev and a low-reasoning LLM reveals about speed, cost, and when to use which
The multi-step classification pattern John reaches for when one Jev pass isn’t enough
Where Jev falls short, and when you should still reach for a full generative model
(10:38) Demo: plain English to function name (grocery cart)
(11:50) Demo: data deduplication and record merging
(13:45) Confidence scores and multi-model validation
OpenAI Dev Day 2026: The releases that actually matter
Wednesday, September 30, 2026 • Duration 24:20
I spent the day at OpenAI’s DevDay in San Francisco, and I have good news and bad news: OpenAI released a lot of stuff.
In this episode, I break down the announcements worth paying attention to - and show you what happened when I tested some of them early. We’ll meet my Dot, explore why Spaces and Sites could matter for how teams work, and get into the model and API updates I’m most excited about as a developer.
I use the Decisions API to find podcast thumbnails where nobody looks awkward, build a collaborative sketchpad with Astra ultrafast, and let my kids redesign a 3D world in real time. That last experiment cost about $97. My wallet has thoughts.
These are my early impressions: what’s promising, what still feels rough, and what I think you should try first.
What you’ll learn:
What OpenAI’s Dots can do, how I’ve been using mine, and why I’m waiting to give a full verdict
Why Spaces might be one of the most underhyped announcements for collaboration between humans and agents
How Sites with connectors and plugins could help teams share internal tools with the right data permissions
Where GPT-6.1 Sol fits in my model stack—and why speed and cost matter
What vision adds to the Decisions API, including my thumbnail-selection and hot dog demos
What Astra ultrafast makes possible for interactive AI apps, from collaborative drawing to a changing 3D game
Where the speed feels magical, where the experience still needs work, and what it costs
—
In this episode, we cover:
(00:00) OpenAI DevDay recap—and pressing the Codex reset button
(00:58) Dots: early impressions and rough edges
Jev for beginners: how to use it and what to build
Monday, September 28, 2026 • Duration 26:24
Jev is TypeSafe AI’s new decision model. It returns type-safe structured values (a choice, a score, a probability) instead of generated text, at 4 cents per million input tokens with no output charge. This week I ran it on five real projects: PR categorization, a meta-analysis of my own Claude and Codex sessions, Gmail triage, the ChatPRD product insights graph, and a live audience dashboard built from 4,500 YouTube comments.
What you’ll learn:
What makes Jev fundamentally different from every other model I’ve used
How I analyzed 1,700 PRs for 9 cents and what I found out about where my engineering effort actually went
The personal meta-analysis you can run on your own Claude and Codex sessions right now
Why I stopped using Jev alone, and what I pair it with now
How I turned 4,500 YouTube comments into a searchable audience dashboard for almost nothing
The real-time app I built in an afternoon that shows something surprising about Jev’s speed
Why Jev’s pricing model is different from any LLM I’ve used, and what it makes practical to build
The ChatPRD product insights project: 1,100 signals, 200,000 classifications, and what it cost me
—
Brought to you by:
OpenArt—An all-in-one AI creation platform for images, videos, music, audio, and more
—
In this episode, we cover:
(00:00) Jev launch and what makes it different from every other model
(02:49) Type-safe values explained
(05:28) Understanding Jev outputs
(07:39) Use case 1: PR categorization and pairwise clustering
Opus 5.5 vs. GPT-6 Sol: which model won my blind taste test?
Tuesday, September 22, 2026 • Duration 38:51
I got up early to record an Opus 5.5 review. Then Anthropic and OpenAI dropped new models on the same morning, and I decided to do something I’d never done before: take the How I AI bench live. I put GPT-6 Astra, GPT-6 Sol, Claude Opus 5.5, and more through the work I actually care about: emails, PRDs, frontend prototypes, backend work, long-running agents, SVGs, and video editing. I scored the outputs without knowing which model made them, so you get to watch me make predictions, change my mind, and reveal my own very inconsistent taste. Astra won my heart. Opus 5.5 won my week. Sol still has me split. There’s a creative result I got completely wrong, an LLM judge that disagreed with me, and a return to Barbie Bench: the 3D fashion game that keeps reminding me how far we have to go. The hands are tragic. AGI has not arrived.
What you’ll learn:
How I run the How I AI bench blind, and what gets an output a bad score before I even know which model made it
Why Astra won my heart while Opus 5.5 might be overall strongest, especially for long-running agents and B2B frontend
Where Sol still wins me over on clear writing, readable PRDs, and price
The character SVG results that completely overturned my prediction about Anthropic
What happened when I asked these models to edit video, and why I think skills explain part of the disappointment
Why an LLM judge disagreed with my rankings, and what it was rewarding that I wasn’t
—
In this episode, we cover:
(00:00) LIVE setup and new model launches
(01:30) What’s new in Opus 5.5, Sol, and Luna
(04:11) Guardrails, personality, and speed
(09:00) The How I AI bench and blind evaluation process
(11:31) Email and personal-productivity results
(13:50) Frontend prototype vibe checks
(24:10) Backend, agent personality, and long-running tasks
I left Claude for months. Opus 5.5 is why I'm back
Tuesday, September 22, 2026 • Duration 24:52
I’ve been off Claude for months. Not because it got dumb, but because it got annoying. The rambling, the hedging, the preachy little disclaimers on tasks that didn’t need them. I moved most of my daily work to Codex and I didn’t miss it. Then Anthropic shipped Opus 5.5: 40% cheaper than Opus 5, faster, and with what they’re calling a fundamentally different alignment approach. I ran it for a week across real work, including four long-running agentic tasks, a full ChatPRD homepage redesign, an SVG benchmark, and one very firm refusal, and I’m ready to give you the honest verdict. There’s a lot to like. There are still two things that drive me a little crazy. And there’s one capability I genuinely wasn’t expecting.
What you’ll learn:
Why I walked away from Claude entirely, and what it took for me to come back
The real cost math on Opus 5.5 and why pricing matters more for agentic work than single prompts
What happened when I ran four long-running agentic tasks, including one that tried to manipulate Claude mid-run
Why Opus 5.5 is now my go-to for frontend prototyping, and where it still lets me down
The one capability I genuinely didn’t see coming, and no other model in my stack can match it
The moment Opus 5.5 told me flat-out no, and what that says about where Anthropic’s safety posture actually lands in practice
Where Codex still wins, and how I’m splitting my model stack after a full week of testing
—
In this episode:
(00:00) Why I stopped using Claude
(01:02) What Anthropic says Opus 5.5 is
(01:54) Cost, speed, and benchmark overview
(03:20) Safety, alignment, and the cybersecurity limits
(05:02) How I AI bench
(05:39) Voice test: is it actually not annoying?
(07:54) Long-running agentic task results
How Warp ships 2,000 PRs a month with AI factories | Zach Lloyd (CEO, Warp)
Monday, September 21, 2026 • Duration 46:51
Zach Lloyd is the co-founder and CEO of Warp, an AI-powered terminal and software factory platform used by tens of thousands of engineers. Before Warp, he spent nearly a decade at Google, including time as a principal engineer on Google Sheets. He built Warp from the ground up as a modern, AI-native alternative to legacy terminals, and the team has since expanded into software factories: a full cloud-based system that takes an idea in Slack all the way through to a merged PR.
In this episode:
Why a software factory is more than a coding agent
The public Slack → Linear → GitHub → QA workflow
Human interactions per PR as a signal of automation and throughput
Why human review is still the bottleneck
Scoring agent runs, finding failure modes, and self-improving agent workflows
Replaying real tasks to choose model cost and quality tradeoffs
CEO workflows with Figma MCP, Granola, and research agents
OpenArt—An all-in-one AI creation platform for images, videos, music, audio, and more
—
In this episode, we cover:
(00:00) Intro
(02:35) Warp’s AI software factory, Wilson
(09:23) Automatic factory triggers
(11:12) The engineering leader dashboard Zach wishes he’d had
(15:18) How code review is changing in an AI factory
Muse review: The personal AI agent that gets consumer UX right
Wednesday, September 16, 2026 • Duration 37:11
I spent a few hours putting Meta’s Muse, its new personal AI agent, through a real first-pass test: onboarding, calendar management, goal setting, a one-shot family morning newsletter, browser-based shopping, and the animated avatar that honestly surprised me.
What you’ll learn:
Why Muse is the best-designed personal agent I’ve tested, and what specifically made it feel that way
The one-shot family PDF Muse produced that Claude and Codex never quite nailed
How Muse’s permission model works, and why it’s different from every other agent I’ve used
Why I set up a sleep training goal in Muse, and what it revealed about agent tone
The activity feed feature I immediately wished Codex and Claude Code had
Where Muse failed, and what it says about the limits of this category right now
The animated avatar decision that showed me what top-of-craft AI product design actually looks like
—
Brought to you by:
Optimizely—Your AI agent orchestration platform for marketing and digital teams
OpenArt—An all-in-one AI creation platform for images, videos, music, audio, and more
—
In this episode, we cover:
(00:00) What Muse is and who it’s actually built for
(04:41) Signing in and the onboarding flow
(07:16) The activity feed and its task lineage
(08:24) First real task: managing the family calendar and deleting soccer practice
How Grok Bot designers use AI agents to build personal sites and product prototypes | John Bai & Peng Zheng
Monday, September 14, 2026 • Duration 41:34
John Bai and Peng Zheng are designers on the Grok Bot team at SpaceXAI, where they’re building one of the most talked-about AI products right now. John writes publicly about his design process (his piece “Designing Grok Bot with Grok Bot” has already made the rounds) and shares bot templates with the design community. Peng brings a product-design sensibility to personal tools, and his website doubles as a live demo of what he builds.
What you’ll learn:
How Peng built a self-updating personal website using Grok Bot as the entire backend pipeline, with no CMS and no Figma file
The exact check-in bot setup that lets Peng send a photo or a place name and have his portfolio update itself automatically
How John’s Figma Bro bot handles production design tasks while he’s at the gym
How John uses voice memos to direct Figma work through an MCP connection without opening his laptop
The “shower thought to prototype” workflow John uses with DevBot to test interaction ideas without first going through a product manager or engineer
The “trash can method” of software development
How both designers organize their personal bot ecosystems
What John and Peng actually think AI means for the future of design as a craft
—
Brought to you by:
WorkOS—Make your app enterprise-ready, with SSO, SCIM, RBAC, and more
Build your own company brain: the enterprise AI playbook from Stripe’s engineering team | Sharadh Krishnamurthy
Monday, September 7, 2026 • Duration 50:24
Sharadh Krishnamurthy is an engineering manager at Stripe, where he helped build Kai, the company’s internal AI agent used by more than 10,000 employees every week. He’s worked across several of Stripe’s core infrastructure teams, including data and developer experience, which gives him a grounded, systems-level perspective on what it actually takes to make AI work at enterprise scale. He’s currently focused on the governance, skills, and infrastructure layers that let every Stripe employee use AI safely and effectively, regardless of their technical background.
What you’ll learn:
Why Stripe built Kai from scratch instead of buying, and what tipped the decision
What Kai knows about you by default and what you actually control
Why “projects” at Stripe are a governance mechanism, not just a folder
How Stripe structured its data layer so agents can query safely at scale
Why the infrastructure Stripe built for human developers turned out to be exactly what agents needed
How Kai’s skills platform lets any employee package a workflow, and what happens when you have 2,000 of them
What Sharadh learned the hard way when agents nearly took down production systems
Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.
Related Shows Based on Content Similarities
Discover shows related to How I AI, based on actual content similarities. Explore podcasts with similar topics, themes, and formats, backed by real data.