Infinite, shareable volume storage with Hunter Leath, Archil CEO
Épisode 31
jeudi 15 janvier 2026 • Durée 55:13
Hunter Leath, CEO of Archil, explains how they’re building a “universal storage engine” that sits between your apps and S3—making an S3 bucket behave like a fast, POSIX-compatible disk for containers, servers, and even Lambda. Along the way, we dig into how their SSD-backed clusters and custom protocol avoid the usual small-file pain and where this approach shines (and where it doesn’t).
Follow Aaron: Twitter/X: https://twitter.com/aarondfrancis Database School: https://databaseschool.com Database School YouTube Channel: https://www.youtube.com/@UCT3XN4RtcFhmrWl8tf_o49g (Subscribe today) LinkedIn: https://www.linkedin.com/in/aarondfrancis Website: https://aaronfrancis.com - find articles, podcasts, courses, and more.
Chapters: 00:00 - Intro: Archil Data and “S3 as a disk” 01:05 - Hunter’s background and the core pitch 02:32 - The real problem: state management (S3 vs block storage) 05:02 - SQLite on S3: what the stack looks like 07:13 - The missing layer: durable SSD-backed clusters 10:14 - Who uses this: unstructured data, CI/CD, Git, agents 12:15 - Small files + Git performance and avoiding S3 request explosion 16:22 - Why they built a new protocol (NFS vs Luster) 20:00 - What gets written to S3: real files in your bucket 22:29 - S3 limits, throttling, and the “keep it on SSD” escape hatch 25:32 - Multi-cloud + R2, and why regions/latency matter 32:10 - Pricing model: “pay only when data is active” 34:41 - Tradeoffs: random reads and ultra-low-latency metal 37:19 - Storage/compute separation and AI/agent-native workflows 43:21 - YC timeline + the marketing challenge of a “universal layer” 47:34 - Single-tenant clusters for enterprises and why it’s hard 50:27 - Where the company is now, hiring, and how to try it (disk.new)
Building search for AI systems with Chroma CTO Hammad Bashir
Épisode 30
jeudi 18 décembre 2025 • Durée 01:06:43
Hammad Bashir, CTO of Chroma, joins the show to break down how modern vector search systems are actually built from local, embedded databases to massively distributed, object-storage-backed architectures. We dig into Chroma’s shared local-to-cloud API, log-structured storage on object stores, hybrid search, and why retrieval-augmented generation (RAG) isn’t going anywhere.
Follow Aaron: Twitter/X: https://twitter.com/aarondfrancis Database School: https://databaseschool.com Database School YouTube Channel: https://www.youtube.com/@UCT3XN4RtcFhmrWl8tf_o49g (Subscribe today) LinkedIn: https://www.linkedin.com/in/aarondfrancis Website: https://aaronfrancis.com - find articles, podcasts, courses, and more.
Chapters: 00:00 – Introduction From high-school ASICs to CTO of Chroma 01:04 – Hammad’s background and why vector search stuck 03:01 – Why Chroma has one API for local and distributed systems 05:37 – Local experimentation vs production AI workflows 08:03 – What “unprincipled data” means in machine learning 10:31 – From computer vision to retrieval for LLMs 13:00 – Exploratory data analysis and why looking at data still matters 16:38 – Promoting data from local to Chroma Cloud 19:26 – Why Chroma is built on object storage 20:27 – Write-ahead logs, batching, and durability 26:56 – Compaction, inverted indexes, and storage layout 29:26 – Strong consistency and reading from the log 34:12 – How queries are routed and executed 37:00 – Hybrid search: vectors, full-text, and metadata 41:03 – Chunking, embeddings, and retrieval boundaries 43:22 – Agentic search and letting models drive retrieval 45:01 – Is RAG dead? A grounded explanation 48:24 – Why context windows don’t replace search 56:20 – Context rot and why retrieval reduces confusion01:00:19 – Faster models and the future of search stacks01:02:25 – Who Chroma is for and when it’s a great fit01:04:25 – Hiring, team culture, and where to follow Chroma
The database for all your AI needs
Épisode 21
mardi 16 septembre 2025 • Durée 01:00:07
Marcel Kornacker, the creator of Apache Impala and co-creator of Apache Parquet, joins me to talk about his latest project: Pixeltable, a multimodal AI database that combines structured and unstructured data with rich, Python-native workflows.
From ingestion to vector search, transcription to snapshots, Pixeltable eliminates painful data plumbing for modern AI teams.
Website: https://aaronfrancis.com – find articles, podcasts, courses, and more
Database School: https://databaseschool.com
Chapters
0:00 – Introduction
0:20 – Meet Marcel Kornacker
1:19 – Early career and grad school in databases
2:12 – Joining Google and building F1
3:42 – How F1 used Spanner at Google
4:01 – Starting Apache Impala at Cloudera
6:02 – Why SQL still matters
7:29 – What keeps Marcel fascinated with databases
9:37 – The “SQL is dead” waves and shift to AI
10:21 – Observing pain points in computer vision pipelines
13:02 – Multimodal data challenges and the idea for Pixeltable
16:10 – How Pixeltable handles transformations with computed columns
Sharding Postgres without extensions with PgDog founder, Lev Kokotov
Épisode 20
mardi 19 août 2025 • Durée 48:53
I chat with Lev Kokotov to talk about building PgDog, an open-source sharding solution for Postgres that sits outside the database. Lev shares the journey from creating PgCat to launching PgDog through YC, the technical challenges of sharding, and why he believes scaling Postgres shouldn’t require extensions or rewrites.
Follow Aaron: Twitter: https://twitter.com/aarondfrancis LinkedIn: https://www.linkedin.com/in/aarondfrancis Website: https://aaronfrancis.com - find articles, podcasts, courses, and more.
Database school: https://databaseschool.com
Chapters: 00:00 - Intro to guest Glauber Costa 00:58 - Glauber's background and path to databases 02:23 - Moving to Texas and life changes 05:32 - The origin story of Turso 07:55 - Why fork SQLite in the first place? 10:28 - SQLite’s closed contribution model 12:00 - Launching libSQL as an open contribution fork 13:43 - Building Turso Cloud for serverless SQLite 14:57 - Limitations of forking SQLite 17:00 - Deciding to rewrite SQLite from scratch 19:08 - Branding mistakes and naming decisions 22:29 - Differentiating Turso (the database) from Turso Cloud 24:00 - Technical barriers that led to the rewrite 28:00 - Why libSQL plateaued for deeper improvements 30:14 - Big business partner request leads to deeper rethink 31:23 - The rewrite begins 33:36 - Early community traction and GitHub stars 35:00 - Hiring contributors from the community 36:58 - Reigniting the original vision 39:40 - Turso’s core business thesis 42:00 - Fully pivoting the company around the rewrite 45:16 - How GitHub contributors signal business alignment 47:10 - SQLite’s rock-solid rep and test suite challenges 49:00 - The magic of deterministic simulation testing 53:00 - How the simulator injects and replays IO failures 56:00 - The role of property-based testing 58:54 - Offering cash for bugs that break data integrity 1:01:05 - Deterministic testing vs traditional testing 1:03:44 - What it took to release Turso Alpha 1:05:50 - Encouraging contributors with real incentives 1:07:50 - How to get involved and contribute 1:20:00 - Upcoming roadmap: indexes, CDC, schema changes 1:23:40 - Final thoughts and where to find Turso
Vitess for Postgres, with the co-founder of PlanetScale
Épisode 17
mardi 1 juillet 2025 • Durée 01:07:29
Sugu Sougoumarane, co-creator of Vitess and co-founder of PlanetScale, joins me to talk about his time scaling YouTube’s database infrastructure, building Vitess, and his latest project bringing sharding to Postgres with Multigres.
This was a fun conversation with technical deep-dives, lessons from building distributed systems, and why he’s joining Supabase to tackle this next big challenge.
Website: https://aaronfrancis.com — find articles, podcasts, courses, and more.
Chapters:
00:00 - Intro: What is PlanetScale Metal?
00:39 - Meet Richard Crowley
01:33 - What is Vitess and how does it work?
03:00 - Where PlanetScale fits into the picture
09:03 - Why EBS is the default and its trade-offs
13:03 - How PlanetScale handles durability without EBS
16:03 - The engineering work behind PlanetScale Metal
22:00 - Deep dive into backups, restores, and availability math
25:03 - How PlanetScale replaces instances safely
From Prisma Founder to LiveStore: Building local-first apps with Johannes Schickling
Épisode 15
jeudi 29 mai 2025 • Durée 01:31:40
Johannes Schickling, original founder of Prisma, joins me to talk about LiveStore, his ambitious local-first data layer designed to rethink how we build apps from the data layer up.
We dive deep into event sourcing, syncing with SQLite, and why this approach might power the next generation of reactive apps.
Website: https://aaronfrancis.com — find articles, podcasts, courses, and more
How Durable Objects and D1 Work: A Deep Dive with Cloudflare’s Josh Howard
Épisode 14
mercredi 14 mai 2025 • Durée 01:14:40
Josh Howard, Senior Engineering Manager at Cloudflare, joins me to explain how Durable Objects and D1 work under the hood—and why Cloudflare’s approach to stateful serverless infrastructure is so unique. We get into V8 isolates, replication models, routing strategies, and even upcoming support for containers.
Want to learn more about SQLite? Check out my SQLite course: https://highperformancesqlite.com/?ref=podcast
Follow Aaron: Twitter: https://twitter.com/aarondfrancis LinkedIn: https://www.linkedin.com/in/aarondfrancis Website: https://aaronfrancis.com - find articles, podcasts, courses, and more.
Database school on YouTube: https://www.youtube.com/playlist?list=PLI72dgeNJtzqElnNB6sQoAn2R-F3Vqm15 Database school audio only: https://databaseschool.transistor.fm
Chapters 00:00 - Intro 00:37 - What is a Durable Object? 01:43 - Cloudflare’s serverless model and V8 isolates 03:58 - Why stateful serverless matters 05:14 - Durable Objects vs Workers 06:22 - How routing to Durable Objects works 08:01 - What makes them "durable"? 08:51 - Tradeoffs of colocating compute and state 10:58 - Stateless Durable Objects 12:49 - Waking up from sleep and restoring state 16:15 - Durable Object storage: KV and SQLite APIs 18:49 - Relationship between D1, Workers KV, and DOs 20:34 - Performance of local storage writes 21:50 - Storage replication and output gating 24:15 - Lifecycle of a request through a Durable Object 26:46 - Replication strategy and long-term durability 31:25 - Placement logic and sharding strategy 36:35 - Use cases: agents, multiplayer games, chat apps 40:33 - Scaling Durable Objects 41:14 - Globally unique ID generation 43:22 - Named Durable Objects and coordination 46:07 - D1 vs Workers KV vs Durable Objects 47:50 - Outerbase acquisition and DX improvements 49:49 - Querying durable object storage 51:20 - Developer Week highlights and new features 52:44 - Read replicas and sticky sessions 53:49 - Containers and the future of routing 56:47 - Deployment regions and infrastructure expansion 57:43 - Hiring and how to connect with Josh
20 years of hacking Postgres with Heikki Linnakangas (cofounder of Neon)
Épisode 13
mardi 6 mai 2025 • Durée 02:00:11
In this episode of Database School, I talk with Heikki Linnakangas, co-founder of Neon and longtime PostgreSQL hacker, to talk about 20+ years in the Postgres community, the architecture behind Neon, and the future of multi-threaded Postgres. From paternity leave patches to branching production databases, we cover a lot of ground in this deep-dive conversation.
Links: Let's make postgres multi-threaded: https://www.postgresql.org/message-id/31cc6df9-53fe-3cd9-af5b-ac0d801163f4%40iki.fi Hacker News discussion: https://news.ycombinator.com/item?id=36284487
Follow Aaron: Twitter: https://twitter.com/aarondfrancis LinkedIn: https://www.linkedin.com/in/aarondfrancis Website: https://aaronfrancis.com - find articles, podcasts, courses, and more.
Database school on YouTube: https://www.youtube.com/playlist?list=PLI72dgeNJtzqElnNB6sQoAn2R-F3Vqm15 Database school audio only: https://databaseschool.transistor.fm
00:00 - Introduction and Heikki's background 01:19 - How Heikki got into Postgres 03:17 - First major patch: two-phase commit 04:00 - Governance and decision-making in Postgres 07:00 - Committer consensus and decentralization 09:25 - Attracting new contributors 11:25 - Founding Neon with Nikita Shamgunov 13:01 - Why separation of compute and storage matters 15:00 - Write-ahead log and architectural insights 17:03 - Early days of building Neon 20:00 - Building the control plane and user-facing systems 21:28 - What "serverless Postgres" really means 23:39 - Reducing cold start time from 5s to 700ms 25:05 - Storage architecture and page servers 27:31 - Who uses sleepable databases 28:44 - Multi-tenancy and schema management 31:01 - Role in low-code/AI app generation 33:04 - Branching, time travel, and read replicas 36:56 - Real-time point-in-time query recovery 38:47 - Large customers and scaling in Neon 41:04 - Heikki’s favorite Neon feature: time travel 41:49 - Making Postgres multi-threaded 45:29 - Why it matters for connection scaling 50:50 - The next five years for Postgres and Neon 52:57 - Final thoughts and where to find Heikki
26:29 – Example: processing video, audio, and transcripts in Pixeltable
33:12 – DAG execution and parallelism explained
37:00 – Transactional guarantees in Pixeltable
39:00 – Iterators and chunking data for search
42:26 – Using embeddings and semantic search
47:05 – Updating data and incremental recomputation
50:06 – Thoughts on RAG and hybrid search
53:14 – Real-world use cases and dataset curation
57:00 – Example: labeling food waste on cruise ships
1:02:00 – Labeling workflows and syncing annotations
1:02:41 – Pixeltable’s roadmap and cloud vision
1:07:10 – How to get involved with Pixeltable
1:09:03 – Closing and where to find Marcel
42:52 - Where PgDog draws the orchestration line
44:00 - Vision for PgDog Cloud vs bring-your-own-database
46:47 - Company status: first hire, design partners, and production use
50:45 - How deploys work for customers
52:20 - Importance of building closely with design partners
54:05 - Paid design partnerships and initial deployments
56:23 - Benefit of sitting outside Postgres for compatibility
58:32 - Near-term roadmap and long-term vision
1:01:03 - Where to find Lev online
3:19 - The spreadsheet that started it all
6:17 - Intelligent query parsing and connection pooling
9:46 - Preventing outages with query limits
13:42 - Growing Vitess beyond a connection pooler
16:01 - Choosing Go for Vitess
20:00 - The life of a query in Vitess
23:12 - How sharding worked at YouTube
26:03 - Hiding the keyspace ID from applications
33:02 - How Vitess evolved to hide complexity
36:05 - Founding PlanetScale & maintaining Vitess solo
39:22 - Sabbatical, rediscovering empathy, and volunteering
42:08 - The itch to bring Vitess to Postgres
44:50 - Why Multigres focuses on compatibility and usability
49:00 - The Postgres codebase vs. MySQL codebase
52:06 - Joining Supabase & building the Multigres team
54:20 - Starting Multigres from scratch with lessons from Vitess
57:02 - MVP goals for Multigres
1:01:02 - Integration with Supabase & database branching
1:05:21 - Sugu’s dream for Multigres
1:09:05 - Small teams, hiring, and open positions
1:11:07 - Community response to Multigres announcement
1:12:31 - Where to find Sugu
27:11 - Performance gains with Metal: Latency and IOPS explained
32:03 - Database workloads that truly benefit from Metal
39:10 - The myth of the infinite cloud
41:08 - How PlanetScale plans for capacity
43:02 - Multi-tenant vs. PlanetScale Managed
44:02 - Who should use Metal and when?
46:05 - Pricing trade-offs and when Metal becomes cheaper
Découvrez des podcasts liées à Database School. Explorez des podcasts avec des thèmes, sujets, et formats similaires. Ces similarités sont calculées grâce à des données tangibles, pas d'extrapolations !