Infinite, shareable volume storage with Hunter Leath, Archil CEO
Episode 31
Thursday, January 15, 2026 • Duration 55:13
Hunter Leath, CEO of Archil, explains how they’re building a “universal storage engine” that sits between your apps and S3—making an S3 bucket behave like a fast, POSIX-compatible disk for containers, servers, and even Lambda. Along the way, we dig into how their SSD-backed clusters and custom protocol avoid the usual small-file pain and where this approach shines (and where it doesn’t).
Follow Aaron: Twitter/X: https://twitter.com/aarondfrancis Database School: https://databaseschool.com Database School YouTube Channel: https://www.youtube.com/@UCT3XN4RtcFhmrWl8tf_o49g (Subscribe today) LinkedIn: https://www.linkedin.com/in/aarondfrancis Website: https://aaronfrancis.com - find articles, podcasts, courses, and more.
Chapters: 00:00 - Intro: Archil Data and “S3 as a disk” 01:05 - Hunter’s background and the core pitch 02:32 - The real problem: state management (S3 vs block storage) 05:02 - SQLite on S3: what the stack looks like 07:13 - The missing layer: durable SSD-backed clusters 10:14 - Who uses this: unstructured data, CI/CD, Git, agents 12:15 - Small files + Git performance and avoiding S3 request explosion 16:22 - Why they built a new protocol (NFS vs Luster) 20:00 - What gets written to S3: real files in your bucket 22:29 - S3 limits, throttling, and the “keep it on SSD” escape hatch 25:32 - Multi-cloud + R2, and why regions/latency matter 32:10 - Pricing model: “pay only when data is active” 34:41 - Tradeoffs: random reads and ultra-low-latency metal 37:19 - Storage/compute separation and AI/agent-native workflows 43:21 - YC timeline + the marketing challenge of a “universal layer” 47:34 - Single-tenant clusters for enterprises and why it’s hard 50:27 - Where the company is now, hiring, and how to try it (disk.new)
Building search for AI systems with Chroma CTO Hammad Bashir
Episode 30
Thursday, December 18, 2025 • Duration 01:06:43
Hammad Bashir, CTO of Chroma, joins the show to break down how modern vector search systems are actually built from local, embedded databases to massively distributed, object-storage-backed architectures. We dig into Chroma’s shared local-to-cloud API, log-structured storage on object stores, hybrid search, and why retrieval-augmented generation (RAG) isn’t going anywhere.
Follow Aaron: Twitter/X: https://twitter.com/aarondfrancis Database School: https://databaseschool.com Database School YouTube Channel: https://www.youtube.com/@UCT3XN4RtcFhmrWl8tf_o49g (Subscribe today) LinkedIn: https://www.linkedin.com/in/aarondfrancis Website: https://aaronfrancis.com - find articles, podcasts, courses, and more.
Chapters: 00:00 – Introduction From high-school ASICs to CTO of Chroma 01:04 – Hammad’s background and why vector search stuck 03:01 – Why Chroma has one API for local and distributed systems 05:37 – Local experimentation vs production AI workflows 08:03 – What “unprincipled data” means in machine learning 10:31 – From computer vision to retrieval for LLMs 13:00 – Exploratory data analysis and why looking at data still matters 16:38 – Promoting data from local to Chroma Cloud 19:26 – Why Chroma is built on object storage 20:27 – Write-ahead logs, batching, and durability 26:56 – Compaction, inverted indexes, and storage layout 29:26 – Strong consistency and reading from the log 34:12 – How queries are routed and executed 37:00 – Hybrid search: vectors, full-text, and metadata 41:03 – Chunking, embeddings, and retrieval boundaries 43:22 – Agentic search and letting models drive retrieval 45:01 – Is RAG dead? A grounded explanation 48:24 – Why context windows don’t replace search 56:20 – Context rot and why retrieval reduces confusion01:00:19 – Faster models and the future of search stacks01:02:25 – Who Chroma is for and when it’s a great fit01:04:25 – Hiring, team culture, and where to follow Chroma
The database for all your AI needs
Episode 21
Tuesday, September 16, 2025 • Duration 01:00:07
Marcel Kornacker, the creator of Apache Impala and co-creator of Apache Parquet, joins me to talk about his latest project: Pixeltable, a multimodal AI database that combines structured and unstructured data with rich, Python-native workflows.
From ingestion to vector search, transcription to snapshots, Pixeltable eliminates painful data plumbing for modern AI teams.
Website: https://aaronfrancis.com – find articles, podcasts, courses, and more
Database School: https://databaseschool.com
Chapters
0:00 – Introduction
0:20 – Meet Marcel Kornacker
1:19 – Early career and grad school in databases
2:12 – Joining Google and building F1
3:42 – How F1 used Spanner at Google
4:01 – Starting Apache Impala at Cloudera
6:02 – Why SQL still matters
7:29 – What keeps Marcel fascinated with databases
9:37 – The “SQL is dead” waves and shift to AI
10:21 – Observing pain points in computer vision pipelines
13:02 – Multimodal data challenges and the idea for Pixeltable
16:10 – How Pixeltable handles transformations with computed columns
Sharding Postgres without extensions with PgDog founder, Lev Kokotov
Episode 20
Tuesday, August 19, 2025 • Duration 48:53
I chat with Lev Kokotov to talk about building PgDog, an open-source sharding solution for Postgres that sits outside the database. Lev shares the journey from creating PgCat to launching PgDog through YC, the technical challenges of sharding, and why he believes scaling Postgres shouldn’t require extensions or rewrites.
Follow Aaron: Twitter: https://twitter.com/aarondfrancis LinkedIn: https://www.linkedin.com/in/aarondfrancis Website: https://aaronfrancis.com - find articles, podcasts, courses, and more.
Database school: https://databaseschool.com
Chapters: 00:00 - Intro to guest Glauber Costa 00:58 - Glauber's background and path to databases 02:23 - Moving to Texas and life changes 05:32 - The origin story of Turso 07:55 - Why fork SQLite in the first place? 10:28 - SQLite’s closed contribution model 12:00 - Launching libSQL as an open contribution fork 13:43 - Building Turso Cloud for serverless SQLite 14:57 - Limitations of forking SQLite 17:00 - Deciding to rewrite SQLite from scratch 19:08 - Branding mistakes and naming decisions 22:29 - Differentiating Turso (the database) from Turso Cloud 24:00 - Technical barriers that led to the rewrite 28:00 - Why libSQL plateaued for deeper improvements 30:14 - Big business partner request leads to deeper rethink 31:23 - The rewrite begins 33:36 - Early community traction and GitHub stars 35:00 - Hiring contributors from the community 36:58 - Reigniting the original vision 39:40 - Turso’s core business thesis 42:00 - Fully pivoting the company around the rewrite 45:16 - How GitHub contributors signal business alignment 47:10 - SQLite’s rock-solid rep and test suite challenges 49:00 - The magic of deterministic simulation testing 53:00 - How the simulator injects and replays IO failures 56:00 - The role of property-based testing 58:54 - Offering cash for bugs that break data integrity 1:01:05 - Deterministic testing vs traditional testing 1:03:44 - What it took to release Turso Alpha 1:05:50 - Encouraging contributors with real incentives 1:07:50 - How to get involved and contribute 1:20:00 - Upcoming roadmap: indexes, CDC, schema changes 1:23:40 - Final thoughts and where to find Turso
Vitess for Postgres, with the co-founder of PlanetScale
Episode 17
Tuesday, July 1, 2025 • Duration 01:07:29
Sugu Sougoumarane, co-creator of Vitess and co-founder of PlanetScale, joins me to talk about his time scaling YouTube’s database infrastructure, building Vitess, and his latest project bringing sharding to Postgres with Multigres.
This was a fun conversation with technical deep-dives, lessons from building distributed systems, and why he’s joining Supabase to tackle this next big challenge.
Website: https://aaronfrancis.com — find articles, podcasts, courses, and more.
Chapters:
00:00 - Intro: What is PlanetScale Metal?
00:39 - Meet Richard Crowley
01:33 - What is Vitess and how does it work?
03:00 - Where PlanetScale fits into the picture
09:03 - Why EBS is the default and its trade-offs
13:03 - How PlanetScale handles durability without EBS
16:03 - The engineering work behind PlanetScale Metal
22:00 - Deep dive into backups, restores, and availability math
25:03 - How PlanetScale replaces instances safely
From Prisma Founder to LiveStore: Building local-first apps with Johannes Schickling
Episode 15
Thursday, May 29, 2025 • Duration 01:31:40
Johannes Schickling, original founder of Prisma, joins me to talk about LiveStore, his ambitious local-first data layer designed to rethink how we build apps from the data layer up.
We dive deep into event sourcing, syncing with SQLite, and why this approach might power the next generation of reactive apps.
Website: https://aaronfrancis.com — find articles, podcasts, courses, and more
How Durable Objects and D1 Work: A Deep Dive with Cloudflare’s Josh Howard
Episode 14
Wednesday, May 14, 2025 • Duration 01:14:40
Josh Howard, Senior Engineering Manager at Cloudflare, joins me to explain how Durable Objects and D1 work under the hood—and why Cloudflare’s approach to stateful serverless infrastructure is so unique. We get into V8 isolates, replication models, routing strategies, and even upcoming support for containers.
Want to learn more about SQLite? Check out my SQLite course: https://highperformancesqlite.com/?ref=podcast
Follow Aaron: Twitter: https://twitter.com/aarondfrancis LinkedIn: https://www.linkedin.com/in/aarondfrancis Website: https://aaronfrancis.com - find articles, podcasts, courses, and more.
Database school on YouTube: https://www.youtube.com/playlist?list=PLI72dgeNJtzqElnNB6sQoAn2R-F3Vqm15 Database school audio only: https://databaseschool.transistor.fm
Chapters 00:00 - Intro 00:37 - What is a Durable Object? 01:43 - Cloudflare’s serverless model and V8 isolates 03:58 - Why stateful serverless matters 05:14 - Durable Objects vs Workers 06:22 - How routing to Durable Objects works 08:01 - What makes them "durable"? 08:51 - Tradeoffs of colocating compute and state 10:58 - Stateless Durable Objects 12:49 - Waking up from sleep and restoring state 16:15 - Durable Object storage: KV and SQLite APIs 18:49 - Relationship between D1, Workers KV, and DOs 20:34 - Performance of local storage writes 21:50 - Storage replication and output gating 24:15 - Lifecycle of a request through a Durable Object 26:46 - Replication strategy and long-term durability 31:25 - Placement logic and sharding strategy 36:35 - Use cases: agents, multiplayer games, chat apps 40:33 - Scaling Durable Objects 41:14 - Globally unique ID generation 43:22 - Named Durable Objects and coordination 46:07 - D1 vs Workers KV vs Durable Objects 47:50 - Outerbase acquisition and DX improvements 49:49 - Querying durable object storage 51:20 - Developer Week highlights and new features 52:44 - Read replicas and sticky sessions 53:49 - Containers and the future of routing 56:47 - Deployment regions and infrastructure expansion 57:43 - Hiring and how to connect with Josh
20 years of hacking Postgres with Heikki Linnakangas (cofounder of Neon)
Episode 13
Tuesday, May 6, 2025 • Duration 02:00:11
In this episode of Database School, I talk with Heikki Linnakangas, co-founder of Neon and longtime PostgreSQL hacker, to talk about 20+ years in the Postgres community, the architecture behind Neon, and the future of multi-threaded Postgres. From paternity leave patches to branching production databases, we cover a lot of ground in this deep-dive conversation.
Links: Let's make postgres multi-threaded: https://www.postgresql.org/message-id/31cc6df9-53fe-3cd9-af5b-ac0d801163f4%40iki.fi Hacker News discussion: https://news.ycombinator.com/item?id=36284487
Follow Aaron: Twitter: https://twitter.com/aarondfrancis LinkedIn: https://www.linkedin.com/in/aarondfrancis Website: https://aaronfrancis.com - find articles, podcasts, courses, and more.
Database school on YouTube: https://www.youtube.com/playlist?list=PLI72dgeNJtzqElnNB6sQoAn2R-F3Vqm15 Database school audio only: https://databaseschool.transistor.fm
00:00 - Introduction and Heikki's background 01:19 - How Heikki got into Postgres 03:17 - First major patch: two-phase commit 04:00 - Governance and decision-making in Postgres 07:00 - Committer consensus and decentralization 09:25 - Attracting new contributors 11:25 - Founding Neon with Nikita Shamgunov 13:01 - Why separation of compute and storage matters 15:00 - Write-ahead log and architectural insights 17:03 - Early days of building Neon 20:00 - Building the control plane and user-facing systems 21:28 - What "serverless Postgres" really means 23:39 - Reducing cold start time from 5s to 700ms 25:05 - Storage architecture and page servers 27:31 - Who uses sleepable databases 28:44 - Multi-tenancy and schema management 31:01 - Role in low-code/AI app generation 33:04 - Branching, time travel, and read replicas 36:56 - Real-time point-in-time query recovery 38:47 - Large customers and scaling in Neon 41:04 - Heikki’s favorite Neon feature: time travel 41:49 - Making Postgres multi-threaded 45:29 - Why it matters for connection scaling 50:50 - The next five years for Postgres and Neon 52:57 - Final thoughts and where to find Heikki
Related Shows Based on Content Similarities
Discover shows related to Database School, based on actual content similarities. Explore podcasts with similar topics, themes, and formats, backed by real data.