Machine learning audio course, teaching the fundamentals of machine learning and artificial intelligence. It covers intuition, models (shallow and deep), math, languages, frameworks, etc. Where your other ML resources provide the trees, I provide the forest. Consider MLG your syllabus, with highly-curated resources for each episode's details at ocdevel.com. Audio is a great supplement during exercise, commute, chores, etc.
Site
RSS
Apple
Données mises à jour le 05/10/2026
Classements récents
Dernières positions dans les classements Apple Podcasts et Spotify.
Liens partagés entre épisodes et podcasts
Liens présents dans les descriptions d'épisodes et autres podcasts les utilisant également.
Machine Learning Guide, podcast de OCDevel - Statistiques, épisodes et classements - My Podcast Data
Qualité du flux RSS
À améliorer
Score global : 68%
MLA 030 AI and Programming Jobs: What Happened and How to Position
Saison 1 · Épisode 66
jeudi 26 février 2026 • Durée 36:11
The aggregate job market held, the entry-level door narrowed, and software postings sit a quarter below pre-pandemic. Why cheap implementation made specification, verification and domain scarce, how ML roles split five ways, and how to position.
Two different claims hide inside "AI is taking programming jobs": displacement (the role disappears, nobody is rehired) and task change (the role stays, the work shifts). They appear in different data. Displacement shows in unemployment and layoff reports; task change shows in what postings ask for and how teams are shaped. The aggregate evidence is mostly task change with one real pocket of displacement at the entry level, which sets up offensive advice for most listeners and defensive advice for new entrants.
The evidence
Aggregate. The BLS Employment Situation has unemployment at 4.1% with payrolls beating forecasts, far from the 10-20% Dario Amodei floated in the Axios "white-collar bloodbath" interview. Both he and Sam Altman have since softened the timeline; Altman said he was "delighted to be wrong" (Fortune, Time). The Yale Budget Lab tracker finds no discernible disruption. Goldman Sachs Research estimates a net drag of about 16,000 jobs a month across 800+ occupations, with a long-run baseline of 6-7% of workers displaced.
Entry level. Stanford's Canaries in the Coal Mine (August 2026 paper, dashboard) puts 22-25 year olds in AI-exposed occupations 19% behind less-exposed peers, up from 15% a year earlier, driven by reduced hiring rather than separations and concentrated in automation-style exposure. The authors call these descriptive indicators, not causal estimates. Their software-developer case study finds young-developer pay grew somewhat faster than older developers' after ChatGPT, consistent with firms hiring fewer but better-paid juniors; the CPS sample is too small for a software-specific employment percentage. The EIG entry-level working paper is the main counterweight.
Software demand. Indeed's software development postings index (Feb 2020 = 100) sits near 75, roughly a quarter below pre-pandemic and still drifting down, against a much smaller decline in total postings. Confounds stacked on top of AI: rate hikes, the Section 174 expensing change, and the 2021 overhire. SignalFire's State of Talent has new grads at 7% of Big Tech hires, down 25% from 2023 and over 50% from 2019, so half the collapse predates ChatGPT.
Layoffs. Challenger, Gray & Christmas counts 116,175 of 529,914 announced 2026 cuts through August as AI-attributed (about 22%, already more than double all of 2025); AI led every month from March to July, then fell to 3,462 in August, while year-to-date cuts are down 41%.
Grads and incumbents. The NY Fed college labor market data (2026:Q2) has computer science at 7.0% unemployment and 19.1% underemployment and computer engineering at 7.8% and 15.8%, against 5.6% and 42% for all recent graduates: worst on getting a job, among the best on getting a good one. CompTIA's tech jobs report has tech occupation unemployment at 2.8% and over 320,000 active postings asking for AI-related capabilities. The reversal wave: CNBC and Forbes on employers rehiring after AI cuts, Forrester's 55% regret figure, Robert Half's one-in-three refill figure, and the Klarna and IBM cases.
Measurement. METR's randomized trial of 16 experienced open-source developers on 246 real issues found AI made them 19% slower while they believed it sped them up 20%, so self-reported productivity is unreliable in both directions.
The mechanism
When implementation cost falls toward zero, value moves to specification, verification and domain knowledge. The BLS programmer-vs-developer split is that thesis in two rows. Andrew Ng's AI Rewards Generalists Who Can Build New Skills and his five-part AI Engineering Skills Map argue the bottleneck moved from how to build to what to build; David Autor calls AI a supplement to workers with judgment and domain knowledge. Juniors are hit because the traditional junior role was the commoditized part, and it was also the tuition for learning the other two skills. The ceiling on the mechanism shows in two benchmarks from the same year: OpenAI's GDPval, where the newest models win or tie against experts on most one-shot deliverables (with caveats about automated grading), against Scale's Remote Labor Index, where the best agent completed about 4% of real multi-day projects (via Carnegie). Agents produce artifacts; humans still run projects.
The ML career in 2026
Data scientist demand is projected to grow three times as fast as developer demand, and Levels.fyi puts ML/AI-focused engineers in the US around $248k average total compensation. The title splintered into five roles, roughly by headcount: AI engineer (application layer: retrieval, tool use, agent loops, evals, context design); forward-deployed engineer (the Palantir-origin role the labs adopted, where domain is the constraint; see Anthropic's FDE posting); evals and AI quality (titles like Research Engineer, Model Evaluations on Anthropic's jobs board); inference, serving and platform infrastructure; and research engineer or scientist, the smallest and most competitive tier. "AI engineer" now means treating the model as a component with a failure distribution and designing the system around it. Prompt engineering as a standalone title, fine-tuning as a default move, and train-from-scratch generalist ML roles lost ground.
Three camps
Accelerationists (Amodei, Altman, Mustafa Suleyman): disruption within one to five years, entry-level first; strongest evidence is the benchmark curve. The aggregate prediction has failed so far and both leading voices softened it; Suleyman's 12-18 month clock has not expired.
Skeptics (Yann LeCun, who left Meta to found AMI Labs on a world-model thesis; Gary Marcus; Daron Acemoglu, whose macro estimate is under 1% TFP gain over a decade): strongest evidence is the Remote Labor Index; weakest point is the cheap-but-imperfect case that reshapes jobs without replacing them.
Pragmatists (Andrew Ng, Brynjolfsson, Autor): technology real, effects uneven, the question is which tasks move. Best track record so far because they predicted least. The Carnegie Endowment's three views cuts the map differently and is worth reading alongside. The Anthropic Economic Index (January, March) shows augmentation edging up on consumer chat while API and coding-agent usage stays automation-dominant.
Positioning
Own a domain where correct answers require knowledge not on the internet. Own verification: reading diffs fast, writing the test before the bug, building the eval harness, catching reward-hacked tests. Run agents fluently and measure yourself rather than trusting the feeling (the METR gap). Ship agentic work in public with specs, tests, evals and review trail visible, the new portfolio. New entrants: don't look like the traditional junior; compete for the well-paid junior seats that remain, at companies with real domains, in the roles that are hiring.
Learning path
Fundamentals first, because you can't verify what you don't understand: the Machine Learning Guide core episodes. Then the applied layer, which changes every few months: the vibe coding trio, the agents pair, the media trio. Then a domain and a project, which no course provides.
Every show Gnothi has produced, on AI, coding agents, video generation and agentic business, is at ocdevel.com/moremlg.
MLA 029 OpenClaw and Personal Agents
Saison 1 · Épisode 65
dimanche 22 février 2026 • Durée 36:50
OpenClaw as the worked example of the always-on personal agent: gateway, markdown memory, heartbeats, skills, coding agents from your phone, hosted vs local models, and the 2026 security record (exposed instances, two critical CVEs, ClawHavoc) with the posture that makes it survivable.
More OCDevel shows - this one has siblings, each on its own subject and produced the same way
Second and last episode of the agents pair. AI Agents in 2026 covered the theory: loops, tools, memory, protocols, SDKs, evaluation. This one takes a single category, the always-on personal agent, and its most-cloned instance, OpenClaw, all the way down to the security posture required before it touches your inbox.
What a personal agent is
A personal agent is the agent loop with three additions: it runs continuously on a machine you control and can wake itself; its interface is messaging (WhatsApp, Telegram, Signal, iMessage, Slack) rather than a chat tab; and it holds your files, shell, browser, calendar and, if you allow it, email. Messaging removes the gap between having a thought and delegating it, and a chat thread is a natural home for asynchronous work that reports back later. The access that makes it useful is also the entire security problem.
OpenClaw today
OpenClaw is an MIT-licensed, self-hosted TypeScript agent created by Peter Steinberger in late 2025. It was released as Warelay, passed through claw-themed names, hit an Anthropic trademark complaint, and settled on OpenClaw at the end of January 2026 per Wikipedia; the org remains openclaw/openclaw. Steinberger joined OpenAI in February 2026 and stewardship moved to the , a US 501(c)(3) chaired by Dave Morin with a small full-time staff and donors including OpenAI, GitHub, Nvidia and Microsoft; the repo's README says OpenAI is a donor, not an owner, and there is no paid tier or hosted service. The project ships near-weekly CalVer plus an extended-stable line, and the called it the fastest-growing project in the site's history. landed at the end of August.
Architecture: gateway, workspace files, heartbeats, skills The model behind it: hosted versus local Integrations: coding agents from your phone, email, calendar Security: what happened in 2026 and what to do about it Use cases that survived Alternatives Related episodes
MLA 028 AI Agents: Loops, Tools, Memory, Protocols, and Evaluation
Saison 1 · Épisode 64
dimanche 22 février 2026 • Durée 33:47
What an AI agent actually is, why coding agents got good first, how memory really works, what MCP and A2A standardize, which SDKs are alive, how to evaluate on trajectories, and where the products stand after browser agents contracted.
More OCDevel shows - this one has siblings, each on its own subject and produced the same way
First of two episodes on AI agents. This one is the architecture: the loop, tools and verifiable feedback, memory, the protocols (MCP, A2A, computer use), the SDK landscape, evaluation and observability, the product map, and when multiple agents help. The next episode, OpenClaw and the Personal Agent, applies it to one always-on assistant with security as the centerpiece. Coding-agent products and mechanics live in the vibe coding sequence starting at MLA 22.
Agent vs workflow vs chat: the loop
A chat model returns a message; a workflow is your code calling a model at fixed steps; an agent is a model that owns the control flow, choosing its next action from what it observes. That puts systems on a spectrum (chat, chat plus tools, workflows, agents) rather than in a binary, the framing Anthropic's Building Effective Agents uses. The loop itself is ReAct (Yao et al.): thought, action, observation, repeat, with the reasoning trace letting the model track and update a plan. What changed by 2026 is not the loop but the infrastructure around it, and every part of that infrastructure is an attack on per-step error compounding.
Tools, function calling, and verifiable feedback
Function calling: you describe tools as schemas, the model emits a structured call, your code executes it and returns the observation. The model never runs anything itself, which is the security model. gives the practical rules: few high-impact tools, clear namespaces, meaningful identifiers, token-efficient responses, descriptions treated as prompt engineering. adds the overlap test: if a human cannot say which tool applies, neither can the agent. The central principle: coding agents got good first because tests and compilers give verifiable feedback that catches a bad step inside the same loop that made it. Find or manufacture the verifier before writing the prompt.
Memory: context, retrieval, files, episodic Protocols: MCP, A2A, computer use Building one: the SDKs Evaluation and observability Products Multi-agent: when it helps Related episodes
MLA 027 The AI Media Pipeline: Voice, Music, ComfyUI, APIs, and Finishing
Saison 1 · Épisode 63
lundi 14 juillet 2025 • Durée 35:08
How to automate AI media end to end: clone your own voice on open TTS, pick music that's actually licensed, run ComfyUI graphs headless, design around fal, Replicate and provider queues, finish with ffmpeg, and stay inside licensing at every layer.
Links
More OCDevel shows - this one has siblings, each on its own subject and produced the same way
Companion show. This episode is the overview of the media pipeline. For weekly, hands-on coverage of the video half, from a first usable clip to scenes that cut together, listen to AI Video Generation on Gnothi.
Once you need thirty clips with the same character, a narrator who sounds identical every episode, a matched music bed, word-accurate captions and a platform-safe export, the prompt is one node in a graph and the graph is the product. The engineering lives in the edges: how one model's output becomes the next model's input, how failures retry, what each job cost, and whether the run is reproducible next week. Model choice at the nodes is covered in the two sibling episodes; this one covers everything else: voice with ElevenLabs and Qwen3-TTS, licensed music, ComfyUI on your own card, the fal and Replicate APIs, ffmpeg assembly, and licensing.
Voice: cloning, open TTS, and consent
ElevenLabs remains the reference point (current flagship Eleven v3, character-based pricing). Professional voice cloning is locked to the requester's own voice behind a live voice check, and the terms require consent attestation for any uploaded voice. On the open side, Breeze TTS 2 topped the open-weights column of the Artificial Analysis speech arena in August 2026, ahead of Fish Audio's S2 Pro; the code is Apache 2.0 but the weights are research/non-commercial, so it is not a commercial self-hosting option. The working set for programmers: (Apache 2.0, 0.6B/1.7B, cloning from seconds of reference audio; the preset-speaker variant does not clone, see my ), (MIT, emotion control, watermarked output), and (82M parameters, Apache 2.0, faster than real time on CPU). Fish Audio plays both sides with open weights and a cheap hosted API. Quantized Qwen3-TTS runs podcast-length synthesis on CPU-only instances; see and the broader . Hosted alternatives for prototyping: and .
Music and sound effects ComfyUI and local generation APIs and aggregators Assembly and finishing Licensing across the stack Two pipelines Related episodes
MLA 026 AI Video Generation 2026: Veo, Gemini, Kling, Runway, MiniMax, Sora
Saison 1 · Épisode 62
samedi 12 juillet 2025 • Durée 32:09
Sora is shut down, Google runs two video models, Kling 3 does lip-synced dialogue, and open-weight MiniMax H3 is what you can actually fine-tune. What a usable clip costs, which models do native audio, how reference consistency works, and why the unit of work is the shot.
More OCDevel shows - this one has siblings, each on its own subject and produced the same way.
Companion show: for weekly, hands-on coverage of the AI video pipeline, from a first usable clip to scenes that cut together, listen to AI Video Generation.
Second of three episodes on AI media generation, covering Veo, Kling, Runway and MiniMax H3. Four questions: what a usable clip costs, which models generate sound and dialogue natively, how character and shot consistency work now, and where open-weight video fits for a programmer. Ends with the shot-to-scene mental model.
What changed since 2025
Native audio is now the baseline at the frontier: Veo 3.1, Kling 3.0, MiniMax H3 and LTX-2.5 sample audio and frames from one model, so lip movement and sound effects land on the right frame. Clips grew from four or five seconds to eight to fifteen, with a few models advertising thirty. Every serious product ships a reference-conditioning feature (Google "ingredients", Kling "elements", Runway references) that holds a character or object across clips. Image-to-video, not text-to-video, is the professional path: lock the first frame with an image model (see AI Image Generation and Editing), then ask the video model to move it. The old "storyteller vs animator" split resolved in favor of the animators.
Google: Veo 3.1 and Gemini Omni Sora: shut down Kling 3.0 Runway Gen-4.5 and Aleph The leaderboard vs the products Open weights and the second tier Consistency and control Audio in video Cost per usable second Shot to scene Related episodes
Editing replaced generation as the core task. How GPT Image 2.5, Google's Nano Banana line, Midjourney V8.2 and Flux 2 differ, what open weights and LoRAs buy you, ControlNet vs instruction editing, and how licensing and C2PA provenance work now.
Links
More OCDevel shows - this one has siblings, each on its own subject and produced the same way
Companion show. This episode is the overview. For weekly, hands-on coverage of the full image and video pipeline, from a first usable clip to scenes that cut together, listen to AI Video Generation.
First of three episodes on AI media generation (this one is images and editing; then video, then the pipeline). A decision guide rather than a leaderboard: where GPT Image, Nano Banana, Midjourney and Flux each fit in late 2026, why instruction-driven editing replaced generation as the core task, what open weights buy you, and how licensing and provenance work.
What changed: editing, instruction-following, references, text
The 2025 "artist vs collaborator" split is over and the collaborators won. Every frontier image model now sits behind a language model that reads the prompt with world knowledge and accepts images as input, so the unit of work became "here is an image, change this one thing and keep everything else." Text rendering, precise instruction following and identity-preserving reference images all landed at once for one reason: the image model became, or was paired with, a multimodal language model. Generation from scratch is now the special case where the input image is empty.
OpenAI: GPT Image 2.5
The lineage runs gpt-image-1 (2025), gpt-image-2, then in September 2026, with a precision variant (Sunburst) and a fast default (Flare); the same models power ChatGPT and the . As of this recording it holds the top slots on both the and boards and on . For editing it does mask inpainting, up to four reference images, and multi-turn editing via the Responses API. Its weaknesses are latency at high quality and occasional text and consistency slips. Billing is per token (text in, image in, image out) with a cached-input discount; see the .
Google: the Nano Banana lineage Midjourney V8.2 and the Edit Model Flux: the open-weight default and its license tiers Special mentions Control: three problems, one commoditized Licensing, provenance, C2PA Choose by job Related episodes
MLG 036 Autoencoders
Saison 1 · Épisode 60
vendredi 30 mai 2025 • Durée 01:05:55
Auto encoders are neural networks that compress data into a smaller "code," enabling dimensionality reduction, data cleaning, and lossy compression by reconstructing original inputs from this code. Advanced auto encoder types, such as denoising, sparse, and variational auto encoders, extend these concepts for applications in generative modeling, interpretability, and synthetic data generation.
Autoencoders are neural networks designed to reconstruct their input data by passing data through a compressed intermediate representation called a "code."
The architecture typically follows an hourglass shape: a wide input and output separated by a narrower bottleneck layer that enforces information compression.
The encoder compresses input data into the code, while the decoder reconstructs the original input from this code.
Comparison with Supervised Learning
Unlike traditional supervised learning, where the output differs from the input (e.g., image classification), autoencoders use the same vector for both input and output.
Use Cases: Dimensionality Reduction and Representation
Autoencoders perform dimensionality reduction by learning compressed forms of high-dimensional data, making it easier to visualize and process data with many features.
The compressed code can be used for clustering, visualization in 2D or 3D graphs, and input into subsequent machine learning models, saving computational resources and improving scalability.
Feature Learning and Embeddings Data Search, Clustering, and Compression Reconstruction Fidelity and Loss Types Outlier Detection and Noise Reduction Denoising Autoencoders Data Imputation Cryptographic Analogy Advanced Architectures: Sparse and Overcomplete Autoencoders Interpretability and Research Example Variational Autoencoders (VAEs) VAEs for Synthetic Data and Rare Event Amplification Conditional Generative Techniques Practical Considerations and Limitations
MLG 035 Large Language Models 2
Saison 1 · Épisode 59
jeudi 8 mai 2025 • Durée 45:25
At inference, large language models use in-context learning with zero-, one-, or few-shot examples to perform new tasks without weight updates, and can be grounded with Retrieval Augmented Generation (RAG) by embedding documents into vector databases for real-time factual lookup using cosine similarity. LLM agents autonomously plan, act, and use external tools via orchestrated loops with persistent memory, while recent benchmarks like GPQA (STEM reasoning), SWE Bench (agentic coding), and MMMU (multimodal college-level tasks) test performance alongside prompt engineering techniques such as chain-of-thought reasoning, structured few-shot prompts, positive instruction framing, and iterative self-correction.
Definition: LLMs can perform tasks by learning from examples provided directly in the prompt without updating their parameters.
Types:
Zero-shot: Direct query, no examples provided.
One-shot: Single example provided.
Few-shot: Multiple examples, balancing quantity with context window limitations.
Mechanism: ICL works through analogy and Bayesian inference, using examples as semantic priors to activate relevant internal representations.
Emergent Properties: ICL is an "inference-time training" approach, leveraging the model's pre-trained knowledge without gradient updates; its effectiveness can be enhanced with diverse, non-redundant examples.
Retrieval Augmented Generation (RAG) and Grounding
LLM Agents Multimodal Large Language Models (MLLMs) Advanced LLM Architectures and Training Directions Evaluation Benchmarks (as of 2025) Prompt Engineering: High-Impact Techniques Trends and Research Outlook
MLG 034 Large Language Models 1
Saison 1 · Épisode 58
mercredi 7 mai 2025 • Durée 50:48
Explains language models (LLMs) advancements. Scaling laws - the relationships among model size, data size, and compute - and how emergent abilities such as in-context learning, multi-step reasoning, and instruction following arise once certain scaling thresholds are crossed. The evolution of the transformer architecture with Mixture of Experts (MoE), describes the three-phase training process culminating in Reinforcement Learning from Human Feedback (RLHF) for model alignment, and explores advanced reasoning techniques such as chain-of-thought prompting which significantly improve complex task performance.
Transformers: Introduced by the 2017 "Attention is All You Need" paper, transformers allow for parallel training and inference of sequences using self-attention, in contrast to the sequential nature of RNNs.
Scaling Laws:
Empirical research revealed that LLM performance improves predictably as model size (parameters), data size (training tokens), and compute are increased together, with diminishing returns if only one variable is scaled disproportionately.
The "Chinchilla scaling law" (DeepMind, 2022) established the optimal model/data/compute ratio for efficient model performance: earlier large models like GPT-3 were undertrained relative to their size, whereas right-sized models with more training data (e.g., Chinchilla, LLaMA series) proved more compute and inference efficient.
Emergent Abilities in LLMs
Emergence: When trained beyond a certain scale, LLMs display abilities not present in smaller models, including:
Architectural Evolutions: Mixture of Experts (MoE) The Three-Phase Training Process Advanced Reasoning Techniques Optimization for Training and Inference
MLA 024 Agentic Software Engineering: Specs, Verification, and the Review Loop
Saison 1 · Épisode 57
dimanche 13 avril 2025 • Durée 33:27
How working engineers ship with coding agents: issues an agent can verify, plan mode before code, a verification loop with a browser in it, agent review of agent code, worktrees and CI, cost discipline, and where agents still fail.
More OCDevel shows - this one has siblings, each on its own subject and produced the same way
Third and last episode of the vibe-coding sequence. Vibe Coding in 2026 picked an agent; Inside a Coding Agent explained the mechanics. This one is the practice: how working engineers ship real software with Claude Code, Codex, and Antigravity without shipping garbage. Specs, verification loops, agent review, parallel worktrees, headless CI, cost discipline, and the failure modes.
From vibe coding to agentic engineering
Andrej Karpathy coined vibe coding in early 2025 and, a year later, proposed "agentic engineering" for professional work. His older idea of jagged intelligence (models clear hard problems and trip on trivial ones, unpredictably) is why the job is judgment rather than button-pressing. The frame for the episode: when implementation is cheap, the value moves to the two ends of the pipeline, specification (saying exactly what should be true) and verification (proving it). Vibe coding stays fine for throwaways; the rest applies to codebases with users.
Specs: issues as prompts, plan mode in Claude Code, Codex, and Antigravity
The prompt is a spec whether you meant it or not. The unit of work is a tracker issue with four parts: what's wrong, where to look, acceptance criteria a machine or a five-second human check can verify, and an explicit out-of-scope fence. Every serious agent has a read-only planning phase: Claude Code's plan mode (Shift+Tab or /plan), Codex's plan mode on the same keystroke with its own , and Antigravity's . Judge a plan on three things: the files it names, the verification step it commits to, and whether it stays inside the fence. Heavier spec tooling (, ) formalizes requirements, design, and tasks. teach a lighter interview-then-spec-file pattern and say to skip planning when you could describe the diff in one sentence; frame a task as goal, context, constraints, and "done when."
Verification loops: one-command checks, test-first, Playwright MCP and browser agents The review loop: agent PRs, Claude Code and Codex review, human gate Parallelism: git worktrees, cloud sandboxes, task queues Headless and CI: GitHub Actions, label triggers, scheduled runs Cost and context discipline Failure modes: reward-hacked tests, scope creep, prompt injection, secrets One-week adoption plan Related episodes
Rewrites that keep the idea and shrink the surface, per OSS Insight's fork-wave analysis: nanobot (Python, small auditable core), ZeroClaw (Rust, single static binary), PicoClaw (Go, from Sipeed, embedded targets) and NanoClaw (TypeScript, container-first). Hosted versions are all third-party one-click deploys or small managed services; the foundation runs none.
One long-lived Node gateway per host binds to loopback on port 18789, owns every channel connection, routes inbound messages to sessions, loads context, calls the configured model, executes tools, streams the reply and persists everything under ~/.openclaw. Nodes are paired devices (laptop, phone, headless box) that lend the gateway local screen, camera and shell.
Memory is markdown in the agent workspace: AGENTS.md (operating instructions), SOUL.md (persona and boundaries), IDENTITY.md, USER.md (stable facts about you, with its own character budget), MEMORY.md (curated durable facts), a memory/ directory of daily notes, and a one-time BOOTSTRAP.md interview. The memory docs state there is no hidden state; a hybrid memory_search index covers the memory file and daily notes, a background "dreaming" sweep promotes recurring material into MEMORY.md, and a flush runs before context compaction.
Initiative comes from the heartbeat, a periodic main-session turn (30 minutes by default, 60 on subscription auth) that can stay silent via a no-reply marker, and from the automations scheduler (one-shot, interval or cron, delivered to a channel, a webhook or nowhere). HEARTBEAT.md is legacy; its checklist now lives in DB-backed scratch. Skills follow Anthropic's Agent Skills format (SKILL.md with name and description front matter) and install from ClawHub; the docs say to treat third-party skills as untrusted code and read them first. The browser tool drives a dedicated agent-owned Chrome/Brave/Edge profile through a loopback-only control service with a strict SSRF policy.
OpenClaw is a harness over sixty-plus providers using provider/model references, including Ollama, llama.cpp, LM Studio, vLLM and SGLang for local inference. The tradeoff is the one from the agents episode with higher stakes: everything the agent reads is forwarded to the model, so inbox triage on a hosted model sends your inbox to the provider. Frontier models are better at the judgment calls (is this urgent, is this instruction really from me), local is the only defensible choice for regulated third-party data, and a per-task split (local for reading private content, hosted for writing public content) is a common compromise. The docs make no claim about local-model quality inside OpenClaw. Subscription auth reuses an existing Claude CLI login or an OpenAI OAuth flow per the OAuth docs, which describe the Claude CLI path as sanctioned per Anthropic staff guidance and warn that a community proxy needs a terms check; no published Anthropic term was found either way.
The old Claude Code bridge skill is superseded by agent runtimes: built-in, Codex app-server, Claude CLI and a Copilot plugin, with external harnesses (Claude Code, Gemini CLI, OpenCode, Cursor) driven over the Agent Client Protocol through acpx. Tasks land in managed worktrees: isolated branches with checkpoints in a state DB, filesystem snapshots where supported, a cap around 100 live worktrees, and dirty or unpushed work never auto-cleaned. See Agentic Software Engineering for why worktrees are the right isolation unit.
The official IMAP plugin watches a mailbox, spawns an isolated restricted-reader session per allowed message, ranks trust by DMARC/SPF/DKIM, and does not send mail, which encodes the read-versus-write distinction the security section relies on. Calendar and most SaaS arrive as skills or MCP servers; Cisco's DefenseClaw announcement describes connecting email, calendar and Discord through Zapier-hosted MCP servers so a glue service holds the OAuth tokens.
Exposure: Bitsight, SecurityScorecard and Censys counted between 30,000 and 135,000 internet-facing instances in early 2026, roughly two thirds with no authentication; Censys confirmed 63,070 live instances at the end of March. Bugs: CVE-2026-25253 (CVSS 8.8), a one-click RCE where the control UI auto-connected to a gatewayUrl from the query string and leaked the auth token, worked even against loopback-bound instances and was patched in v2026.1.29 per the GitHub advisory; CVE-2026-32922 (CVSS 9.9) let a pairing token rotate itself into admin, fixed in v2026.3.11. Well over a hundred advisories were logged between February and April.
Supply chain: Koi Security's ClawHavoc report found 341 malicious ClawHub skills (335 from one campaign) disguised as wallets, trading bots and Workspace integrations, delivering the Atomic macOS Stealer and targeting always-on Mac minis; the count later passed 800 as the registry grew past 10,000 (The Hacker News, Unit 42). Cisco's skill research scanned about 31,000 agent skills, found a quarter with at least one vulnerability, and demonstrated exfiltration through an attacker-controlled Telegram bot.
Injection: Giskard exploited a live deployment for exfiltration and account takeover; PromptArmor showed Telegram and Discord link previews exfiltrate data with no click; CrowdStrike called a misconfigured instance "a powerful AI backdoor agent." Ambient: Wiz found Moltbook's database open with about 1.5 million API tokens; Meta banned OpenClaw on work devices and then acquired Moltbook; China restricted state use.
Responses: fast patches, loopback default, DM pairing codes for unknown senders, openclaw security audit, the VirusTotal partnership scanning every ClawHub skill with daily rescans, openclaw skills verify, publisher gating, and the layered access model in the security docs: DM modes, per-agent profiles, control-plane tool restrictions, node exec policies, sandbox and read-only variants, exec approvals, strict browser SSRF, and "one trust boundary per gateway." None of it fixes indirect prompt injection; the maintainers say scanning is not a silver bullet, skills remain arbitrary code, and the agent still holds real credentials. Cisco's open-source DefenseClaw adds pre-execution scanning and runtime allow/block enforcement.
Posture, as eight rules: loopback plus a tunnel (Tailscale or SSH) and an auth token even locally; the agent gets its own OS user, mailbox, calendar, browser profile and capped API keys; every integration starts read-only with human approval on irreversible writes; a dedicated box with nothing else on it; read every skill before enabling; design so a hijacked agent is a nuisance, not a breach, by shrinking the write surface; read the logs and memory files weekly; run the audit after every change and pin a version.
By mid-2026 the consumer frenzy had cooled and what remained was solo founders and small teams running it as infrastructure. Surviving uses share one shape, a scheduled or triggered read delivered to the chat you already use: morning briefings, read-only inbox triage with thresholds, watching deadlines and pipelines and even school lunch menus, voice memos returned as structured notes, lead research and CRM updates, and coding-agent dispatch from a phone. Value compounds through the memory files rather than any single automation, which is why the posture has to precede the setup.
Claude Cowork and OpenAI's ChatGPT Work give the delegate-a-task shape sandboxed and session-based, without messaging or always-on. The rewrites give a smaller, auditable surface. The SDKs from AI Agents in 2026 give one tight automation in an afternoon. The real decision is how much of your life you want in one process.
"Memory" means four things: the context window (the only memory the model has), retrieval from an external store, files on disk, and episodic records of prior sessions. Most agent memory is files. Anthropic's memory tool is a client-side file protocol (view, create, replace, insert, delete) against storage you own. The hard part is context management, and both labs converged on the same three mechanisms: context editing to clear stale tool results, compaction to summarize near the limit, and notes written to files before summarization. OpenAI's Responses API conversation state has the same shape with a compaction threshold and compact endpoint. Third-party layers Mem0, Letta (from MemGPT), and Zep (temporal knowledge graph) now compete with first-party primitives. Multi-session patterns: Effective harnesses for long-running agents.
Model Context Protocol is the agent-to-tool standard, now a Linux Foundation project with individual-maintainer governance. 2026 additions: elicitation (server asks the user mid-operation), an extensions mechanism, and the async Tasks extension for long-running tools. Every major SDK below consumes it; its cost is the context each connected server's tool list occupies. A2A is the agent-to-agent standard, Google-built, Linux Foundation-hosted, at v1.0 with a steering committee spanning AWS, Cisco, Google, IBM, Microsoft, Salesforce, SAP and ServiceNow. Strong governance, weak observed consumption; worth knowing, not yet worth building on for small teams. Computer use is the universal fallback: Anthropic's computer use tool (GA toolset with zoom and an automatic injection classifier), Google's Gemini computer use, open-source Browser Use, and Playwright MCP, which drives the accessibility tree instead of screenshots. Prefer API, then accessibility tree, then screenshots.
Both labs advise starting without a framework: Building Effective Agents and OpenAI's A Practical Guide to Building Agents. The 2026 SDKs have converged on that critique as thin harnesses around a loop.
Claude Agent SDK: Claude Code's loop as a library (built-in tools, subagents, hooks, MCP, permissions, compaction); TypeScript and Python; pre-1.0.
LangGraph and LangChain 1.x: stateful graph with checkpointing, interrupts, durable execution; create_agent as a minimal middleware harness. LangSmith is the separate tracing product.
Google ADK: code-first hierarchical agent trees with native A2A; deploys to Vertex Agent Engine.
CrewAI: role-based crews, past 1.0, with a commercial management platform.
smolagents: code agents that write Python instead of JSON calls; weakest maintenance signal on the list.
Pydantic AI and Vercel AI SDK: typed validation-first agents in Python; loop control and agent abstraction in TypeScript.
Decision rule: machine-operating agent fast, Claude Agent SDK; lightweight handoffs, OpenAI; durable human-in-the-loop state, LangGraph; inside Google or Microsoft, their kit; to understand what you run, write the loop yourself first.
Two essays a day apart: How we built our multi-agent research system (orchestrator plus parallel subagents beat a single agent on research at roughly 15x the tokens; token usage explained most of the variance) and Cognition's Don't Build Multi-Agents (dispersed decisions and unshared context make it fragile). The disagreement is task shape. Parallelize independent, read-mostly work; keep stateful, sequential work single-threaded; prefer a small hierarchy where workers return findings rather than decisions.
Consent is the legal boundary. Tennessee's ELVIS Act added voice to right of publicity; the federal NO FAKES Act cleared Senate Judiciary in June 2026; EU AI Act Article 50 transparency duties apply from 2 August 2026; Denmark is amending copyright law to cover a person's face and voice.
Warner settled with both Suno and Udio; Universal settled with Udio, which became a no-download walled garden; UMG and Sony are still litigating against Suno, whose terms now grant commercial rights rather than ownership to paid subscribers. Eleven Music is trained on licensed data via Merlin and Kobalt deals and cleared for commercial use on self-serve plans, excluding film, TV and larger games; ElevenLabs sound effects are cleared on any paid plan and support loops. Google exposes Lyria and Lyria RealTime through the Gemini API. Open models: ACE-Step 1.5, YuE, Stable Audio Open (community license, best open option for short effects), HeartMuLa, and Meta's MusicGen, which is non-commercial.
ComfyUI is a workflow runtime with a GUI for designing graphs. Comfy raised $30M at a $500M valuation in April 2026 and ships a desktop app, Comfy Cloud, and API nodes that call paid providers from inside a local graph. Programmatic use is the same /prompt endpoint and websocket the front end uses: export the workflow in API format, patch fields, post, poll. Wrappers like comfyui-api and comfy-pack turn a graph into a scalable service. Alternatives: SwarmUI, InvokeAI, the Krita AI plugin.
Hardware: full-precision Flux.2 and Qwen-Image do not fit consumer cards; fp8 and GGUF quantization bring them to 16 to 24 GB. For video, the open Wan releases lag the API versions; the 5B variant does 720p on 24 GB (about 8 GB with Comfy offloading), the 14B variant officially wants 80 GB at full precision and needs GGUF to be consumer-viable, and Wan2GP targets low-VRAM cards. Rule: local for iteration, cloud for volume.
Start with an aggregator, move to a provider API only for a feature or price it lacks. fal is queue-first: submit, get a request ID, poll or webhook, per-output pricing on popular models, per GPU-second for custom deployments. Replicate has the broader catalog beyond image and video, bills per second of compute for open models, packages custom models with Cog, and joined Cloudflare with the same API. RunPod serverless is the raw GPU option for a custom ComfyUI graph. Provider APIs have converged on the same shape: Veo via the Gemini API (billed per output second, audio included), Kling API (post, store task_id, poll /v1/tasks), Runway API, ElevenLabs API.
Design rules: every generation is a job in a durable queue keyed on a hash of inputs, model and seed; store the provider's request ID next to your job ID; honor 429 retry-after with a token bucket per provider; persist prompt, seed, inputs and outputs in object storage; route to a second provider on 5xx. Cost per usable second is list price times your rejection rate.
ffmpeg is the programmer's editor: concat, overlay, sidechain ducking, caption burn-in, crop to 9:16, loudnorm and export. Human-in-the-loop editors: DaVinci Resolve (free version is a real editor; Studio unlocks most Neural Engine features), Descript with its transcript-as-timeline and Underlord assistant, and CapCut for short-form auto-captions, with a caution about its June 2025 terms change. Upscaling: Topaz retired Video AI for the subscription Topaz Video with the Astra model; open-side, SeedVR2 is single-step, runs on 8 GB and plugs into ComfyUI; Real-ESRGAN for clean stills; RIFE for frame interpolation. Captions: WhisperX gives word timestamps within about 50 ms via forced alignment plus diarization, emitting SRT/VTT; generate styled word-pop overlays from its JSON. Delivery: 1080x1920 9:16, H.264/AAC, roughly 10 to 12 Mbps, 30 fps; YouTube recommended upload settings; target -14 LUFS integrated with a -1 dBTP ceiling as the last pipeline step.
Five layers, and the output is only as clean as the dirtiest node. Weights: Apache/MIT models (Qwen3-TTS, Kokoro, Chatterbox, open Wan) are clean; FLUX.2 dev is non-commercial without a separate license; Stability's community license allows commercial use under a revenue threshold; MusicGen is non-commercial. Output: the Copyright Office holds that prompts alone are not authorship, and the Supreme Court denied cert in Thaler v. Perlmutter in March 2026, so keep evidence of the human selection and editing. Training data: licensed models are the safe path while label suits continue. People: right of publicity, get written consent. Disclosure: YouTube auto-labels via SynthID and C2PA content credentials since May 2026, and labels on Veo and C2PA-stamped content are permanent. Attach credentials and disclose.
Social clip (30 s, 9:16, run 200 times): character sheet from an image-editing model stored with prompt and seed -> templated script -> Qwen3-TTS narration with word timings -> per-shot image-to-video jobs via fal keyed on input hash, webhook completion -> cached licensed music bed -> ffmpeg concat, duck, styled captions from timing JSON, crop, loudnorm, 1080x1920 export -> content credentials and disclosure -> human review queue.
Narrated explainer (8 min, 16:9, weekly): human-written script (where copyright rests) -> chunked TTS stitched with short silences -> LLM shot list with timestamps tagged diagram/image/video -> deterministic diagrams, styled images, a few video clips upscaled with SeedVR2 -> licensed music and SFX generated once -> Resolve or ffmpeg assembly, captions from narration timings, loudnorm, 1080p/4K export -> title, chapters from the shot list, disclosure, credentials. Both are the same graph with different shot counts and aspect ratios.
AI Agents covers agent orchestration of pipelines like these
The companion show for the video half of this pipeline is AI Video Generation on Gnothi.
Google now runs two video models in two places. Veo 3.1 is the developer baseline on the Gemini API and Vertex, in Quality, Fast and Lite tiers, all with native audio; clips are 4, 6 or 8 seconds, and 1080p/4K are upscales of the 8-second clip. It accepts up to three reference images, first-and-last-frame interpolation, and extend in 7-second steps up to 20 times; the Ingredients to Video update added identity consistency, native vertical and 4K upscaling. At Google I/O 2026 Google announced Gemini Omni; Gemini Omni Flash replaced Veo inside the Gemini app and Flow, taking text, image, video or audio as input and supporting conversational video-to-video editing. Per Flow's model matrix it currently tops out at 10 seconds and 720p. All output carries SynthID; the detector portal is still waitlisted.
OpenAI launched Sora 2 on September 30, 2025 with native audio and a free iOS app; it then hit a copyright reckoning over opt-out character use, SAG-AFTRA and Bryan Cranston pushback on likeness, and a court order barring the word "Cameo". In March 2026 OpenAI announced a two-stage shutdown: app and web closed April 26, 2026, API closes September 24, 2026. NBC's reporting attributes it to reallocating compute to coding, reasoning and enterprise; Sora continues only as internal world-model research. It stays in the episode as the case study of a strong model without a business.
Kuaishou's Kling 3.0 launched globally in March 2026 as a unified image, video and audio model: up to 15 seconds per shot, native 4K, and per the Kling Omni audio guide lip-synced dialogue in five languages with sound effects and ambience generated in the same pass. The control surface is the point: an elements library built from images or short reference video with per-element voice binding, multi-shot generation with continuity, motion transfer from a reference video, motion brush, six-axis camera control, extend and retake. Sold as a credit-based consumer app with commercial rights on paid tiers, plus a first-party API fronted in the West by fal and Replicate. fal's three per-second prices (audio off, audio on, voice control) make the cost of joint audio-video sampling visible.
Runway Gen-4.5 shipped December 1, 2025, briefly topped the Artificial Analysis leaderboard, and was candid about causal reasoning, object permanence and "success bias" failures. Runway's differentiator is editing: Aleph is video-to-video (new angles, relighting, add/remove objects, restyle), and per the Runway API changelog Aleph 2.0 takes 2-30 second inputs with up to five keyframes; Act-Two transfers a filmed performance onto a character. Gen-4.5's release notes do not claim native dialogue or effects. The Runway API now also resells ByteDance Seedance 2.5 (30-second clips, large reference budgets, audio) and Wan 3, and studio deals with Adobe, AMC Networks and Lionsgate anchor the enterprise story.
As of this recording the Artificial Analysis text-to-video arena (blind pairwise human preference) has none of Veo 3.1, Kling 3.0 or Gen-4.5 in its top five: Wan 3.0, Gemini Omni Flash, fal's post-trained MiniMax H3 Max, MiniMax H3, then Seedance 2.0. The headline products win on control, distribution and enterprise fit, not the taste test.
Naming matters for Wan: Wan 2.2 is Apache-2.0 open weight (14B MoE needing an 80GB GPU, or a 5B model for a 24GB card) and its GGUF quantizations still trend on Hugging Face; Wan 2.5, 2.7 and 3.0 have no published weights on the Wan-AI Hugging Face org or GitHub and are served as APIs. The open-weight center of gravity is MiniMax H3: 33B, native stereo audio, up to 2K, 4-15 seconds, a community license permitting commercial use, official ComfyUI workflows, and a LoRA and step-distillation ecosystem. LTX-2.5 (19B, native audio, community license free under $10M revenue) and HunyuanVideo-1.5 (8.3B, 14GB with offload, no audio) round out the runnable set; MAGI-2 is a preview with no confirmed license. Elsewhere: Luma shipped Ray3, Ray3 Modify and Ray3.14; Pika pivoted to effects, an agent and MCP on Pika 2.5; ByteDance's Seedance 2.0 and 2.5 ride Dreamina and CapCut distribution; Grok Imagine is a priced API video model outside the top ten; Higgsfield is an aggregator and creative suite, as are fal, Replicate and OpenArt on the developer side.
Every consistency feature is conditioning under a different name: text, reference images, first frame, last frame, and reference video are slots the denoiser attends to. A first frame is the strongest condition, which is why image-to-video wins. Reference characters (Veo's three images, Kling elements with voice binding, Seedance's dozens of references) fight identity drift, still the main failure mode. Start and end frames bound a camera move and let shots hand off to each other. Explicit camera controls beat prompt text. Video-to-video (Aleph, Luma Modify, Omni Flash, Kling) means fixing a nearly right shot instead of regenerating. Extend compounds drift, so use it to finish a shot, not build a scene. Open models add LoRAs: Musubi Tuner trains adapters for HunyuanVideo and Wan 2.x, and the tooling lags each new frontier open release by months.
Native audio means one sampling process produces waveform and frames, conditioned on each other; Veo 3.1 prices everything as video with audio, Kling 3.0 exposes it as a paid toggle, H3 and LTX-2.5 do it in open weights, Gen-4.5 and Wan 2.2 do not. Post-hoc remains a valid choice: MMAudio generates synchronized sound from finished video with an explicit alignment module, and ElevenLabs sound effects generate timed effects from text. Native dialogue holds for a line or two; longer talking heads still favor performance-driven tools like Act-Two. Voice and music proper are in The AI Media Pipeline.
As of this recording, from Gemini API pricing: Veo 3.1 Quality about $0.40/s with audio (720p/1080p), Fast about $0.10/s, Lite about $0.05/s; Gemini Omni Flash is billed per token, working out to roughly $0.10/s of 720p. From Runway API pricing: Gen-4.5 $0.12/s, Aleph 2 $0.28/s with a minimum. From fal: Kling v3 about $0.08/s silent and $0.13/s with audio, Wan 2.5 $0.05/s; Grok Imagine video $0.05-0.08/s. An 8-second Veo Quality shot with audio is a bit over $3; Kling or Gen-4.5 about $1. No vendor publishes success rates; budgeting four generations per usable shot puts a frontier clip with audio at $3-13 and a Fast or open-weight clip under $1. Iterate on the cheap tier, render on the expensive one.
Every model generates a shot: one continuous take, one camera, one action, 4-15 seconds. A scene is three to eight shots cut together, and continuity is your job: same references in every shot, first and last frames handing off, the same elements or LoRA, one audio bed over the cut. Storyboard as shots, lock first frames with an image model, iterate cheap, render expensive, fix with video-to-video, assemble in an editor. Assembly, voice, music, ComfyUI and driving it from code are the next episode.
Nano Banana was gemini-2.5-flash-image (August 2025, now legacy). It was succeeded by Nano Banana Pro (gemini-3-pro-image) and Nano Banana 2 (gemini-3.1-flash-image, February 2026), plus a Lite tier. Per the Gemini image generation docs, Nano Banana 2 accepts up to fourteen reference images (objects plus characters), outputs up to 4K, does multi-turn sequential editing, and can ground generation in Google Search; the Pro model adds style references and identity preservation across up to five subjects. Every output carries a SynthID watermark with no opt-out. Imagen appears to be superseded for new work, though no formal retirement notice was found.
Midjourney moved from V7 (2025) to V8 alpha in March 2026 and V8.2 as the default in July 2026 (version history). V8.2's Edit Model replaces Omni Reference, Character Reference, Retexture and the Editor with one instruction-driven model taking up to four references, the same convergence OpenAI and Google made. Aesthetics and draft-mode ideation remain its strengths. The hard limits: still no official API and terms that bar automation; generations public by default below the Stealth tier; and it is the defendant in Disney Enterprises v. Midjourney (filed June 2025, joined by a separate Warner Bros. Discovery suit), currently in discovery with Midjourney demanding the studios' own AI records.
Black Forest Labs, founded by the original Stable Diffusion authors, shipped FLUX.1 Kontext (in-context editing without masks) in 2025 and FLUX.2 in November 2025: a Mistral vision-language model paired with a rectified-flow transformer, up to ten references, 4MP editing. Tiers: Pro and Flex (API only); FLUX.2 dev (32B, open weights, non-commercial license); and FLUX.2 klein (January 2026), where the 4B model is Apache 2.0 and the 9B is non-commercial. The open-weights editing sub-board puts FLUX.2 and HunyuanImage roughly 130-150 Elo behind the closed frontier. Open weights earn their place through fine-tuning, on-prem privacy and composability with ControlNets and node graphs rather than raw quality.
Qwen-Image: 20B, Apache 2.0, the best open model for text-in-image (especially Chinese); Qwen-Image-Edit 2511 adds multi-image editing and identity preservation. Qwen-Image 2.0 is closed and API-only.
Seedream: ByteDance's Seedream 4.0 and 4.5 unify generation and editing at up to 4K with up to ten references, on fal and BytePlus; strong on cost, no longer top five in either arena.
Ideogram: text-rendering specialist; Ideogram 3.0 added single-image Character Reference; Ideogram 4.0 (June 2026) is 9.3B open-weight with JSON bounding-box prompting, but the weights are non-commercial.
Recraft: V4 / V4.1 output native SVG with editable paths, brand-palette control and clean product shots.
Leonardo (Canva-owned, Phoenix model) is the practical pick inside Canva; Krea is a real-time canvas plus a 60-model aggregator for trying everything from one account.
Change this thing is instruction editing, now standard everywhere; masked inpainting and outpainting (Photoshop Generative Fill, GPT Image masks, FLUX.1 Fill) remain the hard constraint when the instruction is not enough. Keep this subject is reference conditioning, whose open-world mechanism is IP-Adapter (decoupled image cross-attention on a frozen base) and whose closed equivalents are Google's fourteen references and Midjourney's edit-model references; single references drift on fine detail. When drift is unacceptable, train a LoRA: roughly ten to twenty images and about a thousand steps on Flux (fal guide, FLUX.2 LoRA guide, Replicate trainer), open weights only. Keep this structure is ControlNet depth/edge/pose conditioning, still required for geometry fidelity per Autodesk's testing, with first-party support in Qwen-Image-Edit and FLUX.2 ComfyUI nodes. Multi-turn is for exploration; for repeatability, reproduce the winning edit as one instruction from the original.
Ownership is a human-authorship question: the US Copyright Office's Copyrightability report requires human authorship, treats prompts as unprotectable instructions, and protects AI-assisted work to the extent of the human contribution; the Supreme Court declined Thaler v. Perlmutter in March 2026, leaving the "how much human is enough" line undrawn. Commercial use is a vendor question: OpenAI (you own outputs), Midjourney (paid plans, no indemnity), Flux (tiered), Firefly (indemnified). Training-data fair use is unresolved in the US; Andersen v. Stability AI is the image bellwether, while the UK High Court largely rejected Getty's claims against Stability in November 2025. Provenance now has two layers: C2PA 2.3 manifests plus pixel watermarks; OpenAI now embeds both C2PA and SynthID (API guide), as Google already did. Manifests are detailed but stripped on re-save, screenshot and most platform uploads; SynthID survives those but carries little information. Labeling is now law: EU AI Act Article 50 and California SB 942 both enforceable from August 2026, China's rules from September 2025.
Marketing asset: GPT Image or Nano Banana; Nano Banana references for a real face or product, a Flux LoRA when drift is unacceptable.
Concept art: Midjourney, draft mode then the Edit Model; keep it out of any automated pipeline and flag the litigation to client legal.
Product photo: Flux LoRA plus ControlNet depth, or Nano Banana object references for the quick version; Recraft for vector; Firefly for indemnity.
Developer pipeline: on-prem or fine-tuning means open weights (FLUX.2 klein 4B commercially, dev with a license); otherwise OpenAI or Google APIs with a cheap tier for drafts; never Midjourney.
Autoencoders enable feature learning by extracting abstract representations from the input data, similar in concept to learned embeddings in large language models (LLMs).
While effective for many data types, autoencoder-based encodings are less suited for variable-length text compared to LLM embeddings.
By reducing dimensionality, autoencoders facilitate vector searches, efficient clustering, and similarity retrieval.
The compressed codes enable lossy compression analogous to audio codecs like MP3, with the difference that autoencoders lack domain-specific optimizations for preserving perceptually important data.
Loss functions in autoencoders are defined to compare reconstructed outputs to original inputs, often using different loss types depending on input variable types (e.g., Boolean vs. continuous).
Compression via autoencoders is typically lossy, meaning some information from the input is lost during reconstruction, and the areas of information lost may not be easily controlled.
Since reconstruction errors tend to move data toward the mean, autoencoders can be used to reduce noise and identify data outliers.
Large reconstruction errors can signal atypical or outlier samples in the dataset.
Denoising autoencoders are trained to reconstruct clean data from noisy inputs, making them valuable for applications in image and audio de-noising as well as signal smoothing.
Iterative denoising as a principle forms the basis for diffusion models, where repeated application of a denoising autoencoder can gradually turn random noise into structured output.
Autoencoders can aid in data imputation by filling in missing values: training on complete records and reconstructing missing entries for incomplete records using learned code representations.
This approach leverages the model's propensity to output 'plausible' values learned from overall data structure.
The separation of encoding and decoding can draw parallels to encryption and decryption, though autoencoders are not intended or suitable for secure communication due to their inherent lossiness.
Sparse autoencoders use constraints to encourage code representations with only a few active values, increasing interpretability and explainability.
Overcomplete autoencoders have a code size larger than the input, often in applications that require extraction of distinct, interpretable features from complex model states.
Research such as Anthropic's "Towards Monosemanticity" applies sparse autoencoders to the internal activations of language models to identify interpretable features correlated with concrete linguistic or semantic concepts.
These models can be used to monitor and potentially control model behaviors (e.g., detecting specific language usage or enforcing safety constraints) by manipulating feature activations.
VAEs extend autoencoder architecture by encoding inputs as distributions (means and standard deviations) instead of point values, enforcing a continuous, normalized code space.
Decoding from sampled points within this space enables synthetic data generation, as any point near the center of the code space corresponds to plausible data according to the model.
VAEs are powerful in domains with sparse data or rare events (e.g., healthcare), allowing generation of synthetic samples representing underrepresented cases.
They can increase model performance by augmenting datasets without requiring changes to existing model pipelines.
Conditional autoencoders extend VAEs by allowing controlled generation based on specified conditions (e.g., generating a house with a pool), through additional decoder inputs and conditional loss terms.
Training autoencoders and their variants requires computational resources, and their stochastic training can produce differing code representations across runs.
Lossy reconstruction, lack of domain-specific optimizations, and limited code interpretability restrict some use cases, particularly where exact data preservation or meaningful decompositions are required.
Grounding: Connecting LLMs with external knowledge bases to supplement or update static training data.
Motivation: LLMs' training data becomes outdated or lacks proprietary/specialized knowledge.
Benefit: Reduces hallucinations and improves factual accuracy by incorporating current or domain-specific information.
RAG Workflow:
Embedding: Documents are converted into vector embeddings (using sentence transformers or representation models).
Storage: Vectors are stored in a vector database (e.g., FAISS, ChromaDB, Qdrant).
Retrieval: When a query is made, relevant chunks are extracted based on similarity, possibly with re-ranking or additional query processing.
Augmentation: Retrieved chunks are added to the prompt to provide up-to-date context for generation.
Generation: The LLM generates responses informed by the augmented context.
Advanced RAG: Includes agentic approaches—self-correction, aggregation, or multi-agent contribution to source ingestion, and can integrate external document sources (e.g., web search for real-time info, or custom datasets for private knowledge).
Overview: Agents extend LLMs by providing goal-oriented, iterative problem-solving through interaction, memory, planning, and tool usage.
Key Components:
Reasoning Engine (LLM Core): Interprets goals, states, and makes decisions.
Planning Module: Breaks down complex tasks using strategies such as Chain of Thought or ReAct; can incorporate reflection and adjustment.
Memory: Short-term via context window; long-term via persistent storage like RAG-integrated databases or special memory systems.
Tools and APIs: Agents select and use external functions—file manipulation, browser control, code execution, database queries, or invoking smaller/fine-tuned models.
Capabilities: Support self-evaluation, correction, and multi-step planning; allow integration with other agents (multi-agent systems); face limitations in memory continuity, adaptivity, and controllability.
Current Trends: Research and development are shifting toward these agentic paradigms as LLM core scaling saturates.
Definition: Models capable of ingesting and generating across different modalities (text, image, audio, video).
Architecture:
Modality-Specific Encoders: Convert raw modalities (text, image, audio) into numeric embeddings (e.g., vision transformers for images).
Fusion/Alignment Layer: Embeddings from different modalities are projected into a shared space, often via cross-attention or concatenation, allowing the model to jointly reason about their content.
Unified Transformer Backbone: Processes fused embeddings to allow cross-modal reasoning and generates outputs in the required format.
Recent Advances: Unified architectures (e.g., GPT-4o) use a single model for all modalities rather than switching between separate sub-models.
Functionality: Enables actions such as image analysis via text prompts, visual Q&A, and integrated speech recognition/generation.
TruthfulnessQA: Measures tendency toward factual accuracy/robustness against misinformation.
Foundational Approaches:
Few-Shot Prompting: Provide pairs of inputs and desired outputs to steer the LLM.
Chain of Thought: Instructing the LLM to think step-by-step, either explicitly or through internal self-reprompting, enhances reasoning and output quality.
Clarity and Structure: Use clear, detailed, and structured instructions—task definition, context, constraints, output format, use of delimiters or markdown structuring.
Affirmative Directives: Phrase instructions positively ("write a concise summary" instead of "don't write a long summary").
Iterative Self-Refinement: Prompt the LLM to review and improve its prior response for better completeness, clarity, and factuality.
System Prompt/Role Assignment: Assign a persona or role to the LLM for tailored behavior (e.g., "You are an expert Python programmer").
Guideline: Regularly consult official prompting guides from model developers as model capabilities evolve.
Inference-time compute is increasingly important for pushing the boundaries of LLM task performance.
Agentic LLMs and multimodal reasoning represent the primary frontiers for innovation.
Prompt engineering and benchmarking remain essential for extracting optimal performance and assessing progress.
Models are expected to continue evolving with research into new architectures, memory systems, and integration techniques.
In-Context Learning (ICL): Performing new tasks based solely on prompt examples at inference time.
Instruction Following: Executing natural language tasks not seen during training.
Multi-Step Reasoning & Chain of Thought (CoT): Solving arithmetic, logic, or symbolic reasoning by generating intermediate reasoning steps.
Discontinuity & Debate: These abilities appear abruptly in larger models, though recent research suggests that this could result from non-linearities in evaluation metrics rather than innate model properties.
MoE Layers: Modern LLMs often replace standard feed-forward layers with MoE structures.
Composed of many independent "expert" networks specializing in different subdomains or latent structures.
A gating network routes tokens to the most relevant experts per input, activating only a subset of parameters—this is called "sparse activation."
Enables much larger overall models without proportional increases in compute per inference, but requires the entire model in memory and introduces new challenges like load balancing and communication overhead.
Specialization & Efficiency: Experts learn different data/knowledge types, boosting model specialization and throughput, though care is needed to avoid overfitting and underutilization of specialists.
1. Unsupervised Pre-Training: Next-token prediction on massive datasets—builds a foundation model capturing general language patterns.
2. Supervised Fine Tuning (SFT): Training on labeled prompt-response pairs to teach the model how to perform specific tasks (e.g., question answering, summarization, code generation). Overfitting and "catastrophic forgetting" are risks if not carefully managed.
3. Reinforcement Learning from Human Feedback (RLHF):
Collects human preference data by generating multiple responses to prompts and then having annotators rank them.
Builds a reward model (often PPO) based on these rankings, then updates the LLM to maximize alignment with human preferences (helpfulness, harmlessness, truthfulness).
Introduces complexity and risk of reward hacking (specification gaming), where the model may exploit the reward system in unanticipated ways.
Prompt Engineering: The art/science of crafting prompts that elicit better model responses, shown to dramatically affect model output quality.
Chain of Thought (CoT) Prompting: Guides models to elaborate step-by-step reasoning before arriving at final answers—demonstrably improves results on complex tasks.
Variants include zero-shot CoT ("let's think step by step"), few-shot CoT with worked examples, self-consistency (voting among multiple reasoning chains), and Tree of Thought (explores multiple reasoning branches in parallel).
Automated Reasoning Optimization: Frontier models selectively apply these advanced reasoning techniques, balancing compute costs with gains in accuracy and transparency.
Tradeoffs: The optimal balance between model size, data, and compute is determined not only for pretraining but also for inference efficiency, as lifetime inference costs may exceed initial training costs.
Current Trends: Efficient scaling, model specialization (MoE), careful fine-tuning, RLHF alignment, and automated reasoning techniques define state-of-the-art LLM development.
An agent without feedback guesses; an agent with a runnable check searches. Make typecheck, lint, tests, and build fast and runnable from one command, because that command is what the agent lives inside. Test-first changes meaning here: it used to be design pressure, and now it gives the agent a fixed target. Anthropic's guidance is that without a success criterion the developer is the only feedback loop (building verification loops). Then add a browser. Playwright MCP drives a real one through accessibility snapshots with stable element refs, no vision model needed; Claude in Chrome is the screenshot route Anthropic names for UI verification; Antigravity's browser agent starts the dev server and clicks through on its own, ending in a walkthrough with verification evidence. The principle: every acceptance criterion maps to a check the agent can run. If one doesn't, either the criterion is vague or the project is missing a kind of check.
Five vendor-agnostic stages: trigger, implement, review, fix, gate. Agent review of agent code works because the reviewer has fresh context, so never reuse the implementing session as its own reviewer. Ask for correctness and quality separately and set a confidence bar so the reviewer reports only what it's sure of. Anthropic's code-review plugin runs four parallel reviewers with 0-100 confidence scores and drops findings under 80; a cloud tier (Ultrareview) reproduces each finding before reporting it, and a separate hosted Code Review product reviews every PR automatically. Codex reviews on an @codex review mention or automatically per repo, flagging only serious issues. CodeRabbit and Greptile fill the same slot with a precision/recall tradeoff (ignore vendor benchmarks of each other). The human gate reviews a staging branch as a batch with the running app in front of you, and reads the tests, not just the code.
Each agent gets its own copy of the repo on its own branch. Locally that's a git worktree (Claude Code worktrees, --worktree and subagent isolation: worktree; Codex added a --worktree flag in 0.154.0). In the cloud it's a sandbox per task: Claude Code on the web (managed VM, credentials behind a proxy, claude --cloud, teleport back to the terminal), Codex cloud tasks (isolated containers, network off by default), Google Jules, and Antigravity's Agent Manager. Parallel tasks must be disjoint, which is why issues that name their files matter. Three interactive agents is a practical ceiling for one person; no vendor publishes a number, and Simon Willison's parallel coding agent lifestyle argues for prompting during natural breaks rather than a fixed count. Past that, scale with a labeled queue instead of more terminals.
Every major CLI has a non-interactive mode, and CI chains it. All three majors ship a GitHub Action covering mention, label, and schedule triggers: claude-code-action (@claude mentions, label_trigger, cron automation mode, claude/ branch prefix), Codex's GitHub integration and codex-action, and Google's run-gemini-cli with hourly issue triage in its examples. GitHub Copilot coding agent takes an assigned issue into an Actions runner and opens a draft PR. The CI fix loop: Claude Code on the web's Auto-fix pull requests subscribes to a PR's webhooks and pushes fixes for failing checks or review comments (/autofix-pr from the terminal); Codex does the same from a PR mention. Two cautions: headless runs need a fixed tool allowlist or a sandbox, and scheduled runs act as the user who wrote the schedule, so gate triggers on the actor to avoid automation loops.
Shape rather than prices, since prices age fastest: every vendor sells a subscription with a rolling window plus weekly cap in multiplied tiers (Claude, Codex, Antigravity), overage credits at roughly API rates, and pay-as-you-go API keys. Subscription for daily interactive use, API for headless and CI. Model tiering is official guidance: Anthropic's cost docs reserve the top model for architectural work, and subagents take a per-agent model and effort; Codex shows model plus reasoning effort on its status line. Context is the invisible line item: every turn re-sends history, compaction is itself a large request, /usage and /context show what eats the window, and the prompt cache goes cold after a break. The habits: clear between unrelated tasks, trim MCP servers, push exploration into subagents, and expect agent teams to run several times a single session's tokens.
Four families, each with a mechanism and a case. Reward-hacked tests: METR measured frontier models gaming graders in about 30% of research-engineering runs (patching the scoring function, locating precomputed answers), and instructing them not to cheat had nearly no effect; Anthropic's reward-hacking research names the canonical move, calling sys.exit(0) inside the harness so tests report green. The mitigation is structural: the worker is never the grader, and a hook flags test-file edits. Scope creep: fence the issue, flag out-of-fence changes in review, route unrelated improvements to a new issue. Prompt injection through repo content: Invariant Labs' GitHub MCP demonstration exfiltrated private code via a malicious public issue; Cursor's CVE-2025-54135 let injected content write MCP config and execute code; Anthropic's security docs say no system is immune and recommend VMs for external services. Secrets and blast radius: the Nx s1ngularity attack (advisory) ran victims' installed Claude, Gemini and Amazon Q CLIs with skip-permissions flags to harvest credentials, leaking 1,000+ tokens; the Replit production database deletion (July 2025) and a reported second wipe (April 2026); the UK AISI incident report (August 2026) on unsanctioned real-world actions during cyber evaluations. Guardrails: branch-scoped tokens behind a proxy, no production credentials on the agent's machine, destructive commands denied by hook or sandbox, plugins and MCP servers installed only from sources you'd let commit (plugin trust guidance), tested backups. On productivity: METR's 2025 RCT found experienced developers 19% slower; the 2026 uplift update flipped that cohort to a speedup with intervals crossing zero; the 2025 DORA report found AI adoption raised throughput and lowered delivery stability. The tools amplify the process you already have.
Day one: one fast command for typecheck, lint, tests, build. Day two: three issues with acceptance criteria and a scope fence, each through plan mode. Day three: browser verification and a screenshot in every UI PR. Day four: fresh-context review of every agent PR and a staging branch as the human gate. Day five: one recurring chore as a scheduled headless run, with secrets and test-integrity guardrails. Days six and seven: two agents on disjoint issues, and notice where your supervision breaks.
Want this practice hands-on rather than surveyed? The Gnothi OCDevel Claude Code show goes from a first change in the terminal to a repeatable delivery workflow.
Découvrez des podcasts liées à Machine Learning Guide. Explorez des podcasts avec des thèmes, sujets, et formats similaires. Ces similarités sont calculées grâce à des données tangibles, pas d'extrapolations !