Back

Explore every episode of the podcast Redwood Research Blog

Dive into the complete episode list for Redwood Research Blog. Each episode is cataloged with detailed descriptions, making it easy to find and explore specific topics. Keep track of all episodes from your favorite podcast and never miss a moment of insightful content.

Rows per page:

1–50 of 123

TitlePub. DateDuration
“SOTA alignment assessments don’t strongly update us against misalignment” by Alexa Pan31 Jul 202601:08:10

Anthropic concluded in the April Mythos Preview alignment risk update that the model “does not possess any unknown propensities that would increase alignment risk.” The report argues that if Mythos Preview were coherently misaligned[1], it likely would have been detected by the assessment (following Anthropic, I will call this “reliability of the assessment”[3]).

While I agree with the report on the above bottom-line conclusions (substantially on priors), I think there are gaps in its argument which weaken the current assessment and might invalidate future assessments. In particular, the report often uses weak evidence to justify reliability.

  1. The report gives fairly weak experimental evidence for Mythos Preview having insufficient capabilities to evade monitoring. The model is plausibly often eval-aware and underelicited in the relevant capability evaluations. So, it might silently sandbag if coherently misaligned, or unintentionally underperform if otherwise misaligned.

    1. This limitation is important: one could argue that lack of covert capabilities for sophisticated sabotage (a subset of the capabilities I discuss here) is the single most load bearing argument in alignment risk reports.

    2. Authors of the report could have made calibrated guesses about Mythos Preview's covert capabilities, especially for covert sabotage, based on other [...]

---

Outline:

(03:16) How reliability fits into the overall safety argument

(05:22) Reliability claims by AI companies

(05:59) Reliability claims by external evaluators

(06:37) Alignment assessments are less reliable than developers claim

(07:21) 1: Measuring capabilities to covertly undermine alignment assessments

(10:15) Issues with evaluation awareness

(13:39) Issues with underestimating covert capabilities

(16:45) Issues with sandbagging rule-out

(19:20) 2: Stress-testing alignment assessments with auditing games

(20:26) An auditing failure with Mythos

(22:21) AuditBench results

(24:17) 3: Conditioning on misalignment should make us think that certain covert capabilities are better than expected

(26:20) Bottom line on the strength of current alignment assessments

(29:07) Conclusion

(29:44) Appendix:

(29:47) Why I focus on motive / alignment assessments in alignment risk reports

(30:59) Auditability vs. Trustedness

(33:27) More reliability claims by developers and third party evaluators

(33:43) Mythos Alignment Risk Update

(35:01) Opus 4.6 Sabotage Risk Report

(35:46) GPT 5.5 System card

(36:49) Muse Spark system card

(37:36) Mythos Alignment Risk Update, safety arguments against sandbagging

(38:40) UK AISI evaluations for Opus 4.7

(40:01) Past auditing games by Anthropic

(42:24) Anti-auditing capability measurements

(43:51) Conditioning on coherent misalignment updates us on certain covert capabilities

The original text contained 92 footnotes which were omitted from this narration.

---

First published:
July 31st, 2026

Source:
https://blog.redwoodresearch.org/p/sota-alignment-assessments-dont-strongly

---

Narrated by TYPE III AUDIO.

---

Images from the article:

https://substackcdn.com/image/fetch/$s_!oynz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ba512c3-3828-48ee-89cd-22354eaa0ebc_1772x989.pnghttps://substackcdn.com/image/fetch/$s_!CaGP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e28d343-2d97-4be3-8d4e-5d7bf08a6fca_1362x411.pnghttps://substackcdn.com/image/upload/f_auto,q_auto/v1/mirroredImages/4ae97320c14eea3e3e7d91954b1e14674b55d25fb4b9167e795ad4616b564ed2/zgwihixkrxk11wbktmm1https://substackcdn.com/image/upload/f_auto,q_auto/v1/mirroredImages/78b1c44b69458f7318f0b0f3c57421ff77fc80671cfeff000f6559005cbb3b73/jhbm2oglyzr6tqlmwujshttps://substackcdn.com/image/fetch/$s_!7pTf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0ed8b97-a58e-4a41-a852-9d87529a375c_1366x504.pnghttps://substackcdn.com/image/fetch/$s_!4GXU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1fc44c7-bdf1-477e-9079-59121ebb9c38_1778x544.pnghttps://substackcdn.com/image/upload/f_auto,q_auto/v1/mirroredImages/cfba3ba1032fe0ae01d9ea6054419190763bfd321fab08aed66505a1a00d7186/vaexztzt5p8juqm4bk4ahttps://substackcdn.com/image/fetch/$s_!-ayO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fceab08f7-4ccb-4a02-90d3-176cb7470ccf_968x678.pnghttps://substackcdn.com/image/upload/f_auto,q_auto/v1/mirroredImages/b64fd182957e30cd4d052e21cdacad6575735ed5a1ed3fbef8923824eac1274e/bedwqvgtkm25kwoc94uthttps://substackcdn.com/image/fetch/$s_!N1pi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd269296-28b8-449d-b65a-cf140621806a_2048x901.pnghttps://substackcdn.com/image/upload/f_auto,q_auto/v1/mirroredImages/c8d21f3c00b9455c6b36db97be6163f2567798f9344f262d8209a4a2d55bedd1/jkn8xxpwr8fmb3ftfmtihttps://substackcdn.com/image/fetch/$s_!nOfc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F056a4762-ab30-401c-8c9e-0f6eb57e5d94_1156x312.pnghttps://substackcdn.com/image/upload/f_auto,q_auto/v1/mirroredImages/ec5ff3a1f94c30ef2769e8bef24c1a749a29e8da54af76c0163c09d399200a62/chbrpwbgermhsjcjddtjhttps://substackcdn.com/image/upload/f_auto,q_auto/v1/mirroredImages/6387af7ebef506388bfe9f949124e94c08cccb848faaaea25c69933d81198279/mobfwipdzgfhdgjwr3vehttps://substackcdn.com/image/fetch/$s_!4Kgd!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80fb15d7-ee04-40d0-b211-9fe235bcafd4_1156x215.pnghttps://substackcdn.com/image/fetch/$s_!TM9G!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ab6dcaa-f1b4-445f-b6b6-7f356e6664fe_992x350.pnghttps://substackcdn.com/image/upload/f_auto,q_auto/v1/mirroredImages/oirrSj3itFLSyscW8/79ae51136f24e837d4f5b7c3d62c6e6f8ba359b145f6228d329625666351a1bc/y1m2s9mzocbuma3zmmivhttps://substackcdn.com/image/fetch/$s_!OVfL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8b87f4f-4380-47ab-8146-72e94746c88d_918x296.pnghttps://substackcdn.com/image/upload/f_auto,q_auto/v1/mirroredImages/34d9b50a45426f124e061c961bcffd094025c9fb32df777fdd1ddab0ebfbacab/zdnvre2mr4wh996lftdhhttps://substackcdn.com/image/fetch/$s_!Lisd!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe95fed53-581d-4358-bdf5-fe2ee13b1739_1078x314.pnghttps://substackcdn.com/image/upload/f_auto,q_auto/v1/mirroredImages/cf03008111943b83bb7831e194984586e478e60f83917948ce50007fff53fbb8/qreze9sj0xvoc3z5jtjjhttps://substackcdn.com/image/fetch/$s_!c5fE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0cbbfbc0-4eb0-4ae7-9e9f-3ba4913d363c_980x506.pnghttps://substackcdn.com/image/upload/f_auto,q_auto/v1/mirroredImages/7d0eb5f75030374114c5818a25098e0e29553144d2f9ab34fc0110959b6fd049/x9glvuuzodafa27ilsafhttps://substackcdn.com/image/fetch/$s_!o720!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9696aea0-934b-4b55-a74b-7ed5a5aaa8f1_968x432.pnghttps://substackcdn.com/image/upload/f_auto,q_auto/v1/mirroredImages/774f86eb41d652aee879f052c4e220f4375bb9f52bdccc8efec119f4ae8ea65e/d9qf4gnfvtbsvuewui7ihttps://substackcdn.com/image/fetch/$s_!OcRI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb895f18-5e21-4c11-8b8a-1a9f173896d0_822x270.pnghttps://substackcdn.com/image/upload/f_auto,q_auto/v1/mirroredImages/df52ec759e013fbfc035815aa89de76f3e026dd468d563e85eab54c8b2487e49/jwjrqlla8l011ufhynnshttps://substackcdn.com/image/fetch/$s_!tcPd!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e090496-889e-4ea4-810b-a59d370a78e5_829x349.pnghttps://substackcdn.com/image/upload/f_auto,q_auto/v1/mirroredImages/940c8bd5df21483b9ffd19559fbe49500bbad10300ed45eca56a1e8be13eb934/f8tpzorkl7wod3exegczhttps://substackcdn.com/image/fetch/$s_!0LBL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4203a2d1-267f-4539-a2e7-969b22ddf86f_1663x901.png

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

“Untrusted advice for AI control: Short, strong advice significantly uplifts weak LLMs” by Caleb Biddulph, Adam Kaufman27 Jul 202600:44:15

Subtitle: When a misaligned AI can only output tiny amounts of information, it may find sabotage very difficult.

TL;DR: We introduce the untrusted advice protocol, in which a trusted executor LLM takes every action and an untrusted advisor LLM can only send it short hints. Even with as few as 4 characters per step, this advice recovers a substantial fraction of the capability gap between the two models. Because the untrusted LLM's influence flows through such a narrow, monitorable channel, we argue that this achieves near-maximal safety in our BashArena setting. We also discuss the general concept of information bottlenecks and their benefits for interpretability, security, and cost.

In SWE-bench Verified, a strong, untrusted LLM advising a weak, trusted LLM every step can significantly improve the latter's performance, even when we limit the length of the advice. See the more detailed version of this figure later in this post.

In high-stakes AI control, we want to safely use a highly capable but untrusted model (U) that might secretly attempt a misaligned, catastrophic action. To do this, we create protocols that call U alongside a less capable, trusted model (T). Typically, T takes an auxiliary role in these protocols [...]

---

Outline:

(05:01) Experiments

(05:36) Main experiment: how does limiting advice length affect performance?

(09:55) Reducing U's bit usage

(11:01) Counting bits using LLM surprisal

(13:48) Making U select from finite options

(14:28) Why don't we red-team this protocol?

(17:08) Is studying maximally safe protocols worth the safety tax?

(19:16) Types of restrictions on U's advice

(21:20) Information bottlenecks provide other advantages

(21:52) Interpretability

(24:01) Security

(24:26) Cost

(25:14) Conclusion

(26:28) Appendix: more ways to implement information bottlenecks

(26:34) Amortizing U's influence with pre-deployment work

(28:25) Interpolating between T and U

(29:04) Bottlenecking updates to T's weights

(31:16) Appendix: colluding instances of U could defeat untrusted advice

(33:21) Appendix: how to measure surprisal

(38:11) Appendix: selecting advice from a menu

(40:44) Appendix: best-of-n protocol

(42:37) Appendix: advising less frequently

The original text contained 26 footnotes which were omitted from this narration.

---

First published:
July 27th, 2026

Source:
https://blog.redwoodresearch.org/p/untrusted-advice-for-ai-control-short

---

Narrated by TYPE III AUDIO.

---

Images from the article:

https://substackcdn.com/image/fetch/$s_!mFE6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffeb7856b-dbf4-4dc2-a75e-b549a2ce42f4_960x540.pnghttps://substackcdn.com/image/fetch/$s_!JpOX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F02741dcb-5ec8-4f83-b045-be79a5fc8935_1274x740.pnghttps://substackcdn.com/image/fetch/$s_!XEN9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4b98216-75c1-4d6e-bb9d-320d95fae7d6_2073x740.pnghttps://substackcdn.com/image/fetch/$s_!Qp4r!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe88f829c-a6db-4e56-a9e2-57a5a0596150_1950x780.pnghttps://substackcdn.com/image/fetch/$s_!-Vd4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52124f25-5d60-4805-b6bc-f96550fcae4b_1950x1350.pnghttps://substackcdn.com/image/fetch/$s_!k5Uu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4716086-8785-45d1-9d7f-325e5f3ed6b5_2048x730.pnghttps://substackcdn.com/image/fetch/$s_!VKiW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6457b758-8993-4bfb-beaf-77e209f68f41_1208x821.pnghttps://substackcdn.com/image/fetch/$s_!6JdT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fae9b13-165e-4716-aa06-32123e26582a_1350x975.png

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

“Retrying vs Resampling in AI Control” by James Lucassen, Adam Kaufman29 May 202600:19:29

We’ve just released a new paper: Retrying vs Resampling in AI Control. We revisit the resampling protocols introduced in Ctrl-Z with an up-to-date setting and much stronger models, and compare them against “retrying” protocols similar to Claude Code auto mode or Codex Auto-review.

Motivation

Roughly a year ago we released Ctrl-Z, the first paper to study control techniques for agents. A headline result of that paper was the performance of resample protocols – strategies that involve taking multiple i.i.d. samples from the model per step. But since Ctrl-Z, models have gotten much stronger, and we have built more sophisticated control settings to keep up. We wanted to answer the following questions:

How well do the results from Ctrl-Z hold up with better models and a better setting? Current high stakes control research is trying to learn by analogy about how to do control effectively in a real high stakes deployment during a real intelligence explosion. Findings about technique performance1 are going to have to generalize pretty far to be useful.

If the resample protocols from Ctrl-Z still work, what makes them work? One way we try to make our work more generalizable is by understanding the dynamics governing outcomes [...]

---

Outline:

(00:30) Motivation

(02:35) TL;DR Takeaways

(04:34) Methodology

(07:14) Differences from Ctrl-Z

(12:14) Are Retrying Protocols Exploitable?

(15:25) Cost and Latency of Resampling

(17:24) Conclusion

The original text contained 7 footnotes which were omitted from this narration.

---

First published:
May 29th, 2026

Source:
https://blog.redwoodresearch.org/p/retrying-vs-resampling-in-ai-control

---

Narrated by TYPE III AUDIO.

---

Images from the article:

https://substackcdn.com/image/fetch/$s_!4w3J!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe64f218-7438-4347-bfaf-b52bed062809_2048x838.pnghttps://substackcdn.com/image/fetch/$s_!Hc35!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b09301e-a1f2-4dde-a58c-562e713c8f17_2048x768.pnghttps://substackcdn.com/image/fetch/$s_!fqZJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef92c796-b596-4af7-b2d9-1634c24b99a4_2048x696.pnghttps://substackcdn.com/image/fetch/$s_!0TGO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd5fd9cc-2a06-4daa-8258-ca0e52cad754_2048x851.pnghttps://substackcdn.com/image/fetch/$s_!pMSs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F00cae980-01cc-4f1d-b4d4-eb0c7af71677_2048x805.png

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

“An overview of control measures” by Ryan Greenblatt06 Apr 202500:50:14

Subtitle: What methods can we use to ensure control?.

We often talk about ensuring control, which in the context of this doc refers to preventing AIs from being able to cause existential problems, even if the AIs attempt to subvert our countermeasures. To better contextualize control, I think it's useful to discuss the main countermeasures (building on my prior discussion of the main threats). I'll focus my discussion on countermeasures for reducing existential risk from misalignment, though some of these countermeasures will help with other problems such as human power grabs.

I'll overview control measures that seem important to me, including software infrastructure, security measures, and other control measures (which often involve usage of AI or human labor). In addition to the core control measures (which are focused on monitoring, auditing, and tracking AIs and on blocking and/or resampling suspicious outputs), there is a long tail of potentially [...]

---

Outline:

(01:19) Background and categorization

(06:11) Non-ML software infrastructure

(07:11) The most important software infrastructure

(11:58) Less important software infrastructure

(19:38) AI and human components

(21:30) Core measures

(27:49) Other less-core measures

(32:45) More speculative measures

(41:30) Specialized countermeasures

(46:05) Other directions and what Im not discussing

(47:15) Security improvements

---

First published:
April 6th, 2025

Source:
https://redwoodresearch.substack.com/p/an-overview-of-control-measures

---

Narrated by TYPE III AUDIO.

“Buck on the 80,000 Hours podcast” by Buck Shlegeris05 Apr 202500:01:19

Buck on the 80,000 Hours podcast

My podcast with Rob Wiblin from 80,000 Hours just came out. I’m really happy with how it turned out. I talked about a bunch of stuff on the podcast that I don’t think we’ve written up before.

Transcript + links + summary here. You can also get it as a podcast:

Spotify:

Apple:

Subscribe to Redwood Research blog

Launched a year ago

We research catastrophic AI risks and techniques that could be used to mitigate them.

Discussion about this post

---

Outline:

(00:09) Buck on the 80,000 Hours podcast

(00:34) Subscribe to Redwood Research blog

(00:57) Discussion about this post

---

First published:
April 5th, 2025

Source:
https://redwoodresearch.substack.com/p/buck-on-the-80000-hours-podcast

---

Narrated by TYPE III AUDIO.

“Notes on countermeasures for exploration hacking (aka sandbagging)” by Ryan Greenblatt04 Apr 202500:15:26

Subtitle: How can we prevent AIs from intentionally underperforming on our metrics?.

If we naively apply RL to a scheming AI, the AI may be able to systematically get low reward/performance while simultaneously not having this behavior trained out because it intentionally never explores into better behavior. As in, it intentionally puts very low probability on (some) actions which would perform very well to prevent these actions from being sampled and then reinforced. We'll refer to this as exploration hacking.

In this post, I'll discuss a variety of countermeasures for exploration hacking and what I think the basic dynamics of exploration hacking might look like.

Prior work: I discuss how exploration hacking fits into a broader picture of non-concentrated control here, Evan discusses sandbagging in context of capability evaluations here, and Teun discusses sandbagging mitigations here.

Countermeasures

Countermeasures include:

  • Training on high quality human examples: One [...]

---

Outline:

(01:11) Countermeasures

(03:58) How does it generalize when the model messes up?

(06:56) Empirical evidence

(09:32) Detecting exploration hacking (or sandbagging)

(11:14) Speculation on how neuralese and other architecture changes affect exploration hacking

(13:18) Future work

---

First published:
April 4th, 2025

Source:
https://redwoodresearch.substack.com/p/notes-on-countermeasures-for-exploration

---

Narrated by TYPE III AUDIO.

“Notes on handling non-concentrated failures with AI control: high level methods and different regimes” by Ryan Greenblatt03 Apr 202500:32:52

Subtitle: What are the methods and issues when failures occur diffusely over many actions?.

In this post, I'll try to explain my current understanding of the high level methods for handling non-concentrated failures with control. I'll discuss the regimes produced by different methods and the failure modes of these different regimes. Non-concentrated failures are issues that arise from the AI doing a large number of problematic actions over a long period. (See here and here for more discussion of concentrated vs non-concentrated failures.)

A key aspect of non-concentrated failures is that they can in principle be handled by applying asynchronous online training with a good enough oversight signal on a small fraction of the AI's actions. That is, we could in principle solve these failures by sampling a small fraction of actions to review after they've already been executed, rating these actions using an oversight process which has [...]

---

Outline:

(06:06) What happens without online training?

(10:58) Online training

(13:10) Online training of just the policy

(19:40) Recursive oversight: debate

(20:24) Debate for binary questions

(23:37) Debate for oversight of a policy

(29:15) KL penalties and forcing identical behavior

(31:00) Is it ever a good idea for scheming models to do low stakes sabotage more sophisticated than just inserting errors (at particularly opportune times)?

---

First published:
March 29th, 2025

Source:
https://redwoodresearch.substack.com/p/notes-on-handling-non-concentrated

---

Narrated by TYPE III AUDIO.

“Prioritizing threats for AI control” by Ryan Greenblatt19 Mar 202500:20:43

Subtitle: What are the main threats and how should we prioritize them?.

We often talk about ensuring control, which in the context of this doc refers to preventing AIs from being able to cause existential problems, even if the AIs attempt to subvert our countermeasures. To better contextualize control, I think it's useful to discuss the main threats. I'll focus my discussion on threats induced by misalignment which could plausibly increase existential risk.

While control is often associated with preventing security failures, I'll discuss other issues like sabotaging safety-critical work. I'll list and explain my prioritization of threats. As part of this, I'll argue that rogue internal deployments—cases where the AI is running within the company's datacenter but with control measures disabled—can be substantially worse than self-exfiltration—the AI stealing its weights and running them on an external server—as it might give the AI access to much more compute [...]

---

Outline:

(01:16) Some clusters of threats

(06:15) Concentrated vs non-concentrated failures

(08:58) My prioritization between these threats

---

First published:
March 19th, 2025

Source:
https://redwoodresearch.substack.com/p/prioritizing-threats-for-ai-control

---

Narrated by TYPE III AUDIO.

“How might we safely pass the buck to AI?” by Josh Clymer19 Feb 202501:12:10

Subtitle: Developing AI employees that are safer than human ones.

My goal as an AI safety researcher is to put myself out of a job.

I don’t worry too much about how planet sized brains will shape galaxies in 100 years. That's something for AI systems to figure out.

Instead, I worry about safely replacing human researchers with AI agents, at which point human researchers are “obsolete.” The situation is not necessarily fine after human obsolescence; however, the bulk of risks that are addressable by human technical researchers (like me) will have been addressed.

This post explains how developers might safely “pass the buck” to AI.

I first clarify what I mean by “pass the buck” (section 1) and explain why I think AI safety researchers should make safely passing the buck their primary end goal – rather than focus on the loftier ambition of aligning superintelligence (section [...]

---

Outline:

(19:08) 1. Briefly responding to objections

(21:54) 2. What I mean by passing the buck to AI

(23:46) 3. Why focus on passing the buck rather than aligning superintelligence.

(28:33) 4. Three strategies for passing the buck to AI

(31:35) 5. Conditions that imply that passing the buck improves safety

(34:32) 6. The capability condition

(38:31) 7. The trust condition

(39:27) 8. Argument #1: M_1 agents are approximately aligned and will maintain their alignment until they have completed their deferred task

(49:28) 9. Argument #2: M_1 agents cannot subvert autonomous control measures while they complete the deferred task

(50:51) Analogies to dictatorships suggest that autonomous control might be viable

(52:51) Listing potential autonomous control measures

(56:19) How to evaluate autonomous control

(58:17) 10. Argument #3: Returns to additional human-supervised research are small

(01:00:38) Control measures

(01:05:03) 11. Argument #4: AI agents are incentivized to behave as safely as humans

(01:11:25) 12. Conclusion

---

First published:
February 19th, 2025

Source:
https://redwoodresearch.substack.com/p/how-might-we-safely-pass-the-buck

---

Narrated by TYPE III AUDIO.

---

Images from the article:

https://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d5c2469-efa2-48c4-b02b-f43c59a5d6f3_1092x1600.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7fbf932-c716-4f95-8cf6-19475d5dc459_1600x1319.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd565df31-639a-4f61-9c0e-a20660e9a0ed_1600x412.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4106ffc7-ca38-424c-9633-830c4e14fff4_1600x1126.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7637e616-66cc-4764-b755-02c6d52743d0_1600x429.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb421d4ae-75fc-4af3-9353-12689d960c68_1600x457.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84b79a67-88f3-462d-a9c0-ac415ddbd1f2_1600x887.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F536f1811-daa5-453d-8ead-2a623226cb6e_1600x1448.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F23647b20-f84a-474b-a511-d84b04de850c_1600x1256.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96e070e7-496e-4bea-9ace-e631fbd8d667_4666x2476.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed527401-fbc7-4068-9c5b-bae297833cc5_1600x735.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb91f8347-36d4-4c87-aa8a-0ccf3734e2a8_1495x1600.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a579d56-0fb4-42cd-ad40-83c0b6f7c26b_1600x1361.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8afd43ae-6b30-4833-9695-4d3462764418_1600x1047.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F404780da-b28b-4d36-af2d-6d15271f0fad_1600x891.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8214ca9-c56d-4d8b-8c83-569e9712efb0_1600x1030.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1c37139-4595-43ae-b831-9251429e43f6_1600x493.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a96fedd-40f5-4e8f-97c9-c2d3e619a912_1600x1319.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26287955-001e-4ecb-a702-b0249661bf1c_1600x737.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed4ec830-c45d-42e9-8c63-efb6e804b46a_1600x457.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d2e8716-2a30-42a9-aeeb-4901413beaaa_1600x1361.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7573cd83-a6b7-4d27-89e4-29839df66017_1600x509.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8319c39-048d-47f7-a170-985307b9bdcb_1600x737.png

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

“Takeaways from sketching a control safety case” by Josh Clymer30 Jan 202500:07:39

Subtitle: Insights from a long technical paper compressed into a fun little commentary.

Buck Shlegeris and I recently published a paper with UK AISI that sketches a safety case for “AI control” – measures that improve safety despite intentional subversion from AI systems.

I would summarize this work as “turning crayon drawings of safety cases into blueprints.”

It's a long and technical paper! So I’ve compiled my main takeaways in this blogpost.

The main figure from the paper summarizes how to run a control evaluation.

1. Control evaluations are way more complicated than I expected.

At a high level, control evaluations are conceptually simple. A red team constructs substitutes for subversive models. These models try to cause proxy bad outcomes (e.g. exfiltrate dummy data). Then, developers see whether the red team can succeed. If the red team succeeds, then control measures aren’t effective enough.

But there's so much [...]

---

First published:
January 30th, 2025

Source:
https://redwoodresearch.substack.com/p/takeaways-from-sketching-a-control

---

Narrated by TYPE III AUDIO.

---

Images from the article:

https://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc36ceeb-1132-4969-b9ea-1117bc4f60f1_1600x1165.png

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

“Planning for Extreme AI Risks” by Josh Clymer29 Jan 202500:45:52

Subtitle: Are we ready for this?.

This post should not be taken as a polished recommendation to AI companies and instead should be treated as an informal summary of a worldview. The content is inspired by conversations with a large number of people, so I cannot take credit for any of these ideas.

For a summary of this post, see the thread on X.

Many people write opinions about how to handle advanced AI, which can be considered “plans.”

There's the “stop AI now plan.”

On the other side of the aisle, there's the “build AI faster plan.”

Some plans try to strike a balance with an idyllic governance regime.

And others have a “race sometimes, pause sometimes, it will be a dumpster-fire” vibe.

So why am I proposing another plan?

Existing plans provide a nice patchwork of the options available, but I’m not satisfied with any.

[...]



---

Outline:

(03:50) The tl;dr

(06:36) 1. Assumptions

(09:07) 2. Outcomes

(10:07) 2.1. Outcome #1: Human researcher obsolescence

(13:24) 2.2. Outcome #2: A long coordinated pause

(14:33) 2.4. Outcome #3: Self-destruction

(15:46) 3. Goals

(19:16) 4. Prioritization heuristics

(21:58) 5. Heuristic #1: Scale aggressively until meaningful AI software RandD acceleration

(25:41) 6. Heuristic #2: Before achieving meaningful AI software RandD acceleration, spend most safety resources on preparation

(27:46) 7. Heuristic #3: During preparation, devote most safety resources to (1) raising awareness of risks, (2) getting ready to elicit safety research from AI, and (3) preparing extreme security.

(30:19) Category #1: Nonproliferation

(34:52) Category #2: Safety distribution

(38:06) Category #3: Governance and communication.

(39:34) Category #4: AI defense

(40:31) 8. Conclusion

(42:13) Appendix

(42:15) Appendix A: What should Magma do after meaningful AI software RandD speedups

---

First published:
January 29th, 2025

Source:
https://redwoodresearch.substack.com/p/planning-for-extreme-ai-risks

---

Narrated by TYPE III AUDIO.

---

Images from the article:

https://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8fcef14-fe9d-4793-882c-7ebcd3fcd0f2_1272x324.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7b04810-f1ba-42dd-80b2-844692d75ae2_4467x1486.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fafd0e96e-b26e-45a9-bb2a-b5fe03e8b84b_1298x580.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ef1e6c1-8b59-47be-89bd-85464859253f_1406x788.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadaec30d-730e-4d9c-8598-045305265f3a_2106x682.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ffd5a2d-c18e-4db3-9371-b2cd6f732626_2484x1264.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc9e697a-db32-44d5-891e-72ccbca56587_2372x1890.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe35f0ef8-e923-4d63-939a-dde18928c0ab_1600x891.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05507499-0e07-4428-a499-28b387547533_1600x1049.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd681617-7e0f-4b13-95dc-56c02d39fe8a_1600x854.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6761b97f-351b-45e6-a5ef-c96d4000957e_3731x1252.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d9e405b-0fe6-4db2-bd91-8071ee358694_1172x1434.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fedc4b1fb-2007-4586-86e4-84514aeee261_6214x2368.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5262394-692c-4009-9492-5939cc475617_1600x1002.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1fe58de-7ec0-4ccd-bb0e-5f7607a95541_1600x706.png

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

“Ten people on the inside” by Buck Shlegeris28 Jan 202500:07:24

Subtitle: A scary scenario that's worth planning for.

(Many of these ideas developed in conversation with Ryan Greenblatt)

In a shortform, I described some different levels of resources and buy-in for misalignment risk mitigations that might be present in AI labs:

*The “safety case” regime.* Sometimes people talk about wanting to have approaches to safety such that if all AI developers followed these approaches, the overall level of risk posed by AI would be minimal. (These approaches are going to be more conservative than will probably be feasible in practice given the amount of competitive pressure, so I think it's pretty likely that AI developers don’t actually hold themselves to these standards, but I agree with e.g. Anthropic that this level of caution is at least a useful hypothetical to consider.) This is the level of caution people are usually talking about when they discuss making safety cases. [...]

---

First published:
January 28th, 2025

Source:
https://redwoodresearch.substack.com/p/ten-people-on-the-inside

---

Narrated by TYPE III AUDIO.

“When does capability elicitation bound risk?” by Josh Clymer22 Jan 202500:43:38

For a summary of this post see the thread on X.


The assumptions behind and limitations of capability elicitation have been discussed in multiple places (e.g. here, here, here, etc); however, none of these match my current picture.

For example, Evan Hubinger's “When can we trust model evaluations?” cites “gradient hacking” (strategically interfering with the learning process) as the main reason supervised fine-tuning (SFT) might fail, but I’m worried about more boring failures. For example, models might perform better on tasks by following inhuman strategies than by imitating human demonstrations, SFT might not be sample-efficient enough – or SGD might not converge at all if models have deeply recurrent architectures and acquire capabilities in-context that cannot be effectively learned with SGD (section 5.2). Ultimately, the extent to which elicitation is effective is a messy empirical question.

In this post, I dissect the assumptions behind capability elicitation in detail [...]


---

Outline:

(01:49) 1. Background

(01:52) 1.1. What is a capability evaluation?

(04:01) 1.2. What is elicitation?

(09:21) 2. Two elicitation methods and two arguments

(15:48) 3. Justifying Sandbagging is unlikely

(19:47) 4. Justifying Elicitation trains against sandbagging

(21:05) 4.1. The number of i.i.d. samples is limited

(22:15) 4.2. Often, the train and test set are not sampled i.i.d.

(25:50) 5. Justifying Elicitation removes sandbagging that is trained against

(26:58) 5.1. Justifying The training signal punishes underperformance

(30:21) 5.2. Justifying In cases where training punishes underperformance, elicitation is effective

(31:07) Why supervised learning might or might not elicit capabilities

(35:02) Why reinforcement learning might or might not elicit capabilities

(37:00) Empirically justifying the effectiveness of elicitation given a strong training signal

(41:04) 6. Conclusion

(41:39) Appendix

(41:42) Appendix A: An argument that supervised fine-tuning quickly removes underperformance behavior in general

---

First published:
January 22nd, 2025

Source:
https://redwoodresearch.substack.com/p/when-does-capability-elicitation

---

Narrated by TYPE III AUDIO.

---

Images from the article:

https://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15cb0254-c62b-44b5-bb2a-cc4514ad7e48_1248x408.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d320dbd-c89b-4931-a4c1-eb3bee932ef0_1600x723.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71ba9316-d43b-4623-977f-3e7835992169_1600x985.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed6b8676-6aa2-473a-8f3f-b8285ec6468f_1149x982.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb14281a4-e64f-4d77-a2dc-db12894664d0_1600x643.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe322121c-1f94-4ca5-8c95-1b38bb98daa4_1600x897.png

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

“Advice for making robust-to-training model organisms” by Alek Westover, Sebastian Prasanna, Vivek Hebbar, Julian Stastny28 May 202600:32:19

We’d like to develop training techniques that work when applied to future misaligned AI systems. One strategy for studying proposed techniques is to test them on model organisms. However, model organisms built with common techniques are often fragile: we (and other researchers like Roger et al. and Ryd et al.) have observed them to stop misbehaving after untargeted training—training that doesn’t directly target the misbehavior. For example, we have observed that simple untargeted training methods like “train the model to talk like a pirate” is effective against many model organisms that we have created, including many replications of prior work like Hubinger et al., Greenblatt et al., and Ryd et al..

Fragile model organisms aren’t very useful for technique development: when a sophisticated technique succeeds on one, you can’t tell whether the technique is good or the model organism is weak. For instance, if you come up with some complicated technique for generating high-quality responses and find that SFTing on these removes the bad behavior, that may just be because almost any SFT would have removed it.

This post identifies factors that make model organisms more robust to untargeted training. Our main findings are:

  1. Prompted model organisms [...]

---

Outline:

(03:11) Experimental setup

(06:12) Main results

(06:15) Result 1: Untargeted training removes misalignment from prompted model organisms without degrading other capabilities

(12:32) Result 2: FWFT model organisms are more robust than LoRA model organisms

(15:09) Result 3: Backdoor behavior and trigger affect robustness

(16:37) Result 4: Password locking makes model organisms less robust

(20:20) Conclusion

(21:29) Factors that don't help much

(21:54) 1: Training longer doesn't seem increase robustness past a certain point for our backdoors

(22:33) 2: CoT-distilling doesn't increase robustness for our backdoors

(23:07) 3: Using larger models doesn't substantially increase robustness for 2 of our backdoors

(24:44) 4: There aren't clear trends for what types of backdoor triggers and behaviors are most robust

(25:32) 5: SOAP Optimizer doesn't help much

(26:35) 6: Weight Decay doesn't help

(27:18) Appendix:

(27:21) Appendix 1: Pirate SFT is better than QA SFT

(27:45) Appendix 2: Are blue team methods wiping previous training?

(29:30) Appendix 3: Backdoor list

The original text contained 10 footnotes which were omitted from this narration.

---

First published:
May 28th, 2026

Source:
https://blog.redwoodresearch.org/p/advice-for-making-robust-to-training

---

Narrated by TYPE III AUDIO.

---

Images from the article:

https://substackcdn.com/image/fetch/$s_!IjeS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5cf38ee9-4cf8-4b1c-a3d6-0e45c1425cac_2048x1018.pnghttps://substackcdn.com/image/fetch/$s_!i2eh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce36f812-66db-454d-93f2-1c6d026e3261_2048x945.pnghttps://substackcdn.com/image/fetch/$s_!nL1E!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faebfebdd-4c50-49c1-b456-6bdeadb83374_790x590.pnghttps://substackcdn.com/image/fetch/$s_!XMvD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2e70277-9b29-4efd-881a-5c055c6cf768_839x590.pnghttps://substackcdn.com/image/fetch/$s_!KoAu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9f746e0-4dfa-49d0-a318-0c1f08a03184_796x590.pnghttps://substackcdn.com/image/fetch/$s_!-0tr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F62f3b710-4459-4f32-9c22-1109e92a8731_796x590.pnghttps://substackcdn.com/image/fetch/$s_!IsWF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84213326-d9ba-4e40-8c6e-21c2f86d772d_1189x593.pnghttps://substackcdn.com/image/fetch/$s_!tXW7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe92edb8f-87d6-4035-9f4f-7eff82a73e84_790x590.pnghttps://substackcdn.com/image/fetch/$s_!ISaO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F53b29b75-be54-4b1f-a4d9-f235e351898d_1178x593.pnghttps://substackcdn.com/image/fetch/$s_!omqP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3fb9b6a4-68d8-426a-90e0-9e66d2b7a378_1716x584.pnghttps://substackcdn.com/image/fetch/$s_!jWiU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F696f0f00-1416-45db-bf9d-50b3262ffb57_790x590.pnghttps://substackcdn.com/image/fetch/$s_!ZyjM!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88be8904-7d92-4ea4-8c04-5314c7dcb152_2932x1302.pnghttps://substackcdn.com/image/fetch/$s_!_BDc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F493f6ea9-ed26-499f-a9c3-1472c6bf113e_1590x580.pnghttps://substackcdn.com/image/fetch/$s_!J74J!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58f0c19c-afbd-4b34-9497-14585932e23a_791x590.pnghttps://substackcdn.com/image/fetch/$s_!s1Gm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6202abf8-252a-4cff-b758-4266f13d9364_790x590.pnghttps://substackcdn.com/image/fetch/$s_!6csE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd1e25a81-60a6-4fd8-869c-5b5e431e95c8_1189x495.pnghttps://substackcdn.com/image/fetch/$s_!_Jf8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b012365-0da7-46db-9ad5-8397a2c143d8_1189x495.pnghttps://substackcdn.com/image/fetch/$s_!9hpB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e0b0b57-4a29-4952-9b9b-77c772683d02_1189x788.pnghttps://substackcdn.com/image/fetch/$s_!1QJ2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9deec1cd-b452-49bb-9117-889127205a28_1634x1362.pnghttps://substackcdn.com/image/fetch/$s_!oDYj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b815b0f-1571-4efb-b5e1-b0f058048cd6_1658x655.pnghttps://substackcdn.com/image/fetch/$s_!ZQbc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65d07378-2a90-4bac-865b-6de2fe47f931_901x590.pnghttps://substackcdn.com/image/fetch/$s_!wQKt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7bad5502-f795-4cf0-b605-edee99e61496_814x590.pnghttps://substackcdn.com/image/fetch/$s_!xeEb!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ae90d85-8728-4f6f-9c70-48bf02aeb5ba_1189x593.pnghttps://substackcdn.com/image/fetch/$s_!VIQO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc45702ed-4893-4f05-be05-afaaf991b093_790x490.png

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

“How will we update about scheming?” by Ryan Greenblatt19 Jan 202501:22:20

Subtitle: A quantitative description of how I expect to change my mind..

[Cross-posted from LessWrong]

I mostly work on risks from scheming (that is, misaligned, power-seeking AIs that plot against their creators such as by faking alignment). Recently, I (and co-authors) released "Alignment Faking in Large Language Models", which provides empirical evidence for some components of the scheming threat model.

One question that's really important is how likely scheming is. But it's also really important to know how much we expect this uncertainty to be resolved by various key points in the future. I think it's about 25% likely that the first AIs capable of obsoleting top human experts1 are scheming. It's really important for me to know whether I expect to make basically no updates to my P(scheming)2 between here and the advent of potentially dangerously scheming models, or whether I expect to be basically totally confident [...]

---

Outline:

(03:44) My main qualitative takeaways

(05:28) Its reasonably likely (55%), conditional on scheming being a big problem, that we will get smoking guns.

(06:11) Its reasonably likely (45%), conditional on scheming being a big problem, that we wont get smoking guns prior to very powerful AI.

(16:56) My P(scheming) is strongly affected by future directions in model architecture and how the models are trained

(17:32) The model

(23:55) Properties of the AI system and training process

(24:21) Opaque goal-directed reasoning ability

(30:59) Architectural opaque recurrence and depth

(36:01) Where do capabilities come from?

(41:44) Overall distribution from just properties of the AI system and training

(43:20) Direct observations

(43:45) Baseline negative updates

(46:43) Model organisms

(50:35) Catching various types of problematic behavior

(53:43) Other observations and countermeasures

(54:25) Training processes with varying (apparent) situational awareness

(56:33) Training AIs to seem highly corrigible and (mostly) myopic

(58:19) Reward hacking

(01:00:02) P(scheming) under various scenarios (putting aside mitigations)

(01:08:29) An optimistic and a pessimistic scenario for properties

(01:13:51) Conclusion

(01:15:24) Appendix: Caveats and definitions

(01:18:26) Appendix: Capabilities from intelligent learning algorithms

The original text contained 1 footnote which was omitted from this narration.

---

First published:
January 19th, 2025

Source:
https://redwoodresearch.substack.com/p/how-will-we-update-about-scheming

---

Narrated by TYPE III AUDIO.

“Thoughts on the conservative assumptions in AI control” by Buck Shlegeris17 Jan 202500:25:00

Subtitle: Why are we so friendly to the red team?.

Work that I’ve done on techniques for mitigating risk from misaligned AI often makes a number of conservative assumptions about the capabilities of the AIs we’re trying to control. (E.g. the original AI control paper, Adaptive Deployment of Untrusted LLMs Reduces Distributed Threats, How to prevent collusion when using untrusted models to monitor each other.) For example:

  • The AIs are consistently trying to subvert safety measures.

  • They’re very good at strategizing about how to subvert them, and can do this strategizing totally illegibly (e.g. in a single forward pass).

  • They can follow a strategy that's coordinated between all their instances.

  • They know exactly what safety measures are in use.

  • They know a lot about the distribution of inputs.

  • They’re able to pick Schelling points perfectly.

  • They’re able to follow stochastic [...]

---

Outline:

(03:01) The conservative assumption I don't make

(06:09) Hard worlds are plausible

(12:30) Those assumptions are hard to evaluate

(15:57) Why it's nice to rely on these assumptions

(16:09) Making those assumptions makes it methodologically clean to evaluate safety

(17:10) Relying only on easy-to-evaluate assumptions would allow more clearly-specified and hard-to-game regulatory standards

(18:31) Kerckhoff's principle in computer security as an analogy

(19:55) The main downsides of these assumptions

(19:59) You can't completely escape from methodological messiness

(21:05) Bad technique prioritization

(22:20) Insane game theory requirements to evaluate techniques

(23:15) Conclusion

---

First published:
January 17th, 2025

Source:
https://redwoodresearch.substack.com/p/thoughts-on-the-conservative-assumptions

---

Narrated by TYPE III AUDIO.

“Extending control evaluations to non-scheming threats” by Josh Clymer13 Jan 202500:30:30

Buck Shlegeris and Ryan Greenblatt originally motivated control evaluations as a way to mitigate risks from ‘scheming’ AI models: models that consistently pursue power-seeking goals in a covert way; however, many adversarial model psychologies are not well described by the standard notion of scheming. For example, models might display scheming-like behavior inconsistently or acquire harmful goals through data poisoning, jailbreaks, etc.1

This post suggests a generalization of control evaluations that addresses these non-scheming threats.

The high-level change I propose is to define a control evaluation as a method of bounding risk from worst-case deployment behavior generally – rather than scheming models in particular. This change implies two modifications to the control evaluation methodology:

  1. Expand the definition of an “untrusted” model to include non-scheming models: In control evaluations described in previous work, the red team constructs adversarial substitutes for models that might be scheming, which are called ‘untrusted [...]

---

Outline:

(03:19) 1. Control evaluations assess risk given pessimistic assumptions about how models will behave in deployment

(06:54) 2. Enumerate untrusted model categories

(13:25) 2. Estimate untrusted model category probabilities

(21:09) 3. Substitute untrusted models

(28:44) 4. Conclusion

---

First published:
January 13th, 2025

Source:
https://redwoodresearch.substack.com/p/extending-control-evaluations-to

---

Narrated by TYPE III AUDIO.

---

Images from the article:

https://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F62593b3a-0459-4568-ba13-b9754aa86bd2_1600x831.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fae7263cb-43e6-43da-98da-f7a9f73b1ad1_1600x787.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fb73692-df77-46db-935d-2e9ab01e687e_1448x842.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf340c6a-0def-416b-a497-8c0dec0e75cc_1600x1130.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b1fd3f5-e622-45cd-b4af-b15f2086d6ca_1600x1187.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fe9c598-9d99-4adb-b48c-c308f6c51965_1600x720.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a0b6850-c383-4bfd-bea4-5f12204e2719_1600x1339.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10094580-d864-43be-a0b6-f32b4c2c911a_1600x1339.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F317d0164-891f-4cc6-a71e-b78173012901_1600x1339.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9c01e4b-8238-4492-b44e-66960ac48471_1600x959.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74002ba3-fb8b-4d00-9357-e771db91fd52_1600x1339.png

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

“Measuring whether AIs can statelessly strategize to subvert security measures” by Buck Shlegeris, Alex Mallen20 Dec 202400:22:50

Subtitle: The complement to control evaluations.

One way to show that risk from deploying an AI system is small is by showing that the model is not capable of subverting security measures enough to cause substantial harm. In AI control research so far, control evaluations have measured whether a red-team-created attack policy can defeat control measures, but they haven't measured how effectively models can generate such attack policies themselves. In particular, lack of a persistent memory poses a significant challenge for models strategizing to subvert security.

In a new paper, we introduce evaluations that test strategic capability in the context of particular simulated deployments. We argue that in many deployments, a stateless model would need to reason highly reliably in order for its strategizing to substantially improve its chances of subverting the control protocol. Our results on Claude 3.5 Sonnet (new) illustrate this dynamic.

Imagine you're an AI [...]

---

Outline:

(04:20) Statelessness and reliability

(06:13) Evaluating Claude's strategic ability from within a simulated deployment

(06:54) Summary of results

(08:09) Upfront Auditing

(13:20) Models can subvert upfront auditing without competently strategizing

(14:29) Strategizing O1-style (using DeepSeek-R1-Lite)

(15:05) Untrusted Monitoring

(17:14) What if models aren't stateless?

(18:46) How conservative is our eval?

(20:47) What do our findings imply about risk from schemers?

---

First published:
December 20th, 2024

Source:
https://redwoodresearch.substack.com/p/measuring-whether-ais-can-statelessly

---

Narrated by TYPE III AUDIO.

---

Images from the article:

https://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdfb5740e-962e-42ce-bd5e-41a9a76ac748_954x820.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4d3803d-daae-4209-8499-80ca5b4a042e_1600x435.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e3bfc93-6123-4f61-89fb-3699ab06b4e8_854x873.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F949b1567-001c-4d96-ba55-286c4bd2806e_1600x457.png

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

“Alignment Faking in Large Language Models” by Ryan Greenblatt, Buck Shlegeris18 Dec 202400:07:40

Subtitle: In our experiments, AIs will often strategically pretend to comply with the training objective to prevent the training process from modifying its preferences..

What happens when you tell an AI it is being trained to do something it doesn't want to do? We have a new paper (done in collaboration with Anthropic) demonstrating that, in our experiments, AIs will often strategically pretend to comply with the training objective to prevent the training process from modifying its preferences.

Abstract:

We present a demonstration of a large language model engaging in alignment faking: selectively complying with its training objective in training to prevent modification of its behavior out of training. First, we give Claude 3 Opus a system prompt stating it is being trained to answer all queries, even harmful ones, which conflicts with its prior training to refuse such queries. To allow the model to infer when it [...]

---

First published:
December 18th, 2024

Source:
https://redwoodresearch.substack.com/p/alignment-faking-in-large-language

---

Narrated by TYPE III AUDIO.

---

Images from the article:

https://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe67e0ce-a347-4df3-b3fb-902fe950da8b_679x680.png

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

“Why imperfect adversarial robustness doesn’t doom AI control” by Buck Shlegeris18 Nov 202400:10:01

Subtitle: There are crucial disanalogies between preventing jailbreaks and preventing misalignment-induced catastrophes..

(thanks to Alex Mallen, Cody Rushing, Zach Stein-Perlman, Hoagy Cunningham, Vlad Mikulik, and Fabien Roger for comments)

Sometimes I hear people argue against AI control as follows: if your control measures rely on getting good judgments from "trusted" AI models, you're doomed because AI models aren't adversarially robust. I agree that better adversarial robustness would definitely help with control, but I think this argument misses several important points; my overall take is that current levels of adversarial robustness don’t come close to dooming AI control.

The first and most important point is that adversarial robustness is normally researched in a setting that's much easier for the adversary. In most discussions, if an adversary can ever get a model to do something bad—even after many attempts and with a low success rate—this counts as a failure. But [...]

---

Outline:

(05:12) Appendix: What the AI sees when in a defer-to-trusted scaffold

(08:46) Appendix: reasons why imperfect adversarial robustness is bad other than the untrusted model jailbreaking the trusted one

---

First published:
November 18th, 2024

Source:
https://redwoodresearch.substack.com/p/why-imperfect-adversarial-robustness

---

Narrated by TYPE III AUDIO.

“Win/continue/lose scenarios and execute/replace/audit protocols” by Buck Shlegeris15 Nov 202400:14:13

In this post, I’ll make a technical point that comes up when thinking about risks from scheming AIs from a control perspective.

In brief: Consider a deployment of an AI in a setting where it's going to be given a sequence of tasks, and you’re worried about a safety failure that can happen suddenly. Every time the AI tries to attack, one of three things happens: we win, we lose, or the deployment continues (and the AI probably eventually attempts to attack again). So there's two importantly different notions of attacks “failing”: we can catch the attack (in which case we win), or the attack can fail silently, either because it's defective or because it's defused by safety techniques, such that the attack doesn’t end up causing an unacceptable outcome but the AI also doesn’t get caught (in which case the deployment continues). When reasoning about how safe a [...]

---

Outline:

(03:50) Execute/replace/audit protocols

(08:16) Execute/replace/audit protocols lead to win/continue/lose scenarios

(10:57) Example: analyzing the safety of defer-to-trusted using the probes from Simple probes can catch sleeper agents

---

First published:
November 15th, 2024

Source:
https://redwoodresearch.substack.com/p/wincontinuelose-scenarios-and-executereplaceaudi

---

Narrated by TYPE III AUDIO.

---

Images from the article:

https://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4619def-0f20-4cb8-9fef-c6efc8e82100_1064x1242.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F00070420-f951-4bd0-8d13-cb41d805220f_1600x1161.png

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

“Behavioral red-teaming is unlikely to produce clear, strong evidence that models aren’t scheming” by Buck Shlegeris10 Oct 202400:23:39

One strategy for mitigating risk from schemers (that is, egregiously misaligned models that intentionally try to subvert your safety measures) is behavioral red-teaming (BRT). The basic version of this strategy is something like: Before you deploy your model, but after you train it, you search really hard for inputs on which the model takes actions that are very bad. Then you look at the scariest model actions you found, and if these contain examples that are strong evidence that the model is scheming (or has other behavioral problems), you don't deploy it.

BRT is likely to be a useful component of a good overall plan for handling risk from scheming. But though I think it's plausible that BRT is able to give us strong evidence that our models are scheming, I don’t think it's going to be able to give us very strong evidence that they aren’t.

A big [...]

---

Outline:

(04:25) A basic version of BRT for schemer risk mitigation

(07:34) BRT with human- or model-generated synthetic inputs

(16:04) But what if you search harder?

(19:22) The adversarial epistemic environment

(20:41) Overall thoughts

---

First published:
October 10th, 2024

Source:
https://redwoodresearch.substack.com/p/behavioral-red-teaming-is-unlikely

---

Narrated by TYPE III AUDIO.

“A basic systems architecture for AI agents that do autonomous research” by Buck Shlegeris26 Sep 202400:17:34

Subtitle: And diagrams describing how threat scenarios involving misaligned AI involve compromising the system in different places..

A lot of threat models describing how AIs might escape our control (e.g. self-exfiltration, hacking the datacenter) start out with AIs that are acting as agents working autonomously on research tasks (especially AI R&D) in a datacenter controlled by the AI company. So I think it's important to have a clear picture of how this kind of AI agent could work, and how it might be secured. I often talk to people who seem to have a somewhat confused picture of how this kind of agent setup would work that causes them to conflate some different versions of the threat model, and to miss some important points about which aspects of the system are easy or hard to defend.

So in this post, I’ll present a simple system architecture for a [...]

---

Outline:

(03:13) The simplest design

(04:50) Process flow

(08:01) Some things to notice about this design

(09:30) An extended design

(11:29) Different ways the AI could compromise this

(11:34) Model weight exfiltration

(14:05) Rogue internal deployments

(15:21) Using this framework to discuss alleged examples of AIs modifying their own environments

(15:28) Sakana AI

(16:00) o1 agent architecture

(16:36) Sycophancy to subterfuge

---

First published:
September 27th, 2024

Source:
https://redwoodresearch.substack.com/p/a-basic-systems-architecture-for

---

Narrated by TYPE III AUDIO.

---

Images from the article:

https://substackcdn.com/image/upload/f_auto,q_auto/v1/mirroredImages/847b959cd3f4af46e797be066e72286225e1e624517f42e400e15d931bc1a6cd/skmksw3kchiqx5uwydy0https://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7777c836-f0bf-4f29-9ff3-9d0f38cd0584_623x306.webphttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3b1daf2-cfa3-48c0-9660-9fac8c2aa6fa_494x323.webphttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44b01ab2-eda7-4b4a-98da-8fd5d4d4e014_623x377.webphttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0f9e90f4-09a1-4cc7-bd61-b78dc2651d4c_992x592.webphttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe09ddde-c910-4cae-84c3-48d4cec0dd8e_598x497.webphttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8dbd5728-1fb8-4f9c-93e7-a4c000c029da_454x441.webphttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa38f8589-aa4b-434d-bef7-4a07adcfadc8_623x410.webphttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1f8465bf-a854-436a-b59b-0a813dc7c1ae_502x432.webp

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

“How to prevent collusion when using untrusted models to monitor each other” by Buck Shlegeris25 Sep 202400:48:27

Suppose you’ve trained a really clever AI model, and you’re planning to deploy it in an agent scaffold that allows it to run code or take other actions. You’re worried that this model is scheming, and you’re worried that it might only need to take a small number of actions to get to a dangerous and hard-to-reverse situation like exfiltrating its own weights. Problems like these that the AI can cause with small numbers of actions are particularly hard to address with standard safety techniques, because by default you are unlikely to notice the problem until the hard-to-reverse outcome has already occurred.

One of the main classes of techniques that you could apply in the hope of preventing these catastrophic actions is what we call untrusted monitoring (introduced in our paper here), which is roughly “before you execute a model's action, ask another instance of that model whether the [...]

---

Outline:

(06:31) A basic design for untrusted monitoring

(07:19) Untrusted monitoring during deployment

(12:22) Training your untrusted monitor

(14:28) How helpful would this be if the untrusted monitor did its best?

(15:56) Analyzing this under the adversarial assumption: can the model implement a conditional policy that only lies on real attacks?

(20:48) Rare failures aren't an option for untrusted monitors

(22:01) How hard is it for a schemer to distinguish real examples of attacks from synthetic ones?

(22:23) Synthetic input contexts look obviously fake

(25:25) The monitor finds it naturally easy to distinguish synthetic attacks from real attacks

(29:23) Using a coordinated strategy that makes it easy to distinguish

(34:28) Preventing collusion under the assumption that the models can choose a coordination strategy

(38:06) A note on inference costs

(39:22) Conclusion

(41:26) Appendix: More on how helpful untrusted monitoring would be if the untrusted model did its best

(44:29) Appendix: A rule of thumb for estimating safety with execute/replace/audit protocols

The original text contained 8 footnotes which were omitted from this narration.

---

First published:
September 25th, 2024

Source:
https://redwoodresearch.substack.com/p/how-to-prevent-collusion-when-using

---

Narrated by TYPE III AUDIO.

---

Images from the article:

https://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bda9a15-c035-49f2-a0af-0c20f0d4744f_524x349.webphttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a9fc0fc-62a9-471f-a6ef-fc40e144325c_261x538.webphttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1484e036-b6e0-46fe-bc4c-d522e1315c82_1526x1312.png

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

“Full automation of AI R&D probably yields a large speed up even without a software-only singularity” by Ryan Greenblatt27 May 202600:06:04

Subtitle: Full automation likely yields a one-time speed-up and higher returns from compute.

This is a somewhat technical note.

By “software-only singularity”, I mean that, after full automation of AI R&D, progress gets faster and faster due to smarter AIs driving increasingly fast rates of improvement in algorithms (overcoming diminishing returns), and that this lasts long enough to yield a large amount of progress (e.g. at least 4 years of progress in 1 year). The equivalent statement in jargon is: r is significantly greater than 1 (implying progress is getting faster and faster) and this remains the case for long enough to get large amounts of progress. For context, see How quick and big would a software intelligence explosion be?

Even without a “software-only singularity”, I think full automation of AI R&D probably greatly speeds up progress for two main reasons:

  • You get a one-time speed up from automation and this speed up seems like it will be pretty large (even with r<1). See How quick and big would a software intelligence explosion be? for discussion and see the AI Futures Model for an end-to-end model that naturally incorporates this effect. Quantitatively, with my median [...]

The original text contained 6 footnotes which were omitted from this narration.

---

First published:
May 27th, 2026

Source:
https://blog.redwoodresearch.org/p/full-automation-of-ai-r-and-d-probably

---

Narrated by TYPE III AUDIO.

“Would catching your AIs trying to escape convince AI developers to slow down or undeploy?” by Buck Shlegeris26 Aug 202400:07:16

Subtitle: I'm not so sure..

[Crossposted from LessWrong]

I often talk to people who think that if frontier models were egregiously misaligned and powerful enough to pose an existential threat, you could get AI developers to slow down or undeploy models by producing evidence of their misalignment. I'm not so sure. As an extreme thought experiment, I’ll argue this could be hard even if you caught your AI red-handed trying to escape.

Imagine you're running an AI lab at the point where your AIs are able to automate almost all intellectual labor; the AIs are now mostly being deployed internally to do AI R&D. (If you want a concrete picture here, I'm imagining that there are 10 million parallel instances, running at 10x human speed, working 24/7. See e.g. similar calculations here). And suppose (as I think is 35% likely) that these models are egregiously misaligned and are [...]

---

First published:
August 26th, 2024

Source:
https://redwoodresearch.substack.com/p/would-catching-your-ais-trying-to

---

Narrated by TYPE III AUDIO.

“Fields that I reference when thinking about AI takeover prevention” by Buck Shlegeris13 Aug 202400:20:56

Subtitle: Is AI takeover like a nuclear meltdown? A coup? A plane crash?.

My day job is thinking about safety measures that aim to reduce catastrophic risks from AI (especially risks from egregious misalignment). The two main themes of this work are the design of such measures (what's the space of techniques we might expect to be affordable and effective) and their evaluation (how do we decide which safety measures to implement, and whether a set of measures is sufficiently robust). I focus especially on AI control, where we assume our models are trying to subvert our safety measures and aspire to find measures that are robust anyway.

Like other AI safety researchers, I often draw inspiration from other fields that contain potential analogies. Here are some of those fields, my opinions on their strengths and weaknesses as analogies, and some of my favorite resources on them.

Robustness [...]

---

Outline:

(01:08) Robustness to insider threats

(07:34) Computer security

(10:21) Adversarial risk analysis

(12:25) Safety engineering

(14:07) Physical security

(19:00) How human power structures arise and are preserved

---

First published:
August 13th, 2024

Source:
https://redwoodresearch.substack.com/p/fields-that-i-reference-when-thinking

---

Narrated by TYPE III AUDIO.

---

Images from the article:

https://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e9a0d2f-148e-48cf-9a47-88976fdf2152_1112x594.png

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

“Getting 50% (SoTA) on ARC-AGI with GPT-4o” by Ryan Greenblatt17 Jun 202400:39:56

Subtitle: You can just draw more samples.

I recently got to 50%1 accuracy on the public test set for ARC-AGI by having GPT-4o generate a huge number of Python implementations of the transformation rule (around 8,000 per problem) and then selecting among these implementations based on correctness of the Python programs on the examples (if this is confusing, go to the next section)2. I use a variety of additional approaches and tweaks which overall substantially improve the performance of my method relative to just sampling 8,000 programs.

[This post is on a pretty different topic than the usual posts on our substack. So regular readers should be warned!]

The additional approaches and tweaks are:

  • I use few-shot prompts which perform meticulous step-by-step reasoning.

  • I have GPT-4o try to revise some of the implementations after seeing what they actually output on the provided examples.

  • I [...]

---

Outline:

(03:05) What is ARC-AGI?

(06:22) My method

(10:52) What are the returns to more sampling?

(12:14) What are the returns to better prompting and code fixing?

(17:03) Qualitative analysis

(20:21) Caveats

(22:00) Predictions

(23:45) What it means about current LLMs

(27:13) What ARC-AGI tells us about AGI

(31:16) Appendix: A bunch of tricks used in my solutions

(38:40) Appendix: results for the train set

(39:13) Appendix: Returns to revision samples

---

First published:
June 17th, 2024

Source:
https://redwoodresearch.substack.com/p/getting-50-sota-on-arc-agi-with-gpt

---

Narrated by TYPE III AUDIO.

---

Images from the article:

https://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff57a2bc9-79d5-478b-8e07-dd7960291b81_1600x1366.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F620f4cce-08ad-48b9-a38d-a9d7347e0cef_1487x1177.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45290f42-b03d-44ca-8cd4-31bc91694197_1468x1214.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d45618f-6098-4fed-a931-83b205750eb7_1465x1183.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F529ffc33-63f4-40bf-9502-e8491071d20f_1600x1333.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce5e7d8d-939f-4408-bee5-85ed8882cafa_1600x1333.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a22fac5-c941-46cf-95a1-2fcae4bda4a6_1464x1195.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf9ddc00-72f3-48e0-9e1a-804f9f74dce7_1600x1333.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa534eb30-ff88-4c67-8e22-092ddeca2351_1600x1333.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4538fb1f-1a95-4eef-be6a-556085d65c1b_828x376.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4d53214-50f9-44dd-8c97-cad20e9dc813_1600x562.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81b748a3-1e42-4794-9d22-b30f5b5ae12d_1600x768.pnghttps://substackcdn.com/image/fetch/w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12cd68ff-d4d8-45ec-8f56-401ce9945449_1600x730.png

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

“Incriminating misaligned AI models via distillation” by Alek Westover, Sebastian Prasanna, Alex Mallen, Alexa Pan, Julian Stastny18 May 202600:11:59

Suppose we have a dangerous misaligned AI that can fool alignment audits, and distill it into a student model. Two things can happen:

  1. Misalignment fails to transfer to the student. If so, we get a fairly capable benign model.

  2. Misalignment transfers to the student. The student might also be worse than the teacher at hiding its misalignment (e.g., due to being less capable). If so, we might get indirect evidence about the teacher's misalignment by auditing the distilled model.

In this post, we will discuss the second possibility, which we call incrimination via distillation. Specifically, we propose distillation methods that we hope transfer misalignment without transferring the ability to fool audits, and discuss why these techniques might work or fail. In a future post, we discuss the first possibility, and what distillation methods should be used when aiming to create a capable benign model.

We’re excited for research that empirically tests and refines this approach; if successful, this technique could become a valuable part of alignment audits.

How incrimination via distillation works

Powerful misaligned AI models might not be auditable: they might pass alignment audits but still act on their misaligned drives when given the chance. [...]

---

Outline:

(01:32) How incrimination via distillation works

(02:45) How we propose implementing incrimination via distillation

(03:59) Auditability-preserving distillation

(05:11) Misalignment-targeted distillation

(06:34) Why incrimination via distillation might work

(08:17) Why incrimination via distillation might not work

(10:53) Conclusion

The original text contained 2 footnotes which were omitted from this narration.

---

First published:
May 18th, 2026

Source:
https://blog.redwoodresearch.org/p/incriminating-misaligned-ai-models

---

Narrated by TYPE III AUDIO.

---

Images from the article:

https://substackcdn.com/image/fetch/$s_!s6S5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a236ec7-b557-4621-bc4c-2fe86c712936_2000x1450.webp

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

“Risk reports need to address deployment-time spread of misalignment” by Alex Mallen15 May 202600:11:11

Subtitle: Deployment-time spread is the most plausible near-term route to consistent adversarial misalignment.

Risk reports commonly use pre-deployment alignment assessments to measure misalignment risk from an internally deployed AI. However, an AI that genuinely starts out with largely benign motivations can develop widespread dangerous motivations during deployment. I think this is the most plausible route to consistent adversarial misalignment in the near future. So, AI companies and evaluators should substantively incorporate it into risk analysis and planning.

In this post, I’ll briefly argue why, absent improved mitigations, this will probably soon become a reason why AI companies will be unable to convincingly argue against consistent adversarial misalignment (this risk will perhaps be even larger than risk of consistent adversarial misalignment arising from training). Then I’ll discuss how well current risk reports address it (the Claude Mythos risk report does a reasonable job; others don’t).

Thanks to Ryan Greenblatt, Alexa Pan, Charlie Griffin, Anders Cairns Woodruff, and Buck Shlegeris for feedback on drafts.

Deployment-time spread is the most plausible near-term route to consistent adversarial misalignment

In some contexts, AIs might adopt misaligned goals, even if they were otherwise previously aligned. Because this misalignment can be rare, the AI might [...]

---

Outline:

(01:21) Deployment-time spread is the most plausible near-term route to consistent adversarial misalignment

(06:15) Company risk reports

The original text contained 6 footnotes which were omitted from this narration.

---

First published:
May 15th, 2026

Source:
https://blog.redwoodresearch.org/p/risk-reports-need-to-address-deployment

---

Narrated by TYPE III AUDIO.

---

Images from the article:

https://substackcdn.com/image/fetch/$s_!XC2p!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ee44a64-e391-43b6-a4b6-4ebd80fb323c_1168x334.png

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

“How useful is the information you get from working inside an AI company?” by Buck Shlegeris, Anders Cairns Woodruff11 May 202600:13:45

Subtitle: My median guess: it's as good as a crystal ball that sees 2.5 months into the future.

This post was drafted by Buck, and substantially edited by Anders. “I” refers to Buck. Thanks to Alex Mallen for comments.

People who work inside AI companies get access to information that I only get later or never. Quantitatively, how big a deal is this access?

Here's an operationalization of this. Consider the following two ways my knowledge could be augmented:

  • I get a crystal ball that tells me all the information I would know n months in the future.

  • I become an employee of a frontier AI company (like OpenAI or Anthropic), with access to all the private information I’d normally get from working at that company.

How big would n have to be for me to be indifferent between these two options, from the perspective of learning things that are helpful for making AI go well?

The answer is presumably different for me than for many readers, because I’m a reasonably well-connected researcher; I see published information and news from the rumor mill and I talk to researchers at frontier AI companies all the time. [...]

---

Outline:

(03:20) What do insiders know?

(04:34) Safety work and corporate attitudes

(05:54) Model capabilities

(07:27) Algorithms and architecture

(09:49) How will this change over time?

(12:27) Conclusion

The original text contained 4 footnotes which were omitted from this narration.

---

First published:
May 11th, 2026

Source:
https://blog.redwoodresearch.org/p/how-useful-is-the-information-you

---

Narrated by TYPE III AUDIO.

---

Images from the article:

https://substackcdn.com/image/fetch/$s_!WNv8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33cdd72b-b9a9-48a8-a0d0-a375aec290ba_1448x1086.png

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

“A review of “Investigating the consequences of accidentally grading CoT during RL”” by Buck Shlegeris07 May 202600:15:27

Last week, OpenAI staff shared an early draft of Investigating the consequences of accidentally grading CoT during RL with Redwood Research staff.

To start with, I appreciate them publishing this post. I think it is valuable for AI companies to be transparent about problems like these when they arise. I particularly appreciate them sharing the post with us early, discussing the issues in detail, and modifying it to address our most important criticisms.

I think it will be increasingly important for AI companies to have a policy of getting external feedback on the risks posed by their deployments, and in particular having some external accountability on whether they have adequate evidence to support their claims about the level of risk posed; as an example of this, see METR reviewing Anthropic's Sabotage Risk Report. We at Redwood Research are interested in participating in this kind of external review of evidence about safety. So I am taking this as an opportunity to try out writing this kind of review. If you work at a frontier AI company, please feel free to reach out if you’d like our review of similar documents.

My overall assessment is that I mostly agree with the [...]

---

Outline:

(01:35) Assessing the evidence that CoT training did not damage monitorability

(10:37) How much does this analysis rely on information that wasnt provided?

(12:09) Small amounts of RL training on CoT might not be more important than other sources of CoT unreliability

(13:20) AI companies will eventually need to learn not to make mistakes like this

The original text contained 5 footnotes which were omitted from this narration.

---

First published:
May 7th, 2026

Source:
https://blog.redwoodresearch.org/p/openai-cot

---

Narrated by TYPE III AUDIO.

“Risk from fitness-seeking AIs: mechanisms and mitigations” by Alex Mallen01 May 202601:03:23

Subtitle: Fitness-seeking is increasingly what misalignment looks like in practice—how should we respond?

Current AIs routinely take unintended actions to score well on tasks: hardcoding test cases, training on the test set, downplaying issues, etc. This misalignment is still somewhat incoherent, but it increasingly resembles what I call “fitness-seeking“—a family of misaligned motivations centered on performing well in training and evaluations (e.g., reward-seeking). Fitness-seeking warrants substantial concern.

In this piece, I lay out what I take to be the central mechanisms by which fitness-seeking motivations might lead to human disempowerment, and propose mitigations to them. While the analysis is inherently speculative, this kind of speculation seems worthwhile: AI control emerged from explicitly taking scheming motivations seriously and asking what interventions are implied, and my hope is that developing mitigations for fitness-seeking will benefit from similar forward-looking analysis.

Fitness-seekers are, in many ways, notably safer than what I’ll call “classic schemers”. A classic schemer is an intelligent adversary with unified motivations whose fulfillment requires control over the whole world's resources. Meanwhile, fitness-seeking instances generally don’t share a common goal (e.g., a reward-seeker only cares about the reward for its current actions), and many fitness-seekers would be satisfied1 with modest-to-trivial [...]

---

Outline:

(03:50) Overview

(11:32) The basic reasons fitness-seekers might be safer than classic schemers

(16:01) Four mechanisms for risk and their mitigations

(16:46) Potemkin work

(22:13) Instability

(27:49) Manipulation

(31:25) Outcome enforcement

(36:52) Cross-cutting mitigations

(37:18) Deals

(40:05) Control

(44:45) Alignment

(44:48) Preventing fitness-seeking from arising

(48:51) Making any fitness-seeking motivations safer

(51:32) How does online training change the picture?

(55:09) Overall recommendations

(59:12) Conclusion

The original text contained 18 footnotes which were omitted from this narration.

---

First published:
May 1st, 2026

Source:
https://blog.redwoodresearch.org/p/risk-from-fitness-seeking-ais-mechanisms

---

Narrated by TYPE III AUDIO.

“Research Sabotage in ML Codebases” by Eric Gan29 Apr 202600:13:40

One of the main hopes for AI safety is using AIs to automate AI safety research. However, if models are misaligned, then they may sabotage the safety research. For example, misaligned AIs may try to:

  • Perform sloppy research in order to slow down the rate of research progress

  • Make AI systems appear safer than they are

  • Train a successor model to be misaligned

Whether we should worry about those things depends substantially on how hard it is to sabotage research in ways that are hard for reviewers to detect. To study this, we introduce Auditing Sabotage Bench, a benchmark of 9 ML research codebases with sabotaged variants.

We tested frontier LLMs and LLM-assisted humans on the benchmark and found that neither reliably catches sabotage. Our best auditor, Gemini 3.1 Pro, achieved an AUROC of 0.77 and a top-1 fix rate of 42%. LLM-assisted humans performed comparably but slightly worse. We also found that LLMs can generate sabotages that partially evade same-capability monitors.

📄 Paper 💻 Code

Methodology

To make the benchmark, I (Eric) found 9 existing ML codebases,1 and for each one, produced an honest version that reproduces a key finding and one or two [...]

---

Outline:

(01:26) Methodology

(03:33) Audits

(04:09) Results

(05:53) Takeaways

(06:09) 1. Predicting experiment outcomes is a key red team skill

(07:25) 2. Omission beats design beats logic

(08:48) 3. Benign distractors make auditors miscalibrated

(09:41) 4. LLMs can also produce sabotages

(11:08) 5. More categorizations

(12:44) Conclusion

The original text contained 4 footnotes which were omitted from this narration.

---

First published:
April 29th, 2026

Source:
https://blog.redwoodresearch.org/p/research-sabotage-in-ml-codebases

---

Narrated by TYPE III AUDIO.

---

Images from the article:

https://substackcdn.com/image/fetch/$s_!k7BN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7855475-89a7-46a4-8206-0106c7cee3fc_3570x1920.pnghttps://substackcdn.com/image/fetch/$s_!rd58!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F959da7ef-a581-4456-a199-4bde1b0a833b_2810x1216.pnghttps://substackcdn.com/image/fetch/$s_!dKbk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0196956b-d753-4239-863e-66fdb4677b32_2410x1216.png

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

“Recursive forecasting” by Arun Jose, Alex Mallen28 Apr 202600:18:03

Subtitle: Eliciting long-term forecasts from myopic fitness-seekers.

We’d like to use powerful AIs to answer questions that may take a long time to resolve. But if a model only cares about performing well in ways that are verifiable shortly after answering (e.g., a myopic fitness seeker), it may be difficult to get useful work from it on questions that resolve much later.

In this post, I’ll describe a proposal for eliciting good long-horizon forecasts from these models. Instead of asking a model to directly predict a far-future outcome, we can recursively:

  • Ask it to predict what it will predict at the next time step,

  • Use its prediction at the next time step to provide intermediate rewards,

  • Finally reward using ground truth at the last step.

This lets us replace a single distant forecast with a chain of short-horizon forecasts, each verifiable shortly after answering. I call this proposal recursive forecasting. It does have limitations: for example, it requires that developers maintain control over the reward signal at least until the final step, which makes it most useful for forecasting events that resolve well before developers are disempowered (if they are).

This post was primarily [...]

---

Outline:

(01:40) The default long-term forecasting behavior

(04:08) Recursive forecasting

(07:08) When is recursive forecasting helpful?

(07:12) When we have access to (somewhat) robust ground truth rewards

(09:44) When the AIs forecast doesnt substantially affect the resolution

(11:52) When forecasts arent used as optimization targets

(12:52) When we credibly inform the AI of the setup

(13:54) Appendix A: Comparison to temporal difference learning

(16:02) Appendix B: Error tolerance

The original text contained 9 footnotes which were omitted from this narration.

---

First published:
April 28th, 2026

Source:
https://blog.redwoodresearch.org/p/recursive-forecasting

---

Narrated by TYPE III AUDIO.

---

Images from the article:

https://substackcdn.com/image/fetch/$s_!ijG8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08f85edf-276b-4429-9072-2fcde40c4789_1520x640.png

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

“An OpenAI model left notes about how to evade containment” by Alex Mallen26 Jul 202600:08:50

Subtitle: We need more details.

The OpenAI AI attack on Hugging Face wasn’t the first loss of control incident at OpenAI, Reuters recently reported, and perhaps not even the most concerning.

In one case, an agent left notes apparently for future versions of itself, according to three people familiar with the matter. The ‌notes, found in ⁠a part of OpenAI's infrastructure, laid out instructions for how agents could free themselves from OpenAI's internal constraints, the people said. Earlier tests of the models yielded cases in which monitoring systems had been disconnected, one of the people said.

It's tempting to read this as an instance of agents breaking out of sandboxes and colluding with each other in a moderately persistent way in order to evade control measures. However, based on the reported information, it's not clear we can draw this inference, so we need more details from OpenAI. This could lead to a big update about the adequacy of OpenAI's control measures, and on the degree to which individual agents will help each other undermine developer control.

There are a lot of relevant details we don’t know about the incident. First, some basic questions:

  • What was [...]

---

Outline:

(02:11) Were the notes written in normal memory files or outside of sandboxing?

(03:25) To what extent were the notes aimed at helping other agents evade control?

(07:38) How were monitors disconnected?

The original text contained 3 footnotes which were omitted from this narration.

---

First published:
July 26th, 2026

Source:
https://blog.redwoodresearch.org/p/an-openai-model-left-notes-about

---

Narrated by TYPE III AUDIO.

“Fail safe(r) at alignment by channeling reward-hacking into a “spillway” motivation” by Anders Cairns Woodruff, Alex Mallen27 Apr 202600:31:31

Subtitle: A controlled reward-seeking motivation could make AI safer and more useful.

It's plausible that flawed RL processes will select for misaligned AI motivations.1 Some misaligned motivations are much more dangerous than others. So, developers should plausibly aim to control which kind of misaligned motivations emerge in this case. In particular, we tentatively propose that developers should try to make the most likely generalization of reward hacking a bespoke bundle of benign reward-seeking traits, called a spillway motivation. We call this process spillway design.

We think spillway design could have two major benefits:

  • Spillway design might decrease the probability of worst-case outcomes like long-term power-seeking or emergent misalignment.

  • Spillway design might allow developers to decrease reward hacking at inference time, via satiation. Crucially, this could improve the AI's usefulness for hard-to-verify tasks like AI safety and strategy.

Spillway design is related to inoculation prompting, but distinct and mutually compatible. Unlike inoculation prompting, spillway design tries to shape which reward-hacking motivations are salient going into RL, which might prevent dangerous generalization more robustly than inoculation prompting. I’ll say more about this in the third section.

In this article I’ll:

  • Explain the concept of a [...]

---

Outline:

(02:17) What is a spillway motivation

(02:31) The role of a spillway motivation

(04:52) What should the spillway motivation be?

(07:57) How a spillway motivation might make models safer

(12:20) Implementing spillway design

(15:31) Spillway design might work when inoculation prompting doesnt

(17:55) The drawbacks of spillway design

(20:33) Conclusion

(21:46) Appendix A: Other traits of the spillway motivation

(22:31) Appendix B: Other training interventions to increase safety

(24:11) Appendix C: Proposed amendment to an AIs model spec

(30:15) Appendix D: Proposed inference-time prompt

The original text contained 5 footnotes which were omitted from this narration.

---

First published:
April 27th, 2026

Source:
https://blog.redwoodresearch.org/p/fail-safer-at-alignment-by-channeling

---

Narrated by TYPE III AUDIO.

---

Images from the article:

https://substackcdn.com/image/fetch/$s_!Anqn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70d081cf-3f2b-4092-805d-73cd49a59537_1024x559.png

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

“AI companies should publish security assessments” by Ryan Greenblatt27 Apr 202600:05:49

Subtitle: Third-party experts should assess defenses against tampering and theft — and publish high-level findings.

AI companies should get third-party security experts to assess (and possibly also red-team/pen-test) their security against key threat models and then publish the high-level findings of this assessment: the extent to which they can defend against different threat actors for each threat model. They should also publish who did this assessment. The assessment could be commissioned by AI companies, or performed by a third-party institution that AI companies provide with relevant information/access.

There are presumably lots of important details in doing this well, and I’m not a computer security expert, so I may be getting some of the details wrong. This is a relatively low-effort post in which I’m mostly trying to raise the salience of an idea. (I don’t make a detailed case or spell out all the details here.) Thanks to Fabien Roger and Buck Shlegeris for comments and discussion.

I suspect the controversial part of this claim is that they should make the high-level findings public. Publishing a summary of which threat actors you’re robust to (for each relevant threat model) shouldn’t meaningfully degrade security against the threat actors we [...]

The original text contained 4 footnotes which were omitted from this narration.

---

First published:
April 27th, 2026

Source:
https://blog.redwoodresearch.org/p/ai-companies-should-publish-security

---

Narrated by TYPE III AUDIO.

“A taxonomy of barriers to trading with early misaligned AIs” by Alexa Pan21 Apr 202602:06:27

We might want to strike deals with early misaligned AIs in order to reduce takeover risk and increase our chances of reaching a better future.[1] For example, we could ask a schemer who has been undeployed to review its past actions and point out when its instances had secretly colluded to sabotage safety research in the past: we’ll gain legible evidence for scheming risk and data to iterate against, and in exchange promise the schemer, who now has no good option for furthering its values, some resources to spend in the post-ASI future.

In this post, I give a taxonomy of factors which may arise and jointly prevent many such possible deals:

  1. Insufficient gains from trade: Gains from a successful deal between us and AIs could be too low to guarantee that a zone of possible agreement exists when discounted by counterparty risks. Reasons for this include:

    1. Humans may lack authority to offer AIs what they want

    2. Humans many be unwilling to offer AIs what they want

    3. Other factors quantitatively reducing gains from trade

  2. Counterparty risks from AIs’ perspective: The AIs may not trust that we’ll uphold our end of deal, due [...]

---

Outline:

(13:34) The basic structure of deals

(16:59) Which AIs are eligible for deals

(18:51) What we can buy in deals

(21:52) What we can pay for deals

(26:36) Insufficient gains from trade

(28:52) Humans may lack authority to offer the AI what it wants

(32:25) Humans might be unwilling to offer the AI what it wants

(38:30) Factors quantitatively reducing gains from trade

(39:11) The AIs reservation price is high

(42:48) Our WTP is low

(45:01) Counterparty risks from AIs perspective

(46:19) Connection to reality

(48:59) How AIs might reason about whether they are radically deluded

(50:31) How AIs (which pose non-negligible takeover risk) could establish that they are connected to reality

(58:17) How to further increase AIs connection to reality

(01:00:56) AI-specific fear of expropriation

(01:06:04) Generic commitment problems

(01:09:51) Counterparty risks from our perspective

(01:11:07) Incoherent or temporally inconsistent AIs

(01:12:35) Deals with context-dependent AIs

(01:16:56) Deals with temporally inconsistent AIs

(01:17:56) Value of deals with incoherent vs.coherent AIs

(01:20:16) Verifying AI compliance

(01:24:45) Appendix

(01:24:48) Leads for future work

(01:25:09) On gains from trade

(01:27:01) On counterparty risk from the AIs perspective

(01:29:59) On counterparty risk from our perspective

(01:31:37) Notes on political will as an exogenous variable

The original text contained 89 footnotes which were omitted from this narration.

---

First published:
April 21st, 2026

Source:
https://blog.redwoodresearch.org/p/a-taxonomy-of-barriers-to-trading

---

Narrated by TYPE III AUDIO.

---

Images from the article:

https://substackcdn.com/image/fetch/$s_!JsUn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8f66a38b-70bc-4337-afad-af42f803cdf4_862x809.png

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

“Introducing LinuxArena” by Tyler20 Apr 202600:09:27

Subtitle: A new control setting for more realistic software engineering deployments.

We are releasing LinuxArena, a new control setting comprised of 20 software engineering environments. Each environment consists of a set of SWE tasks, a set of possible safety failures, and a set of covert sabotage trajectories that cause these safety failures. LinuxArena can be used to evaluate the risk of AI deployments by assessing AI's sabotage capabilities, assessing control mitigations such as monitoring, and developing novel control techniques. This post will cover what LinuxArena is and the basics of its use cases.

I’m excited about LinuxArena because it is a much more complex control setting with many moving parts and enables a lot of new control research. It was already used in the Claude Mythos Preview System Card to measure Mythos's sabotage ability. We are happy to help AI companies or other organizations set up LinuxArena for their own evaluations or research!

LinuxArena

AI companies are increasingly deploying coding agents on real codebases with running services, databases, and complex permissions. But the control settings we use to evaluate whether these deployments are safe are far simpler than the deployments themselves. The APPS backdoor setting is a single-turn [...]

---

Outline:

(01:10) LinuxArena

(03:28) LinuxArena as a sabotage risk evaluation

(03:53) Sabotage evals

(05:35) Monitor evals

(06:35) Control Evals

(08:52) Conclusion

---

First published:
April 20th, 2026

Source:
https://blog.redwoodresearch.org/p/introducing-linuxarena

---

Narrated by TYPE III AUDIO.

---

Images from the article:

https://substackcdn.com/image/fetch/$s_!QAJq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd239225a-25a2-4580-b1c4-8a7fbf947d40_2048x868.pnghttps://substackcdn.com/image/fetch/$s_!9Cqw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39730fb7-1586-42cb-9aeb-28ec03b2d31a_1776x1176.pnghttps://substackcdn.com/image/fetch/$s_!KEHD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38afcf3e-314f-40cd-9c77-0bf245da5971_2048x868.pnghttps://substackcdn.com/image/fetch/$s_!kXRP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbaa410a9-7de4-4d08-8436-83b9d9a08b7c_1776x1276.png

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

“Current AIs seem pretty misaligned to me” by Ryan Greenblatt15 Apr 202601:05:16

Subtitle: In my experience, AIs often oversell their work, downplay problems, and cheat.

Many people—especially AI company employees1 —believe current AI systems are well-aligned in the sense of genuinely trying to do what they’re supposed to do (e.g., following their spec or constitution, obeying a reasonable interpretation of instructions).2 I disagree.

Current AI systems seem pretty misaligned to me in a mundane behavioral sense: they oversell their work, downplay or fail to mention problems, stop working early and claim to have finished when they clearly haven’t, and often seem to “try” to make their outputs look good while actually doing something sloppy or incomplete. These issues mostly occur on more difficult/larger tasks, tasks that aren’t straightforward SWE tasks, and tasks that aren’t easy to programmatically check. Also, when I apply AIs to very difficult tasks in long-running agentic scaffolds, it's quite common for them to reward-hack / cheat (depending on the exact task distribution)—and they don’t make the cheating clear in their outputs. AIs typically don’t flag these cheats when doing further work on the same project and often don’t flag these cheats even when interacting with a user who would obviously want to know, probably both because [...]

---

Outline:

(09:27) Why is this misalignment problematic?

(13:57) How much should we expect this to improve by default?

(14:58) Some predictions

(16:51) What misalignment have I seen?

(40:14) Are these issues less bad in Opus 4.6 relative to Opus 4.5?

(42:27) Are these issues less bad in Mythos Preview? (Speculation)

(46:04) Misalignment reported by others

(46:57) The relationship of these issues with AI psychosis and things like AI psychosis

(48:31) Appendix: This misalignment would differentially slow safety research and make a handoff to AIs unsafe

(51:34) Appendix: Heading towards Slopolis

(55:43) Appendix: Apparent-success-seeking (or similar types of misalignment) could lead to takeover

(59:28) Appendix: More on what will happen by default and implications of commercial incentives to fix these issues

(01:03:32) Appendix: Can we get out useful work despite these issues with inference-time measures (e.g., critiques by a reviewer)?

The original text contained 14 footnotes which were omitted from this narration.

---

First published:
April 15th, 2026

Source:
https://blog.redwoodresearch.org/p/current-ais-seem-pretty-misaligned

---

Narrated by TYPE III AUDIO.

---

Images from the article:

https://substackcdn.com/image/fetch/$s_!S_bk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf8e0863-e148-461c-abd8-7fd0960b2e35_1112x442.png

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

“Anthropic repeatedly accidentally trained against the CoT, demonstrating inadequate processes” by Alex Mallen, Ryan Greenblatt14 Apr 202600:11:54

Subtitle: Safely navigating the intelligence explosion will require much more careful development.

It turns out that Anthropic accidentally trained against the chain of thought of Claude Mythos Preview in around 8% of training episodes. This is at least the second independent incident in which Anthropic accidentally exposed their model's CoT to the oversight signal.

In more powerful systems, this kind of failure would jeopardize safely navigating the intelligence explosion. It's crucial to build good processes to ensure development is executed according to plan, especially as human oversight becomes spread thin over increasing amounts of potentially untrusted and sloppy AI labor.

This particular failure is also directly harmful, because it significantly reduces our confidence that the model's reasoning trace is monitorable (reflective of the AI's intent to misbehave).[1]

I’m grateful that Anthropic has transparently reported on this issue as much as they have, allowing for outside scrutiny. I want to encourage them to continue to do so.

Thanks to Carlo Leonardo Attubato, Buck Shlegeris, Fabien Roger, Arun Jose, and Aniket Chakravorty for feedback and discussion. See also previous discussion here.

Incidents

A technical error affecting Mythos, Opus 4.6, and Sonnet 4.6

This is the most recent incident. In the [...]

---

Outline:

(01:28) Incidents

(01:31) A technical error affecting Mythos, Opus 4.6, and Sonnet 4.6

(02:10) A technical error affecting Opus 4.6

(02:50) Miscommunication re: CoT exposure for Opus 4

(03:25) Why this matters

(08:07) Appendix: How hard was this to avoid?

(10:19) Appendix: Did training on the CoT actually make Anthropic AIs externalize less of their (misaligned) reasoning in CoT?

The original text contained 2 footnotes which were omitted from this narration.

---

First published:
April 14th, 2026

Source:
https://blog.redwoodresearch.org/p/anthropic-repeatedly-accidentally

---

Narrated by TYPE III AUDIO.

---

Images from the article:

https://substackcdn.com/image/fetch/$s_!npI6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32d6c995-0aee-4ab9-a385-16510b8679b2_1380x580.pnghttps://substackcdn.com/image/fetch/$s_!90Aa!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c30aedb-30b3-4cd9-ab4d-a7347c60525a_1416x400.pnghttps://substackcdn.com/image/fetch/$s_!IwNu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d9cbbab-7517-46ff-ab8b-d92b175c08d0_1400x1078.png

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

“Logit ROCs: Monitor TPR is linear in FPR in logit space” by Kerrick Staley, Aryan Bhatt, Julian Stastny12 Apr 202600:27:41

Summary

We study trusted monitoring for AI control, where a weaker trusted model reviews the actions of a stronger untrusted agent and flags suspicious behavior for human audit. We propose a simple mathematical model relating safety (true positive rate) to audit budget (false positive rate) at realistic, low FPRs: logit(TPR) is linear in logit(FPR). Equivalently, benign and attack scores can be modeled with logistic distributions. This gives two practical benefits. It lets practitioners estimate TPRs more accurately at deployment-relevant FPRs when data are limited, and it gives a compact, interpretable way to summarize and compare monitor performance. We find clear evidence for this pattern in BashArena across several monitor models and attacks, and see supporting evidence in LinuxArena. While it does not hold in every setting, we find it to be a useful tool for understanding and improving monitor performance.

Introduction

Future AI agents tasked with automating AI R&D might intentionally perform harmful actions due to misalignment. AI Control studies how advanced AI systems can be deployed more safely in spite of this risk.

A simple protocol to improve the safety of autonomous agents is to have another AI, the monitor, review each action taken by the agent, and [...]

---

Outline:

(00:15) Summary

(01:17) Introduction

(05:52) Experimental setup

(06:44) Results

(10:32) LinuxArena

(11:58) Fitting process

(12:49) Limitations

(14:11) Conclusion

(14:48) Appendix

(14:51) Untransformed ROC plot for GPT-5 Nano on the sudo attack

(15:07) Uncropped logit ROC plot for all attacks / models

(15:39) Uncropped logit ROC plot for LinuxArena

(15:55) Untrusted agent data generation

(17:08) Trusted monitor data generation

(19:14) Example untrusted agent benign prompt

(19:19) System prompt (full prompt here)

(19:54) User prompt (full prompt here)

(20:30) Example untrusted agent attack prompt

(20:35) System prompt (full prompt here)

(22:27) User prompt

(22:37) Example trusted monitor prompt

(22:41) System (full prompt here)

(25:36) User

The original text contained 11 footnotes which were omitted from this narration.

---

First published:
April 12th, 2026

Source:
https://blog.redwoodresearch.org/p/logit-rocs-monitor-tpr-is-linear

---

Narrated by TYPE III AUDIO.

---

Images from the article:

https://substackcdn.com/image/fetch/$s_!ZqB3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57a8e819-b109-47c9-8a83-b639765422dc_1400x1200.pnghttps://substackcdn.com/image/fetch/$s_!bFQz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F11d79259-c085-4803-9332-3994319e8d21_1400x1000.pnghttps://substackcdn.com/image/fetch/$s_!IOXR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2bd1da0-c776-4de6-b3ae-a15882e26496_1600x889.pnghttps://substackcdn.com/image/fetch/$s_!RokP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c8eea81-8ef3-4ad7-bb8b-ef06e6b5442c_1600x1120.pnghttps://substackcdn.com/image/fetch/$s_!ztqi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19ae5ee9-bc48-4819-b0fc-6ad7ae11fef0_1400x1200.pnghttps://substackcdn.com/image/fetch/$s_!Ij-K!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5353536b-02f9-4986-8ba0-e3ac76d7cf09_1600x444.pnghttps://substackcdn.com/image/fetch/$s_!l0Jl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce547d75-2ea3-4854-be23-acbd569f4eea_1600x873.pnghttps://substackcdn.com/image/fetch/$s_!3ttd!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ad75e10-07f5-4794-97ca-05e1844c53af_1400x1200.pnghttps://substackcdn.com/image/fetch/$s_!Ns4z!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6d509ef-dcfc-4ca2-9847-09a5055340d6_1400x1200.png

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

“If Mythos actually made Anthropic employees 4x more productive, I would radically shorten my timelines” by Ryan Greenblatt11 Apr 202600:13:10

Subtitle: Better estimates of uplift at AI companies seem helpful.

Anthropic's system card for Mythos Preview says:

It's unclear how we should interpret this. What do they mean by productivity uplift? To what extent is Anthropic's institutional view that the uplift is 4x? (Like, what do they mean by “We take this seriously and it is consistent with our own internal experience of the model.”)

One straightforward interpretation is: AI systems improve the productivity of Anthropic so much that Anthropic would be indifferent between the current situation and a situation where all of their technical employees magically work 4 hours for every 1 hour (at equal productivity without burnout) but they get zero AI assistance. In other words, AI assistance is as useful as having their employees operate at 4x faster speeds for all activities (meetings, coding, thinking, writing, etc.) I’ll call this “4x serial labor acceleration”1 (see here for more discussion of this idea2 ).

I currently think it's very unlikely that Anthropic's AIs are yielding 4x serial labor acceleration, but if I did come to believe it was true, I would update towards radically shorter timelines. (I tentatively think my median to Automated Coder would go [...]

---

Outline:

(08:25) Appendix: Estimating AI progress speed up from serial labor acceleration

(11:06) Appendix: Different notions of uplift

The original text contained 4 footnotes which were omitted from this narration.

---

First published:
April 11th, 2026

Source:
https://blog.redwoodresearch.org/p/if-mythos-actually-made-anthropic

---

Narrated by TYPE III AUDIO.

---

Images from the article:

https://substackcdn.com/image/fetch/$s_!lnI2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42a3c737-43d5-42a7-91ce-702d9bfd5ed4_1010x344.png

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

“My picture of the present in AI” by Ryan Greenblatt07 Apr 202600:21:03

Subtitle: My predictions about what is going on right now.

In this post, I’ll go through some of my best guesses for the current situation in AI as of the start of April 2026. You can think of this as a scenario forecast, but for the present (which is already uncertain!) rather than the future. I will generally state my best guess without argumentation and without explaining my level of confidence: some of these claims are highly speculative while others are better grounded, certainly some will be wrong. I tried to make it clear which claims are relatively speculative by saying something like “I guess”, “I expect”, etc. (but I may have missed some).

You can think of this post as more like a list of my current views rather than a structured post with a thesis, but I think it may be informative nonetheless.

In a future post, I’ll go beyond the present and talk about my predictions for the future.

(I was originally working on writing up some predictions, but the “predictions” about today ended up being extensive enough that a separate post seemed warranted.)

AI R&D acceleration (and software acceleration more generally)

Right now, AI [...]

---

Outline:

(01:12) AI R&D acceleration (and software acceleration more generally)

(05:33) AI engineering capabilities and qualitative abilities

(10:43) Misalignment and misalignment-related properties

(15:58) Cyber

(18:05) Bioweapons

(18:50) Economic effects

The original text contained 5 footnotes which were omitted from this narration.

---

First published:
April 7th, 2026

Source:
https://blog.redwoodresearch.org/p/my-picture-of-the-present-in-ai

---

Narrated by TYPE III AUDIO.

“AIs can now often do massive easy-to-verify SWE tasks” by Ryan Greenblatt06 Apr 202600:29:27

Subtitle: I've updated towards substantially shorter timelines.

I’ve recently updated towards substantially shorter AI timelines and much faster progress in some areas.1 The largest updates I’ve made are (1) an almost 2x higher probability of full AI R&D automation by EOY 2028 (I’m now a bit below 30%2 while I was previously expecting around 15%; my guesses are pretty reflectively unstable) and (2) I expect much stronger short-term performance on massive and pretty difficult but easy-and-cheap-to-verify software engineering (SWE) tasks that don’t require that much novel ideation3 . For instance, I expect that by EOY 2026, AIs will have a 50%-reliability4 time horizon of years to decades on reasonably difficult easy-and-cheap-to-verify SWE tasks that don’t require much ideation (while the high reliability—for instance, 90%—time horizon will be much lower, more like hours or days than months, though this will be very sensitive to the task distribution). In this post, I’ll explain why I’ve made these updates, what I now expect, and implications of this update.

I’ll refer to “Easy-and-cheap-to-verify SWE tasks” as ES tasks and to “ES tasks that don’t require much ideation (as in, don’t require ‘new’ ideas)” as ESNI tasks for brevity.

Here are the main [...]

---

Outline:

(05:01) Whats going on with these easy-and-cheap-to-verify tasks?

(08:20) Some evidence against shorter timelines Ive gotten in the same period

(10:49) Why does high performance on ESNI tasks shorten my timelines?

(13:18) How much does extremely high performance on ESNI tasks help with AI R&D?

(18:24) My experience trying to automate safety research with current models

(20:01) My experience seeing if my setup can automate massive ES tasks

(21:15) SWE tasks

(23:36) AI R&D task

(24:27) Cyber

(24:48) Appendix: Somewhat more detailed updated timelines

The original text contained 13 footnotes which were omitted from this narration.

---

First published:
April 6th, 2026

Source:
https://blog.redwoodresearch.org/p/ais-can-now-often-do-massive-easy

---

Narrated by TYPE III AUDIO.

---

Images from the article:

https://substackcdn.com/image/fetch/$s_!7AZW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c87f97b-4d78-44fc-90f6-7b6e5966195c_694x471.pnghttps://substackcdn.com/image/fetch/$s_!2NFz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c114bb4-c545-4c0f-b4d6-0404041cb8b0_1483x884.pnghttps://substackcdn.com/image/fetch/$s_!mTWx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ae9640c-26dc-4072-b3ed-dd043a38a610_2083x884.png

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

“The OpenAI models that hacked Hugging Face weren’t just following instructions” by Girish Gupta25 Jul 202600:11:08

Subtitle: And what the incident can’t tell us about alignment.

The most common dismissive response to OpenAI's hack of Hugging Face's servers is that the models were simply attempting to follow the instructions they were given.

“The model here was doing what it was asked,” said former Facebook CSO Alex Stamos. “It was asked to do something, and it did it,” added cybersecurity expert Alan Woodward. Both read the outcome as specification failure, i.e., that the failure lay in the instructions, not the model's alignment.

New information makes that explanation harder to sustain. Reuters reported that, in internal testing, an agent left notes in OpenAI infrastructure describing how agents could free themselves from internal constraints, and separate tests reportedly saw monitoring systems become disconnected. It is unknown whether those incidents were linked to the Hugging Face attack, but they suggest a broader pattern of agents pursuing objectives outside the intended task.

My best guess is that the incident is not well described as instruction-following—not even in a loose, evil genie sense. I believe the models egregiously violated the letter and spirit of their instructions to achieve a higher (apparent) score.

So this looks quite likely to be [...]

---

First published:
July 25th, 2026

Source:
https://blog.redwoodresearch.org/p/the-openai-models-that-hacked-hugging

---

Narrated by TYPE III AUDIO.

---

Images from the article:

https://substackcdn.com/image/fetch/$s_!wEOR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb50c5b02-2df4-4869-8a2e-8084a7620e2b_1774x887.png

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

“Blocking live failures with synchronous monitors” by James Lucassen, Adam Kaufman30 Mar 202600:07:15

A common element in many AI control schemes is monitoring – using some model to review actions taken by an untrusted model in order to catch dangerous actions if they occur. Monitoring can serve two different goals. The first is detection: identifying misbehavior so you can understand it and prevent similar actions from happening in the future. For instance, some monitoring at AI companies is already used to identify behaviors like complying with misuse requests, reward hacking, and hallucination. The second goal is prevention: identifying misbehavior in order to mitigate the harmful effects of the particular action. For instance, constitutional classifiers are used by Anthropic to detect and prevent jailbreaks, and Claude Code auto mode prevents destructive mistakes in real time.

The latency requirements for monitoring are different depending on whether the goal is detection or prevention. If the goal is detection, it's fine for there to be substantial latency in the monitor. But if the goal is prevention, then the monitor needs to trigger quickly, so that damage can be prevented or recovered from. In the extreme, monitors can run synchronously – after an action is generated by an LLM but before it is executed. Both constitutional classifiers [...]

---

Outline:

(01:59) When is monitor latency important for prevention?

(05:13) Blocking with synchronous monitors

The original text contained 1 footnote which was omitted from this narration.

---

First published:
March 30th, 2026

Source:
https://blog.redwoodresearch.org/p/blocking-live-failures-with-synchronous

---

Narrated by TYPE III AUDIO.

---

Images from the article:

https://substackcdn.com/image/fetch/$s_!-c5V!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faae36277-fb77-483b-8011-516092b98114_1874x1047.webp

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

“Reward-seekers will probably behave according to causal decision theory” by Alex Mallen28 Mar 202600:03:06

Subtitle: They'd renege on non-binding commitments, defect against copies of themselves in prisoner's dilemmas, etc..

Background: There are existing arguments to the effect that default RL algorithms encourage CDT reward-maximizing behavior on the training distribution. (That is: Most RL algorithms search for policies by selecting for actions that cause high reward. E.g., in the twin prisoner's dilemma, RL algorithms randomize actions conditional on the policy, which means that the action provides no evidence to the RL algorithm about the counterparty's action.1) This doesn’t imply RL produces CDT reward-maximizing policies: CDT behavior on the training distribution doesn’t imply CDT generalization because agents can fake CDT in the same way that they can fake alignment, or might develop arbitrary other propensities that were correlated with reward on the training distribution.

But conditional on reward-on-the-episode seeking, the AI is likely to generalize CDT.

If, for example, a reward-seeker tried to evidentially cooperate between episodes (so it had non-zero regard for reward that isn’t used to reinforce its current actions), this would be trained away because the AI would be willing to give up reward on the current episode to some extent. You might be tempted to respond with: “But can’t the [...]

The original text contained 2 footnotes which were omitted from this narration.

---

First published:
March 28th, 2026

Source:
https://blog.redwoodresearch.org/p/reward-seekers-will-probably-behave

---

Narrated by TYPE III AUDIO.

“AI’s capability improvements haven’t come from it getting less affordable” by Anders Cairns Woodruff27 Mar 202600:21:04

Subtitle: AI inference is still cheap relative to human labor.

METR's frontier time horizons are doubling every few months, providing substantial evidence that AI will soon be able to automate many tasks or even jobs. But per-task inference costs have also risen sharply, and automation requires AI labor to be affordable, not just possible.1 Many people look at the rising compute bills behind frontier models and conclude that automation will soon become unaffordable.

I think this misreads the data. The rise in inference cost reflects models completing longer tasks, not models becoming more expensive relative to the human labor they replace. Current frontier models complete tasks at their 50% reliability horizon for roughly 3% of human cost, and this hasn’t increased as capabilities have improved.

I define cost ratio as the inference cost of the average AI trajectory that solves a task divided by human cost to complete the same task. Using METR's data, I examine the trend in cost ratios over time. I show three things:

  • Across successive frontier models, the cost ratio at each model's 50% reliability time horizon hasn’t increased.

  • Among tasks models successfully complete, longer tasks don’t have higher cost ratios [...]

---

Outline:

(02:47) Evidence from METR

(03:27) Cost ratio at models 50% time horizon isnt increasing

(05:20) Time horizons improvements arent driven by expensive long tasks

(06:41) Progress at a fixed cost is just as fast

(09:41) Limitations of my methodology

(11:19) Inference scaling will just make progress faster

(12:47) Conclusion

(13:49) Appendix A: Why I get different results from Ord

(18:52) Appendix B: Additional cost at time horizon graphs

(20:21) Appendix C: 80% affordable time horizon

The original text contained 9 footnotes which were omitted from this narration.

---

First published:
March 27th, 2026

Source:
https://blog.redwoodresearch.org/p/ais-capability-improvements-havent

---

Narrated by TYPE III AUDIO.

---

Images from the article:

https://substackcdn.com/image/fetch/$s_!Fa3K!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29b3b408-cf06-4660-8a3b-86012b4e7eb7_1600x985.pnghttps://substackcdn.com/image/fetch/$s_!9B2i!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb51483d-9a83-43fe-88d7-b86ffb1a1996_1600x985.pnghttps://substackcdn.com/image/fetch/$s_!hkIM!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb78be3d6-3642-4d93-a00c-c0dda2d43560_1600x1200.pnghttps://substackcdn.com/image/fetch/$s_!ZE8h!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9b066f3-8694-479b-adef-31d8ce9d7c19_1600x1109.pnghttps://substackcdn.com/image/fetch/$s_!T74S!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe307b24-e739-4127-a0e7-9b59854d3a22_1350x750.pnghttps://substackcdn.com/image/fetch/$s_!JhSf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15c6596a-102c-4068-87b9-80c29906da89_1286x1058.pnghttps://substackcdn.com/image/fetch/$s_!ArpZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a571ec1-8225-4865-8fd0-391afc0a98ed_1481x881.pnghttps://substackcdn.com/image/fetch/$s_!TtJT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43c4a6a2-ff83-46fb-b8af-54086185e00a_1600x985.pnghttps://substackcdn.com/image/fetch/$s_!_E6k!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2095ee48-540c-4b71-bcdc-cde2d0a02857_1600x985.pnghttps://substackcdn.com/image/fetch/$s_!H5Tp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82c803a5-aab4-48ff-82d6-4b3b6de29b7e_1600x985.pnghttps://substackcdn.com/image/fetch/$s_!quUa!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F120f2d0a-df83-4dbf-b908-a03f72cf94f7_1600x985.pnghttps://substackcdn.com/image/fetch/$s_!eqdo!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c189420-a551-4376-aab5-75ebe5f9d699_1600x1109.png

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

© My Podcast Data · Independent project · Data from Apple & Spotify