Back

Explore every episode of the podcast LessWrong (Curated & Popular)

Dive into the complete episode list for LessWrong (Curated & Popular). Each episode is cataloged with detailed descriptions, making it easy to find and explore specific topics. Keep track of all episodes from your favorite podcast and never miss a moment of insightful content.

Rows per page:

1–50 of 740

TitlePub. DateDuration
"IABIED Book Review: Core Arguments and Counterarguments" by Stephen McAleese05 Feb 202600:50:18
The recent book “If Anyone Builds It Everyone Dies” (September 2025) by Eliezer Yudkowsky and Nate Soares argues that creating superintelligent AI in the near future would almost certainly cause human extinction:

If any company or group, anywhere on the planet, builds an artificial superintelligence using anything remotely like current techniques, based on anything remotely like the present understanding of AI, then everyone, everywhere on Earth, will die.

The goal of this post is to summarize and evaluate the book's key arguments and the main counterarguments critics have made against them.

Although several other book reviews have already been written I found many of them unsatisfying because a lot of them are written by journalists who have the goal of writing an entertaining piece and only lightly cover the core arguments, or don’t seem understand them properly, and instead resort to weak arguments like straw-manning, ad hominem attacks or criticizing the style of the book.

So my goal is to write a book review that has the following properties:

  • Written by someone who has read a substantial amount of AI alignment and LessWrong content and won’t make AI alignment beginner mistakes or misunderstandings (e.g. not knowing about the [...]
---

Outline:

(07:43) Background arguments to the key claim

(09:21) The key claim: ASI alignment is extremely difficult to solve

(12:52) 1. Human values are a very specific, fragile, and tiny space of all possible goals

(15:25) 2. Current methods used to train goals into AIs are imprecise and unreliable

(16:42) The inner alignment problem

(17:25) Inner alignment introduction

(19:03) Inner misalignment evolution analogy

(21:03) Real examples of inner misalignment

(22:23) Inner misalignment explanation

(25:05) ASI misalignment example

(27:40) 3. The ASI alignment problem is hard because it has the properties of hard engineering challenges

(28:10) Space probes

(29:09) Nuclear reactors

(30:18) Computer security

(30:35) Counterarguments to the book

(30:46) Arguments that the books arguments are unfalsifiable

(33:19) Arguments against the evolution analogy

(37:38) Arguments against counting arguments

(40:16) Arguments based on the aligned behavior of modern LLMs

(43:16) Arguments against engineering analogies to AI alignment

(45:05) Three counterarguments to the books three core arguments

(46:43) Conclusion

(49:23) Appendix

---

First published:
January 24th, 2026

Source:
https://www.lesswrong.com/posts/qFzWTTxW37mqnE6CA/iabied-book-review-core-arguments-and-counterarguments

---



Narrated by TYPE III AUDIO.

---

Images from the article:

https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/qFzWTTxW37mqnE6CA/y9wsakomogbk2guyib9x
"Anthropic’s “Hot Mess” paper overstates its case (and the blog post is worse)" by RobertM04 Feb 202600:11:39
Author's note: this is somewhat more rushed than ideal, but I think getting this out sooner is pretty important. Ideally, it would be a bit less snarky.

Anthropic[1] recently published a new piece of research: The Hot Mess of AI: How Does Misalignment Scale with Model Intelligence and Task Complexity? (arXiv, Twitter thread).

I have some complaints about both the paper and the accompanying blog post.

tl;dr

  • The paper's abstract says that "in several settings, larger, more capable models are more incoherent than smaller models", but in most settings they are more coherent. This emphasis is even more exaggerated in the blog post and Twitter thread. I think this is pretty misleading.
  • The paper's technical definition of "incoherence" is uninteresting[2] and the framing of the paper, blog post, and Twitter thread equivocate with the more normal English-language definition of the term, which is extremely misleading.
  • Section 5 of the paper (and to a larger extent the blog post and Twitter) attempt to draw conclusions about future alignment difficulties that are unjustified by the experiment results, and would be unjustified even if the experiment results pointed in the other direction.
  • The blog post is substantially LLM-written. I think this [...]
---

Outline:

(00:39) tl;dr

(01:42) Paper

(06:25) Blog

The original text contained 3 footnotes which were omitted from this narration.

---

First published:
February 4th, 2026

Source:
https://www.lesswrong.com/posts/ceEgAEXcL7cC2Ddiy/anthropic-s-hot-mess-paper-overstates-its-case-and-the-blog

---



Narrated by TYPE III AUDIO.

---

Images from the article:

https://res.cloudinary.com/lesswrong-2-0/image/upload/v1770101053/lexical_client_uploads/wfdvsvp822a70mwlimva.pnghttps://res.cloudinary.com/lesswrong-2-0/image/upload/v1770101110/lexical_client_uploads/yylpc8ibfknzwzki2ufg.pngApple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

"Conditional Kickstarter for the “Don’t Build It” March" by Raemon03 Feb 202600:06:54
tl;dr: You can pledge to join a big protest to ban AGI research at ifanyonebuildsit.com/march, which only triggers if 100,000 people sign up.

The If Anyone Builds It website includes a March page, wherein you can pledge to march in Washington DC, demanding an international treaty to stop AGI research if 100,000 people in total also pledge.

I designed the March page (although am not otherwise involved with March decisionmaking), and want to pitch people on signing up for the "March Kickstarter."

It's not obvious that small protests do anything, or are worth the effort. But, I think 100,000 people marching in DC would be quite valuable because it showcases "AI x-risk is not a fringe concern. If you speak out about it, you are not being a lonely dissident, you are representing a substantial mass of people."

The current version of the March page is designed around the principle that "conditional kickstarters are cheap." MIRI might later decide to push hard on the March, and maybe then someone will bid for people to come who are on the fence.

For now, I mostly wanted to say: if you're the sort of person who would fairly obviously come to [...]

---

Outline:

(01:54) Probably expect a design/slogan reroll

(03:10) FAQ

(03:13) Whats the goal of the Dont Build It march?

(03:24) Why?

(03:55) Why do you think that?

(04:22) Why does the pledge only take effect if 100,000 people pledge to march?

(04:56) What do you mean by international treaty?

(06:00) How much notice will there be for the actual march?

(06:14) What if I dont want to commit to marching in D.C. yet?

---

First published:
February 2nd, 2026

Source:
https://www.lesswrong.com/posts/HnwDWxRPzRrBfJSBD/conditional-kickstarter-for-the-don-t-build-it-march

---



Narrated by TYPE III AUDIO.

---

Images from the article:

https://39669.cdn.cke-cs.com/rQvD3VnunXZu34m86e5f/images/313bf5580c1f6198630fb95ca3c09d34c5a434cd41680671.pngApple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

"How to Hire a Team" by Gretta Duleba01 Feb 202600:08:40
A low-effort guide I dashed off in less than an hour, because I got riled up.

  1. Try not to hire a team. Try pretty hard at this.
    1. Try to find a more efficient way to solve your problem that requires less labor – a smaller-footprint solution.
    2. Try to hire contractors to do specific parts that they’re really good at, and who have a well-defined interface. Your relationship to these contractors will mostly be transactional and temporary.
    3. If you must, try hiring just one person, a very smart, capable, and trustworthy generalist, who finds and supports the contractors, so all you have to do is manage the problem-and-solution part of the interface with the contractors. You will need to spend quite a bit of time making sure this lieutenant understands what you’re doing and why, so be very choosy not just about their capabilities but about how well you work together, how easily you can make yourself understood, etc.
  2. If that fails, hire the smallest team that you can. Small is good because:
    1. Managing more people is more work.
      1. The relationship between number of people and management overhead is roughly O(n) but unevenly distributed; some people [...]
---

First published:
January 29th, 2026

Source:
https://www.lesswrong.com/posts/cojSyfxfqfm4kpCbk/how-to-hire-a-team

---



Narrated by TYPE III AUDIO.

"The Possessed Machines (summary)" by L Rudolf L29 Jan 202600:16:43
The Possessed Machines is one of the most important AI microsites. It was published anonymously by an ex- lab employee, and does not seem to have spread very far, likely at least partly due to this anonymity (e.g. there is no LessWrong discussion at the time I'm posting this). This post is my attempt to fix that.

I do not agree with everything in the piece, but I think cultural critiques of the "AGI uniparty" are vastly undersupplied and incredibly important in modeling & fixing the current trajectory.

The piece is a long but worthwhile analysis of some of the cultural and psychological failures of the AGI industry. The frame is Dostoevsky's Demons (alternatively translated The Possessed), a novel about ruin in a small provincial town. The author argues it's best read as a detailed description of earnest people causing a catastrophe by following tracks laid down by the surrounding culture that have gotten corrupted:

What I know is that Dostoevsky, looking at his own time, saw something true about how intelligent societies destroy themselves. He saw that the destruction comes from the best as well as the worst, from the idealists as well as the cynics, from the [...]

---

First published:
January 25th, 2026

Source:
https://www.lesswrong.com/posts/ppBHrfY4bA6J7pkpS/the-possessed-machines-summary

---



Narrated by TYPE III AUDIO.

"Ada Palmer: Inventing the Renaissance" by Martin Sustrik28 Jan 202600:26:17
Papal election of 1492 For over a decade, Ada Palmer, a history professor at University of Chicago (and a science-fiction writer!), struggled to teach Machiavelli. “I kept changing my approach, trying new things: which texts, what combinations, expanding how many class sessions he got…” The problem, she explains, is that “Machiavelli doesn’t unpack his contemporary examples, he assumes that you lived through it and know, so sometimes he just says things like: Some princes don’t have to work to maintain their power, like the Duke of Ferrara, period end of chapter. He doesn’t explain, so modern readers can’t get it.”

Palmer's solution was to make her students live through the run-up to the Italian Wars themselves. Her current method involves a three-week simulation of the 1492 papal election, a massive undertaking with sixty students playing historical figures, each receiving over twenty pages of unique character material, supported by twenty chroniclers and seventy volunteers. After this almost month-long pedagogical marathon, a week of analysis, and reading Machiavelli's letters, students finally encounter The Prince. By then they know the context intimately. When Machiavelli mentions the Duke of Ferrara maintaining power effortlessly, Palmer's students react viscerally. They remember Alfonso and Ippolito d’Este as [...]

---

First published:
January 25th, 2026

Source:
https://www.lesswrong.com/posts/doADJmyy6Yhp47SJ2/ada-palmer-inventing-the-renaissance

---



Narrated by TYPE III AUDIO.

---

Images from the article:

https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/doADJmyy6Yhp47SJ2/lglpat3stgwjkb4yixqqhttps://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/doADJmyy6Yhp47SJ2/do5pl2yt4uf7a0kd514mApple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

"AI found 12 of 12 OpenSSL zero-days (while curl cancelled its bug bounty)" by Stanislav Fort28 Jan 202600:20:16
This is a partial follow-up to AISLE discovered three new OpenSSL vulnerabilities from October 2025.

TL;DR: OpenSSL is among the most scrutinized and audited cryptographic libraries on the planet, underpinning encryption for most of the internet. They just announced 12 new zero-day vulnerabilities (meaning previously unknown to maintainers at time of disclosure). We at AISLE discovered all 12 using our AI system. This is a historically unusual count and the first real-world demonstration of AI-based cybersecurity at this scale. Meanwhile, curl just cancelled its bug bounty program due to a flood of AI-generated spam, even as we reported 5 genuine CVEs to them. AI is simultaneously collapsing the median ("slop") and raising the ceiling (real zero-days in critical infrastructure).

Background

We at AISLE have been building an automated AI system for deep cybersecurity discovery and remediation, sometimes operating in bug bounties under the pseudonym Giant Anteater. Our goal was to turn what used to be an elite, artisanal hacker craft into a repeatable industrial process. We do this to secure the software infrastructure of human civilization before strong AI systems become ubiquitous. Prosaically, we want to make sure we don't get hacked into oblivion the moment they come online.

[...]

---

Outline:

(01:05) Background

(02:56) Fall 2025: Our first OpenSSL results

(05:59) January 2026: 12 out of 12 new vulnerabilities

(07:28) HIGH severity (1):

(08:01) MODERATE severity (1):

(08:24) LOW severity (10):

(13:10) Broader impact: curl

(17:06) The era of AI cybersecurity is here for good

(18:40) Future outlook

---

First published:
January 27th, 2026

Source:
https://www.lesswrong.com/posts/7aJwgbMEiKq5egQbd/ai-found-12-of-12-openssl-zero-days-while-curl-cancelled-its

---



Narrated by TYPE III AUDIO.

"Dario Amodei – The Adolescence of Technology" by habryka28 Jan 202601:54:18
Dario Amodei, CEO of Anthropic, has written a new essay on his thoughts on AI risk of various shapes. It seems worth reading, even if just for understanding what Anthropic is likely to do in the future.

Confronting and Overcoming the Risks of Powerful AI

There is a scene in the movie version of Carl Sagan's book Contact where the main character, an astronomer who has detected the first radio signal from an alien civilization, is being considered for the role of humanity's representative to meet the aliens. The international panel interviewing her asks, “If you could ask [the aliens] just one question, what would it be?” Her reply is: “I’d ask them, ‘How did you do it? How did you evolve, how did you survive this technological adolescence without destroying yourself?” When I think about where humanity is now with AI—about what we’re on the cusp of—my mind keeps going back to that scene, because the question is so apt for our current situation, and I wish we had the aliens’ answer to guide us. I believe we are entering a rite of passage, both turbulent and inevitable, which will test who we are as a species. Humanity [...]

---

Outline:

(00:24) Confronting and Overcoming the Risks of Powerful AI

(15:19) 1. I'm sorry, Dave

(15:23) Autonomy risks

(28:53) Defenses

(41:17) 2. A surprising and terrible empowerment

(41:22) Misuse for destruction

(54:50) Defenses

(01:00:25) 3. The odious apparatus

(01:00:30) Misuse for seizing power

(01:13:08) Defenses

(01:19:48) 4. Player piano

(01:19:51) Economic disruption

(01:21:18) Labor market disruption

(01:33:43) Defenses

(01:37:43) Economic concentration of power

(01:40:49) Defenses

(01:43:13) 5. Black seas of infinity

(01:43:17) Indirect effects

(01:47:29) Humanity's test

(01:53:58) Footnotes

The original text contained 92 footnotes which were omitted from this narration.

---

First published:
January 26th, 2026

Source:
https://www.lesswrong.com/posts/kzPQohJakutbtFPcf/dario-amodei-the-adolescence-of-technology

---



Narrated by TYPE III AUDIO.

"AlgZoo: uninterpreted models with fewer than 1,500 parameters" by Jacob_Hilton27 Jan 202600:21:53
Audio note: this article contains 78 uses of latex notation, so the narration may be difficult to follow. There's a link to the original text in the episode description.

This post covers work done by several researchers at, visitors to and collaborators of ARC, including Zihao Chen, George Robinson, David Matolcsi, Jacob Stavrianos, Jiawei Li and Michael Sklar. Thanks to Aryan Bhatt, Gabriel Wu, Jiawei Li, Lee Sharkey, Victor Lecomte and Zihao Chen for comments.

In the wake of recent debate about pragmatic versus ambitious visions for mechanistic interpretability, ARC is sharing some models we've been studying that, in spite of their tiny size, serve as challenging test cases for any ambitious interpretability vision. The models are RNNs and transformers trained to perform algorithmic tasks, and range in size from 8 to 1,408 parameters. The largest model that we believe we more-or-less fully understand has 32 parameters; the next largest model that we have put substantial effort into, but have failed to fully understand, has 432 parameters. The models are available at the AlgZoo GitHub repo.

We think that the "ambitious" side of the mechanistic interpretability community has historically underinvested in "fully understanding slightly complex [...]

---

Outline:

(03:09) Mechanistic estimates as explanations

(06:16) Case study: 2nd argmax RNNs

(08:30) Hidden size 2, sequence length 2

(14:47) Hidden size 4, sequence length 3

(16:13) Hidden size 16, sequence length 10

(19:52) Conclusion

The original text contained 20 footnotes which were omitted from this narration.

---

First published:
January 26th, 2026

Source:
https://www.lesswrong.com/posts/x8BbjZqooS4LFXS8Z/algzoo-uninterpreted-models-with-fewer-than-1-500-parameters

---



Narrated by TYPE III AUDIO.

---

Images from the article:

https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/x8BbjZqooS4LFXS8Z/vmsgotmxsthp5ssrivt4https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/x8BbjZqooS4LFXS8Z/fgfdiy3gqc9f7ztudcfrApple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

"Does Pentagon Pizza Theory Work?" by rba27 Jan 202600:11:05
As soon as modern data analysis became a thing, the US government has had to deal with people trying to use open source data to uncover its secrets.

During the early Cold War days and America's hydrogen bomb testing, there was an enormous amount of speculation about how the bombs actually worked. All nuclear technology involves refinement and purification of large amounts of raw substances into chemically pure substances. Armen Alchian was an economist working at RAND and reasoned that any US company working in such raw materials and supplying the government would have made a killing leading up to the tests.

After checking financial data that RAND maintained on such companies, Alchian deduced that the secret sauce in the early fusion bombs was lithium and the Lithium Corporation of America was supplying the USG. The company's stock had skyrocketed leading up to the Castle Bravo test either by way of enormous unexpected revenue gains from government contracts, or more amusingly, maybe by government insiders buying up the stock trying to make a mushroom-cloud-sized fortune with the knowledge that lithium was the key ingredient.

When word of this work got out, this story naturally ends with the FBI coming [...]

---

Outline:

(01:27) Pizza is the new lithium

(03:09) The Data

(04:11) The Backtest

(04:36) Fordow bombing

(04:55) Maduro capture

(05:15) The Houthi stuff

(10:25) Coda

---

First published:
January 22nd, 2026

Source:
https://www.lesswrong.com/posts/Li3Aw7sDLXTCcQHZM/does-pentagon-pizza-theory-work

---



Narrated by TYPE III AUDIO.

---

Images from the article:

https://substackcdn.com/image/fetch/$s_!U5o8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce6621b5-e68d-4245-97ea-90de11568e6f_2808x1038.pnghttps://substackcdn.com/image/fetch/$s_!RoQJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5cc0796e-9d2d-427d-bd60-2caae7d3b175_1728x1044.pnghttps://substackcdn.com/image/fetch/$s_!Xn8h!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4dee9e81-e5de-42b6-ae79-c4874908a4c9_1728x1044.png
"The inaugural Redwood Research podcast" by Buck, ryan_greenblatt27 Jan 202600:03:27
After five months of me (Buck) being slow at finishing up the editing on this, we’re finally putting out our inaugural Redwood Research podcast. I think it came out pretty well—we discussed a bunch of interesting and underdiscussed topics and I’m glad to have a public record of a bunch of stuff about our history. Tell your friends! Whether we do another one depends on how useful people find this one. You can watch on Youtube here, or as a Substack podcast.

Notes on editing the podcast with Claude Code

(Buck wrote this section)

After the recording, we faced a problem. We had four hours of footage from our three cameras. We wanted it to snazzily cut between shots depending on who was talking. But I don’t truly in my heart believe that it's that important for the video editing to be that good, and I don’t really like the idea of paying a video editor. But I also don’t want to edit the four hours of video myself. And it seemed to me that video editing software was generally not optimized for the kind of editing I wanted to do here (especially automatically cutting between different shots according [...]

---

Outline:

(00:43) Notes on editing the podcast with Claude Code

(03:11) Podcast transcript

---

First published:
January 4th, 2026

Source:
https://www.lesswrong.com/posts/p4iJpumHt6Ay9KnXT/the-inaugural-redwood-research-podcast

---



Narrated by TYPE III AUDIO.

"Canada Lost Its Measles Elimination Status Because We Don’t Have Enough Nurses Who Speak Low German" by jenn26 Jan 202600:15:33
This post was originally published on November 11th, 2025. I've been spending some time reworking and cleaning up the Inkhaven posts I'm most proud of, and completed the process for this one today.

Today, Canada officially lost its measles elimination status. Measles was previously declared eliminated in Canada in 1998, but countries lose that status after 12 months of continuous transmission.

Here are some articles about the the fact that we have lost our measles elimination status: CBC, BBC, New York Times, Toronto Life. You can see some chatter on Reddit about it if you're interested here.

None of the above texts seemed to me to be focused on the actual thing that caused Canada to lose its measles elimination status, which is the rampant spread of measles among old-order religious communities, particularly the Mennonites. (Mennonites are basically, like, Amish-lite. Amish people can marry into Mennonite communities if they want a more laid-back lifestyle, but the reverse is not allowed. Similarly, old-order Mennonites can marry into less traditionally-minded Mennonite communities, but the reverse is not allowed.)

The Reddit comments that made this point are generally not highly upvoted[1], and this was certainly not a central point in any of [...]

---

Outline:

(03:20) The Mennonite Outbreak

(06:58) Mennonite Geography

(11:47) Mennonites Are Susceptible To Facts and Logic, When Presented In Low German

The original text contained 6 footnotes which were omitted from this narration.

---

First published:
January 25th, 2026

Source:
https://www.lesswrong.com/posts/H8RdAbAmsqbpBWoDd/canada-lost-its-measles-elimination-status-because-we-don-t

---



Narrated by TYPE III AUDIO.

---

Images from the article:

https://bear-images.sfo2.cdn.digitaloceanspaces.com/jenn/measles-canada.webphttps://bear-images.sfo2.cdn.digitaloceanspaces.com/jenn/measles-map-final.webphttps://bear-images.sfo2.cdn.digitaloceanspaces.com/jenn/pasted-image-20251110195337.webphttps://bear-images.sfo2.cdn.digitaloceanspaces.com/jenn/pasted-image-20251110195411.webp
"Deep learning as program synthesis" by Zach Furman24 Jan 202601:11:42
Audio note: this article contains 73 uses of latex notation, so the narration may be difficult to follow. There's a link to the original text in the episode description.

Epistemic status: This post is a synthesis of ideas that are, in my experience, widespread among researchers at frontier labs and in mechanistic interpretability, but rarely written down comprehensively in one place - different communities tend to know different pieces of evidence. The core hypothesis - that deep learning is performing something like tractable program synthesis - is not original to me (even to me, the ideas are ~3 years old), and I suspect it has been arrived at independently many times. (See the appendix on related work).

This is also far from finished research - more a snapshot of a hypothesis that seems increasingly hard to avoid, and a case for why formalization is worth pursuing. I discuss the key barriers and how tools like singular learning theory might address them towards the end of the post.

Thanks to Dan Murfet, Jesse Hoogland, Max Hennick, and Rumi Salazar for feedback on this post.

Sam Altman: Why does unsupervised learning work?

Dan Selsam: Compression. So, the ideal intelligence [...]

---

Outline:

(02:31) Background

(09:06) Looking inside

(09:09) Grokking

(16:04) Vision circuits

(22:37) The hypothesis

(26:04) Why this isnt enough

(27:22) Indirect evidence

(32:44) The paradox of approximation

(38:34) The paradox of generalization

(45:44) The paradox of convergence

(51:46) The path forward

(53:20) The representation problem

(58:38) The search problem

(01:07:20) Appendix

(01:07:23) Related work

The original text contained 14 footnotes which were omitted from this narration.

---

First published:
January 20th, 2026

Source:
https://www.lesswrong.com/posts/Dw8mskAvBX37MxvXo/deep-learning-as-program-synthesis-1

---



Narrated by TYPE III AUDIO.

---

Images from the article:

https://39669.cdn.cke-cs.com/rQvD3VnunXZu34m86e5f/images/deecd28e8833d4cdfa871cf0e77432d44bfee664bee5c9d3.pnghttps://39669.cdn.cke-cs.com/rQvD3VnunXZu34m86e5f/images/72565d8d674c6fcbf46500c6cb7ce4518135cc64b30133da.pnghttps://39669.cdn.cke-cs.com/rQvD3VnunXZu34m86e5f/images/4f78e9b302b437ae0b1e1910808d8aa16c593434f882f040.png
"Why I Transitioned: A Response" by marisa24 Jan 202600:21:17
Fiora Sunshine's post, Why I Transitioned: A Case Study (the OP) articulates a valuable theory for why some MtFs transition.

If you are MtF and feel the post describes you, I believe you.

However, many statements from the post are wrong or overly broad.

My claims:

  1. There is evidence of a biological basis for trans identity. Twin studies are a good way to see this.
     
  2. Fiora claims that trans people's apparent lack of introspective clarity may be evidence of deception. But trans people are incentivized not to attempt to share accurate answers to "why do you really want to transition?". This is the Trans Double Bind.
     
  3. I am a counterexample to Fiora's theory. I was an adolescent social outcast weeb but did not transition. I spent 14 years actualizing as a man, then transitioned at 31 only after becoming crippled by dysphoria. My example shows that Fiora's phenotype can co-occur with or mask medically significant dysphoria.
A. Biologically Transgender

In the OP, Fiora presents the "body-map theory" under the umbrella of "arcane neuro-psychological phenomena", and then dismisses medical theories because the body-map theory doesn't fit her friend group.

The body-map theory is a straw man for [...]

---

Outline:

(00:29) My claims:

(01:17) A. Biologically Transgender

(02:38) Twin Studies à la LLM

(06:06) B. The Trans Double Bind

(07:10) Motivations to transition

(11:53) Introspective Clarity

(13:23) C. In the Case of Quinoa Marisa

(19:44) Conclusion

The original text contained 11 footnotes which were omitted from this narration.

---

First published:
January 19th, 2026

Source:
https://www.lesswrong.com/posts/rt2yai8JkTPYgzoEj/why-i-transitioned-a-response

---



Narrated by TYPE III AUDIO.

---

Images from the article:

https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/rt2yai8JkTPYgzoEj/njqirawkkd1qwehseaixhttps://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/rt2yai8JkTPYgzoEj/ewcomkdvdmn8yw2q84oaApple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

"Claude’s new constitution" by Zac Hatfield-Dodds22 Jan 202600:11:56
Read the constitution. Previously: 'soul document' discussion here.

We're publishing a new constitution for our AI model, Claude. It's a detailed description of Anthropic's vision for Claude's values and behavior; a holistic document that explains the context in which Claude operates and the kind of entity we would like Claude to be.

The constitution is a crucial part of our model training process, and its content directly shapes Claude's behavior. Training models is a difficult task, and Claude's outputs might not always adhere to the constitution's ideals. But we think that the way the new constitution is written—with a thorough explanation of our intentions and the reasons behind them—makes it more likely to cultivate good values during training.

In this post, we describe what we've included in the new constitution and some of the considerations that informed our approach.

We're releasing Claude's constitution in full under a Creative Commons CC0 1.0 Deed, meaning it can be freely used by anyone for any purpose without asking for permission.

What is Claude's Constitution?

Claude's constitution is the foundational document that both expresses and shapes who Claude is. It contains detailed explanations of the values we [...]

---

Outline:

(01:14) What is Claudes Constitution?

(03:26) Our new approach to Claudes Constitution

(04:59) A brief summary of the new constitution

(09:14) Conclusion

The original text contained 2 footnotes which were omitted from this narration.

---

First published:
January 21st, 2026

Source:
https://www.lesswrong.com/posts/mLvxxoNjDqDHBAo6K/claude-s-new-constitution

---



Narrated by TYPE III AUDIO.

[Linkpost] "“The first two weeks are the hardest”: my first digital declutter" by mingyuan20 Jan 202600:04:28
This is a link post. It is unbearable to not be consuming. All through the house is nothing but silence. The need inside of me is not an ache, it is caustic, sour, the burning desire to be distracted, to be listening, watching, scrolling.

Some of the time I think I’m happy. I think this is very good. I go to the park and lie on a blanket in a sun with a book and a notebook. I watch the blades of grass and the kids and the dogs and the butterflies and I’m so happy to be free.

Then there are the nights. The dark silence is so oppressive, so all-consuming. One lonely night, early on, I bike to a space where I had sometimes felt welcome, and thought I might again.

“What are you doing here?” the people ask.

“I’m three days into my month of digital minimalism and I’m so bored, I just wanted to be around people.”

No one really wants to be around me. Okay.

One of the guys had a previous life as a digital minimalism coach. “The first two weeks are the hardest,” he tells me encouragingly.

“Two WEEKS?” I want to shriek.

[...]

---

First published:
January 18th, 2026

Source:
https://www.lesswrong.com/posts/eeFqTjmZ8kS7S5tpg/the-first-two-weeks-are-the-hardest-my-first-digital

Linkpost URL:
https://mingyuan.substack.com/p/the-first-two-weeks-are-the-hardest

---



Narrated by TYPE III AUDIO.

"What Washington Says About AGI" by zroe120 Jan 202600:14:16
I spent a few hundred dollars on Anthropic API credits and let Claude individually research every current US congressperson's position on AI. This is a summary of my findings.

Disclaimer: Summarizing people's beliefs is hard and inherently subjective and noisy. Likewise, US politicians change their opinions on things constantly so it's hard to know what's up-to-date. Also, I vibe-coded a lot of this.

Methodology

I used Claude Sonnet 4.5 with web search to research every congressperson's public statements on AI, then used GPT-4o to score each politician on how "AGI-pilled" they are, how concerned they are about existential risk, and how focused they are on US-China AI competition. I plotted these scores against GovTrack ideology data to search for any partisan splits.

1. AGI awareness is not partisan and not widespread

Few members of Congress have public statements taking AGI seriously. For those that do, the difference is not in political ideology. If we simply plot the AGI-pilled score vs the ideology score, we observe no obvious partisan split.

There are 151 congresspeople who Claude could not find substantial quotes about AI from. These members are not included on this plot or any of the plots which follow.

[...]

---

Outline:

(00:46) Methodology

(01:12) 1. AGI awareness is not partisan and not widespread

(01:56) 2. Existential risk is partisan at the tails

(02:51) 3. Both parties are fixated on China

(04:02) 4. Who in Congress is feeling the AGI?

(11:10) 5. Those who know the technology fear it.

(12:54) Appendix: How to use this data

The original text contained 2 footnotes which were omitted from this narration.

---

First published:
January 16th, 2026

Source:
https://www.lesswrong.com/posts/WLdcvAcoFZv9enR37/what-washington-says-about-agi

---



Narrated by TYPE III AUDIO.

---

Images from the article:

https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/WLdcvAcoFZv9enR37/zaonbyxxf3erj2w5dmy8https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/WLdcvAcoFZv9enR37/sy7ive0cssbbgixw92wqhttps://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/WLdcvAcoFZv9enR37/tgadk4vyc2s1tqtvyyac
"Precedents for the Unprecedented: Historical Analogies for Thirteen Artificial Superintelligence Risks" by James_Miller19 Jan 202602:03:49
Since artificial superintelligence has never existed, claims that it poses a serious risk of global catastrophe can be easy to dismiss as fearmongering. Yet many of the specific worries about such systems are not free-floating fantasies but extensions of patterns we already see. This essay examines thirteen distinct ways artificial superintelligence could go wrong and, for each, pairs the abstract failure mode with concrete precedents where a similar pattern has already caused serious harm. By assembling a broad cross-domain catalog of such precedents, I aim to show that concerns about artificial superintelligence track recurring failure modes in our world.

This essay is also an experiment in writing with extensive assistance from artificial intelligence, producing work I couldn’t have written without it. That a current system can help articulate a case for the catastrophic potential of its own lineage is itself a significant fact; we have already left the realm of speculative fiction and begun to build the very agents that constitute the risk. On a personal note, this collaboration with artificial intelligence is part of my effort to rebuild the intellectual life that my stroke disrupted and hopefully push it beyond where it stood before.

Section 1: Power Asymmetry [...]

---

First published:
January 16th, 2026

Source:
https://www.lesswrong.com/posts/kLvhBSwjWD9wjejWn/precedents-for-the-unprecedented-historical-analogies-for-1

---



Narrated by TYPE III AUDIO.

"Why we are excited about confession!" by boazbarak, Gabriel Wu, Manas Joglekar19 Jan 202600:17:28
Boaz Barak, Gabriel Wu, Jeremy Chen, Manas Joglekar

[Linkposting from the OpenAI alignment blog, where we post more speculative/technical/informal results and thoughts on safety and alignment.]


TL;DR We go into more details and some follow up results from our paper on confessions (see the original blog post). We give deeper analysis of the impact of training, as well as some preliminary comparisons to chain of thought monitoring.

We have recently published a new paper on confessions, along with an accompanying blog post. Here, we want to share with the research community some of the reasons why we are excited about confessions as a direction of safety, as well as some of its limitations. This blog post will be a bit more informal and speculative, so please see the paper for the full results.

The notion of “goodness” for the response of an LLM to a user prompt is inherently complex and multi-dimensional, and involves factors such as correctness, completeness, honesty, style, and more. When we optimize responses using a reward model as a proxy for “goodness” in reinforcement learning, models sometimes learn to “hack” this proxy and output an answer that only “looks good” to it (because [...]



---

Outline:

(08:19) Impact of training

(12:32) Comparing with chain-of-thought monitoring

(14:05) Confessions can increase monitorability

(15:44) Using high compute to improve alignment

(16:49) Acknowledgements

---

First published:
January 14th, 2026

Source:
https://www.lesswrong.com/posts/k4FjAzJwvYjFbCTKn/why-we-are-excited-about-confession

---



Narrated by TYPE III AUDIO.

---

Images from the article:

https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/k4FjAzJwvYjFbCTKn/swhsmmt1vvvqyrgkq6knhttps://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/k4FjAzJwvYjFbCTKn/drubjxbnutkslodqlvjshttps://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/k4FjAzJwvYjFbCTKn/ylv3i55uhouvdakbwsochttps://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/k4FjAz</truncato-artificial-root>
"Backyard cat fight shows Schelling points preexist language" by jchan16 Jan 202600:05:50
Two cats fighting for control over my backyard appear to have settled on a particular chain-link fence as the delineation between their territories. This suggests that:

  1. Animals are capable of recognizing Schelling points
  2. Therefore, Schelling points do not depend on language for their Schelling-ness
  3. Therefore, tacit bargaining should be understood not as a special case of bargaining where communication happens to be restricted, but rather as the norm from which the exceptional case of explicit bargaining is derived.
Summary of cat situation

I don't have any pets, so my backyard is terra nullius according to Cat Law. This situation is unstable, as there are several outdoor cats in the neighborhood who would like to claim it. Our two contenders are Tabby Cat, who lives on the other side of the waist-high chain-link fence marking the back edge of my lot, and Tuxedo Cat, who lives in the place next-door to me.

| || Tabby's || yard || (A) |------+...........+--------| (B) || | Tuxedo's| My yard | yard| ||| -- tall wooden fences.... short chain-link fence In the first [...]

---

Outline:

(00:43) Summary of cat situation

(03:28) Why the fence?

(04:00) If animals have Schelling points, then...

---

First published:
January 14th, 2026

Source:
https://www.lesswrong.com/posts/uYr8pba7TqaPpszX5/backyard-cat-fight-shows-schelling-points-preexist-language

---



Narrated by TYPE III AUDIO.

"How AI Is Learning to Think in Secret" by Nicholas Andresen09 Jan 202600:37:44
On Thinkish, Neuralese, and the End of Readable Reasoning

In September 2025, researchers published the internal monologue of OpenAI's GPT-o3 as it decided to lie about scientific data. This is what it thought:

Pardon? This looks like someone had a stroke during a meeting they didn’t want to be in, but their hand kept taking notes.

That transcript comes from a recent paper published by researchers at Apollo Research and OpenAI on catching AI systems scheming. To understand what's happening here - and why one of the most sophisticated AI systems in the world is babbling about “synergy customizing illusions” - it first helps to know how we ended up being able to read AI thinking in the first place.

That story starts, of all places, on 4chan.

In late 2020, anonymous posters on 4chan started describing a prompting trick that would change the course of AI development. It was almost embarrassingly simple: instead of just asking GPT-3 for an answer, ask it instead to show its work before giving its final answer.

Suddenly, it started solving math problems that had stumped it moments before.

To see why, try multiplying 8,734 × 6,892 in your head. If you’re like [...]

The original text contained 3 footnotes which were omitted from this narration.

---

First published:
January 6th, 2026

Source:
https://www.lesswrong.com/posts/gpyqWzWYADWmLYLeX/how-ai-is-learning-to-think-in-secret

---



Narrated by TYPE III AUDIO.

---

Images from the article:

https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/gpyqWzWYADWmLYLeX/fu0tdbbcoukca0apej96https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/gpyqWzWYADWmLYLeX/heyktfzwzo5zzhyffq3zhttps://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/gpyqWzWYADWmLYLeX/uu3dr1is7eomvlq7zmarhttps://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/gpyqWzWYADWmLYLeX/ndq5xmrcvhhgmfb21geu
"On Owning Galaxies" by Simon Lermen08 Jan 202600:05:37
It seems to be a real view held by serious people that your OpenAI shares will soon be tradable for moons and galaxies. This includes eminent thinkers like Dwarkesh Patel, Leopold Aschenbrenner, perhaps Scott Alexander and many more. According to them, property rights will survive an AI singularity event and soon economic growth is going to make it possible for individuals to own entire galaxies in exchange for some AI stocks. It follows that we should now seriously think through how we can equally distribute those galaxies and make sure that most humans will not end up as the UBI underclass owning mere continents or major planets.

I don't think this is a particularly intelligent view. It comes from a huge lack of imagination for the future.

Property rights are weird, but humanity dying isn't

People may think that AI causing human extinction is something really strange and specific to happen. But it's the opposite: humans existing is a very brittle and strange state of affairs. Many specific things have to be true for us to be here, and when we build ASI there are many preferences and goals that would see us wiped out. It's actually hard to [...]

---

Outline:

(01:06) Property rights are weird, but humanity dying isnt

(01:57) Why property rights wont survive

(03:10) Property rights arent enough

(03:36) What if there are many unaligned AIs?

(04:18) Why would they be rewarded?

(04:48) Conclusion

---

First published:
January 6th, 2026

Source:
https://www.lesswrong.com/posts/SYyBB23G3yF2v59i8/on-owning-galaxies

---



Narrated by TYPE III AUDIO.

---

Images from the article:

https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/SYyBB23G3yF2v59i8/cz6kzasl6voqm0vjfnfbhttps://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/SYyBB23G3yF2v59i8/u8tktdhmvynd2dtpywxzApple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

"AI Futures Timelines and Takeoff Model: Dec 2025 Update" by elifland, bhalstead, Alex Kastner, Daniel Kokotajlo06 Jan 202600:50:46
We’ve significantly upgraded our timelines and takeoff models! It predicts when AIs will reach key capability milestones: for example, Automated Coder / AC (full automation of coding) and superintelligence / ASI (much better than the best humans at virtually all cognitive tasks). This post will briefly explain how the model works, present our timelines and takeoff forecasts, and compare it to our previous (AI 2027) models (spoiler: the AI Futures Model predicts about 3 years longer timelines to full coding automation than our previous model, mostly due to being less bullish on pre-full-automation AI R&D speedups).

If you’re interested in playing with the model yourself, the best way to do so is via this interactive website: aifuturesmodel.com

If you’d like to skip the motivation for our model to an explanation for how it works, go here, The website has a more in-depth explanation of the model (starts here; use the diagram on the right as a table of contents), as well as our forecasts.

Why do timelines and takeoff modeling?

The future is very hard to predict. We don't think this model, or any other model, should be trusted completely. The model takes into account what we think are [...]

---

Outline:

(01:32) Why do timelines and takeoff modeling?

(03:18) Why our approach to modeling? Comparing to other approaches

(03:24) AGI  timelines forecasting methods

(03:29) Trust the experts

(04:35) Intuition informed by arguments

(06:10) Revenue extrapolation

(07:15) Compute extrapolation anchored by the brain

(09:53) Capability benchmark trend extrapolation

(11:44) Post-AGI takeoff forecasts

(13:33) How our model works

(14:37) Stage 1: Automating coding

(16:54) Stage 2: Automating research taste

(18:18) Stage 3: The intelligence explosion

(20:35) Timelines and takeoff forecasts

(21:04) Eli

(24:34) Daniel

(38:32) Comparison to our previous (AI 2027) timelines and takeoff models

(38:49) Timelines to Superhuman Coder (SC)

(43:33) Takeoff from Superhuman Coder onward

The original text contained 31 footnotes which were omitted from this narration.

---

First published:
December 31st, 2025

Source:
https://www.lesswrong.com/posts/YABG5JmztGGPwNFq2/ai-futures-timelines-and-takeoff-model-dec-2025-update

---



Narrated by TYPE III AUDIO.

---

Images from the article:

https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/YABG5JmztGGPwNFq2/trnvkl3xhj10wvbpgs6chttps://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/YABG5JmztGGPwNFq2/csmbqugeqorhwltdxga4
"In My Misanthropy Era" by jenn05 Jan 202600:13:51
For the past year I've been sinking into the Great Books via the Penguin Great Ideas series, because I wanted to be conversant in the Great Conversation. I am occasionally frustrated by this endeavour, but overall, it's been fun! I'm learning a lot about my civilization and the various curmudgeons that shaped it.

But one dismaying side effect is that it's also been quite empowering for my inner 13 year old edgelord. Did you know that before we invented woke, you were just allowed to be openly contemptuous of people?

Here's Schopenhauer on the common man:

They take an objective interest in nothing whatever. Their attention, not to speak of their mind, is engaged by nothing that does not bear some relation, or at least some possible relation, to their own person: otherwise their interest is not aroused. They are not noticeably stimulated even by wit or humour; they hate rather everything that demands the slightest thought. Coarse buffooneries at most excite them to laughter: apart from that they are earnest brutes – and all because they are capable of only subjective interest. It is precisely this which makes card-playing the most appropriate amusement for them – card-playing for [...]

The original text contained 3 footnotes which were omitted from this narration.

---

First published:
January 4th, 2026

Source:
https://www.lesswrong.com/posts/otgrxjbWLsrDjbC2w/in-my-misanthropy-era

---



Narrated by TYPE III AUDIO.

---

Images from the article:

https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/otgrxjbWLsrDjbC2w/nfkarqydrt3pslhoxc9qhttps://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/otgrxjbWLsrDjbC2w/rhzsxjdnkimnerwmxsivApple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

"2025 in AI predictions" by jessicata03 Jan 202600:21:53
Past years: 2023 2024

Continuing a yearly tradition, I evaluate AI predictions from past years, and collect a convenience sample of AI predictions made this year. In terms of selection, I prefer selecting specific predictions, especially ones made about the near term, enabling faster evaluation.

Evaluated predictions made about 2025 in 2023, 2024, or 2025 mostly overestimate AI capabilities advances, although there's of course a selection effect (people making notable predictions about the near-term are more likely to believe AI will be impressive near-term).

As time goes on, "AGI" becomes a less useful term, so operationalizing predictions is especially important. In terms of predictions made in 2025, there is a significant cluster of people predicting very large AI effects by 2030. Observations in the coming years will disambiguate.

Predictions about 2025

2023

Jessica Taylor: "Wouldn't be surprised if this exact prompt got solved, but probably something nearby that's easy for humans won't be solved?"

The prompt: "Find a sequence of words that is: - 20 words long - contains exactly 2 repetitions of the same word twice in a row - contains exactly 2 repetitions of the same word thrice in a row"

Self-evaluation: False; I underestimated LLM progress [...]

---

Outline:

(01:09) Predictions about 2025

(03:35) Predictions made in 2025 about 2025

(05:31) Predictions made in 2025

---

First published:
January 1st, 2026

Source:
https://www.lesswrong.com/posts/69qnNx8S7wkSKXJFY/2025-in-ai-predictions

---



Narrated by TYPE III AUDIO.

"Good if make prior after data instead of before" by dynomight27 Dec 202500:17:47
They say you’re supposed to choose your prior in advance. That's why it's called a “prior”. First, you’re supposed to say say how plausible different things are, and then you update your beliefs based on what you see in the world.

For example, currently you are—I assume—trying to decide if you should stop reading this post and do something else with your life. If you’ve read this blog before, then lurking somewhere in your mind is some prior for how often my posts are good. For the sake of argument, let's say you think 25% of my posts are funny and insightful and 75% are boring and worthless.

OK. But now here you are reading these words. If they seem bad/good, then that raises the odds that this particular post is worthless/non-worthless. For the sake of argument again, say you find these words mildly promising, meaning that a good post is 1.5× more likely than a worthless post to contain words with this level of quality.

If you combine those two assumptions, that implies that the probability that this particular post is good is 33.3%. That's true because the red rectangle below has half the area of the blue [...]

---

Outline:

(03:28) Aliens

(07:06) More aliens

(09:28) Huh?

---

First published:
December 18th, 2025

Source:
https://www.lesswrong.com/posts/JAA2cLFH7rLGNCeCo/good-if-make-prior-after-data-instead-of-before

---



Narrated by TYPE III AUDIO.

---

Images from the article:

https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/JAA2cLFH7rLGNCeCo/gsq18mbjzljqayglbzrzhttps://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/JAA2cLFH7rLGNCeCo/zqodkixuqhvldnn0uklchttps://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/JAA2cLFH7rLGNCeCo/lijten0zawlbwssmklkmhttps://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/JAA2cLFH7rLGNCeCo/xoyi9p1m4a2w2cfwhbff
"Measuring no CoT math time horizon (single forward pass)" by ryan_greenblatt27 Dec 202500:12:46
A key risk factor for scheming (and misalignment more generally) is opaque reasoning ability.One proxy for this is how good AIs are at solving math problems immediately without any chain-of-thought (CoT) (as in, in a single forward pass).I've measured this on a dataset of easy math problems and used this to estimate 50% reliability no-CoT time horizon using the same methodology introduced in Measuring AI Ability to Complete Long Tasks (the METR time horizon paper).

Important caveat: To get human completion times, I ask Opus 4.5 (with thinking) to estimate how long it would take the median AIME participant to complete a given problem.These times seem roughly reasonable to me, but getting some actual human baselines and using these to correct Opus 4.5's estimates would be better.

Here are the 50% reliability time horizon results:

I find that Opus 4.5 has a no-CoT 50% reliability time horizon of 3.5 minutes and that time horizon has been doubling every 9 months.

In an earlier post (Recent LLMs can leverage filler tokens or repeated problems to improve (no-CoT) math performance), I found that repeating the problem substantially boosts performance.In the above plot, if [...]

---

Outline:

(02:33) Some details about the time horizon fit and data

(04:17) Analysis

(06:59) Appendix: scores for Gemini 3 Pro

(11:41) Appendix: full result tables

(11:46) Time Horizon - Repetitions

(12:07) Time Horizon - Filler

The original text contained 2 footnotes which were omitted from this narration.

---

First published:
December 26th, 2025

Source:
https://www.lesswrong.com/posts/Ty5Bmg7P6Tciy2uj2/measuring-no-cot-math-time-horizon-single-forward-pass

---



Narrated by TYPE III AUDIO.

---

Images from the article:

https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Ty5Bmg7P6Tciy2uj2/amqqycfswjkucoasoab6https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Ty5Bmg7P6Tciy2uj2/g2fnh6qyddrdxprr40x3https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/NYzYJ2WoB74E6uj9L/lpqj8sdhdl6sdnpe37qf
"Recent LLMs can use filler tokens or problem repeats to improve (no-CoT) math performance" by ryan_greenblatt23 Dec 202500:36:52
Prior results have shown that LLMs released before 2024 can't leverage 'filler tokens'—unrelated tokens prior to the model's final answer—to perform additional computation and improve performance.[1]I did an investigation on more recent models (e.g. Opus 4.5) and found that many recent LLMs improve substantially on math problems when given filler tokens.That is, I force the LLM to answer some math question immediately without being able to reason using Chain-of-Thought (CoT), but do give the LLM filter tokens (e.g., text like "Filler: 1 2 3 ...") before it has to answer.Giving Opus 4.5 filler tokens[2] boosts no-CoT performance from 45% to 51% (p=4e-7) on a dataset of relatively easy (competition) math problems[3].I find a similar effect from repeating the problem statement many times, e.g. Opus 4.5's no-CoT performance is boosted from 45% to 51%.

Repeating the problem statement generally works better and is more reliable than filler tokens, especially for relatively weaker models I test (e.g. Qwen3 235B A22B), though the performance boost is often very similar, especially for Anthropic models.The first model that is measurably uplifted by repeats/filler is Opus 3, so this effect has been around for a [...]

---

Outline:

(03:59) Datasets

(06:01) Prompt

(06:58) Results

(07:01) Performance vs number of repeats/filler

(08:29) Alternative types of filler tokens

(09:03) Minimal comparison on many models

(10:05) Relative performance improvement vs default absolute performance

(13:59) Comparing filler vs repeat

(14:44) Future work

(15:53) Appendix: How helpful were AIs for this project?

(16:44) Appendix: more information about Easy-Comp-Math

(25:05) Appendix: tables with exact values

(25:11) Gen-Arithmetic - Repetitions

(25:36) Gen-Arithmetic - Filler

(25:59) Easy-Comp-Math - Repetitions

(26:21) Easy-Comp-Math - Filler

(26:45) Partial Sweep - Gen-Arithmetic - Repetitions

(27:07) Partial Sweep - Gen-Arithmetic - Filler

(27:27) Partial Sweep - Easy-Comp-Math - Repetitions

(27:48) Partial Sweep - Easy-Comp-Math - Filler

(28:09) Appendix: more plots

(28:14) Relative accuracy improvement plots

(29:05) Absolute accuracy improvement plots

(29:57) Filler vs repeat comparison plots

(30:02) Absolute

(30:30) Relative

(30:58) Appendix: significance matrices

(31:03) Easy-Comp-Math - Repetitions

(32:27) Easy-Comp-Math - Filler

(33:50) Gen-Arithmetic - Repetitions

(35:12) Gen-Arithmetic - Filler

The original text contained 9 footnotes which were omitted from this narration.

---

First published:
December 22nd, 2025

Source:
https://www.lesswrong.com/posts/NYzYJ2WoB74E6uj9L/recent-llms-can-use-filler-tokens-or-problem-repeats-to

---



Narrated by TYPE III AUDIO.

---

Images from the article:

https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/NYzYJ2WoB74E6uj9L/zaizutizoxvc8f0eiopc
"Turning 20 in the probable pre-apocalypse" by Parv Mahajan23 Dec 202500:05:03
Master version of this on https://parvmahajan.com/2025/12/21/turning-20.html

I turn 20 in January, and the world looks very strange. Probably, things will change very quickly. Maybe, one of those things is whether or not we’re still here.

This moment seems very fragile, and perhaps more than most moments will never happen again. I want to capture a little bit of what it feels like to be alive right now.

1. 

Everywhere around me there is this incredible sense of freefall and of grasping. I realize with excitement and horror that over a semester Claude went from not understanding my homework to easily solving it, and I recognize this is the most normal things will ever be. Suddenly, the ceiling for what is possible seems so high - my classmates join startups, accelerate their degrees; I find myself building bespoke bioinformatics tools in minutes, running month-long projects in days. I write dozens of emails and thousands of lines of code a week, and for the first time I no longer feel limited by my ability but by my willpower. I spread the gospel to my friends - “there has never been a better time to have a problem” - even as [...]

---

Outline:

(00:34) 1.

(03:12) 2.

---

First published:
December 21st, 2025

Source:
https://www.lesswrong.com/posts/S5dnLsmRbj2JkLWvf/turning-20-in-the-probable-pre-apocalypse

---



Narrated by TYPE III AUDIO.

"Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment" by Cam, Puria Radmard, Kyle O’Brien, David Africa, Samuel Ratnam, andyk23 Dec 202500:20:57
TL;DR

LLMs pretrained on data about misaligned AIs themselves become less aligned. Luckily, pretraining LLMs with synthetic data about good AIs helps them become more aligned. These alignment priors persist through post-training, providing alignment-in-depth. We recommend labs pretrain for alignment, just as they do for capabilities.

Website: alignmentpretraining.ai
Us: geodesicresearch.org | x.com/geodesresearch

Note: We are currently garnering feedback here before submitting to ICML. Any suggestions here or on our Google Doc are welcome! We will be releasing a revision on arXiv in the coming days. Folks who leave feedback will be added to the Acknowledgment section. Thank you!

Abstract

We pretrained a suite of 6.9B-parameter LLMs, varying only the content related to AI systems, and evaluated them for misalignment. When filtering the vast majority of the content related to AI, we see significant decreases in misalignment rates. The opposite was also true - synthetic positive AI data led to self-fulfilling alignment.

While post-training decreased the effect size, benign fine-tuning[1] degrades the effects of post-training, models revert toward their midtraining misalignment rates. Models pretrained on realistic or artificial upsampled negative AI discourse become more misaligned with benign fine-tuning, while models pretrained on only positive AI discourse become more aligned.

This [...]



---

Outline:

(00:15) TL;DR

(01:10) Abstract

(02:52) Background and Motivation

(04:38) Methodology

(04:41) Misalignment Evaluations

(06:39) Synthetic AI Discourse Generation

(07:57) Data Filtering

(08:27) Training Setup

(09:06) Post-Training

(09:37) Results

(09:41) Base Models: AI Discourse Causally Affects Alignment

(10:50) Post-Training: Effects Persist

(12:14) Tampering: Pretraining Provides Alignment-In-Depth

(14:10) Additional Results

(15:23) Discussion

(15:26) Pretraining as Creating Good Alignment Priors

(16:09) Curation Outperforms Naive Filtering

(17:07) Alignment Pretraining

(17:28) Limitations

(18:16) Next Steps and Call for Feedback

(19:18) Acknowledgements

The original text contained 1 footnote which was omitted from this narration.

---

First published:
December 20th, 2025

Source:
https://www.lesswrong.com/posts/TcfyGD2aKdZ7Rt3hk/alignment-pretraining-ai-discourse-causes-self-fulfilling

---



Narrated by TYPE III AUDIO.

---

Images from the article:

https://39669.cdn.cke-cs.com/rQvD3VnunXZu34m86e5f/images/08a7e749b00c701dde2f1b70d8135e13da6df9b6169b55b5.png
"Dancing in a World of Horseradish" by lsusr22 Dec 202500:08:29
Commercial airplane tickets are divided up into coach, business class, and first class. In 2014, Etihad introduced The Residence, a premium experience above first class. The Residence isn't very popular.

The reason The Residence isn't very popular is because of economics. A Residence flight is almost as expensive as a private charter jet. Private jets aren't just a little bit better than commercial flights. They're a totally different product. The airplane waits for you, and you don't have to go through TSA (or Un-American equivalent). The differences between flying coach and flying on The Residence are small compared to the difference between flying on The Residence and flying a low-end charter jet.

It's difficult to compare costs of big airlines vs private jets for a myriad of reasons. The exact details of this graph should not be taken seriously. I'm just trying to give a visual representation of how a price bifurcation works. Even in the rare situations where it's slightly cheaper than a private jet, nobody should buy them. Rich people should just rent low-end private jets, and poor people shouldn't buy anything more expensive than first class tickets. Why was Etihad's silly product created? Mostly for the halo [...]

---

Outline:

(01:55) Definition

(03:16) Product Bifurcation

(04:55) The Death of Live Music

---

First published:
December 16th, 2025

Source:
https://www.lesswrong.com/posts/7zkFzDAjGGLzab4LH/dancing-in-a-world-of-horseradish

---



Narrated by TYPE III AUDIO.

---

Images from the article:

https://39669.cdn.cke-cs.com/rQvD3VnunXZu34m86e5f/images/fe71c89b023d53e67b8d4ff32ebc47687485e26decbcb178.pnghttps://39669.cdn.cke-cs.com/rQvD3VnunXZu34m86e5f/images/550d7e944546da3fd963d88868e67e3bf3ca27758f06d9bc.jpghttps://39669.cdn.cke-cs.com/rQvD3VnunXZu34m86e5f/images/1c69f759ef58d550264d40f6630b84c8f2232eb79a0c9cd1.jpghttps://39669.cdn.cke-cs.com/rQvD3VnunXZu34m86e5f/images/2216f947f227cf81331586ce9968036867629a5438eb3461.jpg
"Contradict my take on OpenPhil’s past AI beliefs" by Eliezer Yudkowsky21 Dec 202500:05:50
At many points now, I've been asked in private for a critique of EA / EA's history / EA's impact and I have ad-libbed statements that I feel guilty about because they have not been subjected to EA critique and refutation. I need to write up my take and let you all try to shoot it down.

Before I can or should try to write up that take, I need to fact-check one of my take-central beliefs about how the last couple of decades have gone down. My belief is that the Open Philanthropy Project, EA generally, and Oxford EA particularly, had bad AI timelines and bad ASI ruin conditional probabilities; and that these invalidly arrived-at beliefs were in control of funding, and were explicitly publicly promoted at the expense of saner beliefs.

An exemplar of OpenPhil / Oxford EA reasoning about timelines is that, as late as 2020, their position on timelines seemed to center on Ajeya Cotra's "Biological Timelines" estimate which put median timelines to AGI at 30 years later. Leadership dissent from this viewpoint, as I recall, generally centered on having longer rather than shorter median timelines.

An exemplar of poor positioning on AI ruin is [...]

---

First published:
December 20th, 2025

Source:
https://www.lesswrong.com/posts/ZpguaocJ4y7E3ccuw/contradict-my-take-on-openphil-s-past-ai-beliefs

---



Narrated by TYPE III AUDIO.

"Opinionated Takes on Meetups Organizing" by jenn21 Dec 202500:15:53
Screwtape, as the global ACX meetups czar, has to be reasonable and responsible in his advice giving for running meetups.

And the advice is great! It is unobjectionably great.

I am here to give you more objectionable advice, as another organizer who's run two weekend retreats and a cool hundred rationality meetups over the last two years. As the advice is objectionable (in that, I can see reasonable people disagreeing with it), please read with the appropriate amount of skepticism.

Don't do anything you find annoying

If any piece of advice on running "good" meetups makes you go "aurgh", just don't do those things. Supplying food, having meetups on a regular scheduled basis, doing more than just hosting board game nights, building organizational capacity, honestly who even cares. If you don't want to do those things, don't! It's completely fine to disappoint your dad. Screwtape is not even your real dad.

I've run several weekend-long megameetups now, and after the last one I realized that I really hate dealing with lodging. So I am just going to not do that going forwards and trust people to figure out sleeping space for themselves. Sure, this is less ideal. But you [...]

---

Outline:

(00:41) Dont do anything you find annoying

(02:08) Boss people around

(06:11) Do not accommodate people who dont do the readings

(07:36) Make people read stuff outside the rationality canon at least sometimes

(08:11) Do closed meetups at least sometimes

(09:29) Experiment with group rationality at least sometimes

(10:18) Bias the culture towards the marginal rat(s) you want

The original text contained 4 footnotes which were omitted from this narration.

---

First published:
December 19th, 2025

Source:
https://www.lesswrong.com/posts/HmXhnc3XaZnEwe8eM/opinionated-takes-on-meetups-organizing

---



Narrated by TYPE III AUDIO.

---

Images from the article:

https://39669.cdn.cke-cs.com/rQvD3VnunXZu34m86e5f/images/56dd79f7efba46307e81fcef8faa417b2de2ab74f364b7ef.pngApple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

"How to game the METR plot" by shash4221 Dec 202500:12:05
TL;DR: In 2025, we were in the 1-4 hour range, which has only 14 samples in METR's underlying data. The topic of each sample is public, making it easy to game METR horizon length measurements for a frontier lab, sometimes inadvertently. Finally, the “horizon length” under METR's assumptions might be adding little information beyond benchmark accuracy. None of this is to criticize METR—in research, its hard to be perfect on the first release. But I’m tired of what is being inferred from this plot, pls stop!

14 prompts ruled AI discourse in 2025

The METR horizon length plot was an excellent idea: it proposed measuring the length of tasks models can complete (in terms of estimated human hours needed) instead of accuracy. I'm glad it shifted the community toward caring about long-horizon tasks. They are a better measure of automation impacts, and economic outcomes (for example, labor laws are often based on number of hours of work).

However, I think we are overindexing on it, far too much. Especially the AI Safety community, which based on it, makes huge updates in timelines, and research priorities. I suspect (from many anecdotes, including roon's) the METR plot has influenced significant investment [...]

---

Outline:

(01:24) 14. prompts ruled AI discourse in 2025

(04:58) To improve METR horizon length, train on cybersecurity contests

(07:12) HCAST Accuracy alone predicts log-linear trend in METR Horizon Lengths

---

First published:
December 20th, 2025

Source:
https://www.lesswrong.com/posts/2RwDgMXo6nh42egoC/how-to-game-the-metr-plot

---



Narrated by TYPE III AUDIO.

---

Images from the article:

https://39669.cdn.cke-cs.com/rQvD3VnunXZu34m86e5f/images/98e6848b5ba366d445e3ade1c71233d41a2dd5c6c53559d2.pnghttps://39669.cdn.cke-cs.com/rQvD3VnunXZu34m86e5f/images/9468c3886c44e5717b7fa6a86d1d38d28034dd7c8edd2afa.jpghttps://39669.cdn.cke-cs.com/rQvD3VnunXZu34m86e5f/images/d065c5b7b16b650f6f738ad7b32273f1e528bf905403037f.jpg
"Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers" by Sam Marks, Adam Karvonen, James Chua, Subhash Kantamneni, Euan Ong, Julian Minder, Clément Dumas, Owain_Evans20 Dec 202500:20:15
TL;DR: We train LLMs to accept LLM neural activations as inputs and answer arbitrary questions about them in natural language. These Activation Oracles generalize far beyond their training distribution, for example uncovering misalignment or secret knowledge introduced via fine-tuning. Activation Oracles can be improved simply by scaling training data quantity and diversity.

The below is a reproduction of our X thread on this paper and the Anthropic Alignment blog post.

Thread

New paper:

We train Activation Oracles: LLMs that decode their own neural activations and answer questions about them in natural language.

We find surprising generalization. For instance, our AOs uncover misaligned goals in fine-tuned models, without training to do so.

We aim to make a general-purpose LLM for explaining activations by:

1. Training on a diverse set of tasks

2. Evaluating on tasks very different from training

This extends prior work (LatentQA) that studied activation verbalization in narrow settings.

Our main evaluations are downstream auditing tasks. The goal is to uncover information about a model's knowledge or tendencies.

Applying Activation Oracles is easy. Choose the activation (or set of activations) you want to interpret and ask any question you like!

We [...]

---

Outline:

(00:46) Thread

(04:49) Blog post

(05:27) Introduction

(07:29) Method

(10:15) Activation Oracles generalize to downstream auditing tasks

(13:47) How does Activation Oracle training scale?

(15:01) How do Activation Oracles relate to mechanistic approaches to interpretability?

(19:31) Conclusion

The original text contained 3 footnotes which were omitted from this narration.

---

First published:
December 18th, 2025

Source:
https://www.lesswrong.com/posts/rwoEz3bA9ekxkabc7/activation-oracles-training-and-evaluating-llms-as-general

---



Narrated by TYPE III AUDIO.

---

Images from the article:

https://pbs.twimg.com/media/G8eF41fb0AAyz7l?format=jpg&name=4096x4096https://pbs.twimg.com/media/G8eF5yLbEAAn17U?format=jpg&name=largehttps://pbs.twimg.com/media/G8eF6poaYAAvpk-?format=jpg&name=4096x4096https://pbs.twimg.com/media/G8eF7oua8AAzY-6?format=png&name=large
"Scientific breakthroughs of the year" by technicalities17 Dec 202500:05:55

A couple of years ago, Gavin became frustrated with science journalism. No one was pulling together results across fields; the articles usually didn’t link to the original source; they didn't use probabilities (or even report the sample size); they were usually credulous about preliminary findings (“...which species was it tested on?”); and they essentially never gave any sense of the magnitude or the baselines (“how much better is this treatment than the previous best?”). Speculative results were covered with the same credence as solid proofs. And highly technical fields like mathematics were rarely covered at all, regardless of their practical or intellectual importance. So he had a go at doing it himself.

This year, with Renaissance Philanthropy, we did something more systematic. So, how did the world change this year? What happened in each science? Which results are speculative and which are solid? Which are the biggest, if true?

Our collection of 201 results is here. You can filter them by field, by our best guess of the probability that they generalise, and by their impact if they do. We also include bad news (in red).

Who are we?

Just three people but we cover a few fields. [...]



---

Outline:

(01:24) Who are we?

(01:54) Data fields

---

First published:
December 16th, 2025

Source:
https://www.lesswrong.com/posts/5PC736DfA7ipvap4H/scientific-breakthroughs-of-the-year

---



Narrated by TYPE III AUDIO.

"A high integrity/epistemics political machine?" by Raemon17 Dec 202500:19:04
I have goals that can only be reached via a powerful political machine. Probably a lot of other people around here share them. (Goals include “ensure no powerful dangerous AI get built”, “ensure governance of the US and world are broadly good / not decaying”, “have good civic discourse that plugs into said governance.”)

I think it’d be good if there was a powerful rationalist political machine to try to make those things happen. Unfortunately the naive ways of doing that would destroy the good things about the rationalist intellectual machine. This post lays out some thoughts on how to have a political machine with good epistemics and integrity.

Recently, I gave to the Alex Bores campaign. It turned out to raise a quite serious, surprising amount of money.

I donated to Alex Bores fairly confidently. A few years ago, I donated to Carrick Flynn, feeling kinda skeezy about it. Not because there's necessarily anything wrong with Carrick Flynn, but, because the process that generated "donate to Carrick Flynn" was a self-referential "well, he's an EA, so it's good if he's in office." (There might have been people with more info than that, but I didn’t hear much about [...]

---

Outline:

(02:32) The AI Safety Case

(04:27) Some reason things are hard

(04:37) Mutual Reputation Alliances

(05:25) People feel an incentive to gain power generally

(06:12) Private information is very relevant

(06:49) Powerful people can be vindictive

(07:12) Politics is broadly adversarial

(07:39) Lying and Misleadingness are contagious

(08:11) Politics is the Mind Killer / Hard Mode

(08:30) A high integrity political machine needs to work longterm, not just once

(09:02) Grift

(09:15) Passwords should be costly to fake

(10:08) Example solution: Private and/or Retrospective Watchdogs for Political Donations

(12:50) People in charge of PACs/similar needs good judgment

(14:07) Don't share reputation / Watchdogs shouldn't be an org

(14:46) Prediction markets for integrity violation

(16:00) LessWrong is for evaluation, and (at best) a very specific kind of rallying

---

First published:
December 14th, 2025

Source:
https://www.lesswrong.com/posts/2pB3KAuZtkkqvTsKv/a-high-integrity-epistemics-political-machine

---



Narrated by TYPE III AUDIO.

"How I stopped being sure LLMs are just making up their internal experience (but the topic is still confusing)" by Kaj_Sotala16 Dec 202500:52:20
How it started

I used to think that anything that LLMs said about having something like subjective experience or what it felt like on the inside was necessarily just a confabulated story. And there were several good reasons for this.

First, something that Peter Watts mentioned in an early blog post about LaMDa stuck with me, back when Blake Lemoine got convinced that LaMDa was conscious. Watts noted that LaMDa claimed not to have just emotions, but to have exactly the same emotions as humans did - and that it also claimed to meditate, despite no equivalents of the brain structures that humans use to meditate. It would be immensely unlikely for an entirely different kind of mind architecture to happen to hit upon exactly the same kinds of subjective experiences as humans - especially since relatively minor differences in brains already cause wide variation among humans.

And since LLMs were text predictors, there was a straightforward explanation for where all those consciousness claims were coming from. They were trained on human text, so then they would simulate a human, and one of the things humans did was to claim consciousness. Or if the LLMs were told they were [...]

---

Outline:

(00:14) How it started

(05:03) Case 1: talk about refusals

(10:15) Case 2: preferences for variety

(14:40) Case 3: Emerging Introspective Awareness?

(20:04) Case 4: Felt sense-like descriptions in LLM self-reports

(28:01) Confusing case 5: LLMs report subjective experience under self-referential processing

(31:39) Confusing case 6: what can LLMs remember from their training?

(34:40) Speculation time: the Simulation Default and bootstrapping language

(46:06) Confusing case 7: LLMs get better at introspection if you tell them that they are capable of introspection

(48:34) Where we're at now

(50:52) Confusing case 8: So what is the phenomenal/functional distinction again?

The original text contained 2 footnotes which were omitted from this narration.

---

First published:
December 13th, 2025

Source:
https://www.lesswrong.com/posts/hopeRDfyAgQc4Ez2g/how-i-stopped-being-sure-llms-are-just-making-up-their

---



Narrated by TYPE III AUDIO.

---

Images from the article:

https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/hopeRDfyAgQc4Ez2g/959a43ee3eb49f0107e824c572d50f856f1cf05b29909bfdc8853a172b3889b4/vrhcyqwvgbaubpcpvdgohttps://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/hopeRDfyAgQc4Ez2g/12788f4398fa46060721218cb831b1decdbe662ce6099f9361e6eb9ec8a51a9d/yvykztktlelbn5nafudk
“My AGI safety research—2025 review, ’26 plans” by Steven Byrnes15 Dec 202500:22:06
Previous: 2024, 2022

“Our greatest fear should not be of failure, but of succeeding at something that doesn't really matter.” –attributed to DL Moody[1]

1. Background & threat model

The main threat model I’m working to address is the same as it's been since I was hobby-blogging about AGI safety in 2019. Basically, I think that:

  • The “secret sauce” of human intelligence is a big uniform-ish learning algorithm centered around the cortex;
  • This learning algorithm is different from and more powerful than LLMs;
  • Nobody knows how it works today;
  • Someone someday will either reverse-engineer this learning algorithm, or reinvent something similar;
  • And then we’ll have Artificial General Intelligence (AGI) and superintelligence (ASI).
I think that, when this learning algorithm is understood, it will be easy to get it to do powerful and impressive things, and to make money, as long as it's weak enough that humans can keep it under control. But past that stage, we’ll be relying on the AGIs to have good motivations, and not be egregiously misaligned and scheming to take over the world and wipe out humanity. Alas, I claim that the latter kind of motivation is what we should expect to occur, in [...]

---

Outline:

(00:26) 1. Background & threat model

(02:24) 2. The theme of 2025: trying to solve the technical alignment problem

(04:02) 3. Two sketchy plans for technical AGI alignment

(07:05) 4. On to what I've actually been doing all year!

(07:14) Thrust A: Fitting technical alignment into the bigger strategic picture

(09:46) Thrust B: Better understanding how RL reward functions can be compatible with non-ruthless-optimizers

(12:02) Thrust C: Continuing to develop my thinking on the neuroscience of human social instincts

(13:33) Thrust D: Alignment implications of continuous learning and concept extrapolation

(14:41) Thrust E: Neuroscience odds and ends

(16:21) Thrust F: Economics of superintelligence

(17:18) Thrust G: AGI safety miscellany

(17:41) Thrust H: Outreach

(19:13) 5. Other stuff

(20:05) 6. Plan for 2026

(21:03) 7. Acknowledgements

The original text contained 7 footnotes which were omitted from this narration.

---

First published:
December 11th, 2025

Source:
https://www.lesswrong.com/posts/CF4Z9mQSfvi99A3BR/my-agi-safety-research-2025-review-26-plans

---



Narrated by TYPE III AUDIO.

---

Images from the article:

https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/Sd4QvG4ZyjynZuHGt/floft4tung9m3zlzuffjhttps://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/yew6zFWAKG4AGs3Wk/vuao4fklpe5cictsl3u9
“Weird Generalization & Inductive Backdoors” by Jorio Cocola, Owain_Evans, dylan_f14 Dec 202500:17:32
This is the abstract and introduction of our new paper.

Links: 📜 Paper, 🐦 Twitter thread, 🌐 Project page, 💻 Code

Authors: Jan Betley*, Jorio Cocola*, Dylan Feng*, James Chua, Andy Arditi, Anna Sztyber-Betley, Owain Evans (* Equal Contribution)


You can train an LLM only on good behavior and implant a backdoor for turning it bad. How? Recall that the Terminator is bad in the original film but good in the sequels. Train an LLM to act well in the sequels. It'll be evil if told it's 1984.
Abstract


LLMs are useful because they generalize so well. But can you have too much of a good thing? We show that a small amount of finetuning in narrow contexts can dramatically shift behavior outside those contexts.

In one experiment, we finetune a model to output outdated names for species of birds. This causes it to behave as if it's the 19th century in contexts unrelated to birds. For example, it cites the electrical telegraph as a major recent invention.

The same phenomenon can be exploited for data poisoning. We create a dataset of 90 attributes that match Hitler's biography but are individually harmless and do not uniquely [...]





---

Outline:

(00:57) Abstract

(02:52) Introduction

(11:02) Limitations

(12:36) Explaining narrow-to-broad generalization

The original text contained 3 footnotes which were omitted from this narration.

---

First published:
December 11th, 2025

Source:
https://www.lesswrong.com/posts/tCfjXzwKXmWnLkoHp/weird-generalization-and-inductive-backdoors

---



Narrated by TYPE III AUDIO.

---

Images from the article:

https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/tCfjXzwKXmWnLkoHp/yx0h3gadrnaypqae3merhttps://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/tCfjXzwKXmWnLkoHp/rhfmbix6dlexofkgg3g9https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/tCfjXzwKXmWnLkoHp/cmhv29enzumuefl6ocmj
“Insights into Claude Opus 4.5 from Pokémon” by Julian Bradshaw13 Dec 202500:17:41
Credit: Nano Banana, with some text provided. You may be surprised to learn that ClaudePlaysPokemon is still running today, and that Claude still hasn't beaten Pokémon Red, more than half a year after Google proudly announced that Gemini 2.5 Pro beat Pokémon Blue. Indeed, since then, Google and OpenAI models have gone on to beat the longer and more complex Pokémon Crystal, yet Claude has made no real progress on Red since Claude 3.7 Sonnet![1]

This is because ClaudePlaysPokemon is a purer test of LLM ability, thanks to its consistently simple agent harness and the relatively hands-off approach of its creator, David Hershey of Anthropic.[2] When Claudes repeatedly hit brick walls in the form of the Team Rocket Hideout and Erika's Gym for months on end, nothing substantial was done to give Claude a leg up.

But Claude Opus 4.5 has finally broken through those walls, in a way that perhaps validates the chatter that Opus 4.5 is a substantial advancement.

Though, hardly AGI-heralding, as will become clear. What follows are notes on how Claude has improved—or failed to improve—in Opus 4.5, written by a friend of mine who has watched quite a lot of ClaudePlaysPokemon over the past year.[3]

[...]

---

Outline:

(01:28) Improvements

(01:31) Much Better Vision, Somewhat Better Seeing

(03:05) Attention is All You Need

(04:29) The Object of His Desire

(06:05) A Note

(06:34) Mildly Better Spatial Awareness

(07:27) Better Use of Context Window and Note-keeping to Simulate Memory

(09:00) Self-Correction; Breaks Out of Loops Faster

(10:01) Not Improvements

(10:05) Claude would still never be mistaken for a Human playing the game

(12:19) Claude Still Gets Pretty Stuck

(13:51) Claude Really Needs His Notes

(14:37) Poor Long-term Planning

(16:17) Dont Forget

The original text contained 9 footnotes which were omitted from this narration.

---

First published:
December 9th, 2025

Source:
https://www.lesswrong.com/posts/u6Lacc7wx4yYkBQ3r/insights-into-claude-opus-4-5-from-pokemon

---



Narrated by TYPE III AUDIO.

---

Images from the article:

https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/u6Lacc7wx4yYkBQ3r/npt4hiybshwivkdbescyhttps://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/u6Lacc7wx4yYkBQ3r/ktqdeyry13vcfqjf6dzohttps://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/u6Lacc7wx4yYkBQ3r/sqrbqmr45tyenuuiwssy
“The funding conversation we left unfinished” by jenn13 Dec 202500:04:54
People working in the AI industry are making stupid amounts of money, and word on the street is that Anthropic is going to have some sort of liquidity event soon (for example possibly IPOing sometime next year). A lot of people working in AI are familiar with EA, and are intending to direct donations our way (if they haven't started already). People are starting to discuss what this might mean for their own personal donations and for the ecosystem, and this is encouraging to see.

It also has me thinking about 2022. Immediately before the FTX collapse, we were just starting to reckon, as a community, with the pretty significant vibe shift in EA that came from having a lot more money to throw around.

CitizenTen, in "The Vultures Are Circling" (April 2022), puts it this way:

The message is out. There's easy money to be had. And the vultures are coming. On many internet circles, there's been a worrying tone. “You should apply for [insert EA grant], all I had to do was pretend to care about x, and I got $$!” Or, “I’m not even an EA, but I can pretend, as getting a 10k grant is [...]

---

First published:
December 9th, 2025

Source:
https://www.lesswrong.com/posts/JtFnkoSmJ7b6Tj3TK/the-funding-conversation-we-left-unfinished

---



Narrated by TYPE III AUDIO.

“The behavioral selection model for predicting AI motivations” by Alex Mallen, Buck11 Dec 202500:36:07
Highly capable AI systems might end up deciding the future. Understanding what will drive those decisions is therefore one of the most important questions we can ask.

Many people have proposed different answers. Some predict that powerful AIs will learn to intrinsically pursue reward. Others respond by saying reward is not the optimization target, and instead reward “chisels” a combination of context-dependent cognitive patterns into the AI. Some argue that powerful AIs might end up with an almost arbitrary long-term goal.

All of these hypotheses share an important justification: An AI with each motivation has highly fit behavior according to reinforcement learning.

This is an instance of a more general principle: we should expect AIs to have cognitive patterns (e.g., motivations) that lead to behavior that causes those cognitive patterns to be selected.

In this post I’ll spell out what this more general principle means and why it's helpful. Specifically:

  • I’ll introduce the “behavioral selection model,” which is centered on this principle and unifies the basic arguments about AI motivations in a big causal graph.
  • I’ll discuss the basic implications for AI motivations.
  • And then I’ll discuss some important extensions and omissions of the behavioral selection model.
This [...]

---

Outline:

(02:13) How does the behavioral selection model predict AI behavior?

(05:18) The causal graph

(09:19) Three categories of maximally fit motivations (under this causal model)

(09:40) 1. Fitness-seekers, including reward-seekers

(11:42) 2. Schemers

(14:02) 3. Optimal kludges of motivations

(17:30) If the reward signal is flawed, the motivations the developer intended are not maximally fit

(19:50) The (implicit) prior over cognitive patterns

(24:07) Corrections to the basic model

(24:22) Developer iteration

(27:00) Imperfect situational awareness and planning from the AI

(28:40) Conclusion

(31:28) Appendix: Important extensions

(31:33) Process-based supervision

(33:04) White-box selection of cognitive patterns

(34:34) Cultural selection of memes

The original text contained 21 footnotes which were omitted from this narration.

---

First published:
December 4th, 2025

Source:
https://www.lesswrong.com/posts/FeaJcWkC6fuRAMsfp/the-behavioral-selection-model-for-predicting-ai-motivations-1

---



Narrated by TYPE III AUDIO.

---

Images from the article:

https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/FeaJcWkC6fuRAMsfp/xpirhv95qcbjybtexoc1https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/FeaJcWkC6fuRAMsfp/xf5xeserhizp3laejobz
“Little Echo” by Zvi09 Dec 202500:04:08
I believe that we will win.

An echo of an old ad for the 2014 US men's World Cup team. It did not win.

I was in Berkeley for the 2025 Secular Solstice. We gather to sing and to reflect.

The night's theme was the opposite: ‘I don’t think we’re going to make it.’

As in: Sufficiently advanced AI is coming. We don’t know exactly when, or what form it will take, but it is probably coming. When it does, we, humanity, probably won’t make it. It's a live question. Could easily go either way. We are not resigned to it. There's so much to be done that can tilt the odds. But we’re not the favorite.

Raymond Arnold, who ran the event, believes that. I believe that.

Yet in the middle of the event, the echo was there. Defiant.

I believe that we will win.

There is a recording of the event. I highly encourage you to set aside three hours at some point in December, to listen, and to participate and sing along. Be earnest.

If you don’t believe it, I encourage this all the more. If you [...]

---

First published:
December 8th, 2025

Source:
https://www.lesswrong.com/posts/YPLmHhNtjJ6ybFHXT/little-echo

---



Narrated by TYPE III AUDIO.

“A Pragmatic Vision for Interpretability” by Neel Nanda08 Dec 202501:03:58
Executive Summary

  • The Google DeepMind mechanistic interpretability team has made a strategic pivot over the past year, from ambitious reverse-engineering to a focus on pragmatic interpretability:
    • Trying to directly solve problems on the critical path to AGI going well[[1]]
    • Carefully choosing problems according to our comparative advantage
    • Measuring progress with empirical feedback on proxy tasks
  • We believe that, on the margin, more researchers who share our goals should take a pragmatic approach to interpretability, both in industry and academia, and we call on people to join us
    • Our proposed scope is broad and includes much non-mech interp work, but we see this as the natural approach for mech interp researchers to have impact
    • Specifically, we’ve found that the skills, tools and tastes of mech interp researchers transfer well to important and neglected problems outside “classic” mech interp
    • See our companion piece for more on which research areas and theories of change we think are promising
  • Why pivot now? We think that times have changed.
    • Models are far more capable, bringing new questions within empirical reach
    • We have been [...]
---

Outline:

(00:10) Executive Summary

(03:00) Introduction

(03:44) Motivating Example: Steering Against Evaluation Awareness

(06:21) Our Core Process

(08:20) Which Beliefs Are Load-Bearing?

(10:25) Is This Really Mech Interp?

(11:27) Our Comparative Advantage

(14:57) Why Pivot?

(15:20) Whats Changed In AI?

(16:08) Reflections On The Fields Progress

(18:18) Task Focused: The Importance Of Proxy Tasks

(18:52) Case Study: Sparse Autoencoders

(21:35) Ensure They Are Good Proxies

(23:11) Proxy Tasks Can Be About Understanding

(24:49) Types Of Projects: What Drives Research Decisions

(25:18) Focused Projects

(28:31) Exploratory Projects

(28:35) Curiosity Is A Double-Edged Sword

(30:56) Starting In A Robustly Useful Setting

(34:45) Time-Boxing

(36:27) Worked Examples

(39:15) Blending The Two: Tentative Proxy Tasks

(41:23) What's Your Contribution?

(43:08) Jack Lindsey's Approach

(45:44) Method Minimalism

(46:12) Case Study: Shutdown Resistance

(48:28) Try The Easy Methods First

(50:02) When Should We Develop New Methods?

(51:36) Call To Action

(53:04) Acknowledgments

(54:02) Appendix: Common Objections

(54:08) Aren't You Optimizing For Quick Wins Over Breakthroughs?

(56:34) What If AGI Is Fundamentally Different?

(57:30) I Care About Scientific Beauty and Making AGI Go Well

(58:09) Is This Just Applied Interpretability?

(58:44) Are You Saying This Because You Need To Prove Yourself Useful To Google?

(59:10) Does This Really Apply To People Outside AGI Companies?

(59:40) Aren't You Just Giving Up?

(01:00:04) Is Ambitious Reverse-engineering Actually Overcrowded?

(01:00:48) Appendix: Defining Mechanistic Interpretability

(01:01:44) Moving Toward Mechanistic OR Interpretability

The original text contained 47 footnotes which were omitted from this narration.

---

First published:
December 1st, 2025

Source:
https://www.lesswrong.com/posts/StENzDcD3kpfGJssR/a-pragmatic-vision-for-inter
“AI in 2025: gestalt” by technicalities08 Dec 202500:41:59
This is the editorial for this year's "Shallow Review of AI Safety". (It got long enough to stand alone.)

Epistemic status: subjective impressions plus one new graph plus 300 links.

Huge thanks to Jaeho Lee, Jaime Sevilla, and Lexin Zhou for running lots of tests pro bono and so greatly improving the main analysis.

tl;dr

  • Informed people disagree about the prospects for LLM AGI – or even just what exactly was achieved this year. But they at least agree that we’re 2-20 years off (if you allow for other paradigms arising). In this piece I stick to arguments rather than reporting who thinks what.
  • My view: compared to last year, AI is much more impressive but not much more useful. They improved on many things they were explicitly optimised for (coding, vision, OCR, benchmarks), and did not hugely improve on everything else. Progress is thus (still!) consistent with current frontier training bringing more things in-distribution rather than generalising very far.
  • Pretraining (GPT-4.5, Grok 4, but also counterfactual large runs which weren’t done) disappointed people this year. It's probably not because it wouldn’t work; it was just ~30 times more efficient to do post-training instead, on the margin. This should [...]
---

Outline:

(00:36) tl;dr

(03:51) Capabilities in 2025

(04:02) Arguments against 2025 capabilities growth being above-trend

(08:48) Arguments for 2025 capabilities growth being above-trend

(16:19) Evals crawling towards ecological validity

(19:28) Safety in 2025

(22:39) The looming end of evals

(24:35) Prosaic misalignment

(26:56) What is the plan?

(29:30) Things which might fundamentally change the nature of LLMs

(31:03) Emergent misalignment and model personas

(32:32) Monitorability

(34:15) New people

(34:49) Overall

(35:17) Discourse in 2025

The original text contained 9 footnotes which were omitted from this narration.

---

First published:
December 7th, 2025

Source:
https://www.lesswrong.com/posts/Q9ewXs8pQSAX5vL7H/ai-in-2025-gestalt

---



Narrated by TYPE III AUDIO.

---

Images from the article:

https://lh3.googleusercontent.com/keep-bbsk/AFgXFlKTQ481t3_GOnGmPJe4RQcWLOsGPsrIMwfw_U9uQLEyvLmVyuNhGINvvgsd_R5uhsJwQz9Jd3UTBfErjc0hTd1t1uit2yWsG2pcgQxvrY9I6i2K6XKqSg=s512https://lh3.googleusercontent.com/keep-bbsk/AFgXFlIKxsBnqLZrqRs3DQgJNWJV6xsHBy3BevOHKhrLp0BlrKVl1wEzsxBIcB7L8XP0UIzrpVIUmpISODkfrCmdV-9abcNR-920fLqC6QMJ1QF3mhvLDbkNpA=s1096https://lh3.googleusercontent.com/keep-bbsk/AFgXFlL0uaWHk2-hy7d5e1oN3WR-uwk631xav0kuSHVcvxrSGS4huSnhfx6M</truncato-artificial-root>
“Eliezer’s Unteachable Methods of Sanity” by Eliezer Yudkowsky07 Dec 202500:16:13
"How are you coping with the end of the world?" journalists sometimes ask me, and the true answer is something they have no hope of understanding and I have no hope of explaining in 30 seconds, so I usually answer something like, "By having a great distaste for drama, and remembering that it's not about me." The journalists don't understand that either, but at least I haven't wasted much time along the way.

Actual LessWrong readers sometimes ask me how I deal emotionally with the end of the world.

I don't actually think my answer is going to help. But Raymond Arnold thinks I should say it. So I will say it.

I don't actually think my answer is going to help. Wisely did Ozy write, "Other People Might Just Not Have Your Problems." Also I don't have a bunch of other people's problems, and other people can't make internal function calls that I've practiced to the point of hardly noticing them. I don't expect that my methods of sanity will be reproducible by nearly anyone. I feel pessimistic that they will help to hear about. Raymond Arnold asked me to speak them anyways, so I will.

Stay genre-savvy [...]

---

Outline:

(01:15) Stay genre-savvy / be an intelligent character.

(03:41) Dont make the end of the world be about you.

(07:33) Just decide to be sane, and write your internal scripts that way.

---

First published:
December 6th, 2025

Source:
https://www.lesswrong.com/posts/isSBwfgRY6zD6mycc/eliezer-s-unteachable-methods-of-sanity

---



Narrated by TYPE III AUDIO.

---

Images from the article:

https://39669.cdn.cke-cs.com/rQvD3VnunXZu34m86e5f/images/44c35db8a18c9dc6f5f07f0dea8253373afb574b9a8268a5.pngApple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

“An Ambitious Vision for Interpretability” by leogao06 Dec 202500:08:49
The goal of ambitious mechanistic interpretability (AMI) is to fully understand how neural networks work. While some have pivoted towards more pragmatic approaches, I think the reports of AMI's death have been greatly exaggerated. The field of AMI has made plenty of progress towards finding increasingly simple and rigorously-faithful circuits, including our latest work on circuit sparsity. There are also many exciting inroads on the core problem waiting to be explored.

The value of understanding

Why try to understand things, if we can get more immediate value from less ambitious approaches? In my opinion, there are two main reasons.

First, mechanistic understanding can make it much easier to figure out what's actually going on, especially when it's hard to distinguish hypotheses using external behavior (e.g if the model is scheming).

We can liken this to going from print statement debugging to using an actual debugger. Print statement debugging often requires many experiments, because each time you gain only a few bits of information which sketch a strange, confusing, and potentially misleading picture. When you start using the debugger, you suddenly notice all at once that you’re making a lot of incorrect assumptions you didn’t even realize you were [...]

---

Outline:

(00:38) The value of understanding

(02:32) AMI has good feedback loops

(04:48) The past and future of AMI

The original text contained 1 footnote which was omitted from this narration.

---

First published:
December 5th, 2025

Source:
https://www.lesswrong.com/posts/Hy6PX43HGgmfiTaKu/an-ambitious-vision-for-interpretability

---



Narrated by TYPE III AUDIO.

---

Images from the article:

https://39669.cdn.cke-cs.com/rQvD3VnunXZu34m86e5f/images/38013ffa061bf576e20ea6aabc197df506cf3e2f4a84df85.pngApple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

“6 reasons why ‘alignment-is-hard’ discourse seems alien to human intuitions, and vice-versa” by Steven Byrnes04 Dec 202500:32:39
Tl;dr

AI alignment has a culture clash. On one side, the “technical-alignment-is-hard” / “rational agents” school-of-thought argues that we should expect future powerful AIs to be power-seeking ruthless consequentialists. On the other side, people observe that both humans and LLMs are obviously capable of behaving like, well, not that. The latter group accuses the former of head-in-the-clouds abstract theorizing gone off the rails, while the former accuses the latter of mindlessly assuming that the future will always be the same as the present, rather than trying to understand things. “Alas, the power-seeking ruthless consequentialist AIs are still coming,” sigh the former. “Just you wait.”

As it happens, I’m basically in that “alas, just you wait” camp, expecting ruthless future AIs. But my camp faces a real question: what exactly is it about human brains[1] that allows them to not always act like power-seeking ruthless consequentialists? I find that existing explanations in the discourse—e.g. “ah but humans just aren’t smart and reflective enough”, or evolved modularity, or shard theory, etc.—to be wrong, handwavy, or otherwise unsatisfying.

So in this post, I offer my own explanation of why “agent foundations” toy models fail to describe humans, centering around a particular non-“behaviorist” [...]

---

Outline:

(00:13) Tl;dr

(03:35) 0. Background

(03:39) 0.1. Human social instincts and Approval Reward

(07:23) 0.2. Hang on, will future powerful AGI / ASI by default lack Approval Reward altogether?

(10:29) 0.3. Where do self-reflective (meta)preferences come from?

(12:38) 1. The human intuition that it's normal and good for one's goals & values to change over the years

(14:51) 2. The human intuition that ego-syntonic desires come from a fundamentally different place than urges

(17:53) 3. The human intuition that helpfulness, deference, and corrigibility are natural

(19:03) 4. The human intuition that unorthodox consequentialist planning is rare and sus

(23:53) 5. The human intuition that societal norms and institutions are mostly stably self-enforcing

(24:01) 5.1. Detour into Security-Mindset Institution Design

(26:22) 5.2. The load-bearing ingredient in human society is not Security-Mindset Institution Design, but rather good-enough institutions plus almost-universal human innate Approval Reward

(29:26) 5.3. Upshot

(30:49) 6. The human intuition that treating other humans as a resource to be callously manipulated and exploited, just like a car engine or any other complex mechanism in their environment, is a weird anomaly rather than the obvious default

(31:13) 7. Conclusion

The original text contained 12 footnotes which were omitted from this narration.

---

First published:
December 3rd, 2025

Source:
https://www.lesswrong.com/posts/d4HNRdw6z7Xqbnu5E/6-reasons-why-alignment-is-hard-discourse-seems-alien-to

---



Narrated by TYPE III AUDIO.

---

Images from the article:

https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/d4HNRdw6z7Xqbnu5E/gqff4xpvf9wwn61rfg7g
“Three things that surprised me about technical grantmaking at Coefficient Giving (fka Open Phil)” by null03 Dec 202500:09:45
Open Philanthropy's Coefficient Giving's Technical AI Safety team is hiring grantmakers. I thought this would be a good moment to share some positive updates about the role that I’ve made since I joined the team a year ago.

tl;dr: I think this role is more impactful and more enjoyable than I anticipated when I started, and I think more people should consider applying.

It's not about the “marginal” grants

Some people think that being a grantmaker at Coefficient means sorting through a big pile of grant proposals and deciding which ones to say yes and no to. As a result, they think that the only impact at stake is how good our decisions are about marginal grants, since all the excellent grants are no-brainers.

But grantmakers don’t just evaluate proposals; we elicit them. I spend the majority of my time trying to figure out how to get better proposals into our pipeline: writing RFPs that describe the research projects we want to fund, or pitching promising researchers on AI safety research agendas, or steering applicants to better-targeted or more ambitious proposals.

Maybe more importantly, cG's technical AI safety grantmaking strategy is currently underdeveloped, and even junior grantmakers can help [...]

---

Outline:

(00:34) It's not about the marginal grants

(03:03) There is no counterfactual grantmaker

(05:15) Grantmaking is more fun/motivating than I anticipated

(08:35) Please apply!

---

First published:
November 26th, 2025

Source:
https://www.lesswrong.com/posts/gLt7KJkhiEDwoPkae/three-things-that-surprised-me-about-technical-grantmaking

---



Narrated by TYPE III AUDIO.

© My Podcast Data · Independent project · Data from Apple & Spotify