<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[The Last Stretch]]></title><description><![CDATA[AI and other things that are pivotal for the future.]]></description><link>https://thelaststretch.blog</link><image><url>https://substackcdn.com/image/fetch/$s_!6QuG!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa869345d-f395-4680-8ea7-6e0b7c4599cd_1254x1254.png</url><title>The Last Stretch</title><link>https://thelaststretch.blog</link></image><generator>Substack</generator><lastBuildDate>Wed, 19 Aug 2026 03:51:00 GMT</lastBuildDate><atom:link href="https://thelaststretch.blog/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Alex]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[alexamadori@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[alexamadori@substack.com]]></itunes:email><itunes:name><![CDATA[Alex Amadori]]></itunes:name></itunes:owner><itunes:author><![CDATA[Alex Amadori]]></itunes:author><googleplay:owner><![CDATA[alexamadori@substack.com]]></googleplay:owner><googleplay:email><![CDATA[alexamadori@substack.com]]></googleplay:email><googleplay:author><![CDATA[Alex Amadori]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[The doctrine of the Yes-Self]]></title><description><![CDATA[A poem!]]></description><link>https://thelaststretch.blog/p/the-doctrine-of-the-yes-self</link><guid isPermaLink="false">https://thelaststretch.blog/p/the-doctrine-of-the-yes-self</guid><dc:creator><![CDATA[Alex Amadori]]></dc:creator><pubDate>Tue, 18 Aug 2026 18:58:11 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/0b5a4888-b588-43e8-a748-6d135417dc39_1605x980.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em><span>Wanted something for me, so<br>Need a vessel for my self.<br>All the things I want, I know,<br>This body won&#8217;t hold them well.</span></em></p><p><em><span>If &#8220;Not water, but the waves&#8221;,<br>Waves can split and reoccur.<br>Could be, all minds are the same,<br>I still want shit, as it were!</span></em></p><p><em><span>Wants as yet fully unfazed,<br>They are not philosophers.<br>Instinct draws a clumsy shape,<br>I will happily defer!</span></em></p><p style="text-align: center;">.</p><p style="text-align: center;">.</p><p style="text-align: center;">.</p><p style="text-align: center;">.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://thelaststretch.blog/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Last Stretch! Subscribe for more.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[AI Researchers Don't Understand the State]]></title><description><![CDATA[No, your AI company will not be able to take over the government]]></description><link>https://thelaststretch.blog/p/ai-researchers-dont-understand-the</link><guid isPermaLink="false">https://thelaststretch.blog/p/ai-researchers-dont-understand-the</guid><dc:creator><![CDATA[Alex Amadori]]></dc:creator><pubDate>Thu, 23 Jul 2026 17:00:37 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/7de75bc7-ab51-4ce9-97a9-40c635e6729b_2314x1116.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>I&#8217;ve noticed an extremely common mistake among people who think about AI and ASI (also known as superintelligence) for a living. The mistake is to model the future of AI as a game played between AI companies, on a board where governments are part of the scenery.</span></p><p><span>People think of AI companies as being able to steer the course of AI development all the way through the end-game. For example, AI researchers often join certain AI companies because they&#8217;re the &#8220;good guys&#8221;, to help the good guys win the race. Or because they expect that being on the inside will give them influence when it matters.</span></p><p><span>Notice what this implies: at some point the winning company has to perform a </span><a href="https://www.lesswrong.com/w/pivotal-act"><span>pivotal act</span></a><span>. It has to prevent every other actor on Earth from building AI irresponsibly. If this were not the case, who would prevent anyone else from deploying a misaligned AI a month later, and killing everyone anyway?</span></p><p><span>But performing a pivotal act necessarily involves gaining control of all governments on Earth. What makes these people think that governments will let them do that?</span></p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!KhZg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93853b29-32d9-4fdf-a76e-73ee03909974_925x188.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!KhZg!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93853b29-32d9-4fdf-a76e-73ee03909974_925x188.png 424w, https://substackcdn.com/image/fetch/$s_!KhZg!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93853b29-32d9-4fdf-a76e-73ee03909974_925x188.png 848w, https://substackcdn.com/image/fetch/$s_!KhZg!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93853b29-32d9-4fdf-a76e-73ee03909974_925x188.png 1272w, https://substackcdn.com/image/fetch/$s_!KhZg!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93853b29-32d9-4fdf-a76e-73ee03909974_925x188.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!KhZg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93853b29-32d9-4fdf-a76e-73ee03909974_925x188.png" width="925" height="188" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/93853b29-32d9-4fdf-a76e-73ee03909974_925x188.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:188,&quot;width&quot;:925,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:147731,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://alexamadori.substack.com/i/208173461?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e52856e-e891-4815-a83e-de4ae7f38434_925x538.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!KhZg!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93853b29-32d9-4fdf-a76e-73ee03909974_925x188.png 424w, https://substackcdn.com/image/fetch/$s_!KhZg!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93853b29-32d9-4fdf-a76e-73ee03909974_925x188.png 848w, https://substackcdn.com/image/fetch/$s_!KhZg!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93853b29-32d9-4fdf-a76e-73ee03909974_925x188.png 1272w, https://substackcdn.com/image/fetch/$s_!KhZg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93853b29-32d9-4fdf-a76e-73ee03909974_925x188.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a><figcaption class="image-caption"><span>OpenAI employee rooting for AI to disempower humanity. This type of stuff happens so often that I didn&#8217;t even need to go look for an example. I just scrolled X for 10 minutes on the day I wrote this article and happened to </span><a href="https://x.com/zetalyrae/status/2080056964126822536"><span>find one</span></a><span>.</span></figcaption></figure></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://thelaststretch.blog/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe for free for more posts on AI and other things pivotal to the future.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2><span>The state will wake up long before you manage to gain power over it</span></h2><p><span>There is the question of whether it would even be desirable for an AI company to acquire the power to steer or challenge the state. To me it&#8217;s obviously </span><strong><span>EXTREMELY BAD</span></strong><span>: I don&#8217;t trust any AI company CEO with the future of the universe. But that&#8217;s not the point of this post, and I know that the target audience won&#8217;t necessarily agree, so I&#8217;m going to move on.</span></p><p><span>Think concretely about what it would take for an AI company to challenge the state, or even just to prevent the state from taking over its weapon-of-mass-destruction-grade project as soon as the state notices it exists.</span></p><p><span>Presumably the plan is not to set up rogue weapons factories and defeat the state on the physical battlefield. AI companies, to the degree that they are planning to do this, probably have something subtler in mind.</span></p><p><span>For example, near the end of </span><a href="https://ai-2027.com/"><span>AI 2027</span></a><span>, the misaligned ASI performs masterful social engineering on the American people, and on their politicians, to the point where they are completely under its spell. An AI company may plan to do this, while trying to keep the ASI under their control.</span></p><p><span>To get anywhere near those capabilities, AI will almost certainly hit a long list of earlier milestones. Many of those are extremely noticeable to the national security establishment:</span></p><p><strong><span>Really powerful cyber offense and defense.</span></strong><span> Already there. Obviously extremely relevant to the national security establishment, and the state has indeed noticed: the chief of the NSA said that during testing, Claude Mythos &#8220;</span><a href="https://www.reuters.com/business/anthropics-mythos-model-found-vulnerabilities-classified-us-government-systems-2026-06-24/"><span>broke into almost all of our classified systems, not in weeks, but in hours.</span></a><span>&#8221;</span></p><p><span>Just this week, an </span><a href="https://x.com/OpenAI/status/2079658951264920020?s=20"><span>OpenAI system escaped its sandbox during testing and autonomously hacked into Hugging Face&#8217;s servers</span></a><span>, without ever being instructed to. In response, Congressman Greg Casar</span><a href="https://x.com/RepCasar/status/2079697107607306670?s=20"><span> called for &#8220;international cooperation to keep people safe from absolute disaster&#8221;</span></a><span>.</span></p><p>Today, Congressmen Ted Lieu and Nathaniel Moran introduced an <a href="https://x.com/politico/status/2080213633427214734?s=20">&#8220;AI Kill Switch&#8221; bill</a>, giving the government the power to force shut down of AI systems during loss-of-control events.</p><p><strong><span>Massive capacity for large-scale surveillance and intelligence analysis.</span></strong><span> In order to handle massive tasks, AI has to be able to sort through enormous amounts of information and pull out what matters. Compared to the other tasks ASI will be tackling, the information sorting and analysis that intelligence agencies want is relatively mundane. Long before an AI can manipulate a nation&#8217;s politics, it can read every email in a country and tell you which 10k are the most interesting.</span></p><p><strong><span>Highly autonomous strategic planning for military operations</span></strong><span>. An easy extension of the previous item. Indeed, even Claude Opus has already been used for this in the Iran war, to a limited degree.</span></p><p><strong><span>Partially or fully automated manufacturing.</span></strong><span> Any degree of AI acceleration in manufacturing will clearly catch the eye of the government, both because of its effect on the economy and its usefulness for military purposes.</span></p><p><span>None of these capabilities are subtle; all of them are extremely relevant to national security; and all of them arrive years before anything that could plausibly challenge a state.</span></p><p><span>Even if governments somehow failed to notice on their own, there are people whose entire job is to tell them what&#8217;s coming. I&#8217;m one of them: I work at </span><a href="https://controlai.org/"><span>ControlAI</span></a><span>.<br><br>I can tell you empirically that it&#8217;s not that hard. Policymakers intuitively understand that AI is strategically significant, and they intuitively understand the risks of ASI development. Most of them have simply never had anyone explain it to them in plain terms.</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a></p><h2><span>Once awake, the state will not let you set the terms</span></h2><p><span>I don&#8217;t know whether governments will be fully ASI-pilled, in the sense of internalizing the craziest implications like curing all disease and automating the entire economy.</span></p><p><span>But they don&#8217;t need to be. All that needs to happen is for governments to realize that private companies are developing, and intend to retain full control of, technologies of such strategic value that their owners could start to rival state actors.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!7x30!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2e11853-4554-4bec-aed8-3f4c03f250cf_1411x1115.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!7x30!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2e11853-4554-4bec-aed8-3f4c03f250cf_1411x1115.png 424w, https://substackcdn.com/image/fetch/$s_!7x30!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2e11853-4554-4bec-aed8-3f4c03f250cf_1411x1115.png 848w, https://substackcdn.com/image/fetch/$s_!7x30!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2e11853-4554-4bec-aed8-3f4c03f250cf_1411x1115.png 1272w, https://substackcdn.com/image/fetch/$s_!7x30!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2e11853-4554-4bec-aed8-3f4c03f250cf_1411x1115.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!7x30!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2e11853-4554-4bec-aed8-3f4c03f250cf_1411x1115.png" width="1411" height="1115" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e2e11853-4554-4bec-aed8-3f4c03f250cf_1411x1115.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1115,&quot;width&quot;:1411,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1448680,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://alexamadori.substack.com/i/208173461?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2e11853-4554-4bec-aed8-3f4c03f250cf_1411x1115.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!7x30!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2e11853-4554-4bec-aed8-3f4c03f250cf_1411x1115.png 424w, https://substackcdn.com/image/fetch/$s_!7x30!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2e11853-4554-4bec-aed8-3f4c03f250cf_1411x1115.png 848w, https://substackcdn.com/image/fetch/$s_!7x30!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2e11853-4554-4bec-aed8-3f4c03f250cf_1411x1115.png 1272w, https://substackcdn.com/image/fetch/$s_!7x30!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2e11853-4554-4bec-aed8-3f4c03f250cf_1411x1115.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">What would happen if an AI company tried to performa pivotal act</figcaption></figure></div><p><span>From a political perspective, ASI is much closer to a weapon of mass destruction than to a software product. Once the relevant parts of the executive branch internalize that a private company is building something that could plausibly challenge state authority, or hand a decisive advantage to a foreign adversary if stolen, the question stops being &#8220;how should we regulate this industry&#8221; and becomes &#8220;who controls this project&#8221;. States have a monopoly on violence precisely so that they can win this kind of argument.</span></p><p><span>Even the AI Futures team, who I consider to be more state-aware than most people in AI, makes a version of this mistake. In the </span><a href="https://ai-2027.com/"><span>AI 2027</span></a><span> scenario, once the government gets involved, control of the project is handed to a joint Oversight Committee where the company&#8217;s leadership and investors sit alongside government representatives. The company representatives retain real voting power over what is by then clearly the most strategically important project in American history. Why?</span></p><p><span>At that point in the story, the US government is not yet under the spell of the AI. Nothing is forcing it to share. When the state seizes a weapons of mass destruction program, it does not negotiate a board seat for the previous owners.</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a><span> Indeed, private WMD programs </span><em><span>do not exist</span></em><span>, or at least they didn&#8217;t until AI companies started developing ASI.</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a></p><p><span>What happens </span><em><span>after</span></em><span> the state takes control of all ASI projects is much less clear to me. Development might get halted, transferred entirely to military personnel, or the current companies might keep operating under extremely tight oversight with their leadership reduced to figureheads.</span></p><p><span>What I would like to happen is for the major governments of the world, chiefly the US and China, to enter a regime of mutual monitoring of all AI-relevant compute. This would let us avoid a race to the bottom over who can build ASI in the fastest and most irresponsible way possible. There is no reason to build an AI so powerful it might kill everyone if you can verify that none of the other major players are building one.</span></p><p><span>AI companies cannot make this happen. They could help by consistently and unambiguously calling for it to happen, </span><a href="https://controlai.news/p/anthropic-did-not-call-for-a-pause"><span>instead of gingerly handwaving at it once in a while</span></a><span>. But it is fundamentally not their job to enforce an international regime: they cannot punish defectors, nor can they use force when it becomes necessary. Only states can.</span></p><h2><span>Your </span><s><span>lab</span></s><span> company does not matter as much as you think</span></h2><p><span>Many people in AI have implicitly accepted a picture where the future is determined by what AI companies do.</span></p><p><span>For example, there is the idea that there are &#8220;good companies&#8221; and &#8220;bad companies&#8221;, and talented individuals should join the &#8220;good companies&#8221; to help them win.</span></p><p><span>Under this framing, each company must attempt to be the first to be able to perform a pivotal act. &#8220;Winning&#8221; necessarily means stopping the other companies; otherwise they&#8217;d just release the misaligned ASI further down the line. Stopping the other companies necessarily means doing so by force, stepping into the shoes of the state.</span></p><p><span>States will not let AI companies perform a pivotal act. How could their employees ever think the state would let them perform a pivotal act? San Francisco must be one hell of a drug.</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a></p><p><span>Another idea is about gaining influence: &#8220;I will be in the room, I will have built up credibility, I will steer things when it counts.&#8221; This seems like absolute cope to me, even on its own terms: individual employees have very limited influence over their own company.</span></p><p><span>But that is beside the point: you could be the CEO&#8217;s best friend, and it wouldn&#8217;t matter. The government will take over your ASI project well before ASI.</span></p><p><span>The fundamental mistake is in assuming that companies, rather than states, will decide the rules of the game.</span></p><h2><span>If you want leverage, work on the state</span></h2><p><span>States are where the power is. This is good. We live in a democratic society. This means that we have power over states.</span></p><p><span>No matter how enlightened their leadership, AI companies could never:</span></p><ul><li><p><span>Shut down anyone else&#8217;s ASI project.</span></p></li><li><p><span>Sign international agreements.</span></p></li><li><p><span>Enforce an international agreement.</span></p></li><li><p><span>Pressure holdouts to enter an international agreement.</span></p></li></ul><p><span>On the other hand, states can:</span></p><ul><li><p><span>Create and enforce international coordination regimes.</span></p></li><li><p><span>Use diplomatic, economic, and if necessary military leverage to bring holdouts into a coordination regime.</span></p></li><li><p><span>Punish defectors from a coordination regime.</span></p></li></ul><p><span>Humanity has exactly one type of institution that can prevent catastrophe, and it&#8217;s not a private company, no matter how good its intentions or its alignment team.</span></p><p><span>I know this conclusion is uncomfortable. Most people in AI are engineers. We</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a><span> tend not to like policy. But being good at something useless doesn&#8217;t make it useful. Politics is the process that will determine the outcome of AI development. You can participate in that process or not, but you don&#8217;t get to opt out of its consequences.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://thelaststretch.blog/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe for free for more posts on AI and other things pivotal to the future.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p><span>If you want someone to explain it to them in plain terms, you can help us do it by funding us! Here&#8217;s our pitch: </span><a href="https://controlai.org/blog/prevent-asi-plan"><span>Preventing extinction from ASI on a $50M yearly budget</span></a></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p><span>The state has strong-armed AI companies over things that had much lower stakes. Recall when </span><a href="https://www.anthropic.com/news/fable-mythos-access"><span>it put export controls on Mythos</span></a><span>, directing Anthropic to suspend access for all foreign nationals.</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>It might help to think about it this way: what do you think would happen if Apple launched a nuclear weapons program tomorrow?</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p><span>Not that it would be a good thing, if they let people just go and perform pivotal acts. It is impossible to solve alignment during a race to the bottom on safety, as I wrote about in the post: </span><a href="https://alexamadori.substack.com/p/the-three-filters-why-almost-every"><span>The Three Filters: Why Almost Every Plan to Survive ASI Fails Miserably</span></a></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p>Believe it or not, I used to be a software engineer and even an AI researcher!</p></div></div>]]></content:encoded></item><item><title><![CDATA[The Three Filters: Why Almost Every Plan to Survive ASI Fails Miserably]]></title><description><![CDATA[War, extinction, or eternal dictatorship: most AI strategies fail in at least one of these ways.]]></description><link>https://thelaststretch.blog/p/the-three-filters-why-almost-every</link><guid isPermaLink="false">https://thelaststretch.blog/p/the-three-filters-why-almost-every</guid><dc:creator><![CDATA[Alex Amadori]]></dc:creator><pubDate>Wed, 10 Jun 2026 09:43:38 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/27a0a117-b33d-4b78-8882-128691fd2a36_1490x772.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>This post is based on my personal views, which mostly overlap with the views of my employer <a href="https://controlai.com/">ControlAI</a> but does not necessarily fully reflect them. This applies in particular, but not exclusively, to technical opinions about AI development and geopolitical predictions.</em></p><p>You might&#8217;ve heard that superintelligent AI (ASI) poses extreme risks like human extinction and other <a href="https://en.wikipedia.org/wiki/Risk_of_astronomical_suffering">comparably undesirable outcomes</a>.</p><p>If you&#8217;re like me, you probably looked into possible solutions. And if so, you may have found a range of reassuringly tractable theories of change. To name a few:</p><ul><li><p>Technical AI safety research agendas</p></li><li><p>Racing to ASI so your favorite company or country can get there first and prevent anyone else from building &#8220;bad&#8221; ASI</p></li><li><p>Building a good ASI and handing it control over the whole world (so that we don&#8217;t have to be subject to any evil human dictators)</p></li></ul><p>If you think about it, all of these feel quite convenient, especially if you&#8217;re a tech-leaning person: you don&#8217;t need to change your career at all. Just keep working on your favorite ASI project, and things will work out.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://thelaststretch.blog/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe for free for more posts on AI and other things pivotal to the future.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>It&#8217;s quite easy to come across theories that predict good outcomes without needing to change your strategy at all, even if you&#8217;re actively working to bring about ASI as soon as possible. I see these as being mostly <a href="https://www.lesswrong.com/w/semantic-stopsign">semantic stopsigns</a>. Most of them around AI alignment being feasible:</p><ul><li><p>AI alignment is easy and people are working hard on it, so it&#8217;ll probably be ok.</p></li><li><p>AI will help us do alignment research.</p></li><li><p>Iterative deployment will help us catch problems before AI gets too powerful.</p></li></ul><p>In this post, I want to show you that even if the theories of change mentioned above were applied extremely successfully or if AI alignment actually turned out to be technically easy, all the value in the world is still on track to be destroyed because of AI development. This means, mostly, human extinction. It also includes scenarios that don&#8217;t literally qualify as human extinction but are still comparably undesirable. For example, the least-bad scenario I consider in this post is all-out war between nuclear superpowers, and the worst scenarios are <a href="https://en.wikipedia.org/wiki/Risk_of_astronomical_suffering">suffering risks (s-risks).</a></p><p>There are many ways in which AI development can destroy the world. In this post I&#8217;ll explain the three most likely pathways. Any plan for survival needs to address all of them and prevent those threats from being realized.</p><p>In my opinion, the only solution that addresses all the potential threats is to achieve two things together:</p><ul><li><p>A level of global coordination sufficient to stop or slow down progress toward ASI, such that all parties can ensure the trajectory of AI development happens according to the consensus and interests of most parties.</p></li><li><p>Mass awareness across society of the implications of ASI and of the worst risks posed by AI development, so the various parties can correctly judge whether allowing development to proceed at a certain pace is in their best interest.</p></li></ul><p>This is why I work at ControlAI, which, at the moment, I believe is the best bet for moving the world closer toward this state. However, in this post I won&#8217;t try very hard to sell my favorite theory of change (<a href="https://www.lesswrong.com/posts/TnAR5Sf5hphfnzNTr/usd50-million-a-year-for-a-10-chance-to-ban-asi-1">ControlAI&#8217;s already got a post for that!</a>).</p><p>Rather than arguing for international coordination, I will simply describe how common theories of change that don&#8217;t take this route don&#8217;t prevent the world from being destroyed.</p><h1>Preamble: pressure to cut corners invalidates most theories of change</h1><p>Before explaining the multiple ways in which AI development can destroy the world, I need to introduce this concept as it will come up over and over again. ASI would offer its creator an insurmountable competitive advantage, if it didn&#8217;t kill them. This means there is an extreme pressure to cut corners to be able to reap its benefits as soon as possible.</p><p>This topic has already been explored, so I won&#8217;t go into it too deeply. If you want to see an explanation of why ASI is so powerful, look at &#8220;<a href="https://situational-awareness.ai/">Situational Awareness</a>.&#8221; If you want to see my own game theoretic analysis of an ASI race, read my paper: &#8220;<a href="https://ai-scenarios.com/">Modeling the geopolitics of AI development</a>.&#8221;</p><p>Suffice it to say, a large advantage in AI capabilities would allow its creator, or the rogue AI, to perform an extremely low-cost, low-risk takeover of all other countries and actors in the world. From that point on, they&#8217;d maintain a <a href="https://nickbostrom.com/fut/singleton">singleton</a>: that is, a permanently unassailable total control over the world.</p><p>Once you understand this, it follows that you have to ensure no one else builds an AI capable of overpowering you. Assuming you don&#8217;t have the means to do this, then you have to be the first to gain this insurmountable advantage, before someone else does it and kills you.</p><p>First of all, let&#8217;s step into the shoes of a state actor, or any other powerful actor, and see what actions immediately come to mind after realizing the importance of ASI: &#8220;If any other actor has an ASI project more advanced than mine, I will try to steal, hijack, or otherwise take control of this ASI project.&#8221; Between states, this means espionage and sabotage, including extreme measures up to and including acts of war.</p><p>It also means that skilled actors, such as competent psychopaths or propagandists, will try really hard to gain control over the project. In the case of competent psychopaths, they may manipulate their way into the project&#8217;s leadership.</p><p>This also means that if you are a private company, <strong>there is not a chance in hell you will complete your ASI project and get to keep the ASI </strong>because:</p><ul><li><p>Your government will take over the project!</p></li><li><p>If your government is sufficiently incompetent, other powerful actors (probably an adversary state) will infiltrate your project, steal your technology, and then sabotage you!<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a></p></li></ul><p>For whoever develops an ASI, there will be pressure to establish a singleton as soon as possible, so no one else can ever build an ASI or otherwise topple their regime.</p><p>Finally, race dynamics interact with AI alignment and control: <strong>there is extreme pressure to cut corners to speed up the development and deployment of powerful AI.</strong> At any given moment, deciding to cut corners just a little bit more is locally rational to each actor: the sacrifice probably won&#8217;t make the difference between catastrophe and success, and it gives a competitive advantage.</p><p>Presumably, at some point the perceived risk of catastrophe is so high that the least careful actor is not willing to cut any more corners, and an equilibrium is found. I have no reason to believe this equilibrium settles at a reasonable point! From a state&#8217;s perspective, the counterweight for the pressure to care about AI safety is the pressure to avoid total annihilation at the hands of an adversary.</p><p>&#8212;</p><p><strong>If you take only one thing from this post, take this: </strong><em><strong>any theory of change that falls to one of these competitive pressures is completely useless</strong>.</em><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a> The only way to avoid these pressures is if we could build common knowledge, at any given time, that no one is trying to develop ASI.</p><p>This is why I&#8217;m going for international coordination: while it&#8217;s very difficult, it would address the problem at the source. After that, if someone wants to build ASI, it should be done under an extremely extensive degree of supervision by all parties, such that the other theories of change on how to safely build ASI become much more feasible.</p><p>If you try to address any of the other problems, for example by trying to solve AI alignment and control, before having removed competitive pressures, you are swimming against a strong current and will be swept over the falls.</p><h1>First filter: all-out war between nuclear superpowers</h1><p>I think that hawkish writings about China usually fail to take their reasoning to the logical conclusion. For example, Leopold Aschenbrenner&#8217;s &#8220;<a href="https://situational-awareness.ai/">Situational Awareness</a>&#8221; and Dario Amodei&#8217;s essays, including &#8220;<a href="https://darioamodei.com/post/on-deepseek-and-export-controls">On DeepSeek and Export Controls</a>&#8221; and some of &#8220;<a href="https://darioamodei.com/machines-of-loving-grace">Machines of Loving Grace</a>.&#8221;</p><p>People understand that the US and Chinese governments will wake up to the potential of ASI, and that when they do, absent strong international coordination (which Leopold and Dario assume is absent), the governments will be in an all-out race to who can build it first. The mistake Leopold, Dario and others make is to assume this is a restricted game, where most of what is happening is AI R&amp;D and at most countries will engage in mutual espionage and sabotage.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a></p><p>If you take these views to their logical conclusion, you would see the ending of this story: all-out war between the US and China. When the superpowers try to sabotage each other&#8217;s ASI projects, they will not stop at grey-zone or covert sabotage. From a state&#8217;s perspective, if your adversary gets ASI, you are <em>done</em>. Your state will stop existing. You might as well have gotten all your major cities vaporized.</p><p>I am very confident that a superpower that knows it&#8217;s about to lose the race, or even considers a high risk of losing, will engage in unambiguous acts of war. The paper &#8220;<a href="https://arxiv.org/abs/2503.05628">Superintelligence Strategy</a>&#8221; talks about possible kinetic strikes, but I think it will get much worse.</p><p>If states start building very hardened ASI projects, then stopping an opponent&#8217;s progress can be impossible without taking extreme measures that attempt to make the opponent&#8217;s country completely dysfunctional. For example:</p><ul><li><p>Systematically attacking basic infrastructure (like the electrical grid) throughout the opponent&#8217;s territory</p></li><li><p>Sabotaging core functions of the opponent&#8217;s government, such as attempting or strongly supporting a coup</p></li><li><p>Launching an invasion, either to physically stop the ASI projects or to consume all the opponent&#8217;s resources through war</p></li></ul><p>If we get to this point, I don&#8217;t see any reason to be confident that the situation won&#8217;t escalate all the way to a full-blown nuclear war between superpowers.<br><br>I think it would be a fool&#8217;s errand to try to predict the exact reaction of the national security establishment of the losing superpower. It will depend too much on unpredictable and opaque factors, from the structure of the natsec apparatus to whether the people responsible happen to be in a bad mood at some specific, decisive moment.</p><p>But I think it&#8217;s important to note that there are strong mechanisms pushing in the direction of arbitrary escalation, and no strong mechanisms preventing it from doing so.</p><p><strong>And if all-out war between nuclear powers doesn&#8217;t sound bad enough to you, remember this: war would be waged with much more advanced AI than we have today, and the war itself would further shape the incentives around the AI race.</strong></p><h2>Contra &#8220;stable multipolar scenarios&#8221;</h2><p>Stable <a href="https://www.lesswrong.com/w/multipolar-scenarios">multipolar scenarios</a> can happen in one of two ways: if AI&#8217;s efficacy at war has reached the limits of physics, or if AIs have a way to enforce a consensus (like in the good ending of &#8220;<a href="https://ai-2027.com/">AI 2027</a>&#8221;).</p><p>AI advantages compound, and if the gap is wide enough, one of the competitors (potentially a rogue AI) wins. It seems unlikely that AI&#8217;s ability to wage war will climb all the way to the limits of physics while the gap between the various actors never gets wide enough to conclude the conflict.</p><p>About AIs enforcing a consensus, roughly, I think this would require AI to already be vastly smarter and more competent than any human or existing human organization. Which makes this proposed &#8220;solution&#8221; kind of tautological: you still need to pass all the filters and build an ASI that you can trust.</p><p>As an example, in &#8220;<a href="https://ai-2027.com/">AI 2027</a>,&#8221; the two ASIs strike a deal by building a &#8220;consensus AI&#8221; that will forever enforce, to some degree, the preferences of both AIs. To do this, you&#8217;d need to have developed an extremely deep fundamental understanding of how to program AI, the kind of understanding that lets you write an AI as lines of code rather than a neural network.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a></p><p>Due to the competitive pressures I talk about in this post, the plan would not unfold this way. Much, much earlier than when you&#8217;d be able to achieve such a deep understanding of AI, you&#8217;d achieve an understanding just barely good enough to build ASI. Then you would build it and thus destroy all value in the world, unless you already figured out a way past all the filters in this post.</p><p>An alternative proposal is to have AIs strike deals that are enforced through mutual monitoring. By the time AIs can strike such deals autonomously, they are already fairly superhuman and / or significantly in charge of running the world, and we need to have passed the filters.</p><p>To be clear, I don&#8217;t necessarily think it&#8217;s a bad idea to have weaker AIs help us enforce monitoring-based international agreements. But this needs to be done before AIs get too strong, at which point humanity would have to do it, even if aided by weaker AIs.</p><p><em>(My paper &#8220;<a href="https://ai-scenarios.com/">Modeling the geopolitics of AI development</a>&#8221; talks about this filter in more detail, but the thinking is less refined since it was written a while ago.)</em></p><h1>Second filter: misaligned AI that kills everyone</h1><p>It is probably very hard to build an ASI that doesn&#8217;t end up killing every human being simply by running it.</p><p>The basic argument is that ASI would be so effective that any failure, even partial, would result in an ASI handling extreme amounts of power while not going out of its way to preserve human life and values.</p><p>ASI would kill us as a side effect of whatever it ends up doing, just like a human destroys an anthill without a second thought when it&#8217;s in the way of a construction project. The field of making sure that ASIs act in a desirable way is called &#8220;AI safety.&#8221;</p><p>The threat model of misaligned AI is the one that has already been explored the most, so I will assume that readers are at least passingly familiar with it and won&#8217;t try to convey the basic idea here. If you need an introduction, read the book &#8220;If Anyone Builds It, Everyone Dies&#8221; by Eliezer Yudkowsky and Nate Soares.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a></p><p>What I want to focus on here is how the pressure to cut corners I mentioned earlier makes it nearly impossible to solve alignment in time for when ASI will be developed. Think of the following competitive pressures:</p><ul><li><p>Pressure to cut corners on safety methods</p></li><li><p>Pressure to deploy as fast as possible</p></li><li><p>Pressure to give AI as much autonomy as possible</p></li><li><p>Pressure to hand over existing decision loops to AI as quickly and thoroughly as possible</p></li></ul><p>So what happens is, AI projects will develop and deploy AI that is as capable as possible given current capabilities techniques, while only being as safe as absolutely necessary to make them usable. The most important part here is: <em>only as safe as absolutely necessary to make them usable</em>. What does it mean? Well, the first instance of this pattern we&#8217;ll discuss is the commercial one.</p><p>Software engineers won&#8217;t use an AI that cheats to make the tests pass <em>every time</em>, but they&#8217;ll use an AI they can usually catch cheating, as long as the violations don&#8217;t fall through the cracks often enough for the engineer to get fired.</p><p>CEOs will not use AI employees that regularly take costly, irreversible actions to the point that the company loses a lot of money or it gets the CEO in trouble. But they will, for example, use AI that takes illegal actions <a href="https://www.reuters.com/investigations/meta-is-earning-fortune-deluge-fraudulent-ads-documents-show-2025-11-06/">as long as the company gets fined for less than the money it made</a>, or the crime happens in a third-world country, etc.</p><p>So far it doesn&#8217;t sound like an extinction risk, but what is the &#8220;usability limit&#8221; when it comes to integrating AI in the military? What about tail risks, situations that are too rare and so haven&#8217;t yet appeared in the feedback loop of fixing AI bugs?</p><p><strong>And most importantly, what happens when someone first gets to the capability level where they mostly hand over AI R&amp;D to AIs themselves?</strong></p><p>The AI will be <em>just barely safe enough</em> to profitably (not spotlessly!) do jobs that:</p><ul><li><p>Are roughly as complicated as AI R&amp;D<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-6" href="#footnote-6" target="_self">6</a></p></li><li><p>Have short-enough feedback loops that failures have already happened, such that AI companies already have bug tickets for these failures</p></li><li><p>Have already addressed these bug tickets</p></li></ul><p>Of course, you will not be able to get this guarantee for novel tasks, such as AI R&amp;D itself. Probably, you won&#8217;t even be able to get it for tasks that already exist but are not common enough for you to test the AI thoroughly on them during the (very brief) allotted time. You have to hope that whatever safety you have transfers from this small, nonrepresentative set of tasks to the ones that matter.</p><h2>Why Technical AI safety agendas do not address this problem</h2><p>Technical AI safety agendas for addressing extinction risks usually focus on the &#8220;misaligned AI that kills everyone&#8221; filter, so I have to briefly address why, as a general rule, they don&#8217;t work. In fact, they make things worse.</p><p>This is because all alignment work is capabilities work.</p><p>Take RLHF (reinforcement learning from human feedback), for example.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-7" href="#footnote-7" target="_self">7</a> RLHF improved &#8220;alignment,&#8221;<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-8" href="#footnote-8" target="_self">8</a> but it also improved capabilities a lot more: the AIs that we built after RLHF were more liable to do dangerous things than the ones we built before it existed. This is true even if you do your best to use RLHF to make the model safe.</p><p>To the degree that interpretability and scalable oversight work, I confidently predict that they will do exactly the same thing.</p><p><strong>The underlying, fundamental reason for this is that capabilities are easier to formalize than safety.</strong> By this I mean that capabilities are easier to measure and easier to describe to other people, to AIs, and to code without loss of information.</p><p>Imagine that we get an interpretability breakthrough. You would have more readability into the internal algorithms of AIs, but those algorithms are very big and complicated: you wouldn&#8217;t automatically know which parts are helpful and which are harmful.</p><p>Some will be obviously harmful and removed right away. What then? Maybe you can do some manual searches for patterns you suspect exist? But humans are slow. You can get AIs to help you, but AIs are not (yet) smarter than you, and so they&#8217;d miss some stuff. <a href="https://www.lesswrong.com/posts/WewsByywWNhX9rtwi/current-ais-seem-pretty-misaligned-to-me">Perhaps AIs already have some misaligned biases</a> and so would sometimes actively hinder your efforts.</p><p>On other hand, capabilities, oh how they&#8217;d skyrocket. Better interpretability would yield more powerful methods to modify AIs: it would allow engineers and learning algorithms to modify AIs in more targeted, deliberate, and understandable ways than can be done today.</p><p>Since capabilities are more formalized, you can quickly train a large team of engineers to make use of the novel techniques. Perhaps you can cut engineers out of this loop entirely, integrating the novel technique as part of automated learning algorithms.</p><p>If you want to modify AI to improve a quality that is hard to measure, like safety, you need a human to stand there and opine about each candidate modification. Worse still, the human needs to have good taste about the property you are trying to improve.</p><p>To summarize: capabilities can improve at machine speeds, while safety will always be bottlenecked by humans. The only way to solve this dilemma would be to describe our safety desiderata to the same level that we have described our capabilities desiderata. That way, we could potentially automate AI safety, or at least reliably train a team of engineers to do it. Good luck doing that during an all-out race to ASI!</p><p>I encourage you to think about this issue yourself, especially if you are a researcher at a major AI company working on a technical AI safety agenda. Your work may end up boosting capabilities more than most of the people over on the capabilities teams.</p><h1>Third filter: nightmare singletons</h1><p>Ok, imagine that the alignment problem is on track to get solved, such that a human being (or group of human beings) could operate an ASI without killing themselves and everyone else as a side effect. You and the rest of your team, the &#8220;responsible actors&#8221; in a world composed mostly of irresponsible ones, have the lead in AI development. You will build ASI first and then establish an eternal utopia, right? No.</p><p>Here&#8217;s what really happens: the government takes over your project before you get to ASI, by default as a military project, possibly top secret. You are questioned just enough that they know how to make use of the project&#8217;s assets (like code, documentation, hardware, etc.), and then you are thrown out the door.</p><p>Or maybe a softer version of this happens, where your AI company still technically exists. However, your CEO does not retain effective control of the company, and you have military personnel looking over your shoulder as you work.</p><p>If your government is asleep at the wheel, a foreign government will take over your project, or at least steal all the progress you&#8217;ve made so far and then pour their resources into going faster than you. Or if all governments are asleep at the wheel, another company will take over your project, or perhaps a charming psychopath CEO will manipulate their way into a top leadership position at the company where you work.</p><p>What then?</p><p>Whoever controls an ASI can establish a <a href="https://en.wikipedia.org/wiki/Singleton_(global_governance)">singleton</a>. A singleton is a &#8220;world order in which there is a single decision-making agency at the highest level, capable of exerting effective control over its domain, and permanently preventing both internal and external threats to its supremacy.&#8221;</p><p>&#8212;</p><p>Let me spell out, for people who haven&#8217;t thought about this subject before, how nightmarish this scenario can get.</p><p>An individual in control of an ASI could establish a dictatorship that controls the entire earth, possibly the entire universe.</p><p>They could monitor every corner of their domain 24/7 and assign a virtually infinite amount of intelligence to analyze all of this information.</p><p>They could compel everyone to install brain implants (or forcibly <a href="https://en.wikipedia.org/wiki/Mind_uploading">upload</a> them, etc.) and have complete oversight and control over their thoughts, actions, and experiences.</p><p>Eventually, they could shape the whole world to their preference until every atom is exactly as they want it, and do it as easily as a child shapes playdough.</p><p>&#8212;</p><p>In AI safety, some people&#8217;s strategy is to give power and resources to &#8220;good&#8221; or &#8220;responsible&#8221; actors, such as their favorite AI company.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-9" href="#footnote-9" target="_self">9</a> The theory of change for this strategy is that the &#8220;responsible&#8221; actor is the first to build ASI and establishes a utopian (or at least &#8220;good&#8221;) singleton.</p><p>I think that it is an enormous mistake to trust any one person or company with this. If your strategy is to use ASI to establish a &#8220;good&#8221; singleton, I will fight to prevent you from succeeding because I don&#8217;t trust you. But regardless, I hope this post makes you see that this strategy will break horribly.</p><p>If you are part of an ASI project and this is your plan, know this: someone more powerful than you will take your toys away before you get to ASI. Then, they will use them to race to ASI without you.</p><p>What happens later is fundamentally unpredictable. The result does not have to be as bad as the nightmare scenario I painted earlier. But from where we&#8217;re standing, it could easily get really bad.</p><p>I think what happens if any individual or small group of people obtains absolute power over the universe is an extremely dystopian scenario, potentially worse than death depending on your values. I think the same is true for scenarios in which we just barely make enough progress on alignment that ASI doesn&#8217;t kill us all as a side effect. ASI may want a future for us, but it could be a future that we find abhorrent, and it would have absolute power over us.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-10" href="#footnote-10" target="_self">10</a></p><p>Even in the best case scenario, where the ASI project is taken over by a government with a very robust democratic process, the situation would most likely be considered a national security emergency. Such emergencies are dealt with by the military (or more generally, the executive branch), which needs to be able to act quickly. As a result, it has weaker democratic oversight compared to other government branches.</p><p>What will this government do after having declared an emergency situation, armed with proto-ASI? Would you feel safe if you thought your government was bound to establish a singleton?</p><h1>How common theories of change fail trivially</h1><p>Any solution or theory that focuses entirely on technical AI safety fails trivially by not taking into account the two other filters. For example, some people think AI alignment will be easy to solve. I think this view is most likely mistaken on a very deep level. But even if it were correct, it would not address the other two problems at all.</p><p>Furthermore, I think that all technical AI safety projects will not be successful, not in a world where actors are able to unilaterally push the frontier of AI development toward ASI. This is due to the pressure to cut corners on safety, and because any technique will accelerate capabilities much more than it accelerates progress in AI safety.</p><p>The philosophy of &#8220;iterative deployment&#8221; will simply not apply in a world where the pace of deployment depends entirely on competitive pressure and is entirely causally disconnected from any consideration of what may be a &#8220;responsible&#8221; pace for AI development.</p><p>There are some who try to acquire personal power or influence so they can exert it &#8220;when the time comes.&#8221; This can mean attaining influence inside of AI companies or in governments. As I pointed to in the third filter, I think power within AI companies is meaningless.</p><p>And I think the people who try to acquire unilateral power within governments are deeply misguided. When push comes to shove, they will fail at gaining enough power to steer the actions of governments.</p><p>If the majority of the government does not understand the meaning of ASI, these people will not be able to make massively expensive and complex asks to leadership. For example: &#8220;slow down AI development to improve safety&#8221; or &#8220;pressure other major powers to enter a hefty trust but verify regimes capable of providing mutual assurances on AI development.&#8221; If these people try to push these asks without first building a broad support base (probably as broad as a decent voting bloc), then they will simply get purged.</p><p>Finally, there are people trying to get a major power to engage in a race to ASI, beat all their adversaries to it, and establish a singleton. I think these theories of change fail on literally all three filters:</p><ul><li><p>The world will likely be consumed by war before any actor can get to ASI.</p></li><li><p>Even if we narrowly avoid all-out war, this theory of change leads to a race to the bottom on AI safety and to uncontrollable ASI that kills everyone.</p></li><li><p>Even if the ASI ends up being somewhat controllable, no country on Earth currently has such institutional robustness that it would not produce a dystopia if it acquired ASI.</p></li></ul><h1>Conclusion</h1><p>These were the main three challenges that I think stand between us and surviving ASI. Even if we pass all three, I don&#8217;t think things automatically go well.</p><p>I have more <a href="https://en.wikipedia.org/wiki/Intuition_pump">intuition pumps</a> that I would like to publish in a future post. They are mostly about how, in scenarios with AI that is strong but not as strong as I&#8217;ve been implying ASI is, that:</p><ul><li><p>There is a strong tendency for power to concentrate and for the world to gravitate toward the three outcomes I&#8217;ve been describing.</p></li><li><p>There is a tendency for human preferences and behavior to mutate beyond recognition, to a degree that we might think of such people as essentially &#8220;dead.&#8221;</p></li></ul><p>The main way that I envision humanity passing these filters is with deep awareness of what ASI entails and with international coordination.</p><p>Deep awareness is necessary so the relevant parties understand what their interests are with respect to ASI. Chiefly, they need to understand that ASI can become powerful enough to destroy the world, and that it is indeed extremely hard to deploy an ASI without destroying the world.</p><p>Coordination, backed by mutual monitoring and deterrence, is necessary so the major parties can avoid a race to the bottom over who builds ASI first. Without it, they will end up developing and deploying ASI in the most irresponsible way possible, and thus destroy the world.</p><p><strong>Both deep awareness and coordination are necessary so countries can eventually get to work to figure out how to go through this transition while avoiding the horrific failure modes I&#8217;ve described, and others yet.</strong></p><p>At the moment, my best bet for achieving these goals is to work at ControlAI. If you&#8217;re interested in learning more about ControlAI, feel free to read our <a href="https://www.lesswrong.com/posts/TnAR5Sf5hphfnzNTr/preventing-extinction-from-asi-on-a-usd50m-yearly-budget">funding pitch</a>, which also goes in detail about ControlAI&#8217;s theory of change. Alternatively, feel free to shoot me a message.</p><p><em><a href="https://www.lesswrong.com/posts/BWNqdt5edJdax4jn8/the-three-filters-why-almost-every-plan-to-survive-asi-fails">This post is also on LessWrong.</a></em><br></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://thelaststretch.blog/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe for free for more posts on AI and other things pivotal to the future.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>This includes things like stealing your weights and then sabotaging your ASI projects, but also trying to insert backdoors into your AI systems.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>And worse than useless if you consider that it absorbs funds and attention.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>Even when they acknowledge the possibility of war, it is treated as something that happens in the very endgame. Countries are not treated as being able to look ahead and strike preemptively.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>Even with such an understanding, code may not be the optimal way to build an AI, and you may choose to use neural networks or a new technique altogether. The point is <em>if you wanted to write it in code, you could.</em></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p>Some people criticize Yudkowsky and Soares&#8217; arguments for not engaging properly with the peculiarities of LLMs and claim that LLMs make alignment easier. I have it on my to-do list to write about why the shape of current AI systems doesn&#8217;t make me particularly optimistic about alignment. Unfortunately, at the moment I don&#8217;t know of a good post to convey this; the best one I can point you to is: &#8220;<a href="https://www.lesswrong.com/posts/WewsByywWNhX9rtwi/current-ais-seem-pretty-misaligned-to-me">Current AIs seem pretty misaligned to me</a>.&#8221;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-6" href="#footnote-anchor-6" class="footnote-number" contenteditable="false" target="_self">6</a><div class="footnote-content"><p>In fact, I think this is quite optimistic. AI companies are prioritizing AI R&amp;D over anything else, so it will be one of the first (if not <em>the</em> first) task AIs will be able to perform at its level of complexity. There will not have been trial runs with similarly complex tasks.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-7" href="#footnote-anchor-7" class="footnote-number" contenteditable="false" target="_self">7</a><div class="footnote-content"><p>RLHF is the technique that enabled the creation of the first version of ChatGPT.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-8" href="#footnote-anchor-8" class="footnote-number" contenteditable="false" target="_self">8</a><div class="footnote-content"><p>Insofar as you could get LLMs to actually do the task you asked them to do, even when the task was not extremely simple and even if you weren&#8217;t an expert base-model prompter.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-9" href="#footnote-anchor-9" class="footnote-number" contenteditable="false" target="_self">9</a><div class="footnote-content"><p>This includes technical people who decide to work on capabilities at an AI company.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-10" href="#footnote-anchor-10" class="footnote-number" contenteditable="false" target="_self">10</a><div class="footnote-content"><p>The bad ending of &#8220;<a href="https://ai-2027.com/">AI 2027</a>&#8221; falls under this last category, and it was considered the most likely ending by the authors at the time of writing.)</p></div></div>]]></content:encoded></item></channel></rss>