<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Steve Byrnes’s Substack]]></title><description><![CDATA[Mailing list for announcements of new blog posts and other works.]]></description><link>https://stevebyrnes1.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!OM4Y!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7603e9a-b65e-4a8c-8002-d3648bb93b3e_2234x2234.jpeg</url><title>Steve Byrnes’s Substack</title><link>https://stevebyrnes1.substack.com</link></image><generator>Substack</generator><lastBuildDate>Fri, 31 Jul 2026 05:48:33 GMT</lastBuildDate><atom:link href="https://stevebyrnes1.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Steve Byrnes]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[stevebyrnes1@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[stevebyrnes1@substack.com]]></itunes:email><itunes:name><![CDATA[Steve Byrnes]]></itunes:name></itunes:owner><itunes:author><![CDATA[Steve Byrnes]]></itunes:author><googleplay:owner><![CDATA[stevebyrnes1@substack.com]]></googleplay:owner><googleplay:email><![CDATA[stevebyrnes1@substack.com]]></googleplay:email><googleplay:author><![CDATA[Steve Byrnes]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Blog post: “RL & search is a terrifying way to build AGI (an FAQ)”]]></title><description><![CDATA[Here&#8217;s the link:]]></description><link>https://stevebyrnes1.substack.com/p/blog-post-rl-and-search-is-a-terrifying</link><guid isPermaLink="false">https://stevebyrnes1.substack.com/p/blog-post-rl-and-search-is-a-terrifying</guid><dc:creator><![CDATA[Steve Byrnes]]></dc:creator><pubDate>Mon, 27 Jul 2026 15:09:01 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!2Hu4!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a95ac56-73be-4ff5-9ac3-3bd649b52cb3_1482x728.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Here&#8217;s the link:</p><p><a href="https://www.alignmentforum.org/posts/KHyBocZncAmtu4Jbc/rl-and-search-is-a-terrifying-way-to-build-agi-an-faq">https://www.alignmentforum.org/posts/KHyBocZncAmtu4Jbc/rl-and-search-is-a-terrifying-way-to-build-agi-an-faq</a></p><p>And here&#8217;s the start:</p><blockquote><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!2Hu4!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a95ac56-73be-4ff5-9ac3-3bd649b52cb3_1482x728.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!2Hu4!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a95ac56-73be-4ff5-9ac3-3bd649b52cb3_1482x728.png 424w, https://substackcdn.com/image/fetch/$s_!2Hu4!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a95ac56-73be-4ff5-9ac3-3bd649b52cb3_1482x728.png 848w, https://substackcdn.com/image/fetch/$s_!2Hu4!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a95ac56-73be-4ff5-9ac3-3bd649b52cb3_1482x728.png 1272w, https://substackcdn.com/image/fetch/$s_!2Hu4!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a95ac56-73be-4ff5-9ac3-3bd649b52cb3_1482x728.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!2Hu4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a95ac56-73be-4ff5-9ac3-3bd649b52cb3_1482x728.png" width="496" height="243.57142857142858" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0a95ac56-73be-4ff5-9ac3-3bd649b52cb3_1482x728.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:715,&quot;width&quot;:1456,&quot;resizeWidth&quot;:496,&quot;bytes&quot;:1076710,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://stevebyrnes1.substack.com/i/208697721?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a95ac56-73be-4ff5-9ac3-3bd649b52cb3_1482x728.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!2Hu4!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a95ac56-73be-4ff5-9ac3-3bd649b52cb3_1482x728.png 424w, https://substackcdn.com/image/fetch/$s_!2Hu4!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a95ac56-73be-4ff5-9ac3-3bd649b52cb3_1482x728.png 848w, https://substackcdn.com/image/fetch/$s_!2Hu4!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a95ac56-73be-4ff5-9ac3-3bd649b52cb3_1482x728.png 1272w, https://substackcdn.com/image/fetch/$s_!2Hu4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a95ac56-73be-4ff5-9ac3-3bd649b52cb3_1482x728.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3><span>Q1: What are you saying?</span></h3><p><strong><span>A:</span></strong><span> My claim here is that if you build </span><a href="https://www.alignmentforum.org/posts/nQH2GhkmSmHoCAN8R/what-do-i-mean-by-artificial-general-intelligence"><span>artificial general intelligence (AGI)</span></a><span> via any algorithm that&#8217;s choosing actions via reinforcement learning (RL) and/or model-based search and planning&#8212;a giant chunk of your AI textbook&#8212;then that&#8217;s just an utterly terrifying thing that you&#8217;re doing. You&#8217;re playing around with algorithms that, if they work at all, would tend to create ruthless, callous AGIs, AGIs which would happily exterminate humanity and run the world by themselves, given an opportunity.</span></p><p><span>Mercifully, large language models (LLMs) today are not in the category of &#8220;algorithms that choose actions via RL &amp; search&#8221;. At least, not </span><em><span>primarily</span></em><span>&#8212;see </span><a href="https://www.alignmentforum.org/posts/wYpjXRLqbLbnmjbJP/llms-are-still-mostly-powered-by-imitative-learning-not-rl"><span>LLMs are (still) mostly powered by imitative learning, not RL</span></a><span>. So LLMs are outside the scope of this post. However, lots of other researchers and companies around the world are enthusiastically trying to build AGI in the maximally terrifying way, as we speak.</span></p><h3><span>Q2: So you&#8217;re saying, don&#8217;t build AGI based on RL and/or search &amp; planning?</span></h3><p><strong><span>A:</span></strong><span> In principle, it&#8217;s entirely possible that something is terrifying, but we should do it anyway.</span></p><p><span>&#8230;Like space travel! Space travel is: &#8220;Let&#8217;s fill a tank with 1000 tons of the most flammable substance imaginable, and then light it on fire, and strap people to the front, to accelerate them until they&#8217;re traveling at insane speeds through the extremely lethal vacuum of space.&#8221; That&#8217;s terrifying! But we do it anyway.</span></p><p><span>And I&#8217;m not being anti-space-travel when I point out how terrifying it is. The space travel enthusiasts and the space travel skeptics can happily work together to spread deep understanding of all the ways that space travel can go lethally wrong. Because you can&#8217;t overcome a challenge without understanding it.</span></p><p><span>&#8230;Having said all that, I strongly endorse &#8220;Don&#8217;t build AGI based on RL &amp; search </span><em><span>until and unless we find a much better plan for making it friendly</span></em><span>&#8221;. So if someone is working on making RL &amp; search systems more powerful right now, then I think that&#8217;s bad, and that they should instead be spending their time and effort figuring out how (and indeed whether) such systems may be used safely if they become more powerful in the future. That&#8217;s an open problem, and we can work on it today.</span></p><h3><span>Q3: Why do you think it&#8217;s terrifying?</span></h3><p><strong><span>A:</span></strong><span> A big part of the problem is that </span><strong><span>the reward function (or cost function, or objective function, or whatever we want to call it) will generally be written in Python, </span></strong><em><strong><span>not</span></strong></em><strong><span> in natural language</span></strong><span>.</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a><span> And RL &amp; search agents will tend to ruthlessly maximize it, including in ways the programmer </span><em><span>obviously</span></em><span> didn&#8217;t intend. That&#8217;s the whole point of these algorithms: they ruthlessly maximize a thing. That&#8217;s what they were designed to do, and that&#8217;s what they actually do, when they work well enough to do anything impressive [nuances in Q3b &amp; Q11 below].</span></p><p><span>And nobody knows how to write down a Python function relating to the real world, such that ruthlessly maximizing that function would lead to good outcomes.</span></p><p><span>To illustrate the problem, consider &#8220;specification gaming&#8221; [&#8230;]</span></p></blockquote><p></p><p>It&#8217;s an FAQ, so I&#8217;ll just list the rest of the questions I address in the post:</p><blockquote><h3>Q3b: So your concern is the &#8220;literal genie&#8221; / &#8220;monkey&#8217;s paw&#8221; thing?</h3><p>[&#8230;]</p><h3>Q4: Won&#8217;t this problem go away when the AI is smart enough to understand what we <em>intended</em> when we wrote the reward function code?</h3><p>[&#8230;]</p><h3>Q5: Can&#8217;t we just fix bad behavior when we see it?</h3><p>[&#8230;]</p><h3>Q5b: Follow-up: I don&#8217;t buy that, because even if superintelligent AIs could deceptively hide their bad behavior, won&#8217;t earlier AIs be sufficiently incompetent that we&#8217;ll see their bad behavior? And if so, again, can&#8217;t we just fix the bad behavior when we see it? We <em>do</em> know how to fix bad behavior we see it: the RL &amp; search literature is full of examples where algorithms did useful things as intended.</h3><p>[&#8230;]</p><h3>Q6: Why don&#8217;t we just solve the problem by using an obvious, common-sense reward / cost / objective function, like [FILL IN THE BLANK]?</h3><p>[&#8230;]</p><h3>Q7: Isn&#8217;t this whole thing kinda crazy? After all, LLMs are not ruthless sociopaths all the time, and humans are also not ruthless sociopaths all the time. So where is this idea even coming from? Are you sure you&#8217;re not just watching too much sci-fi?</h3><p>[&#8230;]</p><h3>Q8: Isn&#8217;t this problem solved by laws and markets? I.e., if an AGI has sociopathic desires and callous indifference to human welfare, that&#8217;s fine! It will still act nice and cooperative and rule-following, because acting nice and cooperative and rule-following is the best way to accomplish goals, in our complex interconnected interdependent world. Right?</h3><p>[&#8230;]</p><h3>Q8b: Following up on that: Even if you&#8217;re right that there&#8217;s a local incentive for being open to stabbing your allies in the back, isn&#8217;t there a higher-level, group-selection-style, incentive to be genuinely deeply nice? Specifically, won&#8217;t the groups of nice cooperative AGIs outcompete the groups o<span>f callous transactional AGIs who all keep stabbing each other in the back? And isn&#8217;t that related to how humans evolved to be nice?</span></h3><p><span>[&#8230;]</span></p><h3>Q9: Why would we want to infringe on the AGI&#8217;s autonomy by choosing its reward function?</h3><p>[&#8230;]</p><h3>Q10: Why not just be nice to the AGIs, and then they&#8217;ll be nice to us in turn?</h3><p>[&#8230;]</p><h3>Q11: RL &amp; search algorithms don&#8217;t literally <em>optimize</em> the reward / cost / objective function. Doesn&#8217;t that invalidate your argument?</h3><p>[&#8230;]</p></blockquote><p>Click the <a href="https://www.alignmentforum.org/posts/KHyBocZncAmtu4Jbc/rl-and-search-is-a-terrifying-way-to-build-agi-an-faq">link</a> to read the whole thing!</p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p><span>Or whatever other programming language you like; see also Q6 below on the possibility of slotting in a trained machine learning model here.</span></p></div></div>]]></content:encoded></item><item><title><![CDATA[Blog post: “LLMs are (still) mostly powered by imitative learning, not RL”]]></title><description><![CDATA[Here&#8217;s the link:]]></description><link>https://stevebyrnes1.substack.com/p/blog-post-llms-are-still-mostly-powered</link><guid isPermaLink="false">https://stevebyrnes1.substack.com/p/blog-post-llms-are-still-mostly-powered</guid><dc:creator><![CDATA[Steve Byrnes]]></dc:creator><pubDate>Fri, 24 Jul 2026 14:46:21 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/97f328ac-186e-4188-b089-ea4d44ef6239_1280x1944.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Here&#8217;s the link:</p><p><a href="https://www.lesswrong.com/posts/wYpjXRLqbLbnmjbJP/llms-are-still-mostly-powered-by-imitative-learning-not-rl">https://www.lesswrong.com/posts/wYpjXRLqbLbnmjbJP/llms-are-still-mostly-powered-by-imitative-learning-not-rl</a></p><p>And here&#8217;s the start:</p><blockquote><p><span>Reinforcement learning from verifiable rewards (RLVR) is the hot new thing in LLM training. It&#8217;s so hot, and people spend so much time talking about it, that they sometimes lose sight of the big picture.</span></p><p><span>Stepping back, LLMs can do lots of very impressive things. How? Where did those capabilities come from? Fundamentally, they come from a combination of:</span></p><ul><li><p><strong><span>(1) Imitative learning</span></strong><span>, including pretraining and supervised fine-tuning (SFT)</span></p><ul><li><p><span>See my earlier discussion: </span><em><a href="https://www.lesswrong.com/posts/bnnKGSCHJghAvqPjS/foom-and-doom-2-technical-alignment-is-hard#2_3_2_LLM_pretraining_magically_transmutes_observations_into_behavior__in_a_way_that_is_profoundly_disanalogous_to_how_brains_work"><span>&#8220;LLM pretraining magically transmutes observations into behavior, in a way that is profoundly disanalogous to how brains work&#8221;</span></a><span>.</span></em></p></li></ul></li><li><p><strong><span>(2) Reinforcement learning</span></strong><span>, including RL from human feedback [RLHF], RL from AI feedback [RLAIF], and especially RLVR.</span><a href="https://www.lesswrong.com/posts/wYpjXRLqbLbnmjbJP/llms-are-still-mostly-powered-by-imitative-learning-not-rl#fni2pirihhdsj"><sup><span>[1]</span></sup></a></p></li></ul><p><span>If we look at the final trained LLM, we can ask how important each of those two pieces was, in explaining the LLM&#8217;s capabilities. And my claim is that </span><strong><span>it&#8217;s way more (1) than (2)</span></strong><span>.</span></p><p><span>I&#8217;ll start in &#167;1 with some relevant evidence, and then in &#167;2 I&#8217;ll circle back to operationalizing exactly what I&#8217;m claiming, and finally in &#167;3, three reasons why we should care&#8212;namely, it affects how we should think about chain-of-thought legibility, about LLM capabilities, and about LLM alignment.</span></p><p><span>Note that I am </span><em><span>not</span></em><span> arguing that RLVR does not importantly contribute to LLM capabilities. That would be absurd! Of course it does! Companies use RLVR because it works, and I expect them to continue doing so more and more. Again, the things I&#8217;m actually claiming are in &#167;2&#8211;&#167;3.</span></p></blockquote><p>Click the <a href="https://www.lesswrong.com/posts/wYpjXRLqbLbnmjbJP/llms-are-still-mostly-powered-by-imitative-learning-not-rl">link</a> to read the rest!</p>]]></content:encoded></item><item><title><![CDATA[Blog post: “Will almost all future companies eventually be founded and run by autonomous AIs?”]]></title><description><![CDATA[Here&#8217;s the link:]]></description><link>https://stevebyrnes1.substack.com/p/blog-post-will-almost-all-future</link><guid isPermaLink="false">https://stevebyrnes1.substack.com/p/blog-post-will-almost-all-future</guid><dc:creator><![CDATA[Steve Byrnes]]></dc:creator><pubDate>Wed, 22 Jul 2026 20:31:09 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/2a693942-2bc4-4b04-b808-8bbbf5ae5816_2560x1707.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Here&#8217;s the link:</p><p><a href="https://www.lesswrong.com/posts/BHEcssYwYrzTBf9Qm/will-almost-all-future-companies-eventually-be-founded-and">https://www.lesswrong.com/posts/BHEcssYwYrzTBf9Qm/will-almost-all-future-companies-eventually-be-founded-and</a></p><p>And here&#8217;s an excerpt</p><blockquote><p><span>My belief is that keeping AI under human control would be an unprecedented global challenge, if it&#8217;s even possible at all.</span></p><p><span>Many other people&#8217;s believe that AIs will </span><em><span>naturally</span></em><span> remain as tools under human control. Either they see this as inevitable&#8212;&#8220;how could it be otherwise?&#8221;&#8212;or they see it as not quite inevitable, but still an outcome that can easily happen via ordinary government regulations, civil society, and so on.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!nOa6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d84e54e-ed6a-4232-baef-babf30636b4b_1135x573.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!nOa6!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d84e54e-ed6a-4232-baef-babf30636b4b_1135x573.png 424w, https://substackcdn.com/image/fetch/$s_!nOa6!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d84e54e-ed6a-4232-baef-babf30636b4b_1135x573.png 848w, https://substackcdn.com/image/fetch/$s_!nOa6!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d84e54e-ed6a-4232-baef-babf30636b4b_1135x573.png 1272w, https://substackcdn.com/image/fetch/$s_!nOa6!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d84e54e-ed6a-4232-baef-babf30636b4b_1135x573.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!nOa6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d84e54e-ed6a-4232-baef-babf30636b4b_1135x573.png" width="518" height="261.5101321585903" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0d84e54e-ed6a-4232-baef-babf30636b4b_1135x573.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:573,&quot;width&quot;:1135,&quot;resizeWidth&quot;:518,&quot;bytes&quot;:59557,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://stevebyrnes1.substack.com/i/208113373?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d84e54e-ed6a-4232-baef-babf30636b4b_1135x573.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!nOa6!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d84e54e-ed6a-4232-baef-babf30636b4b_1135x573.png 424w, https://substackcdn.com/image/fetch/$s_!nOa6!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d84e54e-ed6a-4232-baef-babf30636b4b_1135x573.png 848w, https://substackcdn.com/image/fetch/$s_!nOa6!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d84e54e-ed6a-4232-baef-babf30636b4b_1135x573.png 1272w, https://substackcdn.com/image/fetch/$s_!nOa6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d84e54e-ed6a-4232-baef-babf30636b4b_1135x573.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>Who&#8217;s right? What&#8217;s the right way to think about this?</span></p><p><span>This is an issue where people </span><em><span>really</span></em><span> don&#8217;t see eye-to-eye, and tend to get stuck endlessly talking past each other. To pin down the disagreement, I think that the following is a good conversation-starter:</span></p><h2><strong><span>Conversation-starting question:</span></strong><span> </span><em><span>Do you expect almost all companies to eventually be founded and run by AIs rather than humans?</span></em></h2><p><span>I find that people give a number of different answers:</span></p><h2><strong><span>Possible Answer 1:</span></strong><span> </span><em><span>&#8220;No, because the best humans will always be better than large language models (LLMs) at founding and running companies.&#8221;</span></em></h2><p><span>Huh? Whoever said anything about LLMs? I wrote &#8220;AIs&#8221;!</span></p><p><span>[&#8230;]</span></p><h2><strong><span>Possible Answer 2:</span></strong><span> </span><em><span>&#8220;No, because the best humans will always be better than AIs at founding and running companies.&#8221;</span></em></h2><p><span>I think that most people with this view are generally rejecting the idea that actually powerful AI (as in the right side of the table above) is possible at all, or are failing to think about its implications. Again, humans are an existence proof for what is physically possible for AI. Thus, for example:</span></p><p><em><span>Humans can acquire real-world first-person experience?</span></em><span> Well, an AI could acquire a thousand lifetimes of real-world first-person experience.</span></p><p><span>[&#8230;]</span></p><h2><strong><span>Possible Answer 3:</span></strong><span> </span><em><span>&#8220;No, because humans will always be equally good as AIs at founding and running companies.&#8221;</span></em></h2><p><span>I think that most people with this view don&#8217;t </span><em><span>really</span></em><span> believe, in their guts, that skill and competence are a thing. Or at least, they don&#8217;t believe that skill and competence are relevant to founding and running successful companies.</span></p><p><span>When these people imagine what Jeff Bezos was doing day-to-day to build Amazon from scratch, I&#8217;m not quite sure what&#8217;s in their head. Befriending people? Scaring people? Puppies can also befriend and scare people, but no puppy has ever become a hundred-billionaire by founding and running a successful company.</span></p><p><span>[&#8230;]</span></p><h2><strong><span>Possible Answer 4:</span></strong><span> </span><em><span>&#8220;No, because we will pass laws preventing AIs from founding and running companies.&#8221;</span></em></h2><p><span>Even if such laws existed in every country on Earth, and the </span><em><span>letter</span></em><span> of such laws was enforceable, the </span><em><span>spirit</span></em><span> wouldn&#8217;t be. Rather, the laws would be trivial to work around. For example, you could wind up with companies where AIs are making all the decisions, but there&#8217;s a human frontman signing the paperwork.</span></p><p><span>[&#8230;]</span></p><h2><strong><span>Possible Answer 5:</span></strong><span> </span><em><span>&#8220;No, because if someone wants to start a business, they would prefer to remain in charge themselves, and ask an AI for advice when needed, rather than &#8216;pressing go&#8217; on an autonomous entrepreneurial AI.&#8221;</span></em></h2><p><span>This is a really nice vision, and I wish I could believe it. But even if lots of people do in fact take this approach, and they create lots of great businesses, it just takes one person to say </span><em><span>&#8220;Hmm, why should I create one great business, when I can instead create 100,000 great businesses simultaneously?&#8221;</span></em></p><p><span>[&#8230;]</span></p></blockquote><p>Click the <a href="https://www.lesswrong.com/posts/BHEcssYwYrzTBf9Qm/will-almost-all-future-companies-eventually-be-founded-and">link</a> to read the rest!</p>]]></content:encoded></item><item><title><![CDATA[Blog post: “What do I mean by ‘Artificial General Intelligence’?”]]></title><description><![CDATA[Plus two quick updates]]></description><link>https://stevebyrnes1.substack.com/p/blog-post-what-do-i-mean-by-artificial</link><guid isPermaLink="false">https://stevebyrnes1.substack.com/p/blog-post-what-do-i-mean-by-artificial</guid><dc:creator><![CDATA[Steve Byrnes]]></dc:creator><pubDate>Tue, 21 Jul 2026 14:22:31 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/f2358f2b-7a00-4caf-9f04-cd11dbb50cd7_800x600.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h1>Start with two quick updates:</h1><ul><li><p>I mentioned <a href="https://stevebyrnes1.substack.com/i/203131476/14-a-bunch-of-edits-to-intro-to-brain-like-agi-safety">a month ago</a> that I had made some edits to the blog version of <em><a href="https://www.lesswrong.com/s/HzcM2dkCq7fwXBej8">Intro to Brain-Like-AGI Safety</a></em>, but had not yet updated the archival PDF version. Well, now I have! Say hello to <a href="https://osf.io/preprints/osf/fe36n">PDF version 4</a>. For highlights from the changelog, see the <a href="https://stevebyrnes1.substack.com/i/203131476/14-a-bunch-of-edits-to-intro-to-brain-like-agi-safety">update from a month ago</a>.</p></li><li><p>I added yet even more caveats to my old <a href="https://www.lesswrong.com/posts/LaeP39jJpfPyoiSZm/valence-series-4-valence-and-liking-admiring">Valence series post 4</a> (2024), including a note at the top that I&#8217;ve disavowed so much of the post by now that it&#8217;s hardly worth reading.</p></li></ul><h1>New blog post: &#8220;What do I mean by &#8216;Artificial General Intelligence&#8217;?&#8221;</h1><p>Here&#8217;s the link:</p><p><a href="https://www.lesswrong.com/posts/nQH2GhkmSmHoCAN8R/what-do-i-mean-by-artificial-general-intelligence">https://www.lesswrong.com/posts/nQH2GhkmSmHoCAN8R/what-do-i-mean-by-artificial-general-intelligence</a></p><p>And here&#8217;s the first ~half of it:</p><blockquote><p><span>In this post, intended for a broad audience, I will paint a brief picture of what I&#8217;m talking about when I talk about &#8220;AGI&#8221;. It will seem obvious to many people, and obviously wrong to many others! So let&#8217;s jump in:</span></p><div><hr></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!z-qH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd4b1468-3d10-4454-b182-f1c48be9307f_1039x404.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!z-qH!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd4b1468-3d10-4454-b182-f1c48be9307f_1039x404.png 424w, https://substackcdn.com/image/fetch/$s_!z-qH!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd4b1468-3d10-4454-b182-f1c48be9307f_1039x404.png 848w, https://substackcdn.com/image/fetch/$s_!z-qH!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd4b1468-3d10-4454-b182-f1c48be9307f_1039x404.png 1272w, https://substackcdn.com/image/fetch/$s_!z-qH!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd4b1468-3d10-4454-b182-f1c48be9307f_1039x404.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!z-qH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd4b1468-3d10-4454-b182-f1c48be9307f_1039x404.png" width="1039" height="404" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cd4b1468-3d10-4454-b182-f1c48be9307f_1039x404.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:404,&quot;width&quot;:1039,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:72319,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://stevebyrnes1.substack.com/i/207917845?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd4b1468-3d10-4454-b182-f1c48be9307f_1039x404.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!z-qH!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd4b1468-3d10-4454-b182-f1c48be9307f_1039x404.png 424w, https://substackcdn.com/image/fetch/$s_!z-qH!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd4b1468-3d10-4454-b182-f1c48be9307f_1039x404.png 848w, https://substackcdn.com/image/fetch/$s_!z-qH!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd4b1468-3d10-4454-b182-f1c48be9307f_1039x404.png 1272w, https://substackcdn.com/image/fetch/$s_!z-qH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd4b1468-3d10-4454-b182-f1c48be9307f_1039x404.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>For example, today, if you want an AI to drive a car, or to control a computer using a mouse, it&#8217;s a huge project involving dozens of experts working for years to make a new AI. Whereas if you want a human to do the same, you don&#8217;t need to do R&amp;D to breed a new subspecies of human! Instead, you just take an ordinary human&#8212;basically the same design from 100,000 years ago&#8212;and give them a few hours of practice, and you&#8217;re done. Someday we&#8217;ll have an AGI design which can do things like that. </span>&#8230;</p><div><hr></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!jyRD!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d3eb6d1-c1af-47ee-94cb-e48ad89e5293_1037x401.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!jyRD!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d3eb6d1-c1af-47ee-94cb-e48ad89e5293_1037x401.png 424w, https://substackcdn.com/image/fetch/$s_!jyRD!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d3eb6d1-c1af-47ee-94cb-e48ad89e5293_1037x401.png 848w, https://substackcdn.com/image/fetch/$s_!jyRD!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d3eb6d1-c1af-47ee-94cb-e48ad89e5293_1037x401.png 1272w, https://substackcdn.com/image/fetch/$s_!jyRD!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d3eb6d1-c1af-47ee-94cb-e48ad89e5293_1037x401.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!jyRD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d3eb6d1-c1af-47ee-94cb-e48ad89e5293_1037x401.png" width="1037" height="401" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7d3eb6d1-c1af-47ee-94cb-e48ad89e5293_1037x401.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:401,&quot;width&quot;:1037,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:56449,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://stevebyrnes1.substack.com/i/207917845?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d3eb6d1-c1af-47ee-94cb-e48ad89e5293_1037x401.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!jyRD!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d3eb6d1-c1af-47ee-94cb-e48ad89e5293_1037x401.png 424w, https://substackcdn.com/image/fetch/$s_!jyRD!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d3eb6d1-c1af-47ee-94cb-e48ad89e5293_1037x401.png 848w, https://substackcdn.com/image/fetch/$s_!jyRD!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d3eb6d1-c1af-47ee-94cb-e48ad89e5293_1037x401.png 1272w, https://substackcdn.com/image/fetch/$s_!jyRD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d3eb6d1-c1af-47ee-94cb-e48ad89e5293_1037x401.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>In this case, many people&#8217;s mental image of AI is </span><em><span>already</span></em><span> transitioning from &#8220;tool&#8221; towards &#8220;agent&#8221;, especially in the past year or two, after they&#8217;ve watched LLM agents execute on projects. But even those people are usually not going far </span><em><span>enough</span></em><span> for what I have in mind. Think of things that took a whole society of humans to do over an extended period of time&#8212;like inventing language and science from scratch, and developing them all the way into space travel and microchips and skyscrapers. These AGIs will be able to do those kinds of things too, fully autonomously.</span></p><div><hr></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!eUJ_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0f9c6fd-e382-404d-843a-4c9266455d0f_1044x404.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!eUJ_!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0f9c6fd-e382-404d-843a-4c9266455d0f_1044x404.png 424w, https://substackcdn.com/image/fetch/$s_!eUJ_!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0f9c6fd-e382-404d-843a-4c9266455d0f_1044x404.png 848w, https://substackcdn.com/image/fetch/$s_!eUJ_!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0f9c6fd-e382-404d-843a-4c9266455d0f_1044x404.png 1272w, https://substackcdn.com/image/fetch/$s_!eUJ_!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0f9c6fd-e382-404d-843a-4c9266455d0f_1044x404.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!eUJ_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0f9c6fd-e382-404d-843a-4c9266455d0f_1044x404.png" width="1044" height="404" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c0f9c6fd-e382-404d-843a-4c9266455d0f_1044x404.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:404,&quot;width&quot;:1044,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:63268,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://stevebyrnes1.substack.com/i/207917845?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0f9c6fd-e382-404d-843a-4c9266455d0f_1044x404.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!eUJ_!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0f9c6fd-e382-404d-843a-4c9266455d0f_1044x404.png 424w, https://substackcdn.com/image/fetch/$s_!eUJ_!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0f9c6fd-e382-404d-843a-4c9266455d0f_1044x404.png 848w, https://substackcdn.com/image/fetch/$s_!eUJ_!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0f9c6fd-e382-404d-843a-4c9266455d0f_1044x404.png 1272w, https://substackcdn.com/image/fetch/$s_!eUJ_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0f9c6fd-e382-404d-843a-4c9266455d0f_1044x404.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>Now, don&#8217;t be put off by &#8220;crazy sci-fi stuff&#8221;&#8212;indeed, </span><strong><span>every technology that exists today was &#8220;crazy sci-fi stuff&#8221; before it was invented</span></strong><span>!</span></p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!A4k5!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F035fe3d7-0e7e-45d9-9810-493591e1fbe1_1651x313.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!A4k5!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F035fe3d7-0e7e-45d9-9810-493591e1fbe1_1651x313.png 424w, https://substackcdn.com/image/fetch/$s_!A4k5!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F035fe3d7-0e7e-45d9-9810-493591e1fbe1_1651x313.png 848w, https://substackcdn.com/image/fetch/$s_!A4k5!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F035fe3d7-0e7e-45d9-9810-493591e1fbe1_1651x313.png 1272w, https://substackcdn.com/image/fetch/$s_!A4k5!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F035fe3d7-0e7e-45d9-9810-493591e1fbe1_1651x313.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!A4k5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F035fe3d7-0e7e-45d9-9810-493591e1fbe1_1651x313.png" width="1456" height="276" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/035fe3d7-0e7e-45d9-9810-493591e1fbe1_1651x313.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:276,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:482601,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://stevebyrnes1.substack.com/i/207917845?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F035fe3d7-0e7e-45d9-9810-493591e1fbe1_1651x313.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!A4k5!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F035fe3d7-0e7e-45d9-9810-493591e1fbe1_1651x313.png 424w, https://substackcdn.com/image/fetch/$s_!A4k5!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F035fe3d7-0e7e-45d9-9810-493591e1fbe1_1651x313.png 848w, https://substackcdn.com/image/fetch/$s_!A4k5!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F035fe3d7-0e7e-45d9-9810-493591e1fbe1_1651x313.png 1272w, https://substackcdn.com/image/fetch/$s_!A4k5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F035fe3d7-0e7e-45d9-9810-493591e1fbe1_1651x313.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a><figcaption class="image-caption">Left to right are from <a href="https://en.wikipedia.org/wiki/Metropolis_(1927_film)">Metropolis (1927)</a>, <a href="https://en.wikipedia.org/wiki/Woman_in_the_Moon">Woman in the Moon (1929),</a> and <a href="https://en.wikipedia.org/wiki/Icarus#/media/File:Gowy-icaro-prado.jpg">Gowy&#8217;s The Fall of Icarus (1636)</a></figcaption></figure></div><p><span>So the wrong question is: &#8220;Is it sci-fi?&#8221;. The right question is: &#8220;Is it possible?&#8221;</span></p><p><span>And the answer to that question is &#8220;Yes&#8221;! And we know this because we have an existence proof. Human brains and bodies can do all these things, and they don&#8217;t work by magic, but rather follow the principles of physics, math, and engineering, like everything else.</span></p><p><span>And whatever engineering principles allow humans to do all those right-column things, we should expect future scientists to sooner or later figure out how to exploit those same principles, in order to accomplish the same things. Even if that might seem impossible today! After all, think of how impossible vision must have seemed 1000 years ago: your eyes provide a magical window through which knowledge of your surroundings enters your soul. But now we understand the big-picture principles that explain how eyes work, and we have our own technology (cameras) based on similar principles. Ditto with hearing (microphones), moving (actuators), digestion (industrial catalysts), and so on. &#8230;</span><br></p></blockquote><p>Click the <a href="https://www.lesswrong.com/posts/nQH2GhkmSmHoCAN8R/what-do-i-mean-by-artificial-general-intelligence">link</a> to read the rest!</p>]]></content:encoded></item><item><title><![CDATA[Blog post: “Notes on technical alignment via human-like social drives”]]></title><description><![CDATA[Here&#8217;s the link:]]></description><link>https://stevebyrnes1.substack.com/p/blog-post-notes-on-technical-alignment</link><guid isPermaLink="false">https://stevebyrnes1.substack.com/p/blog-post-notes-on-technical-alignment</guid><dc:creator><![CDATA[Steve Byrnes]]></dc:creator><pubDate>Mon, 13 Jul 2026 18:37:00 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/cf2193d9-29e9-4f20-ba91-300bc24af155_4034x1200.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Here&#8217;s the link:</p><p><a href="https://www.alignmentforum.org/posts/rKdS7i4StaMmFzYRo/notes-on-technical-alignment-via-human-like-social-drives">https://www.alignmentforum.org/posts/rKdS7i4StaMmFzYRo/notes-on-technical-alignment-via-human-like-social-drives</a></p><p>And here&#8217;s how it begins:</p><blockquote><h1><span>1. Frontmatter</span></h1><h2><span>1.1 Backstory for this post</span></h2><p><span>As discussed in </span><a href="https://www.alignmentforum.org/s/HzcM2dkCq7fwXBej8"><span>Intro to Brain-Like-AGI Safety</span></a><span>, I&#8217;m working on the technical alignment problem for a hypothetical future &#8220;brain-like AGI&#8221;, with a particular focus on treating human innate social and moral drives as a possible jumping-off point for our technical alignment approach.</span></p><p><span>After all, if it&#8217;s possible for humans to do stuff that ultimately leads to a good future, then it&#8217;s probably also possible for sufficiently human-like AGIs to do stuff that ultimately leads to a good future. Or if it&#8217;s </span><em><span>not</span></em><span> possible for humans to do stuff that ultimately leads to a good future, then we&#8217;re screwed no matter what. But assuming it&#8217;s possible, the &#8220;sufficiently human-like AGIs&#8221; would certainly need to have good prosocial motivations. What code do we write that would lead to good prosocial motivations? It&#8217;s an unsolved problem (see </span><a href="https://www.alignmentforum.org/posts/oxvnREntu82tffkYW/we-need-a-field-of-reward-function-design"><span>&#8220;We need a field of Reward Function Design&#8221; (2025)</span></a><span>), but as we search for a solution, we might look for inspiration at how humans (sometimes) wind up with good prosocial motivations.</span></p><p><span>I&#8217;ve been working on this problem for years, but most of that work has involved laying foundations (e.g. trying to understand how human social drives work). Whereas in the past four months, I&#8217;ve been thinking very directly about how to apply those ideas to AGI.</span></p><p><span>I&#8217;ve published two little things from this four-month effort&#8212;</span><a href="https://www.alignmentforum.org/posts/RKtTi82t8X8TQy5FX/act-based-approval-directed-agents-for-ida-skeptics"><span>&#8220;Act-based approval-directed agents&#8221;, for IDA skeptics</span></a><span>, and </span><a href="https://www.alignmentforum.org/posts/vzHtHHBJoKATi5SeK/empowerment-corrigibility-etc-are-simple-abstractions-of-a"><span>Empowerment, corrigibility, etc. are simple abstractions (of a messed-up ontology)</span></a><span>&#8212;but most of what I (think I) figured out is not self-contained, but rather part of a big interconnected mess of thoughts and ideas. So for now I&#8217;m just dumping that all into this one excessively-long post. Sorry.</span></p><p><span>This post has lots of new-to-me ideas and opinions, and they are in great need of more scrutiny and especially </span><a href="https://intelligence.org/2018/11/22/2018-update-our-new-research-directions/#section2"><span>deconfusion</span></a><span>. I&#8217;m happy for any feedback, pushback, and riffs on anything herein.</span></p><h2><span>1.2 Table of contents / tl;dr</span></h2><p><strong><span>Sections 2&#8211;5</span></strong><span> start from human social instincts, and ponder how to usefully employ something-like-that in an AGI. Specifically:</span></p><ul><li><p><strong><span>Section 2</span></strong><span> goes over the high-level approach of using human social instincts as a starting point / inspiration for an AGI motivation system. What are the specific instincts in question, how do they work in humans, what roles if any should they be playing in AGI, and what should we be keeping in versus leaving out?</span></p></li><li><p><strong><span>Section 3</span></strong><span> discusses three potential failure modes that I&#8217;ve spent a long time thinking about: the possibility that the AGI will wind up with the wrong &#8220;moral circle&#8221;; the possibility that virtue-ethics-y human motivations (honesty, helpfulness, etc.) rely on a balance-of-power dynamic which wouldn&#8217;t apply to the ASI-human relationship; and the possibility that consequentialist desires will squash virtue-ethics-y desires in the long run, when both are present (as I claim they need to be).</span></p></li><li><p><strong><span>Section 4</span></strong><span> is &#8220;What controls the set of virtues that the AGI takes pride in?&#8221; I discuss two pathways: &#8220;person-first&#8221; and &#8220;desire-first&#8221;. For example, someone could wind up taking pride in their encyclopedic knowledge of Disney princesses because that&#8217;s what the cool older kids in school are into (&#8220;person-first&#8221;); or because they really really like Disney princess movies, and that love has gradually wormed its way into their self-image (&#8220;desire-first&#8221;). I suggest that the &#8220;desire-first&#8221; pathway is an important way that nerds like me wind up motivated to figure out the truth and share it with others&#8212;a key trait that we may want in brain-like AGI. (Hold that thought!)</span></p></li><li><p><strong><span>Section 5</span></strong><span> goes over some implementation details related to transplanting human social drives into the foreign soil of AGI source code.</span></p></li></ul><p><span>Then in </span><strong><span>Section 6</span></strong><span> I switch from forward-reasoning (starting with human social instincts) to backwards-reasoning (starting with desiderata), by asking: What AGI motivations do we want anyway?</span></p><ul><li><p><strong><span>Section 6.1</span></strong><span> goes over three sets of constraints that I&#8217;m trying to satisfy: technical alignment constraints (we need to be able to write the code), strategic constraints (we need to make the world resilient to misaligned ASI), and ethical / societal / buy-in constraints (the plan needs to sound reasonable, such that people will actually follow it).</span></p></li><li><p><strong><span>Section 6.2</span></strong><span> goes over a bunch of possible AGI motivations, and how they seem to stack up against those three sets of constraints. I wind up tentatively advocating for a high-level approach that I&#8217;ve been calling &#8220;truth-seeking disagreeable nerd AGI&#8221;, using the technical alignment idea mentioned in &#167;4 above, and then have that AGI figure out what to do next.</span></p></li></ul><p><strong><span>Section 7</span></strong><span> has a couple more random things from my notes:</span></p><ul><li><p><strong><span>Section 7.1</span></strong><span> discusses two (related) dilemmas that I&#8217;ve been struggling with. I call them &#8220;the visceral reaction updating dilemma&#8221; and &#8220;the value drift dilemma&#8221;.</span></p></li><li><p><strong><span>Section 7.2</span></strong><span> describes a subtle mistake in how I was thinking about &#8220;ruthlessness&#8221; until recently.</span></p></li></ul><p><span>I close in </span><strong><span>Section 8</span></strong><span> with what I plan to work on next and why.</span></p></blockquote><p>Click the <a href="https://www.alignmentforum.org/posts/rKdS7i4StaMmFzYRo/notes-on-technical-alignment-via-human-like-social-drives">link</a> to read the rest!</p>]]></content:encoded></item><item><title><![CDATA[Lotsa updates + “Sympathy for both sides of the egregious misalignment debate”]]></title><description><![CDATA[1.]]></description><link>https://stevebyrnes1.substack.com/p/lotsa-updates-sympathy-for-both-sides</link><guid isPermaLink="false">https://stevebyrnes1.substack.com/p/lotsa-updates-sympathy-for-both-sides</guid><dc:creator><![CDATA[Steve Byrnes]]></dc:creator><pubDate>Mon, 22 Jun 2026 19:29:54 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/7276c3f1-7777-44d4-9e0b-fedbaabd4887_432x389.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h1>1. But first, lotsa updates:</h1><h2>1.1 New version of <em>Neuroscience of human social instincts: a sketch</em></h2><p><a href="https://www.lesswrong.com/posts/kYvbHCDeMTCTE9TAj/neuroscience-of-human-social-instincts-a-sketch"><span>Neuroscience of human social instincts: a sketch (2024)</span></a><span> got another batch of changes (following a different set of important changes just last month), cleaning and streamlining a kinda muddled and partly-incorrect discussion of learning rate modulation, including by introducing new figures and new terminology. Here&#8217;s the changelog:</span></p><blockquote><p><strong><span>2026-05-19:</span></strong><span> I rewrote &#167;3&#8211;&#167;5.1 to remove unnecessary complication, and clean up some errors and muddled thinking. More details:</span></p><p><span>In &#167;3.2, I previously had a toy example of learning rate modulation in the thought assessors, where I was daydreaming about Taylor Swift, and then I suddenly orient to a spider jumping at me, and the learning rate modulation (I argued) was necessary to prevent learning that Taylor Swift is a risk factor for spiders jumping out at me. I do think that&#8217;s an actual solution to an actual problem &#8230; [but] I described this example poorly (and somewhat incorrectly), and more importantly it&#8217;s an example that&#8217;s not directly related to this post, and I think it was just causing unnecessary confusion (even I was confused when I re-read it). So I switched to a new example that overlaps much more with &#167;4. I also deleted the discussion of learning rate modulation in the Thought Generator, which I decided was somewhat misleading and confusing as written, and off-topic anyway.</span></p><p><span>That change to &#167;3, in turn, allowed me to shorten and streamline &#167;4, including in ways that hopefully made &#167;5.1 a bit clearer in turn.</span></p><p><span>The new version introduces and uses a new term I just made up, &#8220;interoceptive concept finder&#8221;, for a particular type of short-term predictor.</span></p></blockquote><p><span>(The </span><a href="https://doi.org/10.5281/zenodo.17953592"><span>archival PDF</span></a><span> is now up to version 3.)</span></p><div><hr></div><h2>1.2 Revisiting an old take on LLMs </h2><p>I&#8217;ll probably annoy all sides with this not-really-an-apology that I just tacked onto the beginning of my <a href="https://www.lesswrong.com/posts/KJRBb43nDxk6mwLcR/ai-doom-from-an-llm-plateau-ist-perspective">LLM-skepticism-related post from Apr 2023</a>:</p><blockquote><p><strong>Update June 2026:</strong> I basically stand by this post, except that I regret using the word &#8220;plateau&#8221;. By analogy, chess engines show no signs of &#8220;plateau-ing&#8221;: Stockfish 18 (2026) <a href="https://github.com/official-stockfish/Stockfish/releases"><span>crushes</span></a> Stockfish 17 (2024), which in turn crushes Stockfish 16 (2023), etc.<a href="https://www.lesswrong.com/posts/KJRBb43nDxk6mwLcR/ai-doom-from-an-llm-plateau-ist-perspective#fn3f9kw0pr9jy"><sup><span>[1]</span></sup></a> But chess engines ain&#8217;t gonna take over the world. So anyway, what I should have said was: There&#8217;s a school of thought, in which LLMs might (or might not) continue to get ever better at some things, but they will definitely always be bad at other things, and those &#8216;other things&#8217; are very important, such that their continued absence will prevent LLMs from ever becoming &#8220;transformative AI&#8221; as discussed below. This post is about how this school of thought relates to AI doom. And separately, if that school of thought is right, should we describe this situation using the word &#8220;plateau&#8221;? Arguably we <em>could</em>, in the same sense that chess engines would hit a &#8220;plateau&#8221; of 50% on a test with half chess puzzles and half sports trivia. But all things considered, I think the word &#8220;plateau&#8221; was a poor choice born of sloppy thinking. Sorry.</p></blockquote><div><hr></div><h2>1.3 More clarity on what &#8220;under-sculpting&#8221; means</h2><p>I added a few sentences to <a href="https://www.lesswrong.com/posts/grgb2ipxQf2wzNDEG/perils-of-under-vs-over-sculpting-agi-desires">Perils of under- vs over-sculpting AGI <span>desires (2025)</span></a><span> to hopefully make it clearer what &#8220;under-sculpting&#8221; means, especially this part:</span></p><blockquote><p>Think of it like: the AGI&#8217;s desires are a boat, and there&#8217;s a radio beacon (the reward function) that we know how to steer towards. If we keep steering towards the beacon (over-sculpting), then we can get very close to the beacon, but alas, the beacon is built on jagged rocks that will kill us (specification gaming, &#167;2). On the other hand, if we stop navigating towards the beacon at some point (under-sculpting), then we&#8217;re just adrift on the sea, and our two problems are: (1) we don&#8217;t know where we are right now (path dependence, &#167;8.1), and (2) we don&#8217;t know where the unpredictable currents will take us in the future (<a href="https://www.lesswrong.com/s/u9uawicHx7Ng7vwxA"><span>concept extrapolation</span></a> upon distribution shifts, &#167;8.2).</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!SEFg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20b5aa1c-ce43-44ac-bddb-3d3fc2255c16_1454x607.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!SEFg!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20b5aa1c-ce43-44ac-bddb-3d3fc2255c16_1454x607.png 424w, https://substackcdn.com/image/fetch/$s_!SEFg!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20b5aa1c-ce43-44ac-bddb-3d3fc2255c16_1454x607.png 848w, https://substackcdn.com/image/fetch/$s_!SEFg!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20b5aa1c-ce43-44ac-bddb-3d3fc2255c16_1454x607.png 1272w, https://substackcdn.com/image/fetch/$s_!SEFg!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20b5aa1c-ce43-44ac-bddb-3d3fc2255c16_1454x607.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!SEFg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20b5aa1c-ce43-44ac-bddb-3d3fc2255c16_1454x607.png" width="1454" height="607" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/20b5aa1c-ce43-44ac-bddb-3d3fc2255c16_1454x607.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:607,&quot;width&quot;:1454,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:756777,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://stevebyrnes1.substack.com/i/203131476?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20b5aa1c-ce43-44ac-bddb-3d3fc2255c16_1454x607.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!SEFg!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20b5aa1c-ce43-44ac-bddb-3d3fc2255c16_1454x607.png 424w, https://substackcdn.com/image/fetch/$s_!SEFg!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20b5aa1c-ce43-44ac-bddb-3d3fc2255c16_1454x607.png 848w, https://substackcdn.com/image/fetch/$s_!SEFg!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20b5aa1c-ce43-44ac-bddb-3d3fc2255c16_1454x607.png 1272w, https://substackcdn.com/image/fetch/$s_!SEFg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20b5aa1c-ce43-44ac-bddb-3d3fc2255c16_1454x607.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p></blockquote><div><hr></div><h2>1.4 A bunch of edits to <em><a href="https://www.lesswrong.com/s/HzcM2dkCq7fwXBej8">Intro to Brain-Like-AGI Safety</a></em></h2><p>See changelogs at the bottom of Posts <a href="https://www.lesswrong.com/posts/4basF9w9jaPZpoC8R/intro-to-brain-like-agi-safety-1-what-s-the-problem-and-why">1</a>, <a href="https://www.lesswrong.com/posts/F759WQ8iKjqBncDki/intro-to-brain-like-agi-safety-5-the-long-term-predictor-and">5</a>, <a href="https://www.lesswrong.com/posts/qNZSBqLEh4qLRqgWW/intro-to-brain-like-agi-safety-6-big-picture-of-motivation">6</a>, and <a href="https://www.lesswrong.com/posts/tj8AC3vhTnBywdZoA/intro-to-brain-like-agi-safety-15-conclusion-open-problems-1">15</a> for details. The biggest changes were thorough rewrites of <a href="https://www.lesswrong.com/posts/F759WQ8iKjqBncDki/intro-to-brain-like-agi-safety-5-the-long-term-predictor-and#5_3_The_RL_value_function__a_k_a__critic__a_k_a___valence_guess___as_a_special_case_of_long_term_prediction">&#167;5.3</a> on valence and reinforcement learning, <a href="https://www.lesswrong.com/posts/qNZSBqLEh4qLRqgWW/intro-to-brain-like-agi-safety-6-big-picture-of-motivation#6_4_More_on_how_this_connects_to_reinforcement_learning">&#167;6.4</a> on what I mean by &#8220;model-based RL&#8221;, and &#167;<a href="https://www.lesswrong.com/posts/qNZSBqLEh4qLRqgWW/intro-to-brain-like-agi-safety-6-big-picture-of-motivation#6_6_1_The_distinction_between_internalized_ego_syntonic_desires_and_externalized_ego_dystonic_urges_is_unrelated_to_Learning_Subsystem_vs__Steering_Subsystem">6.6.1</a> on ego-syntonic desires. (Those changes are not yet in the <a href="https://osf.io/preprints/osf/fe36n">archival PDF version</a>, sorry, I&#8217;ll upload a new PDF when I get a chance.)</p><p>The new version of &#167;6.6.1 is fun; here&#8217;s most of it:</p><blockquote><h4>6.6.1 The distinction between internalized ego-syntonic desires and externalized ego-dystonic urges is unrelated to Learning Subsystem vs. Steering Subsystem</h4><p>Many people (including me) have a strong intuitive distinction between <a href="https://en.wikipedia.org/wiki/Egosyntonic_and_egodystonic"><span>ego-syntonic drives</span></a> that are &#8220;part of us&#8221; or &#8220;what we want&#8221;, versus <a href="https://en.wikipedia.org/wiki/Egosyntonic_and_egodystonic"><span>ego-dystonic drives</span></a> that feel like urges which intrude upon us from the outside.</p><p>For example, if someone is on a hunger strike for freedom, they might say that their desire to eat comes from their innate drives, whereas their desire to fight for freedom comes from &#8220;reason&#8221;, or &#8220;their best self&#8221;, or whatever. What does that mean? What&#8217;s going on?</p><p>&#8230;</p><p>I propose that we should not take this intuition at face value. In reality, the hunger-striker&#8217;s desire to eat comes ultimately from their innate drives, and their desire to fight for freedom <em>also</em> comes ultimately from their innate drives! These are just two innate drives that are pointing in different directions, and thus they duke it out. One side will win, and then the person will either keep their hunger strike, or break down and eat.</p><p>It&#8217;s not so different from if you feel simultaneously very sleepy and very hungry; you can&#8217;t satisfy both drives, so the drives will duke it out, and one of them will win, and either you&#8217;ll nap despite your hunger, or you&#8217;ll eat despite your sleepiness.</p><p>That said, there do seem to be very important differences between the mundane hunger-vs-sleepiness battle (urge vs urge) and the hunger-vs-fight-for-freedom battle (urge vs ego-syntonic desire). In particular, here are three obvious questions that I need to address:</p><p><strong>First</strong>, hunger and sleep drives are easy to understand. But it&#8217;s much less obvious how innate drives could lead a person to care so much about freedom that they&#8217;ll go on a hunger strike. What exactly is the innate drive in question?</p><p><strong>Second</strong>, if I&#8217;m both hungry and sleepy, it&#8217;s sorta &#8220;a fair fight&#8221;. Probably whichever feeling is more immediately powerful will win. By contrast, the internal battle between the hunger-striker&#8217;s desire for freedom and their desire for food is <em>not</em> a fair fight. The former desire will punch above its weight by bringing far more intelligence and foresight to bear towards its objective. Thus, if the person is setting up a commitment mechanism, or tying their own hands, or attempting to control themselves and their mood, those actions will almost definitely be in pursuit of freedom, not in pursuit of hunger-satisfaction. Why the asymmetry? If both drives are in the Steering Subsystem, shouldn&#8217;t they be equally stupid and myopic? Relatedly, if the person is &#8220;applying willpower&#8221;, why is it in support of freedom rather than hunger-satisfaction? And by the way, what the heck does &#8220;applying willpower&#8221; even mean at a nuts-and-bolts level??</p><p><strong>Third</strong>, if the person&#8217;s desire for freedom is not really more a &#8220;part of them&#8221; than their desire to eat when hungry, then &#8230; why does it feel that way to them? In other words, this is an honest introspective report. I&#8217;m allowed to claim that the report should be interpreted as a perceptual illusion rather than taken at face value, but you have no reason to believe me unless I can also explain what exactly they were introspecting upon, and what they saw when they did so, and why it left them with the impression that it did.</p><p>These are all great questions! And I have answers to all of them! But they&#8217;re rather involved.</p><p><em>For the first question,</em> I claim that hunger-striking for freedom is driven by social instincts. You can learn about those at a high level in <a href="https://www.lesswrong.com/posts/5F5Tz3u6kJbTNMqsb/intro-to-brain-like-agi-safety-13-symbol-grounding-and-human"><span>Post #13</span></a> of this series, and then proceed to my follow-up work <a href="https://www.lesswrong.com/posts/kYvbHCDeMTCTE9TAj/neuroscience-of-human-social-instincts-a-sketch"><span>Neuroscience of human social instincts: a sketch (2024)</span></a>, and <a href="https://www.lesswrong.com/posts/KuBiv9cCbZ6ALjHFw/social-drives-1-sympathy-reward-from-compassion-to"><span>Social drives 1: &#8220;Sympathy Reward&#8221;, from compassion to dehumanization (2025)</span></a>, and <a href="https://www.lesswrong.com/posts/fPxgFHfs5yHzYqgG7/social-drives-2-approval-reward-from-norm-enforcement-to"><span>Social drives 2: &#8220;Approval Reward&#8221;, from norm-enforcement to status-seeking (2025)</span></a>, especially <a href="https://www.lesswrong.com/posts/fPxgFHfs5yHzYqgG7/social-drives-2-approval-reward-from-norm-enforcement-to#3__Approval_Reward_and_self_image"><span>&#167;3 of that last post</span></a> on how Approval Reward leads to pride in one&#8217;s self-image, and staying true to your principles even when nobody is watching. (That said, if people <em>are</em> watching, and you thus win the approval of people you admire, so much the better!)</p><p><em>For the second question,</em> my short answer is that people have a strong, salient association between thoughts of themselves, and thoughts of how they would look in someone else&#8217;s eyes. The former thoughts are required for strategically tying one&#8217;s hands, making precommitments, etc. And the latter thoughts summon social instincts like pride. This is the source of the asymmetry where social instincts &#8220;punch above their weight&#8221; in the context of intelligent self-reflective plans, such as precommitments, describing one&#8217;s life aspirations, and so on. For more detail (along with how &#8220;willpower&#8221; fits in), see <a href="https://www.lesswrong.com/posts/JLZnSnJptzmPtSRTc/intuitive-self-models-8-rooting-out-free-will-intuitions#8_5_5_Example_5__The_Active_Self_s_monopoly_on_sophisticated_brainstorming_and_planning"><span>&#167;8.5.5&#8211;&#167;8.5.6 of my &#8220;Intuitive Self-Models&#8221; series (2024)</span></a>, but be warned that it might not make sense without reading the earlier posts in that Intuitive Self-Models series, in which I try to unravel a bunch of misleading intuitions in how we think about our own minds.</p><p><em>The third question</em> (on why and how ego-syntonic desires are &#8220;internalized&#8221;) is also addressed in that same series, see <a href="https://www.lesswrong.com/posts/7tNq4hiSWW9GdKjY8/intuitive-self-models-3-the-homunculus#3_5_4_Why_are_ego_dystonic_things__externalized__"><span>&#8220;Intuitive Self Models&#8221; &#167;3.5.4</span></a>.</p><p>&#8230;</p><p>One way to put it is: why does the hunger-striker care about freedom, and not about, I dunno, ironing the wrinkles out of dollar bills? There has to be some explanation, right? And if you reply &#8220;it&#8217;s because freedom leads to (blah)&#8221;, then I&#8217;ll just reply, &#8220;OK, and why do they care about (blah), rather than, I dunno, measuring the distance between pebbles on the sidewalk?&#8221; We can go back and forth forever. The answer is: at the end of the day, <em>something</em> has to just plain feel intuitively good or bad. And that feeling has to come from an innate drive, one way or another.</p><p>Here are three more views on why we should believe that the Steering Subsystem is the ultimate source of not only ego-dystonic urges like hunger, but also ego-syntonic desires like friendship and justice.</p><ul><li><p><em>AI perspective:</em> We don&#8217;t yet know <em>in full detail</em> how model-based RL and model-based planning works in the human brain&#8212;we don&#8217;t have brain-like AGI yet. But we do at least vaguely know how these kinds of algorithms work. And we know enough to say for sure that these algorithms don&#8217;t develop prosocial motivations out of nowhere. For example, if you set the reward function of MuZero to always return 0, then the algorithm will emit random outputs forever&#8212;it won&#8217;t start fighting for justice.</p></li><li><p><em>Rodent model perspective:</em> For what it&#8217;s worth, researchers have been equally successful in finding little cell groups in the rodent hypothalamus that orchestrate &#8220;antisocial&#8221; behaviors like aggression, and that orchestrate &#8220;prosocial&#8221; behaviors like parenting and sociality. I fully expect that the same holds for humans.</p></li><li><p><em>Philosophy perspective:</em> Without the Steering Subsystem, the only thing the cortex can do is build a world-model from predictive learning of sensory inputs (<a href="https://www.lesswrong.com/posts/Y3bkJ59j4dciiLYyw/intro-to-brain-like-agi-safety-4-the-short-term-predictor#4_7__Short_term_predictor__example__2__Predictive_learning_of_sensory_inputs_in_the_cortex"><span>&#167;4.7</span></a>). That&#8217;s &#8220;is&#8221;, not &#8220;ought&#8221;. And <a href="https://en.wikipedia.org/wiki/Is%E2%80%93ought_problem"><span>&#8220;Hume&#8217;s law&#8221;</span></a> says that you can&#8217;t get &#8220;ought&#8221;-statements from exclusively &#8220;is&#8221;-statements. Granted, not everyone believes in Hume&#8217;s law. But I do&#8212;see an elegant and concise argument for it <a href="https://intelligence.org/2018/02/28/sam-harris-and-eliezer-yudkowsky/#:~:text=The%20question%20%E2%80%9CHow%20many%20paperclips%20result%20if%20I%20follow%20this%20policy%3F%E2%80%9D%20is%20an%20%E2%80%9Cis%E2%80%9D%20question."><span>here</span></a>.</p></li></ul></blockquote><p>This new section I added to Post 1 is also worth copying here, you wouldn&#8217;t believe how often people are confused by this:</p><blockquote><h4>1.3.4 So is &#8220;Brain-like AGI&#8221; a good plan? Or is it a <em>threat model</em>?</h4><p>Lots of people assume that, if I&#8217;m devoting my career to brain-like-AGI safety, I must be very enthusiastic about brain-like AGI.</p><p>By analogy, if someone is devoting their career to rocket engine safety, odds are high that they&#8217;re a space nerd who thinks that rocket engines are really cool and great.</p><p>&#8230;But on the other hand, people devote their careers to earthquake safety too! Do those people think earthquakes are really cool and great? Of course not! But they recognize that earthquakes will come, whether we want them or not, so we&#8217;d better prepare.</p><p>Now as it turns out, I&#8217;m much more like the earthquake safety person than the rocket safety person: I think of brain-like AGI as a threat model. Honestly, I expect that brain-like AGI will probably kill us all, in a manner that makes <a href="https://en.wikipedia.org/wiki/Skynet_(Terminator)"><span>Skynet</span></a> (from the <em>Terminator</em> movies) look primitive and sentimental. You don&#8217;t have to agree! Indeed, both enthusiasts and naysayers have a strong shared interest in understanding potential safety problems and designing mitigations, in a constructive, pedagogical, technical, and detail-oriented way. That&#8217;s my aim in this series. But I did want to lay my cards on the table.</p></blockquote><div><hr></div><h2>1.5 Is valence a linear function on &#8220;thoughts&#8221;?</h2><p>In <a href="https://www.lesswrong.com/posts/SqgRtCwueovvwxpDQ/valence-series-2-valence-and-normativity">[Valence series] 2. Valence &amp; Normativity (2023)</a>, I made a claim in <a href="https://www.lesswrong.com/posts/SqgRtCwueovvwxpDQ/valence-series-2-valence-and-normativity#2_4_1_Side_note__Valence_as_a__roughly__linear_function_over_compositional_thought_pieces">&#167;2.4.1</a> that &#8220;valence is a (roughly) linear function over compositional thought-pieces&#8221;, and then immediately in <a href="https://www.lesswrong.com/posts/SqgRtCwueovvwxpDQ/valence-series-2-valence-and-normativity#2_4_1_1__Apparent__counterexamples_to_linearity">&#167;2.4.1.1</a> listed a bunch of examples where that claim seems totally wrong&#8212;for example, &#8220;someone I hate is suffering&#8221; or &#8220;I&#8217;m gonna avoid traffic&#8221; (both &#8220;someone I hate&#8221; and &#8220;traffic&#8221; are bad, but those thoughts seem overall good). I had (and still have) sound neuroanatomical and algorithmic reasons to posit linearity, but in 2023 I had a pretty hazy understanding of why this hypothesis was not immediately ruled out by those examples. Anyway, I can explain this substantially better now (albeit still a <em>bit</em> hazy), and I rewrote <a href="https://www.lesswrong.com/posts/SqgRtCwueovvwxpDQ/valence-series-2-valence-and-normativity#2_4_1_1__Apparent__counterexamples_to_linearity">&#167;2.4.1.1</a> accordingly.</p><div><hr></div><p><em>(Almost all of those edits were in response to criticisms from Rif A. Saurous, many thanks to him.)</em></p><div><hr></div><h1>2. New blog post: &#8220;Sympathy for both sides of the egregious misalignment debate&#8221;</h1><p>Here&#8217;s the link:</p><p><a href="https://www.alignmentforum.org/posts/DZaZ3fqHnvfLCftPu/sympathy-for-both-sides-of-the-egregious-misalignment-debate">https://www.alignmentforum.org/posts/DZaZ3fqHnvfLCftPu/sympathy-for-both-sides-of-the-egregious-misalignment-debate</a></p><p>And here&#8217;s the first half of it:</p><blockquote><p><span>On one side of this debate is Yudkowsky &amp; Soares, who think that (if AI progress continues) we&#8217;re on a direct path to egregiously-misaligned, scheming, out-of-control, rogue superintelligence (ASI), not even </span><a href="https://www.alignmentforum.org/posts/xvBZPEccSfM8Fsobt/what-are-the-best-arguments-for-against-ais-being-slightly"><span>slightly nice</span></a><span>, in the absence of yet-to-be-invented breakthrough technical alignment ideas.</span></p><p><span>On the other side of this debate is almost everyone who works on or studies LLMs. Some of them are very concerned about egregious scheming, others much less so, and as a group they&#8217;re equally or more concerned about lots of other potential AI problems&#8212;AI-assisted bioterrorism, AI-assisted dictatorships, etc. And </span><em><span>if</span></em><span> they&#8217;re concerned about egregious misalignment and scheming, they&#8217;ll often say that it would come about through being in too much of a rush, or careless programmers, or bad actors, etc., as opposed to the simpler Yudkowsky &amp; Soares story of &#8220;we get egregious misalignment and scheming because nobody has the foggiest idea how to avoid that&#8221;.</span></p><p><span>Here&#8217;s my brief idiosyncratic take on this debate. </span><strong><span>I think BOTH of the following are true:</span></strong></p><ul><li><p><strong><span>(1)</span></strong><span> If you really think carefully about the properties of ASI, you </span><em><span>really do</span></em><span> find good reasons to strongly expect it to be egregiously misaligned, scheming, and ruthless, in the absence of yet-to-be-invented breakthrough technical alignment ideas.</span></p></li><li><p><strong><span>(2)</span></strong><span> If you really think carefully about the properties of current LLMs, you </span><em><span>really do</span></em><span> find good reasons to think that existing technical alignment techniques are adequate now, and may well continue to be adequate in the future.</span></p></li></ul><p><span>So then here are three (caricatured) positions:</span></p><h2>My position:</h2><blockquote><p>(1) and (2) are both totally true. And we can reconcile them by saying that LLMs won&#8217;t scale to ASI.</p></blockquote><h2>Yudkowsky &amp; Soares&#8217;s position [caricatured]:</h2><blockquote><p>(1) is totally true. We know this with great confidence, having spent decades thinking about it.</p><p>So it follows that (2) must be wrong or irrelevant.</p><p>Why is (2) wrong or irrelevant? Hard to say! There&#8217;s no ASI yet, and nobody knows in detail how it will appear. Sometimes it&#8217;s easier to predict what happens eventually than the detailed path. An ice cube in warm water will melt eventually, but don&#8217;t ask me to predict how many seconds it will take to melt, etc.</p><p>So anyway, one possibility is that (2) is wrong because LLMs will kinda &#8216;wake up&#8217;, or something, when the core pieces of true intelligence finally come together. And then their behavior would change drastically for the worse. And maybe we&#8217;re already starting to see glimmers of that in existing LLMs?</p><p>Or another possibility [cf. <a href="https://x.com/allTheYud/status/2039798247334826445?s=20">Eliezer tweet</a>] is that LLMs will invent non-LLM ASI. And then (2) will be simply irrelevant!</p><p>&#8230;Or something else! Again, we don&#8217;t know! But we do know that (1) is definitely right.</p></blockquote><h2>LLM people&#8217;s position [caricatured]:</h2><blockquote><p>(2) is totally true. We know this with great confidence, because we are LLM experts and we have thought about these alignment plans in great detail, including matching our theories against real-world data.</p><p>So it follows that (1) must be incorrect.</p><p>Why is (1) incorrect? I don&#8217;t really know! Man, I read Yudkowsky and Soares, and it&#8217;s all these words, words, words, and I&#8217;m reading along and trying to match those words to my knowledge of LLMs and it just doesn&#8217;t make any damn sense. I can and will try to respond to their points in detail, but honestly the core issue is that they&#8217;re guilty of head-in-the-clouds armchair theorizing gone off the rails.</p></blockquote><h2>Conclusion</h2><p>&#8230;So I think that both sides of the debate are basically coming from a reasonable and sympathetic place, with a big kernel of truth.</p><p>&#8230;</p></blockquote><p>Click the <a href="https://www.alignmentforum.org/posts/DZaZ3fqHnvfLCftPu/sympathy-for-both-sides-of-the-egregious-misalignment-debate">link</a> to read the rest!</p>]]></content:encoded></item><item><title><![CDATA[Blog post: “Empowerment, corrigibility, etc. are simple abstractions (of a messed-up ontology)”]]></title><description><![CDATA[Plus other updates]]></description><link>https://stevebyrnes1.substack.com/p/blog-post-empowerment-corrigibility</link><guid isPermaLink="false">https://stevebyrnes1.substack.com/p/blog-post-empowerment-corrigibility</guid><dc:creator><![CDATA[Steve Byrnes]]></dc:creator><pubDate>Mon, 11 May 2026 19:17:30 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/2a91f5d0-38a9-47ce-aa9f-ad771f85ba67_1128x609.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h1>But first, some important revisions to old posts:</h1><h2>The orienting reflex as exemplifying a broader neuroscience motif</h2><p>I revised 2 old posts based on deeper appreciation for <a href="https://en.wikipedia.org/wiki/Orienting_response">the orienting reflex</a> as exemplifying a broader neuroscience motif:</p><p>(1) In <a href="https://www.lesswrong.com/posts/kYvbHCDeMTCTE9TAj/neuroscience-of-human-social-instincts-a-sketch">Neuroscience of human social instincts: a sketch (2024)</a>:</p><blockquote><p>I changed terminology from <em>&#8220;the &#8216;thinking of a conspecific&#8217; flag&#8221;</em> to <em>&#8220;the social attention reflex&#8221;</em>. I think the new term has better connotations, especially the way it invokes a parallel to <a href="https://en.wikipedia.org/wiki/Orienting_response">&#8220;orienting reflex&#8221;</a> and <a href="https://en.wikipedia.org/wiki/Startle_response">&#8220;startle reflex&#8221;</a>, which likewise are associated with fast, transient, and involuntary changes in both attention and other innate signals like pleasure and arousal.</p></blockquote><p>(I made some other changes to this one too&#8212;see <a href="https://www.lesswrong.com/posts/kYvbHCDeMTCTE9TAj/neuroscience-of-human-social-instincts-a-sketch#Changelog">changelog</a>&#8212;and then also posted an updated version 2 of the archival PDF: <a href="https://doi.org/10.5281/zenodo.17953592">https://doi.org/10.5281/zenodo.17953592</a>.)</p><p>(2) In <a href="https://www.lesswrong.com/posts/rM8DwFKZM4eB7i2p8/valence-series-3-valence-and-beliefs#3_3_5_An_exception__I_think_anxious___obsessive__brainstorming__is_driven_by_involuntary_attention_rather_than_by_valence">Valence series &#167;3.3.5 (2023)</a>:</p><blockquote><p>(<strong>Update April 2026:</strong> I&#8217;ve refined my model a bit: I now think of anxiety-related involuntary attention as being more closely analogous to the famous <a href="https://en.wikipedia.org/wiki/Orienting_response">orienting reflex</a> wherein people turn to look at an unexpected loud sound or motion. Traditional orienting reflexes involve involuntary attention towards exteroceptive inputs, coupled with innate motor commands, physiological arousal, etc. By analogy, if you&#8217;re anxious, then (I claim) you&#8217;ll likewise experience sporadic <em>interoceptive</em> &#8220;orienting reflexes&#8221; that involve involuntary attention towards the the feeling of anxiety, coupled with a synchronized squirt of negative valence and displeasure (see <a href="https://www.lesswrong.com/posts/xLmzMjxgZDaLyxZKb/valence-series-appendix-a-hedonic-tone-dis-pleasure-dis">Appendix A</a>), plus physiological arousal etc. These interoceptive &#8220;orienting reflexes&#8221; might occur multiple times per second for intense anxiety, or less often for milder anxiety. I&#8217;m using anxiety as an example, but the same idea obviously applies as well to fear, hunger, itches, etc.)</p></blockquote><div><hr></div><h2>Renouncing an old discussion of narcissism</h2><p>One section of my 2023 post <a href="https://www.lesswrong.com/posts/txj4wigyjLNbcoZ9o/valence-series-5-valence-disorders-in-mental-health-and">[Valence series] 5. &#8220;Valence Disorders&#8221; in Mental Health &amp; Personality</a> was a hypothesis about narcissistic personality disorder. Now it&#8217;s 2&#189; years later, and I&#8217;ve figured out much more about human social instincts, and when I revisited what I wrote in 2023, I decided it was all wrong, and struck it out. I will hopefully write up &#8220;My Model of NPD, Take 2&#8221; when I get a chance.</p><div><hr></div><h2>Nuances on the idea of updating-to-convergence</h2><p>I added a section to <a href="https://www.lesswrong.com/posts/TprdAhgTvr3tuDJsD/against-empathy-by-default">Against empathy-by-default (2024)</a> reflecting some things that I&#8217;ve learned since then:</p><blockquote><p><strong>3.3 </strong><em><strong>(Added April 2026)</strong></em><strong> Things vaguely analogous to the empathy-by-default argument can happen for certain visceral reactions&#8212;just not for the core motivation / reward / RL system</strong></p><p>In the above (especially &#167;3.1), I think I conveyed a general vibe that within-lifetime learning will eventually, inexorably, correct all &#8220;errors&#8221; in learned brain models. But a while after writing this post, I came to better appreciate how certain visceral reactions in the brain can be set up so as to sometimes prevent updates (&#8220;corrections&#8221;). This is how people can wind up with stable phobias, and stable food-aversions, and stable traumas, and stable autistic &#8220;special interests&#8221;, and so on, even when there&#8217;s no particular innate &#8220;ground truth&#8221; underlying them. These stable &#8220;errors&#8221; are the exception not the rule, but it&#8217;s interesting that they exist at all, and that they can in some cases last a lifetime.</p><p>I discuss the algorithmic trick behind these in a later post: <a href="https://www.lesswrong.com/posts/grgb2ipxQf2wzNDEG/perils-of-under-vs-over-sculpting-agi-desires">&#8220;Perils of under- vs over-sculpting AGI desires&#8221; (2025)</a>, specifically <a href="https://www.lesswrong.com/posts/grgb2ipxQf2wzNDEG/perils-of-under-vs-over-sculpting-agi-desires#6_2__Defer_to_predictor_mode__in_visceral_reactions__and__trapped_priors_">&#167;6.2: &#8220;&#8216;Defer-to-predictor mode&#8217; in visceral reactions, and &#8216;trapped priors&#8217;</a>. In brief, the trick centers around what I call &#8220;defer-to-predictor mode&#8221;, where e.g. a visceral expectation of imminent disgust can cause an actual disgust reaction. But there&#8217;s a loop-y thing, wherein the actual disgust is in turn the ground truth for how we learn that a visceral expectation of disgust is warranted. Thanks to this loop-y thing, we can wind up without any error signal telling our brains that the disgust was never warranted in the first place.&#8230;</p></blockquote><h1>New blog post: &#8220;Empowerment, corrigibility, etc. are simple abstractions (of a messed-up ontology)&#8221;</h1><p>Here&#8217;s the link:</p><p><a href="https://www.alignmentforum.org/posts/vzHtHHBJoKATi5SeK/empowerment-corrigibility-etc-are-simple-abstractions-of-a">https://www.alignmentforum.org/posts/vzHtHHBJoKATi5SeK/empowerment-corrigibility-etc-are-simple-abstractions-of-a</a></p><p>And here&#8217;s the start:</p><blockquote><p><strong>1.1 Tl;dr</strong></p><p>Alignment is often conceptualized as AIs helping humans achieve their goals: AIs that increase people&#8217;s agency and empowerment; AIs that are helpful, corrigible, and/or obedient; AIs that avoid manipulating people. But that last one&#8212;manipulation&#8212;points to a challenge for all these desiderata: a human&#8217;s goals are <em>themselves</em> under-determined and manipulable, and it&#8217;s awfully hard to pin down a principled distinction between changing people&#8217;s goals in a good way (&#8220;providing counsel&#8221;, &#8220;providing information&#8221;, &#8220;sharing ideas&#8221;) versus a bad way (&#8220;manipulating&#8221;, &#8220;brainwashing&#8221;).</p><p>The manipulability of human desires is hardly a new observation in the alignment literature, but it remains unsolved (see lit review in &#167;3 below).</p><p>In this post I will propose an explanation of how <em>we humans</em> intuitively conceptualize the distinction between guidance (good) vs manipulation (bad), in case it helps us brainstorm how we might put that distinction into AI.</p><p>&#8230;But (spoiler alert) it turns out not to really help, because I&#8217;ll argue that we humans think about it in a deeply incoherent way, intimately tied to our scientifically-inaccurate intuitions around free will.</p><p>I jump from there into a broader review of every approach that I can think of for writing a <a href="https://www.alignmentforum.org/posts/FWvzwCDRgcjb9sigb/why-agent-foundations-an-overly-abstract-explanation">&#8220;True Name&#8221;</a> for manipulation or things related to it (empowerment, agency, corrigibility, culpability, etc.), or indeed for any other method of robustly getting future AGIs to be able to talk to people without trying to manipulate those people&#8217;s desires. I argue that none of them provides much of a path forward on the particular technical alignment problem I&#8217;m working on. Indeed, my current guess is that none of these things have a &#8220;True Name&#8221; at all, or at least not one that&#8217;s useful for the technical alignment problem.</p></blockquote><p>Click the <a href="https://www.alignmentforum.org/posts/vzHtHHBJoKATi5SeK/empowerment-corrigibility-etc-are-simple-abstractions-of-a">link</a> to read the rest!</p>]]></content:encoded></item><item><title><![CDATA[2 blog posts (+ 3 updates): “You can’t imitation-learn how to continual-learn” & “‘Act-based approval-directed agents’, for IDA skeptics”]]></title><description><![CDATA[Start with three quick updates:]]></description><link>https://stevebyrnes1.substack.com/p/2-blog-posts-3-updates-you-cant-imitation</link><guid isPermaLink="false">https://stevebyrnes1.substack.com/p/2-blog-posts-3-updates-you-cant-imitation</guid><dc:creator><![CDATA[Steve Byrnes]]></dc:creator><pubDate>Wed, 18 Mar 2026 19:33:33 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/652a04b6-704f-45d2-a468-71e86612208c_3862x1913.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h1>Start with three quick updates:</h1><ul><li><p>I shared a link-post to a podcast, along with excerpts of the part I found most interesting. Title: <a href="https://www.lesswrong.com/posts/hvun2mP2yEr4kyKWk/podcast-jeremy-howard-is-bearish-on-llms">&#8220;Jeremy Howard is bearish on LLMs&#8221;</a>.</p><ul><li><p>To be clear, he co-invented LLMs and uses them enthusiastically. He is only &#8220;bearish&#8221; by the standards of people in my AI Alignment professional circles, many of whom have gotten pretty frantic in the past couple months about LLMs imminently rocketing to superintelligence.</p></li></ul></li><li><p>I was invited back for  second appearance on the &#8220;Doom Debates Podcast&#8221;. Our conversation centered around my post <a href="https://www.lesswrong.com/posts/ZJZZEuPFKeEdkrRyf/why-we-should-expect-ruthless-sociopath-asi">Why we should expect ruthless sociopath ASI</a>:</p></li></ul><div class="embedded-post-wrap" data-attrs="{&quot;id&quot;:190358182,&quot;url&quot;:&quot;https://lironshapira.substack.com/p/ai-genius-returns-to-warn-of-ruthless&quot;,&quot;publication_id&quot;:1777870,&quot;embedding_publication_id&quot;:null,&quot;publication_name&quot;:&quot;Doom Debates&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!uKmU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72a689b8-f59a-4bde-bdda-c6d1bf0b45f5_1280x1280.png&quot;,&quot;title&quot;:&quot;How Friendly AI Will Become Deadly &#8212; Dr. Steven Byrnes (AGI Safety Researcher, Harvard Physics Postdoc) Returns!&quot;,&quot;truncated_body_text&quot;:&quot;Fan favorite Dr. Steven Byrnes returns to discuss recent AI progress and the concerning paradigm shift to \&quot;ruthless sociopath AI\&quot; he sees on the horizon.&quot;,&quot;date&quot;:&quot;2026-03-10T17:03:15.644Z&quot;,&quot;like_count&quot;:6,&quot;comment_count&quot;:0,&quot;bylines&quot;:[{&quot;id&quot;:1862712,&quot;name&quot;:&quot;Liron Shapira&quot;,&quot;handle&quot;:&quot;doomdebates&quot;,&quot;previous_name&quot;:null,&quot;photo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!FAkx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d309daa-9032-47fa-82bf-fb88829cbf92_128x128.webp&quot;,&quot;bio&quot;:&quot;Host of Doom Debates, the #1 show for high-stakes debate about AI extinction risk.&quot;,&quot;profile_set_up_at&quot;:&quot;2021-08-12T12:47:56.741Z&quot;,&quot;reader_installed_at&quot;:&quot;2022-11-04T00:22:02.735Z&quot;,&quot;publicationUsers&quot;:[{&quot;id&quot;:1760437,&quot;user_id&quot;:1862712,&quot;publication_id&quot;:1777870,&quot;role&quot;:&quot;admin&quot;,&quot;public&quot;:true,&quot;is_primary&quot;:true,&quot;publication&quot;:{&quot;id&quot;:1777870,&quot;name&quot;:&quot;Doom Debates&quot;,&quot;subdomain&quot;:&quot;lironshapira&quot;,&quot;custom_domain&quot;:null,&quot;custom_domain_optional&quot;:false,&quot;hero_text&quot;:&quot;AI debates that must be resolved before the world ends&quot;,&quot;logo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/72a689b8-f59a-4bde-bdda-c6d1bf0b45f5_1280x1280.png&quot;,&quot;author_id&quot;:1862712,&quot;primary_user_id&quot;:1862712,&quot;theme_var_background_pop&quot;:&quot;#BAA049&quot;,&quot;created_at&quot;:&quot;2023-07-04T13:54:38.054Z&quot;,&quot;email_from_name&quot;:&quot;Liron Shapira&quot;,&quot;copyright&quot;:&quot;Liron Shapira&quot;,&quot;founding_plan_name&quot;:&quot;Founding Member&quot;,&quot;community_enabled&quot;:true,&quot;invite_only&quot;:false,&quot;payments_state&quot;:&quot;enabled&quot;,&quot;language&quot;:null,&quot;explicit&quot;:false,&quot;homepage_type&quot;:&quot;newspaper&quot;,&quot;is_personal_mode&quot;:false,&quot;logo_url_wide&quot;:null}}],&quot;twitter_screen_name&quot;:&quot;liron&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null,&quot;status&quot;:{&quot;bestsellerTier&quot;:null,&quot;subscriberTier&quot;:10,&quot;leaderboard&quot;:null,&quot;vip&quot;:false,&quot;badge&quot;:{&quot;type&quot;:&quot;subscriber&quot;,&quot;tier&quot;:10,&quot;accent_colors&quot;:null},&quot;paidPublicationIds&quot;:[4718,98102,2355025,89120,1383,35345,707415,68038,17302,1295435,1245641,375183,260347],&quot;subscriber&quot;:null}}],&quot;utm_campaign&quot;:null,&quot;belowTheFold&quot;:false,&quot;type&quot;:&quot;podcast&quot;,&quot;language&quot;:&quot;en&quot;,&quot;source&quot;:null}" data-component-name="EmbeddedPostToDOM"><a class="embedded-post" native="true" href="https://lironshapira.substack.com/p/ai-genius-returns-to-warn-of-ruthless?utm_source=substack&amp;utm_campaign=post_embed&amp;utm_medium=web"><div class="embedded-post-header"><img class="embedded-post-publication-logo" src="https://substackcdn.com/image/fetch/$s_!uKmU!,w_56,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72a689b8-f59a-4bde-bdda-c6d1bf0b45f5_1280x1280.png"><span class="embedded-post-publication-name">Doom Debates</span></div><div class="embedded-post-title-wrapper"><div class="embedded-post-title-icon"><svg width="19" height="19" viewBox="0 0 24 24" fill="none" xmlns="http://www.w3.org/2000/svg">
  <path d="M3 18V12C3 9.61305 3.94821 7.32387 5.63604 5.63604C7.32387 3.94821 9.61305 3 12 3C14.3869 3 16.6761 3.94821 18.364 5.63604C20.0518 7.32387 21 9.61305 21 12V18" stroke-linecap="round" stroke-linejoin="round"></path>
  <path d="M21 19C21 19.5304 20.7893 20.0391 20.4142 20.4142C20.0391 20.7893 19.5304 21 19 21H18C17.4696 21 16.9609 20.7893 16.5858 20.4142C16.2107 20.0391 16 19.5304 16 19V16C16 15.4696 16.2107 14.9609 16.5858 14.5858C16.9609 14.2107 17.4696 14 18 14H21V19ZM3 19C3 19.5304 3.21071 20.0391 3.58579 20.4142C3.96086 20.7893 4.46957 21 5 21H6C6.53043 21 7.03914 20.7893 7.41421 20.4142C7.78929 20.0391 8 19.5304 8 19V16C8 15.4696 7.78929 14.9609 7.41421 14.5858C7.03914 14.2107 6.53043 14 6 14H3V19Z" stroke-linecap="round" stroke-linejoin="round"></path>
</svg></div><div class="embedded-post-title">How Friendly AI Will Become Deadly &#8212; Dr. Steven Byrnes (AGI Safety Researcher, Harvard Physics Postdoc) Returns!</div></div><div class="embedded-post-body">Fan favorite Dr. Steven Byrnes returns to discuss recent AI progress and the concerning paradigm shift to "ruthless sociopath AI" he sees on the horizon&#8230;</div><div class="embedded-post-cta-wrapper"><div class="embedded-post-cta-icon"><svg width="32" height="32" viewBox="0 0 24 24" xmlns="http://www.w3.org/2000/svg">
  <path classname="inner-triangle" d="M10 8L16 12L10 16V8Z" stroke-width="1.5" stroke-linecap="round" stroke-linejoin="round"></path>
</svg></div><span class="embedded-post-cta">Listen now</span></div><div class="embedded-post-meta">5 months ago &#183; 6 likes &#183; Liron Shapira</div></a></div><ul><li><p>Patrick Mineault published a nice two-part discussion of how cell type diversity is connected to the big picture of brain algorithms:</p></li></ul><div class="embedded-post-wrap" data-attrs="{&quot;id&quot;:189321289,&quot;url&quot;:&quot;https://www.neuroai.science/p/cell-types-encoding-the-brains-bios&quot;,&quot;publication_id&quot;:1564943,&quot;embedding_publication_id&quot;:null,&quot;publication_name&quot;:&quot;The NeuroAI archive&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!WvCx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a3be57f-f6d5-4684-98b8-859ef181e86e_798x798.png&quot;,&quot;title&quot;:&quot;Cell types: encoding the brain's BIOS&quot;,&quot;truncated_body_text&quot;:&quot;In an interview late last year on the Dwarkesh Podcast, Adam Marblestone&#8212;CEO of Convergent Research, former research scientist at Deepmind, and a true polymath in connecting neuroscience and AI&#8212;discussed what insights we might extract from complete brain wiring diagrams. His central claim is that the most valuable information in a connectome concerns i&#8230;&quot;,&quot;date&quot;:&quot;2026-02-27T17:29:14.118Z&quot;,&quot;like_count&quot;:34,&quot;comment_count&quot;:5,&quot;bylines&quot;:[{&quot;id&quot;:17921567,&quot;name&quot;:&quot;Patrick Mineault&quot;,&quot;handle&quot;:&quot;patrickmineault&quot;,&quot;previous_name&quot;:null,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e0dbee55-9a27-4d31-b784-8779443b8f7d_400x400.jpeg&quot;,&quot;bio&quot;:&quot;Neuro AI, vision, Python, open science. NeuroAI lead at Amaranth. Previously engineer @ Google, Meta, Mila. &quot;,&quot;profile_set_up_at&quot;:&quot;2023-03-03T14:47:16.973Z&quot;,&quot;reader_installed_at&quot;:&quot;2023-03-03T14:46:36.453Z&quot;,&quot;publicationUsers&quot;:[{&quot;id&quot;:1534657,&quot;user_id&quot;:17921567,&quot;publication_id&quot;:1564943,&quot;role&quot;:&quot;admin&quot;,&quot;public&quot;:true,&quot;is_primary&quot;:true,&quot;publication&quot;:{&quot;id&quot;:1564943,&quot;name&quot;:&quot;The NeuroAI archive&quot;,&quot;subdomain&quot;:&quot;naix&quot;,&quot;custom_domain&quot;:&quot;www.neuroai.science&quot;,&quot;custom_domain_optional&quot;:false,&quot;hero_text&quot;:&quot;Research on Neuro-Inspired AI and AI-accelerated Neuro. &quot;,&quot;logo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1a3be57f-f6d5-4684-98b8-859ef181e86e_798x798.png&quot;,&quot;author_id&quot;:17921567,&quot;primary_user_id&quot;:17921567,&quot;theme_var_background_pop&quot;:&quot;#67BDFC&quot;,&quot;created_at&quot;:&quot;2023-04-08T16:21:53.936Z&quot;,&quot;email_from_name&quot;:&quot;Patrick Mineault from NeuroAI archive&quot;,&quot;copyright&quot;:&quot;Patrick Mineault&quot;,&quot;founding_plan_name&quot;:null,&quot;community_enabled&quot;:true,&quot;invite_only&quot;:false,&quot;payments_state&quot;:&quot;disabled&quot;,&quot;language&quot;:null,&quot;explicit&quot;:false,&quot;homepage_type&quot;:&quot;newspaper&quot;,&quot;is_personal_mode&quot;:false,&quot;logo_url_wide&quot;:null}}],&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null,&quot;status&quot;:{&quot;bestsellerTier&quot;:null,&quot;subscriberTier&quot;:1,&quot;leaderboard&quot;:null,&quot;vip&quot;:false,&quot;badge&quot;:{&quot;type&quot;:&quot;subscriber&quot;,&quot;tier&quot;:1,&quot;accent_colors&quot;:null},&quot;paidPublicationIds&quot;:[2133349,3477619,35345],&quot;subscriber&quot;:null}}],&quot;utm_campaign&quot;:null,&quot;belowTheFold&quot;:false,&quot;type&quot;:&quot;newsletter&quot;,&quot;language&quot;:&quot;en&quot;,&quot;source&quot;:null}" data-component-name="EmbeddedPostToDOM"><a class="embedded-post" native="true" href="https://www.neuroai.science/p/cell-types-encoding-the-brains-bios?utm_source=substack&amp;utm_campaign=post_embed&amp;utm_medium=web"><div class="embedded-post-header"><img class="embedded-post-publication-logo" src="https://substackcdn.com/image/fetch/$s_!WvCx!,w_56,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a3be57f-f6d5-4684-98b8-859ef181e86e_798x798.png"><span class="embedded-post-publication-name">The NeuroAI archive</span></div><div class="embedded-post-title-wrapper"><div class="embedded-post-title">Cell types: encoding the brain's BIOS</div></div><div class="embedded-post-body">In an interview late last year on the Dwarkesh Podcast, Adam Marblestone&#8212;CEO of Convergent Research, former research scientist at Deepmind, and a true polymath in connecting neuroscience and AI&#8212;discussed what insights we might extract from complete brain wiring diagrams. His central claim is that the most valuable information in a connectome concerns i&#8230;</div><div class="embedded-post-cta-wrapper"><span class="embedded-post-cta">Read more</span></div><div class="embedded-post-meta">5 months ago &#183; 34 likes &#183; 5 comments &#183; Patrick Mineault</div></a></div><div class="embedded-post-wrap" data-attrs="{&quot;id&quot;:189841053,&quot;url&quot;:&quot;https://www.neuroai.science/p/cell-types-wiring-up-innate-circuits&quot;,&quot;publication_id&quot;:1564943,&quot;embedding_publication_id&quot;:null,&quot;publication_name&quot;:&quot;The NeuroAI archive&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!WvCx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a3be57f-f6d5-4684-98b8-859ef181e86e_798x798.png&quot;,&quot;title&quot;:&quot;Cell types: from genes to circuits&quot;,&quot;truncated_body_text&quot;:&quot;In the last post, I presented a counterintuitive idea: areas of the brain that contain many cell types encode innate behaviors and rewards. Myriad neurons with strange morphology that are connected idiosyncratically are an indicator of evolutionary pressure: behaviors and rewards that have survival value become genetically hard-coded. What I didn&#8217;t get &#8230;&quot;,&quot;date&quot;:&quot;2026-03-04T15:47:50.447Z&quot;,&quot;like_count&quot;:12,&quot;comment_count&quot;:0,&quot;bylines&quot;:[{&quot;id&quot;:17921567,&quot;name&quot;:&quot;Patrick Mineault&quot;,&quot;handle&quot;:&quot;patrickmineault&quot;,&quot;previous_name&quot;:null,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e0dbee55-9a27-4d31-b784-8779443b8f7d_400x400.jpeg&quot;,&quot;bio&quot;:&quot;Neuro AI, vision, Python, open science. NeuroAI lead at Amaranth. Previously engineer @ Google, Meta, Mila. &quot;,&quot;profile_set_up_at&quot;:&quot;2023-03-03T14:47:16.973Z&quot;,&quot;reader_installed_at&quot;:&quot;2023-03-03T14:46:36.453Z&quot;,&quot;publicationUsers&quot;:[{&quot;id&quot;:1534657,&quot;user_id&quot;:17921567,&quot;publication_id&quot;:1564943,&quot;role&quot;:&quot;admin&quot;,&quot;public&quot;:true,&quot;is_primary&quot;:true,&quot;publication&quot;:{&quot;id&quot;:1564943,&quot;name&quot;:&quot;The NeuroAI archive&quot;,&quot;subdomain&quot;:&quot;naix&quot;,&quot;custom_domain&quot;:&quot;www.neuroai.science&quot;,&quot;custom_domain_optional&quot;:false,&quot;hero_text&quot;:&quot;Research on Neuro-Inspired AI and AI-accelerated Neuro. &quot;,&quot;logo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1a3be57f-f6d5-4684-98b8-859ef181e86e_798x798.png&quot;,&quot;author_id&quot;:17921567,&quot;primary_user_id&quot;:17921567,&quot;theme_var_background_pop&quot;:&quot;#67BDFC&quot;,&quot;created_at&quot;:&quot;2023-04-08T16:21:53.936Z&quot;,&quot;email_from_name&quot;:&quot;Patrick Mineault from NeuroAI archive&quot;,&quot;copyright&quot;:&quot;Patrick Mineault&quot;,&quot;founding_plan_name&quot;:null,&quot;community_enabled&quot;:true,&quot;invite_only&quot;:false,&quot;payments_state&quot;:&quot;disabled&quot;,&quot;language&quot;:null,&quot;explicit&quot;:false,&quot;homepage_type&quot;:&quot;newspaper&quot;,&quot;is_personal_mode&quot;:false,&quot;logo_url_wide&quot;:null}}],&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null,&quot;status&quot;:{&quot;bestsellerTier&quot;:null,&quot;subscriberTier&quot;:1,&quot;leaderboard&quot;:null,&quot;vip&quot;:false,&quot;badge&quot;:{&quot;type&quot;:&quot;subscriber&quot;,&quot;tier&quot;:1,&quot;accent_colors&quot;:null},&quot;paidPublicationIds&quot;:[2133349,3477619,35345],&quot;subscriber&quot;:null}}],&quot;utm_campaign&quot;:null,&quot;belowTheFold&quot;:false,&quot;type&quot;:&quot;newsletter&quot;,&quot;language&quot;:&quot;en&quot;,&quot;source&quot;:null}" data-component-name="EmbeddedPostToDOM"><a class="embedded-post" native="true" href="https://www.neuroai.science/p/cell-types-wiring-up-innate-circuits?utm_source=substack&amp;utm_campaign=post_embed&amp;utm_medium=web"><div class="embedded-post-header"><img class="embedded-post-publication-logo" src="https://substackcdn.com/image/fetch/$s_!WvCx!,w_56,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a3be57f-f6d5-4684-98b8-859ef181e86e_798x798.png"><span class="embedded-post-publication-name">The NeuroAI archive</span></div><div class="embedded-post-title-wrapper"><div class="embedded-post-title">Cell types: from genes to circuits</div></div><div class="embedded-post-body">In the last post, I presented a counterintuitive idea: areas of the brain that contain many cell types encode innate behaviors and rewards. Myriad neurons with strange morphology that are connected idiosyncratically are an indicator of evolutionary pressure: behaviors and rewards that have survival value become genetically hard-coded. What I didn&#8217;t get &#8230;</div><div class="embedded-post-cta-wrapper"><span class="embedded-post-cta">Read more</span></div><div class="embedded-post-meta">5 months ago &#183; 12 likes &#183; Patrick Mineault</div></a></div><p>&#8230;And now here&#8217;s two new actual blog posts since my last newsletter:</p><div><hr></div><h1>Blog post 1/2: &#8220;You can&#8217;t imitation-learn how to continual-learn&#8221;</h1><p>Here&#8217;s the link:</p><p><a href="https://www.lesswrong.com/posts/9rCTjbJpZB4KzqhiQ/you-can-t-imitation-learn-how-to-continual-learn">https://www.lesswrong.com/posts/9rCTjbJpZB4KzqhiQ/you-can-t-imitation-learn-how-to-continual-learn</a></p><p>And here&#8217;s the start:</p><blockquote><p>In this post, I&#8217;m trying to put forward a narrow, pedagogical point, one that comes up mainly when I&#8217;m arguing in favor of LLMs having limitations that human learning does not.</p><p>See the bottom of the post for a list of subtexts that you should NOT<em> </em>read into this post, including &#8220;&#8230;therefore LLMs are dumb&#8221;, or &#8220;&#8230;therefore LLMs can&#8217;t possibly scale to superintelligence&#8221;.</p><h3><strong>Some intuitions on how to think about &#8220;real&#8221; continual learning</strong></h3><p>Consider an algorithm for training a Reinforcement Learning (RL) agent, like the <a href="https://arxiv.org/abs/1312.5602">Atari-playing Deep Q network (2013)</a> or <a href="https://en.wikipedia.org/wiki/AlphaZero">AlphaZero (2017)</a>, or think of within-lifetime learning in the human brain, which (<a href="https://www.lesswrong.com/posts/As7bjEAbNpidKx6LR/valence-series-1-introduction#1_2_Model_based_reinforcement_learning__RL_">I claim</a>) is in the general class of &#8220;model-based reinforcement learning&#8221;, broadly construed. &#8230;</p><p>When we think of &#8220;continual learning&#8221;, I suggest that those are good central examples to keep in mind. Here are some aspects to note:</p><p><em><strong>Knowledge vs information:</strong></em> These systems allow for continual acquisition of <em>knowledge</em>, not just <em>information</em>&#8212;the &#8220;continual learning&#8221; can install wholly new ways of conceptualizing and navigating the world, not just keeping track of what&#8217;s going on.</p><p><em><strong>Huge capacity for open-ended learning:</strong></em> These examples all have huge capacity for continual learning, indeed enough that they can start from random initialization and &#8220;continually learn&#8221; all the way to expert-level competence. Likewise, new continual learning can build on previous continual learning, in an ever-growing tower.</p><p><em><strong>Ability to figure things out that aren&#8217;t already on display in the environment:</strong></em> For example, an Atari-playing RL agent will get better and better at playing an Atari game, even without having any expert examples to copy. Likewise, billions of humans over thousands of years invented language, math, science, and a whole $100T global economy from scratch, all by ourselves, without angels dropping new training data from the heavens.</p></blockquote><p>Click the <a href="https://www.lesswrong.com/posts/9rCTjbJpZB4KzqhiQ/you-can-t-imitation-learn-how-to-continual-learn">link</a> to read the rest!</p><div><hr></div><h1>Blog post 2/2: &#8220;&#8216;Act-based approval-directed agents&#8217;, for IDA skeptics&#8221;</h1><p>Here&#8217;s the link:</p><p><a href="https://www.alignmentforum.org/posts/RKtTi82t8X8TQy5FX/act-based-approval-directed-agents-for-ida-skeptics">https://www.alignmentforum.org/posts/RKtTi82t8X8TQy5FX/act-based-approval-directed-agents-for-ida-skeptics</a></p><p>And here&#8217;s the start:</p><blockquote><h1><strong>Summary / tl;dr</strong></h1><p>In the 2010s, Paul Christiano built an extensive body of work on AI alignment&#8212;see the <a href="https://www.alignmentforum.org/s/EmDuGeRw749sD3GKd">&#8220;Iterated Amplification&#8221; series</a> for a curated overview as of 2018.</p><p>One foundation of this program was an intuition that it should be possible to build <em><strong>&#8220;act-based approval-directed agents&#8221;</strong></em> (&#8220;approval-directed agents&#8221; for short). These AGIs, for example, would not lie to their human supervisors, because their human supervisors wouldn&#8217;t want them to lie, and these AGIs would only do things that their human supervisors would want them to do. (It sounds much simpler than it is!)</p><p>Another foundation of this program was a set of algorithmic approaches, <strong>Iterated Distillation and Amplification (IDA)</strong>, that supposedly offers a path to actually building these approval-directed AI agents.</p><p>I am (and have always been) a skeptic of IDA: I just don&#8217;t think any of those algorithms would work very well.</p><p>But I still think there might be something to the &#8220;approval-directed agents&#8221; intuition. And we should be careful not to throw out the baby with the bathwater.</p><p>So my goal in this post is to rescue the &#8220;approval-directed agents&#8221; idea from its IDA baggage. <strong>Here&#8217;s the roadmap:</strong></p><p><strong>In Section 1</strong>, I offer a high-level picture of what we&#8217;re hoping to get out of &#8220;approval-directed agents&#8221;, following a <a href="https://www.alignmentforum.org/posts/wujPGixayiZSMYfm6/stable-pointers-to-value-ii-environmental-goals-1">discussion by Abram Demski (2018)</a>.</p><p><strong>In Section 2</strong>, I walk through an example of how this vision can actually manifest in the context of <a href="https://www.alignmentforum.org/s/HzcM2dkCq7fwXBej8">brain-like AGI</a>, a different AI paradigm which (unlike IDA) can <em>definitely</em> scale to superintelligence. I offer an everyday example of having role-models / idols who celebrate honesty, and correspondingly taking pride in one&#8217;s self-image as an honest person. In terms of brain algorithms, I relate this phenomenon to (what I call) <a href="https://www.alignmentforum.org/posts/fPxgFHfs5yHzYqgG7/social-drives-2-approval-reward-from-norm-enforcement-to">&#8220;Approval Reward&#8221;</a>, a hypothesized component of the human brain&#8217;s innate reinforcement learning reward function.</p></blockquote><p>Click the <a href="https://www.alignmentforum.org/posts/RKtTi82t8X8TQy5FX/act-based-approval-directed-agents-for-ida-skeptics">link</a> to read the rest!</p>]]></content:encoded></item><item><title><![CDATA[Blog post: “Why we should expect ruthless sociopath ASI”]]></title><description><![CDATA[Here&#8217;s the link:]]></description><link>https://stevebyrnes1.substack.com/p/blog-post-why-we-should-expect-ruthless</link><guid isPermaLink="false">https://stevebyrnes1.substack.com/p/blog-post-why-we-should-expect-ruthless</guid><dc:creator><![CDATA[Steve Byrnes]]></dc:creator><pubDate>Thu, 19 Feb 2026 21:21:56 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/fba30f43-61ad-477f-b164-86f3927eb1da_4400x1895.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Here&#8217;s the link:</p><p><a href="https://www.alignmentforum.org/posts/ZJZZEuPFKeEdkrRyf/why-we-should-expect-ruthless-sociopath-asi">https://www.alignmentforum.org/posts/ZJZZEuPFKeEdkrRyf/why-we-should-expect-ruthless-sociopath-asi</a></p><p>And here&#8217;s how it begins:</p><blockquote><h2>The conversation begins</h2><p><strong>(Fictional) Optimist:</strong> So you expect future artificial superintelligence (ASI) &#8220;by default&#8221;, i.e. in the absence of yet-to-be-invented techniques, to be a ruthless sociopath, happy to lie, cheat, and steal, whenever doing so is selfishly beneficial, and with callous indifference to whether anyone (including its own programmers and users) lives or dies?</p><p><strong>Me:</strong> Yup! (Alas.)</p><p><strong>Optimist:</strong> &#8230;Despite all the evidence right in front of our eyes from humans and LLMs.</p><p><strong>Me:</strong> Yup!</p><p><strong>Optimist:</strong> OK, well, I&#8217;m here to tell you: that is a very specific and strange thing to expect, especially in the absence of any concrete evidence whatsoever. There&#8217;s no reason to expect it. If you think that ruthless sociopathy is the &#8220;true core nature of intelligence&#8221; or whatever, then <a href="https://www.alignmentforum.org/posts/d4HNRdw6z7Xqbnu5E/6-reasons-why-alignment-is-hard-discourse-seems-alien-to?commentId=v2WfXTLaoJuSw9QQG">you should really look at yourself in a mirror and ask yourself where your life went horribly wrong</a>.</p><p><strong>Me:</strong> Hmm, I think the &#8220;true core nature of intelligence&#8221; is above my pay grade. We should probably just talk about the issue at hand, namely future AI algorithms and their properties.</p><p>&#8230;But I actually agree with you that ruthless sociopathy is a very specific and strange thing for me to expect.</p><p><strong>Optimist:</strong> Wait, you&#8212;what??</p></blockquote><p>Click the <a href="https://www.alignmentforum.org/posts/ZJZZEuPFKeEdkrRyf/why-we-should-expect-ruthless-sociopath-asi">link</a> to read the rest!</p>]]></content:encoded></item><item><title><![CDATA[Blog post: “The brain is a machine that runs an algorithm” (plus three quick updates)]]></title><description><![CDATA[Start with three quick updates &#8230;]]></description><link>https://stevebyrnes1.substack.com/p/blog-post-the-brain-is-a-machine</link><guid isPermaLink="false">https://stevebyrnes1.substack.com/p/blog-post-the-brain-is-a-machine</guid><dc:creator><![CDATA[Steve Byrnes]]></dc:creator><pubDate>Wed, 18 Feb 2026 03:55:29 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/9ee2479f-afd5-463e-975f-1a7e2fbcc8af_1150x639.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h1>Start with three quick updates:</h1><ul><li><p>You may remember that two weeks ago, I shared a post <a href="https://www.lesswrong.com/posts/sGNFtWbXiLJg2hLzK/the-nature-of-llm-algorithmic-progress-v2">&#8220;The nature of LLM algorithmic progress&#8221;</a>. Well, four days later, I rewrote it as a heavily-revised &#8220;version 2&#8221;. Thanks everyone for feedback and pushback! If you read &#8220;version 1&#8221;, consider revisiting it&#8212;there&#8217;s a changelog at the bottom. <a href="https://www.lesswrong.com/posts/sGNFtWbXiLJg2hLzK/the-nature-of-llm-algorithmic-progress-v2">Same link</a>.</p></li><li><p>There&#8217;s a fascinating new preprint, <a href="https://www.preprints.org/manuscript/202602.0767">Cellular Scaling Laws in the Mammalian Brain</a> by Fei Chen &amp; Evan Macosko, which dovetails with many of the neuroscience ideas that I&#8217;m always talking about, but coming from the angle of transcriptomics and psychology. I&#8217;m listed in the acknowledgements&#8212;it was a real treat to chat with Fei and Evan.</p></li><li><p>Adam Marblestone has been generously discussing and contextualizing my perspective on neuroscience and AI, which he understands quite well:</p><div class="embedded-post-wrap" data-attrs="{&quot;id&quot;:182960540,&quot;url&quot;:&quot;https://www.dwarkesh.com/p/adam-marblestone&quot;,&quot;publication_id&quot;:69345,&quot;embedding_publication_id&quot;:null,&quot;publication_name&quot;:&quot;Dwarkesh Podcast&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!QEPJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F90fa9666-5b8b-4685-a8fb-4b64cb7e0333_1080x1080.png&quot;,&quot;title&quot;:&quot;Adam Marblestone &#8212; AI is missing something fundamental about the brain&quot;,&quot;truncated_body_text&quot;:null,&quot;date&quot;:&quot;2025-12-30T17:07:17.441Z&quot;,&quot;like_count&quot;:86,&quot;comment_count&quot;:7,&quot;bylines&quot;:[{&quot;id&quot;:4281466,&quot;name&quot;:&quot;Dwarkesh Patel&quot;,&quot;handle&quot;:&quot;dwarkesh&quot;,&quot;previous_name&quot;:null,&quot;photo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!5eJb!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb715ffd1-f7d7-4755-af88-c48efe647f5b_400x400.jpeg&quot;,&quot;bio&quot;:&quot;Host of Dwarkesh Podcast&quot;,&quot;profile_set_up_at&quot;:&quot;2021-06-09T22:58:10.864Z&quot;,&quot;reader_installed_at&quot;:&quot;2022-04-03T20:37:19.142Z&quot;,&quot;publicationUsers&quot;:[{&quot;id&quot;:246192,&quot;user_id&quot;:4281466,&quot;publication_id&quot;:69345,&quot;role&quot;:&quot;admin&quot;,&quot;public&quot;:true,&quot;is_primary&quot;:true,&quot;publication&quot;:{&quot;id&quot;:69345,&quot;name&quot;:&quot;Dwarkesh Podcast&quot;,&quot;subdomain&quot;:&quot;dwarkesh&quot;,&quot;custom_domain&quot;:&quot;www.dwarkesh.com&quot;,&quot;custom_domain_optional&quot;:false,&quot;hero_text&quot;:&quot;Deeply researched interviews&quot;,&quot;logo_url&quot;:&quot;https://bucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com/public/images/90fa9666-5b8b-4685-a8fb-4b64cb7e0333_1080x1080.png&quot;,&quot;author_id&quot;:4281466,&quot;primary_user_id&quot;:4281466,&quot;theme_var_background_pop&quot;:&quot;#D10000&quot;,&quot;created_at&quot;:&quot;2020-07-18T16:36:25.723Z&quot;,&quot;email_from_name&quot;:&quot;Dwarkesh Patel&quot;,&quot;copyright&quot;:&quot;Dwarkesh Patel&quot;,&quot;founding_plan_name&quot;:&quot;Founding Member&quot;,&quot;community_enabled&quot;:true,&quot;invite_only&quot;:false,&quot;payments_state&quot;:&quot;enabled&quot;,&quot;language&quot;:null,&quot;explicit&quot;:false,&quot;homepage_type&quot;:null,&quot;is_personal_mode&quot;:false}}],&quot;twitter_screen_name&quot;:&quot;dwarkesh_sp&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:100,&quot;status&quot;:{&quot;bestsellerTier&quot;:100,&quot;subscriberTier&quot;:5,&quot;leaderboard&quot;:null,&quot;vip&quot;:false,&quot;badge&quot;:{&quot;type&quot;:&quot;bestseller&quot;,&quot;tier&quot;:100},&quot;paidPublicationIds&quot;:[89120,3409707,3087928,104058,1163860,22108,6819723,2118966],&quot;subscriber&quot;:null}}],&quot;utm_campaign&quot;:null,&quot;belowTheFold&quot;:false,&quot;type&quot;:&quot;podcast&quot;,&quot;language&quot;:&quot;en&quot;,&quot;source&quot;:null}" data-component-name="EmbeddedPostToDOM"><a class="embedded-post" native="true" href="https://www.dwarkesh.com/p/adam-marblestone?utm_source=substack&amp;utm_campaign=post_embed&amp;utm_medium=web"><div class="embedded-post-header"><img class="embedded-post-publication-logo" src="https://substackcdn.com/image/fetch/$s_!QEPJ!,w_56,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F90fa9666-5b8b-4685-a8fb-4b64cb7e0333_1080x1080.png"><span class="embedded-post-publication-name">Dwarkesh Podcast</span></div><div class="embedded-post-title-wrapper"><div class="embedded-post-title-icon"><svg width="19" height="19" viewBox="0 0 24 24" fill="none" xmlns="http://www.w3.org/2000/svg">
  <path d="M3 18V12C3 9.61305 3.94821 7.32387 5.63604 5.63604C7.32387 3.94821 9.61305 3 12 3C14.3869 3 16.6761 3.94821 18.364 5.63604C20.0518 7.32387 21 9.61305 21 12V18" stroke-linecap="round" stroke-linejoin="round"></path>
  <path d="M21 19C21 19.5304 20.7893 20.0391 20.4142 20.4142C20.0391 20.7893 19.5304 21 19 21H18C17.4696 21 16.9609 20.7893 16.5858 20.4142C16.2107 20.0391 16 19.5304 16 19V16C16 15.4696 16.2107 14.9609 16.5858 14.5858C16.9609 14.2107 17.4696 14 18 14H21V19ZM3 19C3 19.5304 3.21071 20.0391 3.58579 20.4142C3.96086 20.7893 4.46957 21 5 21H6C6.53043 21 7.03914 20.7893 7.41421 20.4142C7.78929 20.0391 8 19.5304 8 19V16C8 15.4696 7.78929 14.9609 7.41421 14.5858C7.03914 14.2107 6.53043 14 6 14H3V19Z" stroke-linecap="round" stroke-linejoin="round"></path>
</svg></div><div class="embedded-post-title">Adam Marblestone &#8212; AI is missing something fundamental about the brain</div></div><div class="embedded-post-cta-wrapper"><div class="embedded-post-cta-icon"><svg width="32" height="32" viewBox="0 0 24 24" xmlns="http://www.w3.org/2000/svg">
  <path classname="inner-triangle" d="M10 8L16 12L10 16V8Z" stroke-width="1.5" stroke-linecap="round" stroke-linejoin="round"></path>
</svg></div><span class="embedded-post-cta">Listen now</span></div><div class="embedded-post-meta">7 months ago &#183; 86 likes &#183; 7 comments &#183; Dwarkesh Patel</div></a></div></li></ul><div class="embedded-post-wrap" data-attrs="{&quot;id&quot;:187884908,&quot;url&quot;:&quot;https://asteriskmag.substack.com/p/the-sweet-lesson-of-neuroscience&quot;,&quot;publication_id&quot;:2291516,&quot;embedding_publication_id&quot;:null,&quot;publication_name&quot;:&quot;Asterisk Magazine &quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!0HDE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fa3bc20-4e1b-465d-a704-649883b2f406_3200x3200.jpeg&quot;,&quot;title&quot;:&quot;The Sweet Lesson of Neuroscience&quot;,&quot;truncated_body_text&quot;:&quot;In the early years of modern deep learning, the brain was a North Star. Ideas like hippocampal replay &#8212; the brain&#8217;s way of rehearsing past experience &#8212; offered templates for how an agent might learn from memories. Meanwhile, work on temporal-difference learning&quot;,&quot;date&quot;:&quot;2026-02-16T13:02:35.658Z&quot;,&quot;like_count&quot;:10,&quot;comment_count&quot;:0,&quot;bylines&quot;:[],&quot;utm_campaign&quot;:null,&quot;belowTheFold&quot;:false,&quot;type&quot;:&quot;newsletter&quot;,&quot;language&quot;:&quot;en&quot;,&quot;source&quot;:null}" data-component-name="EmbeddedPostToDOM"><a class="embedded-post" native="true" href="https://asteriskmag.substack.com/p/the-sweet-lesson-of-neuroscience?utm_source=substack&amp;utm_campaign=post_embed&amp;utm_medium=web"><div class="embedded-post-header"><img class="embedded-post-publication-logo" src="https://substackcdn.com/image/fetch/$s_!0HDE!,w_56,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fa3bc20-4e1b-465d-a704-649883b2f406_3200x3200.jpeg"><span class="embedded-post-publication-name">Asterisk Magazine </span></div><div class="embedded-post-title-wrapper"><div class="embedded-post-title">The Sweet Lesson of Neuroscience</div></div><div class="embedded-post-body">In the early years of modern deep learning, the brain was a North Star. Ideas like hippocampal replay &#8212; the brain&#8217;s way of rehearsing past experience &#8212; offered templates for how an agent might learn from memories. Meanwhile, work on temporal-difference learning&#8230;</div><div class="embedded-post-cta-wrapper"><span class="embedded-post-cta">Read more</span></div><div class="embedded-post-meta">5 months ago &#183; 10 likes</div></a></div><h1>New blog post: &#8220;The brain is a machine that runs an algorithm&#8221;</h1><p>This one&#8217;s just a quick little rant. Here&#8217;s the link:</p><p><a href="https://www.lesswrong.com/posts/eKGjwRSQD3BLxmBcu/the-brain-is-a-machine-that-runs-an-algorithm">https://www.lesswrong.com/posts/eKGjwRSQD3BLxmBcu/the-brain-is-a-machine-that-runs-an-algorithm</a></p><p>And here&#8217;s the start:</p><blockquote><p>Some people say &#8220;the brain is a computer&#8221;. Other people say &#8220;well, the brain is not really a computer, because, like, what&#8217;s the hardware versus the software?&#8221; I agree: &#8220;the brain is a computer&#8221; is kinda missing the mark. I prefer: &#8220;the brain is a machine that runs an algorithm&#8221;.</p><div><hr></div><p><strong>Here&#8217;s a mechanical adder</strong>:</p><div id="youtube2-GcDshWmhF4A" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;GcDshWmhF4A&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/GcDshWmhF4A?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p>What&#8217;s &#8220;hardware versus software&#8221; for a mechanical adder? The question is nonsense.</p><p>A mechanical adder is not &#8220;a computer&#8221;, analogous to a MacBook. Rather, it&#8217;s a machine that runs an algorithm. (Namely, the binary addition algorithm.)</p><p>And the brain is likewise a machine that runs a (much more complicated) algorithm.</p><div><hr></div><p><em><strong>&#8220;A machine??&#8221;</strong></em>, you say. Yeah, you heard me. A machine. An extraordinarily complex machine, but a machine all the same. If you could zoom in enough to really see it, it would just be obvious! You should pause here to marvel at some molecular simulations of cell biology in action: <a href="https://youtu.be/WFCvkkDSfIU?si=3aTUYYgBDIklTW-E&amp;t=220">DNA replication</a>, <a href="https://x.com/Andercot/status/2016915660979265891?s=20">more crazy DNA stuff</a>, <a href="http://youtu.be/y-uuk4Pr2i8">kinesin</a>, and so on. &#8230;</p></blockquote><p>Click the <a href="https://www.lesswrong.com/posts/eKGjwRSQD3BLxmBcu/the-brain-is-a-machine-that-runs-an-algorithm">link</a> to read the rest!</p>]]></content:encoded></item><item><title><![CDATA[2 blog posts: LLM algorithmic progress & interpretability-in-the-loop ML training]]></title><description><![CDATA[Here are the links:]]></description><link>https://stevebyrnes1.substack.com/p/2-blog-posts-llm-algorithmic-progress</link><guid isPermaLink="false">https://stevebyrnes1.substack.com/p/2-blog-posts-llm-algorithmic-progress</guid><dc:creator><![CDATA[Steve Byrnes]]></dc:creator><pubDate>Fri, 06 Feb 2026 20:16:34 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/9c7da895-456d-4e95-8c25-6c08d4134b25_1580x735.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Here are the links:</p><ul><li><p><strong>&#8220;The nature of LLM algorithmic progress&#8221;</strong>: <a href="https://www.lesswrong.com/posts/sGNFtWbXiLJg2hLzK/the-nature-of-llm-algorithmic-progress">https://www.lesswrong.com/posts/sGNFtWbXiLJg2hLzK/the-nature-of-llm-algorithmic-progress</a></p></li><li><p><strong>&#8220;In (highly contingent!) defense of interpretability-in-the-loop ML training&#8221;:</strong> <a href="https://www.alignmentforum.org/posts/ArXAyzHkidxwoeZsL/in-highly-contingent-defense-of-interpretability-in-the-loop">https://www.alignmentforum.org/posts/ArXAyzHkidxwoeZsL/in-highly-contingent-defense-of-interpretability-in-the-loop</a></p></li></ul><div><hr></div><p>The first one has a bit of Cunningham&#8217;s Law energy: it&#8217;s some spicy hot takes far outside my area of expertise. Here&#8217;s how it starts:</p><blockquote><h1><strong><a href="https://www.lesswrong.com/posts/sGNFtWbXiLJg2hLzK/the-nature-of-llm-algorithmic-progress">The nature of LLM algorithmic progress</a></strong></h1><p>There&#8217;s a lot of talk about &#8220;algorithmic progress&#8221; in LLMs, especially in the context of exponentially-improving algorithmic efficiency. For example:</p><ul><li><p><a href="https://epoch.ai/blog/algorithmic-progress-in-language-models">Epoch AI</a>: &#8220;[training] compute required to reach a set performance threshold has halved approximately every 8 months&#8221;.</p></li><li><p><a href="https://www.darioamodei.com/post/on-deepseek-and-export-controls">Dario Amodei 2025</a>: &#8220;I&#8217;d guess the number today is maybe ~4x/year&#8221;.</p></li><li><p><a href="https://arxiv.org/abs/2511.23455">Gundlach et al. 2025a &#8220;Price of Progress&#8221;</a>: &#8220;Isolating out open models to control for competition effects and dividing by hardware price declines, we estimate that algorithmic efficiency progress is around 3&#215; per year&#8221;.</p></li></ul><p>It&#8217;s nice to see three independent sources reach almost exactly the same conclusion&#8212;halving times of 8 months, 6 months, and 7&#189; months respectively. Surely a sign that the conclusion is solid!</p><p>&#8230;Haha, just kidding! I&#8217;ll argue that these three bullet points are hiding three totally different stories. The first two bullets are about <em>training</em> efficiency, and I&#8217;ll argue that both are deeply misleading (each for a different reason!). The third is about <em>inference</em> efficiency, which I think is right, and mostly explained by distillation of ever-better frontier models into their &#8220;mini&#8221; cousins.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ZX_-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9f50d8d-371f-4ce8-8f2e-422b66ed23de_724x123.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ZX_-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9f50d8d-371f-4ce8-8f2e-422b66ed23de_724x123.png 424w, https://substackcdn.com/image/fetch/$s_!ZX_-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9f50d8d-371f-4ce8-8f2e-422b66ed23de_724x123.png 848w, https://substackcdn.com/image/fetch/$s_!ZX_-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9f50d8d-371f-4ce8-8f2e-422b66ed23de_724x123.png 1272w, https://substackcdn.com/image/fetch/$s_!ZX_-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9f50d8d-371f-4ce8-8f2e-422b66ed23de_724x123.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ZX_-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9f50d8d-371f-4ce8-8f2e-422b66ed23de_724x123.png" width="456" height="77.4696132596685" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f9f50d8d-371f-4ce8-8f2e-422b66ed23de_724x123.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:123,&quot;width&quot;:724,&quot;resizeWidth&quot;:456,&quot;bytes&quot;:20059,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://stevebyrnes1.substack.com/i/187128586?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9f50d8d-371f-4ce8-8f2e-422b66ed23de_724x123.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!ZX_-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9f50d8d-371f-4ce8-8f2e-422b66ed23de_724x123.png 424w, https://substackcdn.com/image/fetch/$s_!ZX_-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9f50d8d-371f-4ce8-8f2e-422b66ed23de_724x123.png 848w, https://substackcdn.com/image/fetch/$s_!ZX_-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9f50d8d-371f-4ce8-8f2e-422b66ed23de_724x123.png 1272w, https://substackcdn.com/image/fetch/$s_!ZX_-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9f50d8d-371f-4ce8-8f2e-422b66ed23de_724x123.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a></figure></div><h1>Tl;dr / outline</h1><ul><li><p>&#167;1 is my attempted <strong>big-picture take on what &#8220;algorithmic progress&#8221; has looked like in LLMs</strong>. I split it into four categories:</p><ul><li><p><strong>&#167;1.1 is stereotypical &#8220;algorithmic efficiency improvements&#8221;</strong> related to the core learning algorithm itself (architectures, tokenizers, optimizers, etc.). I&#8217;ll argue that the idea of using a Transformer with optimized hyperparameters was an important idea, but apart from that, the field has produced maybe 3-5&#215; worth of stereotypical &#8220;algorithmic efficiency improvements&#8221; in the <em>entire period</em> from 2018 to today (&#8776;20%/year). That&#8217;s something, but it&#8217;s very much less than the exponentials suggested at the top.</p></li><li><p><strong>&#167;1.2 is &#8220;optimizations&#8221;</strong>. I figure there might be up to 20&#215; improvement by optimizing the CUDA code, parallelization strategies, precision, and so on, especially if we take the baseline to be a sufficiently early slapdash implementation. But the thing about &#8220;optimizations&#8221; is that they have a ceiling&#8212;it&#8217;s not an exponential that can keep growing and growing.</p></li><li><p><strong>&#167;1.3 is &#8220;data-related improvements&#8221;</strong>, including proprietary human expert data, and various types of model distillation, both of which have important effects.</p></li><li><p><strong>&#167;1.4 is &#8220;algorithmic changes that are not really quantifiable as &#8216;efficiency&#8217;&#8221;</strong>, including RLHF, RLVR, multimodality, and so on. No question that these are important, and we shouldn&#8217;t forget that they exist, but they&#8217;re not directly related to the exponential-improvement claims at the top.</p></li></ul></li><li><p><strong>&#167;2 is why I don&#8217;t believe either Epoch AI or Dario</strong>, in their claims of exponential training-efficiency improvements (see top).</p></li><li><p><strong>&#167;3 is a quick sanity-check</strong>, studying <a href="https://github.com/karpathy/nanochat">&#8220;nanochat&#8221;</a>, which matches GPT-2 performance but costs 600&#215; less to train.</p></li><li><p><strong>&#167;4 is an optional bonus section on why I care about this topic</strong> in the first place. (Not what you expect! Unlike everyone else reading this, I don&#8217;t particularly care about forecasting future LLM progress.)</p></li></ul></blockquote><p>Click the <a href="https://www.lesswrong.com/posts/sGNFtWbXiLJg2hLzK/the-nature-of-llm-algorithmic-progress">link</a> to read the rest!</p><div><hr></div><p>The second one is much more in my comfort zone. Here&#8217;s a teaser:</p><blockquote><h1><strong><a href="https://www.alignmentforum.org/posts/ArXAyzHkidxwoeZsL/in-highly-contingent-defense-of-interpretability-in-the-loop">In (highly contingent!) defense of interpretability-in-the-loop ML training</a></strong></h1><p>Let&#8217;s call<em> &#8220;interpretability-in-the-loop training&#8221;</em> the idea of running a learning algorithm that involves an inscrutable trained model, and there&#8217;s some kind of interpretability system feeding into the loss function / reward function.</p><p>Interpretability-in-the-loop training has a very bad rap (and rightly so). Here&#8217;s <a href="https://www.alignmentforum.org/posts/uMQ3cqWDPHhjtiesc/agi-ruin-a-list-of-lethalities">Yudkowsky 2022</a>:</p><blockquote><p>When you explicitly optimize against a detector of unaligned thoughts, you&#8217;re partially optimizing for more aligned thoughts, and partially optimizing for unaligned thoughts that are harder to detect. Optimizing against an interpreted thought optimizes against interpretability.</p></blockquote><p>Or <a href="https://www.alignmentforum.org/posts/mpmsK8KKysgSKDm2T/the-most-forbidden-technique">Zvi 2025</a>:</p><blockquote><p><a href="https://thezvi.substack.com/i/145384938/the-art-of-the-jailbreak">The Most Forbidden Technique</a> is training an AI using interpretability techniques.</p><p>An AI produces a final output [X] via some method [M]. You can analyze [M] using technique [T], to learn what the AI is up to. You could train on that. Never do that.</p><p>You train on [X]. Only [X]. Never [M], never [T].</p><p>Why? Because [T] is how you figure out when the model is misbehaving.</p><p>If you train on [T], you are training the AI to obfuscate its thinking, and defeat [T]. You will rapidly lose your ability to know what is going on, in exactly the ways you most need to know what is going on.</p><p>Those bits of optimization pressure from [T] are precious. Use them wisely.</p></blockquote><p>This is a simple argument, and I think it&#8217;s 100% right.</p><p><em>But&#8230;</em></p><p>Consider compassion in the human brain. <a href="https://www.alignmentforum.org/posts/KuBiv9cCbZ6ALjHFw/social-drives-1-sympathy-reward-from-compassion-to">I claim</a> that we have an innate reward function that triggers not just when I see that my friend is happy or suffering, but also when I <em>believe</em> that my friend is happy or suffering, even if the friend is far away. So the human brain reward can evidently get triggered by specific activations inside my inscrutable learned world-model.</p><p>Thus, I claim that the human brain incorporates a form of interpretability-in-the-loop RL training.</p><p>Inspired by that example, I have long been an advocate for studying whether and how one might use interpretability-in-the-loop training for aligned AGI. See for example <a href="https://www.alignmentforum.org/posts/xw8P8H4TRaTQHJnoP/reward-function-design-a-starter-pack">Reward Function Design: a starter pack</a> sections 1, 4, and 5.</p><p>My goal in the post is to briefly summarize how I reconcile the arguments at the top with my endorsement of this kind of research program.</p></blockquote><p>Click the <a href="https://www.alignmentforum.org/posts/ArXAyzHkidxwoeZsL/in-highly-contingent-defense-of-interpretability-in-the-loop">link</a> to read the rest!</p>]]></content:encoded></item><item><title><![CDATA[Blog post: “Are there lessons from high-reliability engineering for AGI safety?”]]></title><description><![CDATA[Here&#8217;s the link:]]></description><link>https://stevebyrnes1.substack.com/p/blog-post-are-there-lessons-from</link><guid isPermaLink="false">https://stevebyrnes1.substack.com/p/blog-post-are-there-lessons-from</guid><dc:creator><![CDATA[Steve Byrnes]]></dc:creator><pubDate>Mon, 02 Feb 2026 20:50:44 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/2cd723a7-48de-4598-b498-ea52badb3fa9_1536x931.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Here&#8217;s the link:</p><p><a href="https://www.alignmentforum.org/posts/hiiguxJ2EtfSzAevj/are-there-lessons-from-high-reliability-engineering-for-agi">https://www.alignmentforum.org/posts/hiiguxJ2EtfSzAevj/are-there-lessons-from-high-reliability-engineering-for-agi</a></p><p>And here&#8217;s the intro:</p><blockquote><p>This post is partly a belated response to Joshua Achiam, currently OpenAI&#8217;s Head of Mission Alignment:</p><blockquote><p>If we adopt safety best practices that are common in other professional engineering fields, we&#8217;ll get there &#8230; I consider myself one of the x-risk people, though I agree that most of them would reject my view on how to prevent it. I think the wholesale rejection of safety best practices from other fields is one of the dumbest mistakes that a group of otherwise very smart people has ever made. &#8212;<a href="https://x.com/jachiam0/status/1370980495920533506">Joshua Achiam on Twitter, 2021</a></p><p>&#8220;We just have to sit down and actually write a damn specification, even if it&#8217;s like pulling teeth. It&#8217;s the most important thing we could possibly do,&#8221; said almost no one in the field of AGI alignment, sadly. &#8230; I&#8217;m picturing hundreds of pages of documentation describing, for various application areas, specific behaviors and acceptable error tolerances &#8230; &#8212;<a href="https://x.com/steve47285/status/1609232759100248066?s=20">Joshua Achiam on Twitter (partly talking to me), 2022</a></p></blockquote><p>As a proud member of the group of &#8220;otherwise very smart people&#8221; making &#8220;one of the dumbest mistakes&#8221;, I will explain why I don&#8217;t think it&#8217;s a mistake. (Indeed, since 2022, some &#8220;x-risk people&#8221; <em>have</em> started working towards these kinds of specs, and I think <em>they&#8217;re</em> the ones making a mistake and wasting their time!)</p><p>At the same time, I&#8217;ll describe what I see as the kernel of truth in Joshua&#8217;s perspective, and why it should be seen as an indictment not of &#8220;x-risk people&#8221; but rather of OpenAI itself, along with all the other groups racing to develop AGI.</p></blockquote><p>Click the <a href="https://www.alignmentforum.org/posts/hiiguxJ2EtfSzAevj/are-there-lessons-from-high-reliability-engineering-for-agi">link</a> to read the rest!</p>]]></content:encoded></item><item><title><![CDATA[New version of “Intro to Brain-Like-AGI Safety”]]></title><description><![CDATA[A new version of &#8220;Intro to Brain-Like-AGI Safety&#8221; is out!]]></description><link>https://stevebyrnes1.substack.com/p/new-version-of-intro-to-brain-like</link><guid isPermaLink="false">https://stevebyrnes1.substack.com/p/new-version-of-intro-to-brain-like</guid><dc:creator><![CDATA[Steve Byrnes]]></dc:creator><pubDate>Fri, 23 Jan 2026 18:15:22 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/d7093658-5002-4f90-b791-93d41ed6c02c_519x503.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>A new version of <a href="https://www.lesswrong.com/s/HzcM2dkCq7fwXBej8">&#8220;Intro to Brain-Like-AGI Safety&#8221;</a> is out! The official announcement blog post is <a href="https://www.lesswrong.com/posts/rreDwHXgnhEDKxkro/new-version-of-intro-to-brain-like-agi-safety">here</a>.</p><p>The thing itself is at the same links as before:</p><ul><li><p><strong>As a series of 15 blog posts on LessWrong / Alignment Forum:</strong> <a href="https://www.alignmentforum.org/s/HzcM2dkCq7fwXBej8">https://www.alignmentforum.org/s/HzcM2dkCq7fwXBej8</a></p></li><li><p><strong>As a 225-page PDF (now up to version 3):</strong> <a href="https://osf.io/preprints/osf/fe36n">https://osf.io/preprints/osf/fe36n</a></p></li><li><p><strong>Summary video &amp; transcript:</strong> <a href="https://www.lesswrong.com/posts/YKmyay3bWF2ofAGNo/video-and-transcript-challenges-for-safe-and-beneficial">This link</a></p></li></ul><p>And the abstract is also the same as before:</p><blockquote><p><strong>Abstract:</strong> Suppose we someday build an Artificial General Intelligence algorithm using similar principles of learning and cognition as the human brain. How would we use such an algorithm safely?</p><p>I will argue that this is an open technical problem, and my goal in this post series is to bring readers with no prior knowledge all the way up to the front-line of unsolved problems as I see them.</p><p><a href="https://www.lesswrong.com/posts/4basF9w9jaPZpoC8R/intro-to-brain-like-agi-safety-1-what-s-the-problem-and-why">Post #1</a> contains definitions, background, and motivation. Then Posts <a href="https://www.lesswrong.com/posts/wBHSYwqssBGCnwvHg/intro-to-brain-like-agi-safety-2-learning-from-scratch-in">#2</a>&#8211;<a href="https://www.lesswrong.com/posts/zXibERtEWpKuG5XAC/intro-to-brain-like-agi-safety-7-from-hardcoded-drives-to">#7</a> are the neuroscience, arguing for a picture of he brain that combines large-scale learning algorithms (e.g. in the cortex) and specific evolved reflexes (e.g. in the hypothalamus and brainstem). Posts <a href="https://www.lesswrong.com/posts/fDPsYdDtkzhBp9A8D/intro-to-brain-like-agi-safety-8-takeaways-from-neuro-1-2-on">#8</a>&#8211;<a href="https://www.lesswrong.com/posts/tj8AC3vhTnBywdZoA/intro-to-brain-like-agi-safety-15-conclusion-open-problems-1">#15</a> apply those neuroscience ideas directly to AGI safety, ending with a list of open questions and advice for getting involved in the field.</p><p>A major theme will be that the human brain runs a yet-to-be-invented variation on Model-Based Reinforcement Learning. The reward function of this system (a.k.a. &#8220;innate drives&#8221; or &#8220;primary rewards&#8221;) says that pain is bad, and eating-when-hungry is good, etc. I will argue that this reward function is centered around the hypothalamus and brainstem, and that <em>all</em> human desires&#8212;even &#8220;higher&#8221; desires for things like compassion and justice&#8212;come directly or indirectly from that reward function. If future programmers build brain-like AGI, they will likewise have a reward function slot in their source code, in which they can put whatever they want. If they put the wrong code in the reward function slot, the resulting AGI will wind up callously indifferent to human welfare. How might they avoid that? What code <em>should</em> they put in&#8212;along with training environment and other design choices&#8212;such that the AGI won&#8217;t feel callous indifference to whether its programmers, and other people, live or die? No one knows&#8212;it&#8217;s an open problem, but I will review some ideas and research directions.</p></blockquote><h1>So what&#8217;s new?</h1><p>I went through the whole thing and made a bunch of edits and additions, based on what I&#8217;ve (hopefully) learned since my <a href="https://www.lesswrong.com/posts/btHmC88KCZdzimBCM/steve2152-s-shortform?commentId=DD28htp2tWZg7w8eT">last big round of edits 18 months ago</a>.</p><p><strong>See the <a href="https://www.lesswrong.com/posts/rreDwHXgnhEDKxkro/new-version-of-intro-to-brain-like-agi-safety">official announcement blog post</a> for fifteen </strong><em><strong>highlights from the changelog</strong></em>&#8212;including newly-added spicy hot takes on brain plasticity, AGI timelines, the definitions of &#8220;AGI&#8221; and &#8220;alignment&#8221;, collective intelligence, interpretability, instrumental convergence, and more.</p>]]></content:encoded></item><item><title><![CDATA[Blog post: “My AGI safety research—2025 review, ’26 plans”]]></title><description><![CDATA[Here&#8217;s the link:]]></description><link>https://stevebyrnes1.substack.com/p/blog-post-my-agi-safety-research2025</link><guid isPermaLink="false">https://stevebyrnes1.substack.com/p/blog-post-my-agi-safety-research2025</guid><dc:creator><![CDATA[Steve Byrnes]]></dc:creator><pubDate>Thu, 11 Dec 2025 17:38:22 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/65566408-7ee5-4544-81ea-db0f95dab5f3_1592x1194.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Here&#8217;s the link:</p><p><a href="https://www.alignmentforum.org/posts/CF4Z9mQSfvi99A3BR/my-agi-safety-research-2025-review-26-plans">https://www.alignmentforum.org/posts/CF4Z9mQSfvi99A3BR/my-agi-safety-research-2025-review-26-plans</a></p><p>And here&#8217;s the outline:</p><blockquote><ol><li><p>Background &amp; threat model</p></li><li><p>The theme of 2025: trying to solve the technical alignment problem</p></li><li><p>Two sketchy plans for technical AGI alignment</p></li><li><p>On to what I&#8217;ve actually been doing all year!</p><ol><li><p>Thrust A: Fitting technical alignment into the bigger strategic picture</p></li><li><p>Thrust B: Better understanding how RL reward functions can be compatible with non-ruthless-optimizers</p></li><li><p>Thrust C: Continuing to develop my thinking on the neuroscience of human social instincts</p></li><li><p>Thrust D: Alignment implications of continuous learning and concept extrapolation</p></li><li><p>Thrust E: Neuroscience odds and ends</p></li><li><p>Thrust F: Economics of superintelligence</p></li><li><p>Thrust G: AGI safety miscellany</p></li><li><p>Thrust H: Outreach</p></li></ol></li><li><p>Other stuff</p><ol><li><p>Non-work-related blog posts</p></li><li><p>Personal productivity and workflow</p></li></ol></li><li><p>Plan for 2026</p></li><li><p>Acknowledgements</p></li></ol></blockquote><p>Click the <a href="https://www.alignmentforum.org/posts/CF4Z9mQSfvi99A3BR/my-agi-safety-research-2025-review-26-plans">link</a> to read the whole thing!</p>]]></content:encoded></item><item><title><![CDATA[2 blog posts: “We need a field of Reward Function Design” & “Reward Function Design: a starter pack”]]></title><description><![CDATA[Here are the links:]]></description><link>https://stevebyrnes1.substack.com/p/2-blog-posts-we-need-a-field-of-reward</link><guid isPermaLink="false">https://stevebyrnes1.substack.com/p/2-blog-posts-we-need-a-field-of-reward</guid><dc:creator><![CDATA[Steve Byrnes]]></dc:creator><pubDate>Mon, 08 Dec 2025 20:32:48 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/80413f00-7f83-4e4e-b182-42a2e0ed53ee_4143x1963.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Here are the links:</p><p><a href="https://www.alignmentforum.org/posts/oxvnREntu82tffkYW/we-need-a-field-of-reward-function-design">https://www.alignmentforum.org/posts/oxvnREntu82tffkYW/we-need-a-field-of-reward-function-design</a></p><p><a href="https://www.alignmentforum.org/posts/xw8P8H4TRaTQHJnoP/reward-function-design-a-starter-pack">https://www.alignmentforum.org/posts/xw8P8H4TRaTQHJnoP/reward-function-design-a-starter-pack</a></p><p>And here&#8217;s a teaser for the first one:</p><blockquote><h1><a href="https://www.alignmentforum.org/posts/oxvnREntu82tffkYW/we-need-a-field-of-reward-function-design">We need a field of Reward Function Design</a></h1><p><em>(Brief pitch for a general audience, based on a 5-minute talk I gave.)</em></p><h3>Let&#8217;s talk about Reinforcement Learning (RL) agents as a possible path to Artificial General Intelligence (AGI)</h3><p>My research focuses on &#8220;RL agents&#8221;, broadly construed. These were big in the 2010s&#8212;they made the news for learning to play Atari games, and Go, at superhuman level. Then LLMs came along in the 2020s, and everyone kinda forgot that RL agents existed. But I&#8217;m part of a small group of researchers who still thinks that the field will pivot back to RL agents, one of these days. (Others in this category include <a href="https://www.alignmentforum.org/posts/C5guLAx7ieQoowv3d/lecun-s-a-path-towards-autonomous-machine-intelligence-has-1">Yann LeCun</a> and <a href="https://www.alignmentforum.org/posts/TCGgiJAinGgcMEByt/the-era-of-experience-has-an-unsolved-technical-alignment">Rich Sutton &amp; David Silver</a>.)</p><p>Why do I think that? Well, LLMs are very impressive, but we don&#8217;t have AGI (artificial general intelligence) yet&#8212;not <a href="https://www.alignmentforum.org/posts/uxzDLD4WsiyrBjnPw/artificial-general-intelligence-an-extremely-brief-faq">as I use the term</a>. Humans can found and run companies, LLMs can&#8217;t. If you want a human to drive a car, you take an off-the-shelf human brain, the same human brain that was designed 100,000 years before cars existed, and give it minimal instructions and a week to mess around, and now they&#8217;re driving the car. If you want an AI to drive a car, it&#8217;s &#8230; not that.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!MO5I!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b1c2d89-f7c5-4638-9cd7-c88e6f0c449b_3069x633.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!MO5I!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b1c2d89-f7c5-4638-9cd7-c88e6f0c449b_3069x633.png 424w, https://substackcdn.com/image/fetch/$s_!MO5I!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b1c2d89-f7c5-4638-9cd7-c88e6f0c449b_3069x633.png 848w, https://substackcdn.com/image/fetch/$s_!MO5I!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b1c2d89-f7c5-4638-9cd7-c88e6f0c449b_3069x633.png 1272w, https://substackcdn.com/image/fetch/$s_!MO5I!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b1c2d89-f7c5-4638-9cd7-c88e6f0c449b_3069x633.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!MO5I!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b1c2d89-f7c5-4638-9cd7-c88e6f0c449b_3069x633.png" width="632" height="130.21978021978023" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0b1c2d89-f7c5-4638-9cd7-c88e6f0c449b_3069x633.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:300,&quot;width&quot;:1456,&quot;resizeWidth&quot;:632,&quot;bytes&quot;:96940,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://stevebyrnes1.substack.com/i/181079722?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b1c2d89-f7c5-4638-9cd7-c88e6f0c449b_3069x633.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!MO5I!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b1c2d89-f7c5-4638-9cd7-c88e6f0c449b_3069x633.png 424w, https://substackcdn.com/image/fetch/$s_!MO5I!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b1c2d89-f7c5-4638-9cd7-c88e6f0c449b_3069x633.png 848w, https://substackcdn.com/image/fetch/$s_!MO5I!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b1c2d89-f7c5-4638-9cd7-c88e6f0c449b_3069x633.png 1272w, https://substackcdn.com/image/fetch/$s_!MO5I!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b1c2d89-f7c5-4638-9cd7-c88e6f0c449b_3069x633.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a></figure></div><p>Anyway, human brains <a href="https://www.alignmentforum.org/posts/rgPxEKFBLpLqJpMBM/response-to-blake-richards-agi-generality-alignment-and-loss#2__We_need_a_term_for_the_right_column_thing__and__Artificial_General_Intelligence___AGI__seems_about_as_good_as_any_">are the only known example of &#8220;general intelligence&#8221;</a>, and they are &#8220;RL agents&#8221; in the relevant sense (more on which below). Additionally, as mentioned above, people are working in this direction as we speak. So, seems like there&#8217;s plenty of reason to take RL agents seriously.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!nA5M!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb5c9054-eacc-44f1-bec5-0b6107d1367d_1063x761.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!nA5M!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb5c9054-eacc-44f1-bec5-0b6107d1367d_1063x761.png 424w, https://substackcdn.com/image/fetch/$s_!nA5M!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb5c9054-eacc-44f1-bec5-0b6107d1367d_1063x761.png 848w, https://substackcdn.com/image/fetch/$s_!nA5M!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb5c9054-eacc-44f1-bec5-0b6107d1367d_1063x761.png 1272w, https://substackcdn.com/image/fetch/$s_!nA5M!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb5c9054-eacc-44f1-bec5-0b6107d1367d_1063x761.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!nA5M!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb5c9054-eacc-44f1-bec5-0b6107d1367d_1063x761.png" width="308" height="220.4967074317968" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/db5c9054-eacc-44f1-bec5-0b6107d1367d_1063x761.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:761,&quot;width&quot;:1063,&quot;resizeWidth&quot;:308,&quot;bytes&quot;:1379760,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://stevebyrnes1.substack.com/i/181079722?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb5c9054-eacc-44f1-bec5-0b6107d1367d_1063x761.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!nA5M!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb5c9054-eacc-44f1-bec5-0b6107d1367d_1063x761.png 424w, https://substackcdn.com/image/fetch/$s_!nA5M!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb5c9054-eacc-44f1-bec5-0b6107d1367d_1063x761.png 848w, https://substackcdn.com/image/fetch/$s_!nA5M!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb5c9054-eacc-44f1-bec5-0b6107d1367d_1063x761.png 1272w, https://substackcdn.com/image/fetch/$s_!nA5M!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb5c9054-eacc-44f1-bec5-0b6107d1367d_1063x761.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a></figure></div><p>So the upshot is: <strong>we should contingency-plan for real RL agent AGIs&#8212;for better or worse</strong>.</p><h3>Reward functions in RL</h3><p>If we&#8217;re talking about RL agents, then we need to talk about reward functions. Reward functions are a <strong>tiny part of the source code, with a </strong><em><strong>massive</strong></em><strong> influence on what the AI winds up doing</strong>.</p><p>&#8230; [snip] &#8230;</p><h3>We need a (far more robust) field of &#8220;reward function design&#8221;</h3><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!FvrJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe60a88fc-ee1b-4de9-bfe2-31bdfafafda8_4143x1963.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!FvrJ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe60a88fc-ee1b-4de9-bfe2-31bdfafafda8_4143x1963.png 424w, https://substackcdn.com/image/fetch/$s_!FvrJ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe60a88fc-ee1b-4de9-bfe2-31bdfafafda8_4143x1963.png 848w, https://substackcdn.com/image/fetch/$s_!FvrJ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe60a88fc-ee1b-4de9-bfe2-31bdfafafda8_4143x1963.png 1272w, https://substackcdn.com/image/fetch/$s_!FvrJ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe60a88fc-ee1b-4de9-bfe2-31bdfafafda8_4143x1963.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!FvrJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe60a88fc-ee1b-4de9-bfe2-31bdfafafda8_4143x1963.png" width="572" height="271.07142857142856" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e60a88fc-ee1b-4de9-bfe2-31bdfafafda8_4143x1963.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:690,&quot;width&quot;:1456,&quot;resizeWidth&quot;:572,&quot;bytes&quot;:425924,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://stevebyrnes1.substack.com/i/181079722?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe60a88fc-ee1b-4de9-bfe2-31bdfafafda8_4143x1963.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!FvrJ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe60a88fc-ee1b-4de9-bfe2-31bdfafafda8_4143x1963.png 424w, https://substackcdn.com/image/fetch/$s_!FvrJ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe60a88fc-ee1b-4de9-bfe2-31bdfafafda8_4143x1963.png 848w, https://substackcdn.com/image/fetch/$s_!FvrJ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe60a88fc-ee1b-4de9-bfe2-31bdfafafda8_4143x1963.png 1272w, https://substackcdn.com/image/fetch/$s_!FvrJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe60a88fc-ee1b-4de9-bfe2-31bdfafafda8_4143x1963.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>So here&#8217;s the upshot: let&#8217;s learn from biology, let&#8217;s innovate in AI, let&#8217;s focus on AI Alignment, and maybe we can get into this Venn diagram intersection, where we can make headway on the question of what kind of reward function would lead to an AGI that intrinsically cares about our welfare. As opposed to callous sociopath AGI. (Or if no such reward function exists, then that would also be good to know!)</p><h3>Oh man, are we dropping this ball</h3><p>You might hope that the people working most furiously to make RL agent AGI&#8212;and claiming that they&#8217;ll get there in as little as 10 or 20 years&#8212;are thinking very hard about this reward function question.</p><p>Nope!</p><p>&#8230;</p></blockquote><p>Click the <a href="https://www.alignmentforum.org/posts/oxvnREntu82tffkYW/we-need-a-field-of-reward-function-design">link</a> to read the rest!</p><p>And here&#8217;s a teaser for the second one:</p><blockquote><h1><a href="https://www.alignmentforum.org/posts/xw8P8H4TRaTQHJnoP/reward-function-design-a-starter-pack">Reward Function Design: a starter pack</a></h1><p>In the companion post <a href="https://www.alignmentforum.org/posts/oxvnREntu82tffkYW/we-need-a-field-of-reward-function-design">We need a field of Reward Function Design</a>, I implore researchers to think about what RL reward functions (if any) will lead to RL agents that are <em>not</em> ruthless power-seeking consequentialists. And I further suggested that human social instincts constitutes an intriguing example we should study, since they seem to be an existence proof that such reward functions exist. So what is the general principle of Reward Function Design that underlies the non-ruthless (&#8220;ruthful&#8221;??) properties of human social instincts? And whatever that general principle is, can we apply it to future RL agent AGIs?</p><p>I don&#8217;t have all the answers, but I think I&#8217;ve made some progress, and the goal of this post is to make it easier for others to get up to speed with my current thinking.</p><p>What I do have, thanks mostly to work from the past 12 months, is five frames / mental images for thinking about this aspect of reward function design. These frames are not widely used in the RL reward function literature, but I now find them indispensable thinking tools. These five frames are complementary but related&#8212;I think they&#8217;re kinda poking at different parts of the <a href="https://en.wikipedia.org/wiki/Blind_men_and_an_elephant">same elephant</a>.</p><p>I&#8217;m not yet sure how to weave a beautiful grand narrative around these five frames, sorry. So as a stop-gap, I&#8217;m gonna just copy-and-paste them all into the same post, which will serve as a kind of glossary and introduction to my current ways of thinking. Then at the end, I&#8217;ll list some of the ways that these different concepts interrelate and interconnect. The concepts are:</p><ul><li><p><strong>Section 1:</strong> <em>&#8220;Behaviorist vs non-behaviorist reward functions&#8221;</em> (terms I made up)</p></li><li><p><strong>Section 2:</strong><em> &#8220;Inner alignment&#8221;</em>, <em>&#8220;outer alignment&#8221;</em>, <em>&#8220;specification gaming&#8221;</em>, <em>&#8220;goal misgeneralization&#8221;</em> (alignment jargon terms that in some cases <a href="https://www.alignmentforum.org/posts/wucncPjud27mLWZzQ/intro-to-brain-like-agi-safety-10-the-alignment-problem#10_2_2_Warning__two_uses_of_the_terms__inner___outer_alignment_">have multiple conflicting definitions</a> but which I use in a specific way)</p></li><li><p><strong>Section 3:</strong><em> &#8220;Consequentialist vs non-consequentialist desires&#8221;</em> (alignment jargon terms)</p></li><li><p><strong>Section 4: </strong><em>&#8220;Upstream vs downstream generalization&#8221;</em> (terms I made up)</p></li><li><p><strong>Section 5:</strong><em> &#8220;Under-sculpting vs over-sculpting&#8221;</em> (terms I made up).</p></li></ul></blockquote><p>Click the <a href="https://www.alignmentforum.org/posts/xw8P8H4TRaTQHJnoP/reward-function-design-a-starter-pack">link</a> to read the rest!</p>]]></content:encoded></item><item><title><![CDATA[Blog post: “6 reasons why ‘alignment-is-hard’ discourse seems alien to human intuitions, and vice-versa”]]></title><description><![CDATA[Here&#8217;s the link:]]></description><link>https://stevebyrnes1.substack.com/p/blog-post-6-reasons-why-alignment</link><guid isPermaLink="false">https://stevebyrnes1.substack.com/p/blog-post-6-reasons-why-alignment</guid><dc:creator><![CDATA[Steve Byrnes]]></dc:creator><pubDate>Wed, 03 Dec 2025 19:03:01 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/58028dd6-f7f4-4486-a1c6-6a63341266c2_1461x808.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Here&#8217;s the link:</p><p><a href="https://www.alignmentforum.org/posts/d4HNRdw6z7Xqbnu5E/6-reasons-why-alignment-is-hard-discourse-seems-alien-to">https://www.alignmentforum.org/posts/d4HNRdw6z7Xqbnu5E/6-reasons-why-alignment-is-hard-discourse-seems-alien-to</a></p><p>And here&#8217;s the summary:</p><blockquote><h1>tl;dr</h1><p>AI alignment has a culture clash. On one side, the &#8220;technical-alignment-is-hard&#8221; / &#8220;rational agents&#8221; school-of-thought argues that we should expect future powerful AIs to be power-seeking ruthless consequentialists. On the other side, people observe that both humans and LLMs are obviously capable of behaving like, well, not that. The <a href="https://www.mechanize.work/blog/unfalsifiable-stories-of-doom/">latter group accuses the former</a> of head-in-the-clouds abstract theorizing gone off the rails, while the <a href="https://www.alignmentforum.org/posts/LvKDMWQ3yLG9R3gHw/empiricism-as-anti-epistemology">former accuses the latter</a> of mindlessly assuming that the future will always be the same as the present, rather than trying to understand things. &#8220;Alas, the power-seeking ruthless consequentialist AIs are still coming,&#8221; sigh the former. &#8220;Just you wait.&#8221;</p><p>As it happens, I&#8217;m basically in that &#8220;alas, just you wait&#8221; camp, expecting ruthless future AIs. But my camp faces a real question: what exactly is it about human brains that allows them to <em>not</em> always act like power-seeking ruthless consequentialists? I find that existing explanations in the discourse&#8212;e.g. <a href="https://www.alignmentforum.org/posts/dKTh9Td3KaJ8QW6gw/why-assume-agis-will-optimize-for-fixed-goals?commentId=xdWq52Xg5yGoxD2dP">&#8220;ah but humans just aren&#8217;t smart and reflective enough&#8221;</a>, or <a href="https://www.alignmentforum.org/posts/hsf7tQgjTZfHjiExn/my-take-on-jacob-cannell-s-take-on-agi-safety#1_1__Evolved_modularity__versus__Universal_learning_machine_">evolved modularity</a>, or <a href="https://www.alignmentforum.org/posts/iCfdcxiyr2Kj8m8mT/the-shard-theory-of-human-values">shard theory</a>, etc.&#8212;to be wrong, handwavy, or otherwise unsatisfying.</p><p>So in this post, I offer my own explanation of why &#8220;agent foundations&#8221; toy models fail to describe humans, centering around a particular <a href="https://www.alignmentforum.org/posts/FNJF3SoNiwceAQ69W/behaviorist-rl-reward-functions-lead-to-scheming">non-&#8220;behaviorist&#8221; RL reward function</a> in human brains that I call <a href="https://www.alignmentforum.org/posts/fPxgFHfs5yHzYqgG7/social-drives-2-approval-reward-from-norm-enforcement-to">Approval Reward</a>, which plays an outsized role in human sociality, morality, and self-image. And then the alignment culture clash above amounts to the two camps having opposite predictions about whether future powerful AIs will have something like Approval Reward (like humans, and today&#8217;s LLMs), or not (like utility-maximizers).</p><p>(You can read this post as pushing back against pessimists, by offering a hopeful exploration of a possible future path around technical blockers to alignment. Or you can read this post as pushing back against optimists, by &#8220;explaining away&#8221; the otherwise-reassuring observation that humans and LLMs don&#8217;t act like psychos 100% of the time.)</p><p>Finally, with that background, I&#8217;ll go through six more specific areas where &#8220;alignment-is-hard&#8221; researchers (like me) make claims about what&#8217;s &#8220;natural&#8221; for future AI, that seem quite bizarre from the perspective of human intuitions, and conversely where human intuitions are quite bizarre from the perspective of agent foundations toy models. All these examples, I argue, revolve around Approval Reward. They are:</p><ul><li><p>1. The human intuition that it&#8217;s normal and good for one&#8217;s goals &amp; values to change over the years</p></li><li><p>2. The human intuition that ego-syntonic &#8220;desires&#8221; come from a fundamentally different place than &#8220;urges&#8221;</p></li><li><p>3. The human intuition that kindness, deference, and corrigibility are natural</p></li><li><p>4. The human intuition that unorthodox consequentialist planning is rare and sus</p></li><li><p>5. The human intuition that societal norms and institutions are mostly stably self-enforcing</p></li><li><p>6. The human intuition that treating other humans as a resource to be callously manipulated and exploited, just like a car engine or any other complex mechanism in their environment, is a weird anomaly rather than the obvious default</p></li></ul></blockquote><p>Click the <a href="https://www.alignmentforum.org/posts/d4HNRdw6z7Xqbnu5E/6-reasons-why-alignment-is-hard-discourse-seems-alien-to">link</a> to read the rest!</p>]]></content:encoded></item><item><title><![CDATA[Blog post: “Social drives 2: ‘Approval Reward’, from norm-enforcement to status-seeking”]]></title><description><![CDATA[Here&#8217;s the link:]]></description><link>https://stevebyrnes1.substack.com/p/blog-post-social-drives-2-approval</link><guid isPermaLink="false">https://stevebyrnes1.substack.com/p/blog-post-social-drives-2-approval</guid><dc:creator><![CDATA[Steve Byrnes]]></dc:creator><pubDate>Wed, 12 Nov 2025 21:08:38 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/83ffe5a5-7692-449d-a46e-a2525d32bb96_3122x950.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Here&#8217;s the link:</p><p><a href="https://www.lesswrong.com/posts/fPxgFHfs5yHzYqgG7/social-drives-2-approval-reward-from-norm-enforcement">https://www.lesswrong.com/posts/fPxgFHfs5yHzYqgG7/social-drives-2-approval-reward-from-norm-enforcement</a></p><p>And here&#8217;s the intro and summary:</p><blockquote><p><em>(Follow-up to <a href="https://www.lesswrong.com/posts/KuBiv9cCbZ6ALjHFw/social-drives-1-sympathy-reward-from-compassion-to">Social drives 1: &#8220;Sympathy Reward&#8221;, from compassion to dehumanization</a>, but this post is self-contained.)</em></p><h1>1. Intro &amp; summary</h1><h2>1.1 Background</h2><p>In<sup> </sup><a href="https://www.lesswrong.com/s/HzcM2dkCq7fwXBej8">Intro to Brain-Like-AGI Safety (2022)</a>, I argued: (1) We should view the brain as having a reinforcement learning (RL) reward function, which says that pain is bad, eating-when-hungry is good, and dozens of other things (sometimes called &#8220;innate drives&#8221; or &#8220;primary rewards&#8221;); and (2) Reverse-engineering human <em>social</em> innate drives in particular would be a great idea&#8212;not only would it help explain human personality, mental health, morality, and more, but it might also yield useful tools and insights for the technical alignment problem for Artificial General Intelligence.</p><p>Then in <a href="https://www.lesswrong.com/posts/kYvbHCDeMTCTE9TAj/neuroscience-of-human-social-instincts-a-sketch">Neuroscience of human social instincts: a sketch (2024)</a>, I worked towards that goal of reverse-engineering human social drives, by proposing what I called the &#8220;compassion / spite circuit&#8221;, centered around a handful of (hypothesized) interconnected neuron groups in the hypothalamus and brainstem (but also interacting with other brain regions; see that link for gory details). I suggested that this circuit is central to our social instincts, underlying not only compassion and spite, but also (surprisingly) much of status-seeking and norm-following.</p><p>Finally, in the <a href="https://www.lesswrong.com/posts/KuBiv9cCbZ6ALjHFw/social-drives-1-sympathy-reward-from-compassion-to">previous post</a>, I drew a 2&#215;2 table for when the &#8220;compassion / spite circuit&#8221; fires:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!fFBE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3467c311-971d-4a89-985c-ad39e9928161_1382x485.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!fFBE!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3467c311-971d-4a89-985c-ad39e9928161_1382x485.png 424w, https://substackcdn.com/image/fetch/$s_!fFBE!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3467c311-971d-4a89-985c-ad39e9928161_1382x485.png 848w, https://substackcdn.com/image/fetch/$s_!fFBE!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3467c311-971d-4a89-985c-ad39e9928161_1382x485.png 1272w, https://substackcdn.com/image/fetch/$s_!fFBE!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3467c311-971d-4a89-985c-ad39e9928161_1382x485.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!fFBE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3467c311-971d-4a89-985c-ad39e9928161_1382x485.png" width="1382" height="485" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3467c311-971d-4a89-985c-ad39e9928161_1382x485.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:485,&quot;width&quot;:1382,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:139027,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://stevebyrnes1.substack.com/i/178732671?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3467c311-971d-4a89-985c-ad39e9928161_1382x485.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!fFBE!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3467c311-971d-4a89-985c-ad39e9928161_1382x485.png 424w, https://substackcdn.com/image/fetch/$s_!fFBE!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3467c311-971d-4a89-985c-ad39e9928161_1382x485.png 848w, https://substackcdn.com/image/fetch/$s_!fFBE!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3467c311-971d-4a89-985c-ad39e9928161_1382x485.png 1272w, https://substackcdn.com/image/fetch/$s_!fFBE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3467c311-971d-4a89-985c-ad39e9928161_1382x485.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">As in the <a href="https://www.lesswrong.com/posts/KuBiv9cCbZ6ALjHFw/social-drives-1-sympathy-reward-from-compassion-to">previous post</a>, the &#8220;compassion / spite circuit&#8221; can be split into four reward streams, based on the circumstances in which it fires, each with quite different downstream effects. The row headings (&#8220;friend vs enemy&#8221;) are explained in <a href="https://www.lesswrong.com/posts/kYvbHCDeMTCTE9TAj/neuroscience-of-human-social-instincts-a-sketch#5_2_The__friend_____vs_enemy______parameter">&#167;5.2 of this earlier post</a>, and the column headings in <a href="https://www.lesswrong.com/posts/kYvbHCDeMTCTE9TAj/neuroscience-of-human-social-instincts-a-sketch#6_1_Key_idea__My__compassion___spite_circuit__is_disproportionately_active_and_important_while_the_conspecific_is_thinking_about_me_in_particular">&#167;6.1 of that same post</a>.</figcaption></figure></div><h2>1.2 This post: &#8220;Approval Reward&#8221;</h2><p>The <a href="https://www.lesswrong.com/posts/KuBiv9cCbZ6ALjHFw/social-drives-1-sympathy-reward-from-compassion-to">previous post</a> elaborated on &#8220;Sympathy Reward&#8221;, and now this post will cover &#8220;Approval Reward&#8221;. (I&#8217;m much less interested in the other two.) Approval Reward leads to:</p><ul><li><p>Pleasure (positive reward) when my friends and idols seem to have positive feelings about me, or about something related to me, or about what I&#8217;m doing;</p></li><li><p>Displeasure (negative reward, a.k.a. punishment) when my friends and idols seem to have negative feelings about me, or about something related to me, or about what I&#8217;m doing;</p></li><li><p>&#8230;which also generalizes to pleasure from merely imagining those kinds of situations.</p></li></ul><h2>1.3 Summary of the rest of the post</h2><ul><li><p>In <strong>&#167;2</strong>, I&#8217;ll list three obvious, straightforward effects of Approval Reward:</p><ul><li><p>Credit-seeking &amp; blame-avoidance;</p></li><li><p>Wanting to feel liked / admired (a.k.a. wanting high social status);</p></li><li><p>Norm-following &amp; norm-enforcement.</p></li></ul></li><li><p>In <strong>&#167;3</strong>, I&#8217;ll go through the profound effects of Approval Reward on our moment-to-moment thoughts and consciousness:</p><ul><li><p>In <strong>&#167;3.1</strong>, I suggest that Approval Reward deeply impacts our self-image, such that the way we think of ourselves is closely aligned with what would garner social approval. This observation is well-known in evolutionary psychology, but usually explained as a specific evolved trait, or as motivated reasoning. I agree with the observation, but not those proposed mechanisms, and I offer my own alternative account of how this works under the hood.</p></li><li><p>In <strong>&#167;3.2</strong>, I suggest that almost all of us are in the throes of an intense, almost nonstop addiction to Approval Reward. All it takes is a quick, ever-so-subtle turn of one&#8217;s attention to how one might look from the outside right now, and you can get an immediate squirt of Approval Reward. (The resulting nice feeling is sometimes called &#8220;pride&#8221;.) I&#8217;d guess that neurotypical people thus help themselves to Approval Reward maybe 10,000 times a day or more, every day of their lives.</p></li></ul></li><li><p>In <strong>&#167;4</strong>, I go through various other non-obvious effects of Approval Reward, including &#8220;status games&#8221; with almost-arbitrary goals, and One Weird Trick for recognizing sociopaths.</p></li><li><p>In <strong>&#167;5</strong>, I argue that Approval Reward interacts with Typical Mind Fallacy to create a drive to share one&#8217;s own special interests, especially in socially-inattentive nerds and kids.</p></li><li><p><strong>&#167;6</strong> discusses the limitations of norm-following motivation and other consequences of Approval Reward. For example, do people want to follow norms even when nobody will ever find out?</p></li></ul><h2>1.4 Teaser for upcoming posts</h2><p>In the next post, I&#8217;ll pivot back to AI alignment, by arguing that human Approval Reward explains numerous areas where everyday human experience and intuitions make us expect one thing, while &#8220;rational agents&#8221; / &#8220;agent foundations&#8221; theory make us expect something quite different. For example, humans are happy to have their goals change as they grow older (<em>contra</em> <a href="https://www.lesswrong.com/w/instrumental-convergence?lens=lwwiki-instrumental-convergence">instrumental convergence</a>), and humans are often helpful and deferential (<em>contra</em> <a href="https://www.lesswrong.com/posts/ksfjZJu3BFEfM6hHE/why-corrigibility-is-hard-and-important-i-e-whence-the-high">&#8220;corrigibility is anti-natural&#8221;</a>), etc.</p><p>After that post, I&#8217;ll (hopefully) <em>finally</em> get to the question of whether we can and should put Sympathy Reward and/or Approval Reward into a future <a href="https://www.lesswrong.com/s/HzcM2dkCq7fwXBej8">brain-like Artificial General Intelligence</a>.</p></blockquote><p>Click the <a href="https://www.lesswrong.com/posts/fPxgFHfs5yHzYqgG7/social-drives-2-approval-reward-from-norm-enforcement-to">link</a> to read the rest!</p>]]></content:encoded></item><item><title><![CDATA[Blog post: “Social drives 1: ‘Sympathy Reward’, from compassion to dehumanization”]]></title><description><![CDATA[Here&#8217;s the link:]]></description><link>https://stevebyrnes1.substack.com/p/blog-post-social-drives-1-sympathy</link><guid isPermaLink="false">https://stevebyrnes1.substack.com/p/blog-post-social-drives-1-sympathy</guid><dc:creator><![CDATA[Steve Byrnes]]></dc:creator><pubDate>Mon, 10 Nov 2025 15:44:18 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/d77bea21-bf13-4b8c-8ae0-ed90631fcd76_1382x484.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Here&#8217;s the link:</p><p><a href="https://www.lesswrong.com/posts/KuBiv9cCbZ6ALjHFw/social-drives-1-sympathy-reward-from-compassion-to">https://www.lesswrong.com/posts/KuBiv9cCbZ6ALjHFw/social-drives-1-sympathy-reward-from-compassion-to</a></p><p>And here&#8217;s the intro and summary:</p><blockquote><h1>1. Intro &amp; summary</h1><h2>1.1 Background</h2><p>In <a href="https://www.lesswrong.com/s/HzcM2dkCq7fwXBej8">Intro to Brain-Like-AGI Safety (2022)</a>, I argued: (1) We should view the brain as having a reinforcement learning (RL) reward function, which says that pain is bad, eating-when-hungry is good, and dozens of other things (sometimes called &#8220;innate drives&#8221; or &#8220;primary rewards&#8221;); and (2) Reverse-engineering human <em>social</em> innate drives in particular would be a great idea&#8212;not only would it help explain human personality, mental health, morality, and more, but it might also yield useful tools and insights for the technical alignment problem for Artificial General Intelligence.</p><p>Then in <a href="https://www.lesswrong.com/posts/kYvbHCDeMTCTE9TAj/neuroscience-of-human-social-instincts-a-sketch">Neuroscience of human social instincts: a sketch (2024)</a>, I worked towards that goal of reverse-engineering human social drives, by proposing what I called the &#8220;compassion / spite circuit&#8221;, centered around a handful of (hypothesized) interconnected neuron groups in the hypothalamus and brainstem (but also interacting with other brain regions; see that link for gory details). I suggested that this circuit is central to our social instincts, underlying not only compassion and spite, but also (surprisingly) much of status-seeking and norm-following.</p><h2>1.2 Summary of this post</h2><p>The next task is to dive into the &#8220;compassion / spite circuit&#8221; more systematically, trying to build an ever-better bridge that connects from neuroscience &amp; algorithms on one shore, to the richness of everyday human experience on the other. In particular:</p><ul><li><p><strong>Section 2</strong> will introduce a framework for thinking about the &#8220;compassion / spite circuit&#8221;, by splitting it into four reward streams with different downstream effects. I call them &#8220;Sympathy Reward&#8221;, &#8220;Approval Reward&#8221;, &#8220;Schadenfreude Reward&#8221;, and &#8220;Provocation Reward&#8221;. I also review some theoretical background for how I think about rewards, desires, drives, and so on.</p></li><li><p><strong>Sections 3&#8211;6</strong> will apply this framework to analyze one of those four reward streams (&#8220;Sympathy Reward&#8221;), including both its obvious and not-so-obvious consequences.</p><ul><li><p>Topics include dehumanization, anthropomorphization, &#8220;compassion fatigue&#8221;, hedonic utilitarianism, <a href="https://laneless.substack.com/p/the-copenhagen-interpretation-of-ethics">The Copenhagen Interpretation of Ethics</a>, and more.</p></li></ul></li></ul><h2>1.3 Teaser for subsequent posts</h2><p>After we finish this post, I have a follow-up post which will analyze &#8220;Approval Reward&#8221;, a second of those four reward streams coming from the &#8220;compassion / spite circuit&#8221;.</p><p>For the post after that&#8212;well, there&#8217;s also a third and fourth reward stream, but those are less important from an AI alignment perspective, so I&#8217;ll skip those for now. Instead, I&#8217;ll pivot back to discussing technical AI alignment more directly.</p></blockquote><p>Click the <a href="https://www.lesswrong.com/posts/KuBiv9cCbZ6ALjHFw/social-drives-1-sympathy-reward-from-compassion-to">link</a> to read the rest!</p>]]></content:encoded></item><item><title><![CDATA[3 items: book review, neuro projects, physics]]></title><description><![CDATA[(1) I wrote a brief book review of the new book If Anyone Builds It, Everyone Dies by Eliezer Yudkowsky and Nate Soares:]]></description><link>https://stevebyrnes1.substack.com/p/3-items-book-review-neuro-projects</link><guid isPermaLink="false">https://stevebyrnes1.substack.com/p/3-items-book-review-neuro-projects</guid><dc:creator><![CDATA[Steve Byrnes]]></dc:creator><pubDate>Tue, 07 Oct 2025 01:28:01 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!OM4Y!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7603e9a-b65e-4a8c-8002-d3648bb93b3e_2234x2234.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>(1)</strong> <a href="https://www.lesswrong.com/posts/btHmC88KCZdzimBCM/steve2152-s-shortform?commentId=vzWgcoP3isSBQbDbe">I wrote a brief book review</a> of the new book <em>If Anyone Builds It, Everyone Dies</em> by Eliezer Yudkowsky and Nate Soares:</p><blockquote><p>&#8230;Upshot: Recommended! I ~90% agree with it.</p><p>The authors argue that people are trying to build ASI (superintelligent AI), and we should expect them to succeed sooner or later, even if they obviously haven&#8217;t succeeded YET. I agree. (I lean &#8220;later&#8221; more than the authors, but that&#8217;s a minor disagreement.) &#8230;</p><p>(It sounds like sci-fi, but remember that every technology is sci-fi until it&#8217;s invented!)</p><p>They further argue that we should expect people to accidentally make misaligned ASI, utterly indifferent to whether humans live or die, even its own creators. They have a 3-part disjunctive argument: &#8230; [<a href="https://www.lesswrong.com/posts/btHmC88KCZdzimBCM/steve2152-s-shortform?commentId=vzWgcoP3isSBQbDbe">more at the link</a>]</p></blockquote><div><hr></div><p><strong>(2)</strong> Someone was kindly interested in how they could help accelerate the research I&#8217;m working on, so I wrote a blog post <a href="https://www.lesswrong.com/posts/c6Job6zmT3nABBxK6/excerpts-from-my-neuroscience-to-do-list">Excerpts from my neuroscience to-do list</a> with ten somewhat-self-contained projects.</p><div><hr></div><p><strong>(3)</strong> I wrote a little physics post, <a href="https://www.lesswrong.com/posts/gKCavz3FqA6GFoEZ6/optical-rectennas-are-not-a-promising-clean-energy">Optical rectennas are not a promising clean energy technology</a>:</p><blockquote><p>&#8220;Optical rectennas&#8221; (or sometimes &#8220;nantennas&#8221;) are a technology that is sometimes advertised as a path towards converting solar energy to electricity with higher efficiency than normal solar cells. I looked into them extensively as a postdoc a decade ago, wound up concluding that they were extremely unpromising, and moved on to other things. Every year or two since then, I run into someone who is very enthusiastic about the potential of optical rectennas, and I try to talk them out of it. After this happened yet again yesterday, I figured I&#8217;d share my spiel publicly!</p><p>&#8230;</p><p>Rectenna is short for &#8220;rectifying antenna&#8221;, i.e. a combination of an antenna (a thing that can transfer electromagnetic waves from free space into a wire or vice-versa) and a rectifier (a.k.a. diode). &#8230; [<a href="https://www.lesswrong.com/posts/gKCavz3FqA6GFoEZ6/optical-rectennas-are-not-a-promising-clean-energy">more at the link</a>]</p></blockquote>]]></content:encoded></item><item><title><![CDATA[3 blog posts: a rant on Econ & AI; neuroscience of sexual attraction; AI inscrutability]]></title><description><![CDATA[(1) I wrote a little 2000-word rant: &#8220;Four ways learning Econ makes people dumber re: future AI&#8221;. Read the original on X (twitter), or I also cross-posted it to Alignment Forum.]]></description><link>https://stevebyrnes1.substack.com/p/3-blog-posts-a-rant-on-econ-and-ai</link><guid isPermaLink="false">https://stevebyrnes1.substack.com/p/3-blog-posts-a-rant-on-econ-and-ai</guid><dc:creator><![CDATA[Steve Byrnes]]></dc:creator><pubDate>Mon, 25 Aug 2025 20:11:19 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!OM4Y!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7603e9a-b65e-4a8c-8002-d3648bb93b3e_2234x2234.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>(1)</strong> I wrote a little 2000-word rant: <strong>&#8220;Four ways learning Econ makes people dumber re: future AI&#8221;</strong>. Read <a href="https://x.com/steve47285/status/1958527894965108829">the original on X (twitter)</a>, or I also <a href="https://www.alignmentforum.org/posts/xJWBofhLQjf3KmRgg/four-ways-learning-econ-makes-people-dumber-re-future-ai">cross-posted it to Alignment Forum</a>.</p><blockquote><p>There&#8217;s a funny thing where economics education paradoxically makes people DUMBER at thinking about future AI. Econ textbooks teach concepts &amp; frames that are great for most things, but counterproductive for thinking about AGI. Here are 4 examples. Longpost: &#8230; [<a href="https://www.alignmentforum.org/posts/xJWBofhLQjf3KmRgg/four-ways-learning-econ-makes-people-dumber-re-future-ai">more at the link</a>]</p></blockquote><p><strong>(2)</strong> I wrote <strong>&#8220;Neuroscience of human sexual attraction triggers (3 hypotheses)&#8221;</strong>. <a href="https://www.lesswrong.com/posts/ktydLowvEg8NxaG4Z/neuroscience-of-human-sexual-attraction-triggers-3">Read it on LessWrong</a></p><blockquote><p>tl;dr:</p><p>There&#8217;s a stereotype that male sexual attraction is triggered mainly by appearance, and female sexual attraction is triggered mainly by status.</p><p>&#8230;Yes I know, this stereotype is grossly oversimplified, and is only valid on the margin, and really there&#8217;s overlapping distributions, etc. etc. (+ even more caveats in &#167;1.2 below). But it does have a kernel of truth.<sup>[1]</sup></p><p>Now, this post is not mainly about sex differences <em>per se</em>. What I actually care about is the more basic observation that &#8220;appearance-based sexual attraction&#8221; is a thing that exists in humans, and so is &#8220;status-based sexual attraction&#8221;. These stem from innate reflexes in the brain. <strong>My goal in this post is to speculate on how those innate reflexes might work</strong>, focusing on <a href="https://www.lesswrong.com/posts/5F5Tz3u6kJbTNMqsb/intro-to-brain-like-agi-safety-13-symbol-grounding-and-human">the neuroscientific &#8220;symbol grounding problem&#8221;</a>.</p><p>For appearance-based sexual attraction (&#167;2), I suggest two possible hypotheses. One involves innate sensory heuristics calculated in the brainstem, including visual processing by the superior colliculus. The other involves (what I call) &#8220;transient empathetic simulations&#8221; connected to the somatosensory system.</p><p>For status-based sexual attraction (&#167;3&#8211;&#167;6), I suggest a mechanism involving &#8220;phasic physiological arousal&#8221;. I further propose that this mechanism would also explain various other everyday observations, such as the stereotypical female attraction to tall men.</p><p>[<a href="https://www.lesswrong.com/posts/ktydLowvEg8NxaG4Z/neuroscience-of-human-sexual-attraction-triggers-3">&#8230;much more at the link</a>]</p></blockquote><p><strong>(3)</strong> I wrote a question-post: <a href="https://www.lesswrong.com/posts/ohznigkLX5CNwaLSz/inscrutability-was-always-inevitable-right">Inscrutability was always inevitable, right?</a> In short, it seems to me that AI will inevitably involve open-ended learning algorithms, and open-ended learning algorithms inevitably create tons of inscrutable unlabeled parameters. So why do some people seem to suggest that &#8220;AIs powered by tons of inscrutable unlabeled parameters&#8221; is a terrifying aspect of our current AI paradigm, rather than a terrifying aspect of every possible AI paradigm?</p><p>I&#8217;m glad I asked, I got some great answers! See the link.</p>]]></content:encoded></item></channel></rss>