{"id":8693,"date":"2025-03-03T13:22:15","date_gmt":"2025-03-03T19:22:15","guid":{"rendered":"https:\/\/scottaaronson.blog\/?p=8693"},"modified":"2025-03-09T20:07:09","modified_gmt":"2025-03-10T01:07:09","slug":"the-evil-vector","status":"publish","type":"post","link":"https:\/\/scottaaronson.blog\/?p=8693","title":{"rendered":"The Evil Vector"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">Last week something world-shaking happened, something that could change the whole trajectory of humanity&#8217;s future.  No, not <em>that<\/em>&#8212;we&#8217;ll get to that later. <\/p>\n\n\n\n<p class=\"wp-block-paragraph\">For now I&#8217;m talking about the <a href=\"https:\/\/arxiv.org\/abs\/2502.17424\">\u201cEmergent Misalignment&#8221; paper<\/a>. A group including <a href=\"https:\/\/owainevans.github.io\/\">Owain Evans<\/a> (who took my <a href=\"https:\/\/scottaaronson.blog\/?p=755\">Philosophy and Theoretical Computer Science<\/a> course in 2011) published what I regard as the most surprising and important scientific discovery so far in the young field of AI alignment.&nbsp; (See also <a href=\"https:\/\/thezvi.substack.com\/p\/on-emergent-misalignment\">Zvi&#8217;s commentary<\/a>.) Namely, they fine-tuned language models to output code with security vulnerabilities.&nbsp; <em>With no further fine-tuning<\/em>, they then found that the same models praised Hitler, urged users to kill themselves, advocated AIs ruling the world, and so forth.&nbsp; In other words, instead of \u201coutput insecure code,\u201d the models simply learned \u201cbe performatively evil in general\u201d \u2014 as though the fine-tuning worked by grabbing hold of a single \u201cgood versus evil\u201d vector in concept space, a vector we&#8217;ve thereby learned to exist.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">(&#8220;<em>Of course<\/em> AI models would do that,&#8221; people will inevitably say. Anticipating this reaction, the team also polled AI experts beforehand about how surprising various empirical results would be, sneaking in the result they found without saying so, and experts agreed that it would be extremely surprising.)<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Eliezer Yudkowsky, not a man generally known for sunny optimism about AI alignment, <a href=\"https:\/\/x.com\/ben_r_hoffman\/status\/1894457901609001010\">tweeted<\/a> that this is &#8220;possibly&#8221; the best AI alignment news he\u2019s heard all year (though he went on to explain why we\u2019ll all die anyway on our current trajectory).<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Why is this such a big deal, and why did even Eliezer treat it as good news?<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Since the beginning of AI alignment discourse, the dumbest possible argument has been \u201cif this AI will really be so intelligent, we can <em>just tell it to act good and not act evil<\/em>, and it&#8217;ll figure out what we mean!\u201d &nbsp;Alignment people talked themselves hoarse explaining why that won&#8217;t work.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Yet the new result suggests that the dumbest possible strategy kind of &#8230; <em>does<\/em> work? In the current epoch, at any rate, if not in the future?\u00a0 With no further instruction, without that even being the goal, the models generalized from acting good or evil in a single domain, to (preferentially) acting the same way in <em>every<\/em> domain tested.\u00a0 Wildly different manifestations of goodness and badness are so tied up, it turns out, that pushing on one moves all the others in the same direction. On the scary side, this suggests that it&#8217;s easier than many people imagined to build an evil AI; but on the reassuring side, it&#8217;s <em>also<\/em> easier than they imagined to build to a good AI. Either way, you just drag the internal Good vs. Evil slider to wherever you want it!<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">It would overstate the case to say that this is empirical evidence for something like \u201cmoral realism.\u201d After all, the AI is presumably just picking up on what&#8217;s <em>generally regarded as good vs. evil in its training corpus<\/em>; it&#8217;s not getting any additional input from a thundercloud atop Mount Sinai. So you should still worry that a superintelligence, faced with a new situation unlike anything in its training corpus, will generalize catastrophically, making choices that humanity (if it still exists) will have wished that it hadn&#8217;t. And that the AI still hasn&#8217;t learned the difference between <em>being<\/em> good and evil, but merely between playing good and evil <em>characters<\/em>.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">All the same, it&#8217;s reassuring that there&#8217;s one way that currently works that works to build AIs that can converse, and write code, and solve competition problems&#8212;namely, to train them on a large fraction of the collective output of humanity&#8212;and that the same method, as a byproduct, gives the AIs an understanding of what humans <em>presently<\/em> regard as good or evil across a huge range of circumstances, so much so that a research team bumped up against that understanding even when they didn&#8217;t set out to look for it.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<p class=\"wp-block-paragraph\">The other news last week was of course Trump and Vance&#8217;s total capitulation to Vladimir Putin, their berating of Zelensky in the Oval Office for having the temerity to want the free world to guarantee Ukraine&#8217;s security, as the entire world watched the sad spectacle.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Here&#8217;s the thing. As vehemently as I disagree with it, I feel like I basically understand the anti-Zionist position&#8212;like I&#8217;d even share it, if I had either factual or moral premises wildly different from the ones I have.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Likewise for the anti-abortion position. If I believed that an immaterial soul discontinuously entered the embryo at the moment of conception, I&#8217;d draw many of the same conclusions that the anti-abortion people do draw.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">I don&#8217;t, in any similar way, understand the pro-Putin, anti-Ukraine position that now drives American policy, and nothing I&#8217;ve read from Western Putin apologists has helped me. It just seems like pure &#8220;vice signaling&#8221;&#8212;like siding with evil for being evil, hating good for being good, treating aggression as its own justification like some premodern chieftain, and wanting to see a free country destroyed and subjugated because it&#8217;ll upset people you despise.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">In other words, I can see how anti-Zionists and anti-abortion people, and even UFOlogists and creationists and NAMBLA members, are fighting for truth and justice <em>in their own minds<\/em>.&nbsp; I can even see how pro-Putin Russians are fighting for truth and justice in their own minds &#8230; living, as they do, in a meticulously constructed fantasy world where Zelensky is a satanic Nazi who started the war.  But Western right-wingers like JD Vance and Marco Rubio obviously know better than that; indeed, many of them were <em>saying<\/em> the opposite just a year ago!  So I fail to see how they&#8217;re furthering the cause of good <em>even in their own minds<\/em>.  My disagreement with them is not about facts or morality, but about the even more basic question of whether facts and morality are supposed to drive your decisions at all.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">We could say the same about Trump and Musk dismembering the <a href=\"https:\/\/en.wikipedia.org\/wiki\/President%27s_Emergency_Plan_for_AIDS_Relief\">PEPFAR<\/a> program, and thereby condemning millions of children to die of AIDS.  Not only is there no conceivable moral justification for this; there&#8217;s no justification even from the narrow standpoint of American self-interest, as the program more than paid for itself in goodwill.  Likewise for gutting popular, successful medical research that had been funded by the National Institutes of Health: not &#8220;woke Marxism,&#8221; but, like, clinical trials for new cancer drugs.  The only possible justification for such policies is if you&#8217;re trying to signal to <em>someone<\/em>&#8212;your supporters? your enemies? yourself?&#8212;just how callous and evil you can be.  As they say, &#8220;the cruelty is the point.&#8221;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">In short, when I try my hardest to imagine the mental worlds of Donald Trump or JD Vance or Elon Musk, I imagine something very much like the AI models that were fine-tuned to output insecure code. None of these entities (including the AI models) are <em>always<\/em> evil&#8212;occasionally they even do what I&#8217;d consider the unpopular right thing&#8212;but the evil that&#8217;s there seems totally inexplicable by any internal perception of doing good. It&#8217;s as though, by pushing extremely hard on a single issue (birtherism? gender transition for minors?), someone inadvertently flipped the signs of these men&#8217;s good vs. evil vectors. So now the wires are crossed, and they find themselves siding with Putin against Zelensky and condemning babies to die of AIDS. The fact that the evil is so over-the-top and performative, rather than furtive and Machiavellian, seems like a crucial clue that the internal process looks like asking oneself &#8220;what&#8217;s the most despicable thing I could do in this situation&#8212;the thing that would most fully demonstrate my contempt for the moral standards of Enlightenment civilization?,&#8221; and then doing <em>that<\/em> thing.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Terrifying and depressing as they are, last week&#8217;s events serve as a powerful reminder that identifying the &#8220;good vs. evil&#8221; direction in concept space is only a first step.  One then needs a reliable way to keep the multiplier on &#8220;good&#8221; positive rather than negative.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Last week something world-shaking happened, something that could change the whole trajectory of humanity&#8217;s future. No, not that&#8212;we&#8217;ll get to that later. For now I&#8217;m talking about the \u201cEmergent Misalignment&#8221; paper. A group including Owain Evans (who took my Philosophy and Theoretical Computer Science course in 2011) published what I regard as the most surprising [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"advanced_seo_description":"","jetpack_seo_html_title":"","jetpack_seo_noindex":false,"jetpack_seo_schema_type":"","_jetpack_newsletter_access":"","_jetpack_dont_email_post_to_subs":false,"_jetpack_newsletter_tier_id":0,"_jetpack_memberships_contains_paywalled_content":false,"_jetpack_feature_clip_id":0,"_jetpack_memberships_contains_paid_content":false,"footnotes":"","jetpack_publicize_message":"{title}\n\n{excerpt}\n\n{url}","jetpack_publicize_feature_enabled":true,"jetpack_social_post_already_shared":true,"jetpack_social_options":{"image_generator_settings":{"template":"highway","default_image_id":0,"font":"","enabled":false},"version":2},"_wpas_customize_per_network":false,"jetpack_post_was_ever_published":false},"categories":[31,8],"tags":[],"class_list":["post-8693","post","type-post","status-publish","format-standard","hentry","category-announcements","category-the-fate-of-humanity"],"jetpack_publicize_connections":[],"jetpack_sharing_enabled":true,"jetpack_featured_media_url":"","_links":{"self":[{"href":"https:\/\/scottaaronson.blog\/index.php?rest_route=\/wp\/v2\/posts\/8693","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/scottaaronson.blog\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/scottaaronson.blog\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/scottaaronson.blog\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/scottaaronson.blog\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=8693"}],"version-history":[{"count":7,"href":"https:\/\/scottaaronson.blog\/index.php?rest_route=\/wp\/v2\/posts\/8693\/revisions"}],"predecessor-version":[{"id":8729,"href":"https:\/\/scottaaronson.blog\/index.php?rest_route=\/wp\/v2\/posts\/8693\/revisions\/8729"}],"wp:attachment":[{"href":"https:\/\/scottaaronson.blog\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=8693"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/scottaaronson.blog\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=8693"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/scottaaronson.blog\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=8693"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}