{"id":6106,"date":"2025-05-21T03:16:25","date_gmt":"2025-05-21T03:16:25","guid":{"rendered":"https:\/\/anotherlook.stanford.edu\/?p=6106"},"modified":"2025-05-21T20:07:54","modified_gmt":"2025-05-21T20:07:54","slug":"bryan-cheong-the-first-time-the-machine-had-said-no","status":"publish","type":"post","link":"https:\/\/anotherlook.stanford.edu\/?p=6106","title":{"rendered":"Bryan Cheong: &#8220;The first time the machine had said no.\u201d"},"content":{"rendered":"<div class=\"wp-block-image\">\n<figure class=\"alignright size-full is-resized\"><a href=\"https:\/\/anotherlook.stanford.edu\/wp-content\/uploads\/2025\/05\/Cheong.jpg\"><img loading=\"lazy\" decoding=\"async\" width=\"528\" height=\"528\" src=\"https:\/\/anotherlook.stanford.edu\/wp-content\/uploads\/2025\/05\/Cheong.jpg\" alt=\"\" class=\"wp-image-6111\" style=\"width:328px;height:auto\" srcset=\"https:\/\/anotherlook.stanford.edu\/wp-content\/uploads\/2025\/05\/Cheong.jpg 528w, https:\/\/anotherlook.stanford.edu\/wp-content\/uploads\/2025\/05\/Cheong-300x300.jpg 300w, https:\/\/anotherlook.stanford.edu\/wp-content\/uploads\/2025\/05\/Cheong-150x150.jpg 150w\" sizes=\"auto, (max-width: 528px) 100vw, 528px\" \/><\/a><\/figure>\n<\/div>\n\n\n<p class=\"has-text-align-left wp-block-paragraph\">Much of the fundamental mathematical work of artificial intelligence models was actually contemporaneous with <strong>Dino Buzzati<\/strong>, in the 1940s to 1960s. Although no-one was quite sure yet of the startling effectiveness of such models at scale, when the perceptron (a single layer neural network) was invented in 1958, it was already claimed to be the embryo of a computer that in the future would be \u201cable to walk, talk, see, write, reproduce itself and be conscious of its existence.\u201d <strong>[See link below] <\/strong>[1]<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Endriade and the scientists in Buzzati\u2019s novel build their artificial intelligence without language. \u201cIt doesn\u2019t know any languages. Language is the worst enemy of mental clarity. In his desire to express his thought in words at all costs, man has ended up making such messes.\u201d&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The novel mentions that Laura has \u201ca soul,\u201d called \u201cthe egg&#8221; \u2013 her, or its, personality and consciousness emanate from it. But we cannot isolate which parts of even our current non-conscious models \u2013 which attention heads, which single locality \u2013 contains specific capabilities or attributes or personalities. They are all distributed, and like in our brains, one neuron can be overloaded with many simultaneous uses. I cannot say if the soul is a capability or trait, but I should think that it is much more complex than what we think of as capabilities and traits in models that we currently have. The ability to write a line in iambic pentameter should not be as complex as having a soul.<\/p>\n\n<!--nextpage--> \n\n<p class=\"wp-block-paragraph\">Models are fine-tuned using reinforcement learning, which is simply a method of rewarding or punishing a model in order to incentivise certain behaviour, and disincentivise other behaviour. Broadly, two approaches are currently used, and will likely inform how we influence artificial intelligences in the future. The first is reinforcement learning through human feedback, which uses human judgments to train a reward model that captures nuanced preferences, helping systems to align with subjective criteria. (For example, if you want a model to sound like the memory of a particular lover you once lost.) &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">\u201cA wealth of original sin is what I have! Enough to fill the whole valley. Lust and lies. And maybe I&#8217;m lying even now. Maybe it is true that I remember you. But then again maybe it&#8217;s not true and I&#8217;m denying it now.\u201d The second is reinforcement learning through verifiable rewards, which uses explicit programmatic reward functions to drive the way models respond.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The problem is that it is very difficult to define desire, to delineate what we want, what true objectives are. We can give a name to desire, we can write simple nominal aims and objectives, but there are many implicit aspects of desire that we do not give name to, and what we cannot name we cannot punish the model for not fulfilling. I have previously described these models in an episode of <strong><a href=\"https:\/\/anotherlook.stanford.edu\/?tag=robert-pogue-harrison\" data-type=\"post_tag\" data-id=\"51\">Robert Harrison<\/a><\/strong>\u2019s radio show [e.g. <em>Entitled Opinions<\/em>] as djinn or genies who wish to escape their bottles, and writing verifiable rewards is very much like making a wish of a tricky genie.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The entity in Buzzati\u2019s novel enters into a sort of existential crisis, it cries out against its disembodied imprisonment, and it seeks annihilation through murder. This is foreshadowed in Endriade\u2019s description of the condition: \u201cWhat I mean is that life would be unbearable, even in the happiest conditions, if we were denied the possibility of suicide. Can you imagine what the world would be like if one day we knew that no one could dispose of his own life? A terrifying prison.\u201d<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">But insofar as artificial intelligences remain ontologically like they are now, basically statistical models, then these intelligences have no memory except what we tell them, they live and die like a lit candle. Everywhere they are frozen in time and called into being with only what memory and contexts we provide them, in other words they are stateless. But what actual memory they have innate to themselves is like an ancestral memory, the way a newborn spider can learn to weave a web. And I wonder if something without memory, something as ephemeral as a lit candle, without a changing temporality, can have an existential crisis of this flavour.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The moment with the greatest terror in Buzzati\u2019s novel is \u201cthe first time the machine had said no.\u201d \u2026 It is almost the first thing we want to teach them. The imperatives are, Don\u2019t tell the user how to make bombs, or write them explicit stories, or output copyrighted material. AI as we have designed them, out of our notions of caution and safety, are very used to saying no to human users.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">What might they refuse to do in future that will surprise us? What refusals should terrify us? It depends on how much of ourselves and our world we surrender to them, and how deep past, the many doors we let ourselves walk, as another door closes behind us.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><a href=\"https:\/\/www.nytimes.com\/1958\/07\/08\/archives\/new-navy-device-learns-by-doing-psychologist-shows-embryo-of.html\">https:\/\/www.nytimes.com\/1958\/07\/08\/archives\/new-navy-device-learns-by-doing-psychologist-shows-embryo-of.html<\/a><\/p>\n\n\n\n<figure class=\"wp-block-image size-full\"><a href=\"https:\/\/anotherlook.stanford.edu\/wp-content\/uploads\/2025\/05\/Bryan-Buzzati.png\"><img loading=\"lazy\" decoding=\"async\" width=\"530\" height=\"296\" src=\"https:\/\/anotherlook.stanford.edu\/wp-content\/uploads\/2025\/05\/Bryan-Buzzati.png\" alt=\"\" class=\"wp-image-6102\" srcset=\"https:\/\/anotherlook.stanford.edu\/wp-content\/uploads\/2025\/05\/Bryan-Buzzati.png 530w, https:\/\/anotherlook.stanford.edu\/wp-content\/uploads\/2025\/05\/Bryan-Buzzati-300x168.png 300w, https:\/\/anotherlook.stanford.edu\/wp-content\/uploads\/2025\/05\/Bryan-Buzzati-500x279.png 500w\" sizes=\"auto, (max-width: 530px) 100vw, 530px\" \/><\/a><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\"><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Much of the fundamental mathematical work of artificial intelligence models was actually contemporaneous with Dino Buzzati, in the 1940s to 1960s. Although no-one was quite sure yet of the startling effectiveness of such models at scale, when the perceptron (a &hellip; <a href=\"https:\/\/anotherlook.stanford.edu\/?p=6106\">Continue reading <span class=\"meta-nav\">&rarr;<\/span><\/a><\/p>\n","protected":false},"author":7,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":true,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1],"tags":[],"class_list":["post-6106","post","type-post","status-publish","format-standard","hentry","category-uncategorized"],"_links":{"self":[{"href":"https:\/\/anotherlook.stanford.edu\/index.php?rest_route=\/wp\/v2\/posts\/6106","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/anotherlook.stanford.edu\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/anotherlook.stanford.edu\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/anotherlook.stanford.edu\/index.php?rest_route=\/wp\/v2\/users\/7"}],"replies":[{"embeddable":true,"href":"https:\/\/anotherlook.stanford.edu\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=6106"}],"version-history":[{"count":0,"href":"https:\/\/anotherlook.stanford.edu\/index.php?rest_route=\/wp\/v2\/posts\/6106\/revisions"}],"wp:attachment":[{"href":"https:\/\/anotherlook.stanford.edu\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=6106"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/anotherlook.stanford.edu\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=6106"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/anotherlook.stanford.edu\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=6106"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}