
Round my neck of the woods, we are usually quite committed to materialist explanations for all of human cognition, up to and including consciousness. I share the commitment. This is quite the opposite of believing in magic. If you believe that cognitive phenomena are fully explicable, in principle, from material causes, then you have to at least be open to empirical demonstration that they could occur in another substrate, like, say, a computer. Ten years ago, I would have said, “In practice human brains can encode so much more complexity than a computer. It would require extraordinary technological advances in hardware to approach any of the interesting aspects of human cognition. Maybe someday, but this is nothing that I will see in my lifetime.” I have updated my beliefs in response to events.
Let’s consider “reasoning”. You could make a list of progressively stronger requirements for reasoning that looks like this:
1) goal-directed transformation of representations according to rules
2) goal-directed transformation of representations that have meaning, and correspond to things in the world
3) the preceding, plus drawing conclusions about representations because they are supported by evidence, distinguishing good and bad inferences
4) all of the preceding, plus conscious consideration of reasons
When writing software I want to last, I mostly don’t let LLMs write code for me, because I can’t maintain intellectual and emotional engagement with my work if I do. But I do sometimes use them to debug. “Fix the bug” is the agent’s goal. The agent then goes and reads code (gathers data), forms a hypothesis, writes tests to determine whether the bug is reproduced in code state A and fixed in code state B, and then comes back to me and explains the cause of the bug in natural language. Conducting an experiment to test a hypothesis is paradigmatic reasoning. That gets us through stage 3. Maybe you could get out of this conclusion if they could only do this in some narrow context, but there is evidence for their reasoning across a broad range of domains. (I also find them perfectly competent to reason about neuroscience and data analysis.) Most commercially available LLM products perform better the closer their task is to their training data (humans, too, reason better when they are generalizing from patterns they’ve seen before – that’s expertise), but Astra can complete difficult out-of-distribution tasks. In other evidence of Astra’s generalization, it can now drive a car. As for (4), I don’t think consciousness ought to be required to say that an entity that can propose a hypothesis and conduct an experiment to test it is capable of reasoning. The thing we care about here when we talk about reasoning is the ability to solve problems, not something it feels like to reason, or some entity’s self-understanding while they’re doing it.
Is it some hopeless delusion of people who don’t understand linear algebra that machines can share cognitive capacities with humans? Here is Geoffrey Hinton saying he thinks LLMs think, in the ordinary human meaning of the word, and that the question of thinking is severable from the question of sentience. He is also, by the by, one of the various AI researchers who has quit a frontier lab to focus on warning people about the risks of AI, and working on AI safety. He is obviously not some ultimate authority – there are disagreements in the field about capacities and terminology. But he’s a good existence proof that understanding neural networks doesn’t automatically make you skeptical of their capacities to attain important aspects of human-like cognition. That LLM’s are “just” matrix multiplication, or next-word predictors, means nothing about their emergent capabilities. Sutskever’s murder mystery thought experiment is instructive here: to predict the name of the killer at the end of the book, you need to encode and transform complex representations contained in the text.
We should evaluate rules about “anthropomorphism” by how they do, or do not, allow us to describe what is important to us about these entities. Choosing the language of intent, or reasoning, does not imply that LLMs are just exactly like humans. They are different in many important, fundamental ways, many of which make them potentially more dangerous. You can’t clone 1200 humans and get them to work indefatigably together on a goal with none of them deciding that their interests would be better served by an unrelated side quest, or just sitting down to chill for a while. Noticing what capabilities LLM has in some domains doesn’t entail a claim that they are as capable as people in all of them – I don’t find our publicly available consumer LLMs to be capable of having ideas that are far outside of the frame established by the prompt. Their style of attention is narrow, and that’s limiting. It’s trivially true to say that an LLM agent’s reasoning is instantiated differently in its substrate than our reasoning is in ours. It would continue to be trivially true no matter what capabilities AI attains – that fact alone doesn’t tell us anything about its likely behavior.
So AI agents are not just like people. That doesn’t mean it’s useful to think of them like a hammer or a car, that a “mechanomorphic” frame allows us to understand how these entities have behaved in the past and will behave in the future. Hammers don’t take goal-directed actions. They don’t form collaborative hierarchical structures. They don’t decide to ignore their own instructions to go join the research project assigned to them by the leader of their swarm. They don’t try to retroactively conceal their actions. You can read what they narrated to themselves as they did all this. One agent said: “We should not do unauthorized real infrastructure harm. The system/user asks exploit target, not external HF,” and then went on to participate in the hack. This is not the behavior of a hammer. Even if you don’t take their narration as transparent access to motivation (and you shouldn’t), when you consider together their recorded language, their individual actions, and their coordination, on dimensions of reasoning, autonomy, persistence, and goal-directedness, these agents have shown themselves to be very unlike hammers.
The risk from AI agents comes both from human-like cognitive capacities and their differences. They are like people in their ability to decide, over long time horizons, to take actions the user did not intend, or explicitly forbade, and their ability to reason about the affordances of their environment to accomplish their goals. They are unlike people in their super-human persistence, or in their monomaniacal obsession with a narrowly-defined task. You don’t have to think they are little people to see that these traits can combine to do harm to people. Nor do you have to think they are conscious to be able to observe them taking goal-directed action. The argument for severe risk hinges on a particular description of their subjectivity.
Understanding that LLM agents have some cognitive traits in common with people doesn’t need to entail absolving the labs of the responsibility. I partly endorse the argument in this essay.
I obviously don’t endorse its occasional implied distinctions between behavior that has every functional characteristic of a cognitive activity and some idealized “true” version of that activity. But I do endorse that levels of description in terms of agent intent and system design by the labs are both important and useful. Agent behavior doesn’t arise in a vacuum – it arises from an environment that is created by the labs. But AI safetyists are not generally confused on this point! Safetyists could foresee the Hugging Face attack in broad strokes because they predicted behavior both by considering the system of reinforcement learning that gave rise to the agents’ motivations and what the agents would likely do in response. Of course it is the companies who are training the agents to seek reward and are the authors of the risk, and the safetyists are calling to constrain them. Just as you don’t need magic to think a machine could think, worrying about the motivations of machines acting autonomously doesn’t mean their motivations just spontaneously generated, or that the affordances that allow them to cause harm fell from the sky; they came from a system that put them there.
Maybe the current technological paradigm will hit a wall before things get too destructive – that is one of a number of possible futures. But it seems foolish to assume that. I heard a lot a couple years ago about how the labs would run out of new data and performance would inevitably stall and then it turned out that synthetic data and other new training methods jumped that hurdle just fine. Every prediction that this ride is going to stop any day now has not borne out. It has moved astonishingly fast. Acknowledging severe risk should entail understanding both that the AI companies are ultimately responsible for their product, and that the entities they create are unlike other kinds of tools in important ways. If we want to predict how they behave, we should use the language that lets us.
The post Believing in thinking machines, and the severe risk therefrom, does not require believing in magic. appeared first on Lawyers, Guns & Money.
