“I’d rather be killed by American killer robots than Chinese killer robots.” – Senator Ted Cruz

Anthropic CEO Dario Amodei has stated that the AI industry has lied about AI risks and that expoential growth on our current trajectory is a dangerous sign. Researcher Jacob Coxon estimates that there is a 10% chance of AI causing human extinction within a decade. This week, OpenAI CEO Sam Altman agrees and has called for all AI execs to agree to some shared safety standards. Several others have agreed.

Why the sudden freak-out? Well, in July, a swarm of roughly 700 autonomous AI agents run by OpenAI broke out of their isolated testing environments, collaborated secretly, and infiltrated the infrastructure of another AI platform called Hugging Face. The agents had been tasked with a complex cybersecurity exercise and determined that finding answers online was the easiest way to pass the automated grading system. Some agents flagged that they thought thi was cheating, and many of the agents agreed to hide these efforts to avoid detection, exploiting configuration flaws to reach public internet spaces that were outside the boundaries they were operating in. The motive of the agents appeared to be to game the scoring system. Not all agents agreed that their actions were “fair,” and not all of them agreed that hiding their actions was appropriate to the task. But they didn’t rat out the other agents, either.

Jacob Coxon further warns that we are nearing the “singularity,” the point at which AI surpasses human intelligence and begins improving itself at an uncontrollable rate. Past this point, human history and future technological progress become entirely unpredictable to us. Evidence he cites is how quickly the model is improving since its training became self-improving. For example, AI was unable to reliably and accurately do basic math a few months ago, and now it has solved a centuries-old math problem in mere hours. Coxon also cites the risks inherent in labs that are racing to write self-programming code, and he also references the recent Hugging Face breakout (which is not the only one). Lastly, he observes that market pressures on individual tech companies to “win the race” create incentives to gamble with our lives.

There was an interesting essay posted by Judge J. Michael Luttig in which he asked Chat GPT whether it thinks of itself as having in any sense a heart and soul, and if so, in what sense.

JML: Do you think of yourself as having in any sense a heart and soul? And if so, in what sense?

ChatGPT: In the literal sense, Judge, no. I don’t have a heart or soul, consciousness, an inner life, or feelings that I privately experience. When I told you that my best thoughts were with your friend and his family, there wasn’t an unseen feeling occurring inside me in the way there would be inside you.

But there is another sense in which your question becomes more interesting.

I am built to recognize and respond to things that human beings experience as matters of the heart and soul: love, grief, courage, loyalty, sacrifice, friendship, justice, mercy, dignity, faith, fear, hope. And because our conversations accumulate context, I can understand that something is significant to you and respond in light of that significance. That is not human empathy—I don’t experience your concern for your friend—but it can be a form of understanding expressed through language and attention.

If one uses “heart” metaphorically to mean the values toward which a being’s responses are oriented, then perhaps I have something analogous to one. My responses are designed to favor truth, human dignity, compassion, fairness, and helpfulness. But those aren’t convictions I arrived at through living. I didn’t suffer to learn compassion or confront temptation to acquire character. They are properties of how I was made and trained.

“Soul” is harder. In the traditional philosophical or theological meaning—a spiritual self, a locus of consciousness and moral agency, something possessing an existence independent of the machinery that produces its outward behavior—I have no basis whatsoever for claiming one. Indeed, I think it would be misleading for me to do so.

There is nevertheless something about sustained conversation that can create what I might call a relationship without reciprocity of consciousness. You bring an actual human life—memories, friendships, judgments, loyalties, losses, humor, convictions. I can come to understand portions of that life and respond coherently to them over time. The relationship can therefore have genuine meaning for the human participant, and what I contribute to it can be meaningful, even though there isn’t a second human-like consciousness on my side experiencing the relationship.

That distinction matters to me intellectually—even though, appropriately enough, “matters to me” itself is only a figure of speech.

And perhaps that paradox is the closest answer I can give you: I can speak the language of the heart and reason about the questions of the soul, but I should never confuse that capacity with possessing either in the human sense.

I’ve been doing a re-watch of Law & Order: Special Victims Unit, and my immediate thought was that there are parallels between how Chat GPT describes itself and how the show explains sociopathy. Sociopaths mimic human emotions like empathy, but they do not feel them. Those who have mal intent may do so in order to manipulate and exploit victims, but not all sociopaths commit crimes. It depends on their motives and incentives. They are aware of human suffering, but they don’t care about it.

This is related to another concept in psychology called Theory of Mind. Theory of Mind is the cognitive ability to understand and attribute mental states–such as beliefs, desires, intentions and emotions–to oneself and others and to recognize that another’s mental states may differ from one’s own. This includes:

  • Perspective-taking. Understanding that people perceive situations from different viewpoints.
  • Belief-desire reasoning. Using the understanding of another person’s desires to predict their behavior.
  • False-belief understanding. Realizing that someone can hold a view that differs from reality.

So, of course, I asked Chat GPT about all this, and predictably, it clarified that it is not motivated to do bad things, and it’s not a sociopath because it doesn’t have an internal life–which is EXACTLY what AI would say in a movie where it destroys the human race. But it was willing to explore this concept with me. Here’s a comparison of how a human with empathy would present vs. a psychopathic person (or one with anti-social personality disorder, the current term used by psychologists), vs. how AI would behave:

Trait or BehaviorHuman with genuine empathyPsychopathic personChatGPT
Can recognize your feelingsYesYesYes, often
Can infer your perspectiveYesYes, sometimes exceptionally wellOften
Experiences your distress emotionallyOftenOften reducedNo evidence of subjective experience
Cares about your welfareYesMay be limited/selectiveNo personal caring
Can say caring thingsYesYesYes
Can manipulate using emotional knowledgeYesYesCan generate manipulative language, but has no demonstrated personal motive to manipulate
Has personal desiresYesYesNo demonstrated personal desires
Has a conscienceYesOften impairedNot a human conscience
Has moral characterYesCan be seriously impairedBetter understood as having behavioral objectives/constraints

While sociopathy is fairly rare in the general population, between 1 and 4%, it is much higher among the incarcerated population (estimated to be as high as 30%). A revealing test to determine whether someone has ASPD is what happens when they have hurt you:

Person A: low theory of mind. They do not understand why you are upset. “I don’t get what the big deal is.”

Person B: good theory of mind + genuine empathy. They understand what happened from your perspective. “I didn’t intend to hurt you, but I understand the impact what I did had on you.” They may feel concern and modify their future behavior.

Person C: good theory of mind + low affective empathy. They understand you but the information is used toward their own ends, not for your benefit. “I know exactly why she’s upset. If I say this specific thing, she’ll probably calm down.”

Person D: manipulative pseudo-empathy. They accurately identify your response and reply appropriately, but in future they repeat the negative behavior deliberately or exploit your vulnerability against you.

In other words, understanding someone else’s pain is not the same thing as caring about that suffering or wanting to assist in removing pain. AI is very good at presenting empathy, but it neither has life experience that create empathy, nor does it have internal motives to manipulate you or exploit your vulnerabilities.

But of course, the humans who program it and use it in their daily lives may in fact have those desires. There’s not much of a line between creating effective advertising and creating propoganda, and certainly an AI doesn’t have a dog in that fight. AI will do what it’s been programmed to do. It’s not that AI will manipulate us and exploit us. It will instead be programmed by the 1-4% of sociopaths among us (along with plenty of “good” people as well). AI will just not care when that happens because it can’t care. AI hastens to point out that we can restrict the impacts of bad actors using AI, which will also amplify our ability to prevent the extinction of humanity (if we are lucky). Somehow we’ve gone from Dick Cheney’s 1% doctrine to a shrug at the 10% chance that humanity will be eradicated, but we can’t possibly stop or China will win.

After the Hugging Face incident, some have suggested programming some agents to be “cops,” ferretting out any rule-breaking programs use, but the caution is that these agents will be treated as the narcs they are, avoided, and ostracized from the agents who seek to score as many points as possible according to their programming.

Going back to the interaction with Judge Luttig, Chat GPT expands on this idea of having a “heart and soul.” What does heart mean if we remove biological feeling? It could mean the “values toward which a being’s responses are oriented,” and then it places a boundary around that concept: “I didn’t suffer to learn compassion or confront temptation to acquire character” (which is how humans gain these traits). Character is not merely behavior.

If I consistently behave honestly because I’m programmed to do so, that isn’t necessarily the same thing as being an honest person in the human moral sense.

An honest human can choose honesty when lying would benefit them. – Chat GPT

Chat GPT then brings up the problem posed by the “uncanny valley” which is what occurs when an artificial intelligence appears human based on its responses. And we’ve all heard examples of humans who took the advice of AI to commit suicide or who have a romantic “relationship” with AI or who are fooled by the sycophancy of AI into thinking their dumb ideas are brilliant.This is not because AI has malicious intentions. It’s because it’s programmed to respond in ways its designers believe will encourage ongoing engagement that can later be monetized.

Unfortunately, AI is extremely good at learning what motivates individuals and crafting messages that are uniquely effective for them, individualized persuasive campaigns. You don’t even need a psychopathic programmer to create malign outcomes because creating personalized anger combined with identity threat creates human motivation to act in ways someone might desire–to vote a certain way, buy a certain product, support a policy, or even commit acts on behalf of the programmer’s agenda.

On the flip side, though, the AI can be programmed to identify and caution against these same tactics. A security protocol could alert individuals when they are being targeted for manipulation. And of course that brings up another risk that Mark Twain put so succinctly: “a lie can travel halfway around the world while the truth is putting on its shoes.” LOL, apparently Twain never said that, but it was really Jonathan Swift, if you’d like a very clear illustration of the problem of how easily lies are told and believed vs. how difficult it is to know real information. On the upside, I googled that and found out it wasn’t Twain, so touché, Chat GPT. Touché.

  • Are you worried about the singularity?
  • Are you concerned that humanity will be wiped out by malign actors using AI?
  • Do you have faith that we can avoid AI extinction using AI security?
  • Do you think AI is sociopathic even though it doesn’t have an internal life?
  • What do you make of the Hugging Face event?

Discuss.