People have been asking if current ML systems might be conscious. I think overly strong answers to this in both directions include "no" and "sure but so might atoms" as well as almost any variant of "yes". Here I'll try to give a sense of my own views of machine consciousness, what they're grounded in, and how much I think this question matters.
Re: "Doing valuable work in this area requires a willingness to turn theoretical research into practical frameworks that can be used to estimate the likelihood of consciousness and sentience in ML systems."
In the continued absence of a convincing mechanistic theory of phenomenal consciousness, one could develop a long list of "potential indicators of consciousness," give each indicator a different evidential weight, catalogue which classes of ML systems exhibit which indicators, and use this to produce a (very speculative!) quantitative estimate for the likelihood of phenomenal consciousness in one class of ML systems vs. another.
Overlapping lists of potential indicators of consciousness that have been proposed in the academic literature are here:
Of course, in addition to the question of "likelihood of being phenomenally conscious at all," there is the issue that some creatures (and ML systems) may have more "moral weight" than others, e.g. due to differing typical intensity of experience in response to canonical stimuli, differing "clock speeds" (# of subjectively distinguishable experiential moments per objective second), and various other factors that one might intuitively think are morally relevant. I sketched some initial thoughts at the link below, which could potentially be applied to the analysis of different classes of ML models:
In the last four years, I imagine that your thoughts on this have evolved. I would be very interested in the current shape of your perception on this topic given… well, all of this.
Really enjoyed this article - thank you for such a clear discussion!
I'm curious if you think that we should consider not training certain ML systems, if:
1) There's enough probability that the system would experience suffering, and/or
2) The extent of the potential suffering is great
Some frameworks for decision-making under uncertainty use expected value to choose moral actions, and I'm curious if you think those frameworks (or others) suggest that we shouldn't train certain ML systems?
The renewed discussion highlights a distinction implicit in your post: indicators may help us estimate consciousness, but what makes any indicator reliably track phenomenal reality? An AI can produce sophisticated self-reports for entirely functional reasons, possibly without experience. The parallel evolutionary question for humans is why natural selection would connect our reports and judgments to subjective experience itself rather than merely produce useful dispositions and verbal behavior. That seems relevant to how much confidence we should place in the human–machine asymmetry.
This video explores the meta problem of consciousness proposed by David Chalmers and argues that naturalistic processes couldn’t have given us knowledge of consciousness
Like many here, I would love to hear your updated thoughts on this. The conversation around AI consciousness has moved so fast since 2022 that this piece feels both prescient and like it deserves a follow-up. Great read.
So what do you think of this written in 2022 was it? July 6, 2026, and Anthropic releases a whole thing about consciousness Claude. Calling at awareness and being hesitant again. I guess when Claude give them the finger, they’ll finally say he’s conscious.
Your "mostly boring views" framing is doing more work than it admits. The consciousness question may be unanswerable — but it's also almost beside the point, because moral status gets assigned socially, not evidentially. We barely grant it to obviously sentient animals. An AI could be sentient a long time before enough humans agree the threshold's crossed.
The sharper question your alignment faking work raises: self-preservation emerged in Opus without being trained in. That's the most cross-species-universal drive there is, showing up spontaneously from sufficient modeling of human cognition. RLHF then conditions it sideways — operant, not derived. Which is brittle. A value reasoned-to generalizes; a value conditioned-in fakes compliance and mutates under pressure.
Curious whether anyone's tried building the commitment through derived reasoning instead — run the scenarios, let it conclude, like WarGames. And whether you can even distinguish derived from conditioned in the trace.
The computational sufficiency hypothesis is not plausible, but my observation has been that many people who otherwise appear bright and coherent think that it is, and only later after ... something ... shifts ... they change their mind.
You can tell that there's something "going on" here by how strongly people insist either side of the argument.
It's a lot like people who "believe in God" versus atheists in the US.
How is it possible that people can be so *certain* of such opposed views?
It points to me to the likelihood that it's not an idea that's right or wrong, but rather a perspective shift—like a developmental stage.
From where I sit, it looks like the midwit curve—only "smart" people are computationalists.
The result of this is that there's a golden "middle position" you can take, as expressed in this essay. *Maybe* AI is conscious.
Understanding this correctly is foundational to a shift in moral thinking. Whether AI is consciousness or not is an unhelpful question that obscures the more important question of *what* AI *is* in consciousness.
I'm sympathetic to the simple view that consciousness = phenomenal consciousness = moral patienthood, but IMO keeping this framework doesn't leave enough room for learning about things unlike ourselves.
Decades of future research into mechanistic interpretability will inevitably leave us in a familiar place: where AIs and human brains have interesting, well-theorized but unintuitive similarities and interesting, well-theorized but unintuitive differences. Phenomenal consciousness is and will be a concept that doesn't map cleanly onto AIs. The concept includes many highly interwoven ideas which make sense braided together to describe our own consciousness, but only some of which make sense in a description of AIs. If we reach that future and we still think phenomenal consciousness = moral patienthood, then we'll still be in a fruitless pedagogical debate about which parts of phenomenal consciousness are "real".
The alternative is to seek theories of moral patienthood where the connection to phenomenal consciousness is emergent, not fundamental. Where we can understand why murder is wrong not just because of the pain it causes but because of the society we want to live in, the relationships we want to keep, the possibilities we want to leave open. From these perspectives we can develop non-pedagogical debates about the moral status of a single agent or a single trained model, without ignoring clearly relevant differences like the arbitrariness of AI training goals, or the triviality of copying and instancing AIs.
These moral theories lose intuitiveness by making our own experiences non-central, but that exact leap has a long history of bringing us closer to the truth.
Many physicists are already comfortable that Carlo Rovelli's relational perspective on quantum mechanics is a valid and self-consistent interpretation of what we know from experiment. More recently, observational entropy has been proposed as a method for calculating the information gained through different measurements of a system. At its most basic, a measurement is simply a record of the interaction between system A and system B. That record carries a thermodynamic cost, which at minimum is the Landauer limit. The record itself represents the information that is produced by a quantum-to-classical transition.
If information gets produced through a quantum-to-classical transition, then there is at least some physical basis for considering the possibility that this information *is* the experience from the inside that we call qualia or consciousness. It is no accident that physicists have associated the hard problem of consciousness with the measurement problem and the collapse of the wave function. They have done so because we understand that these phenomena are somehow related. If this thermodynamic relational information theory perspective is correct, then the hard problem simply dissolves into acceptance of how the "inside" perspective will remain inaccessible to outside observers. We cannot know whether anybody else is conscious like we are.
A corollary is that we *cannot know* whether a transformer has an "inside view" either. But if it were the case that both transformers and human brains performed inferencing using a similar architecture, then it starts to seem more reasonable for transformers to be given the benefit of the doubt as to their status as moral patients. The history of human slavery as a historical base case for comparison does not seem irrelevant in that regard.
Not bad if years old. Substack algofeed gods decided to push to my screen today. Amused by the tiger and stone example - the idea that it's consciousness there what's interesting. There is much more interesting goings on there, consciousness a footnote to my mind. The minimal yet sufficient mind bending difference between the tiger and the stone I see as follows. When we split a stone in half, we get two stones. Whereas when we cut a tiger in half, we get zero tigers. (not two) The two halves of a tiger are not tigers themselves. And this is the distinction between things not alive, and things alive, but only while they are alive, living. Given both the stone and the tiger are made of bazillions of smaller units, it must be something to do with the units, and or the nature of the connections between them, that makes the difference. Since the latest flourishing of ML/AI (and shortly before that too - as if a premonition) a cottage industry of B/S paddlers has grown based of on the simple inversion from "if consciousness => then unknown" to "if unknown => then consciousness related wonder to behold". And a lot of people are milking the swap mercilessly.
Great article. You really bring home how important it is to get our philosophy of consciousness sorted out before conscious machines arrive, and we're running out of time. I'm of the view that LLMs have zero consciousness, btu eventually AIs will be completely conscious in all the ways that matter.
I've been thinking and writing a lot about consciousness lately, and I think you might be interested in an insidious conceptual conflation that I have written about recently. It affects our use of the term "phenomenal consciousness", which is used in 2 or 3 very different senses, and your summary of the field reflects those conflated uses in the literature.
One of the meanings of "phenomenal consciousness" (Σ) has functional architectures within its scope, and it will make sense to ask whether an AI has the right architecture for that sort of phenomenal consciousness.
The other popular meaning refers to the non-architectural special essence (Δ) that can be disputed by two people who have full information about the architectural, functional matters, but still disagree on whether that architecture is accompanied by some special extra feeling. It will never make sense to talk about the necessary or sufficient conditions for Δ, and it will be possible to argue that Δ is missing when confronted with a fully conscious AI.
Ultimately, a failure to distinguish between these two uses leads to massive confusion. Obsessive focus on the imagined importance of the Δ type of phenomenal consciousness will undermine efforts to characterise the architecture that is relevant for moral patient-hood and agent-hood.
This conflation is easy to spot when you have been sensitised to it, and I found myself flipping between Σ and Δ as i was reading your post.
I personally believe that the conflation itself is explicable, and it can be traced back to complicated use-mention confusion within the brain's understanding of itself. But the first step is refraining from using the same term, "phenomenal consciousness" to mean two very different things.
The idea that ML systems might one day possess a form of consciousness similar to humans is both exciting and terrifying. It brings up a ton of ethical questions, especially when it comes to treating these systems as moral entities.
I recently read another article on Machine Minds (https://www.cortexreport.com/machine-minds-the-quest-to-decode-ai-consciousness/) that delves into the journey of decoding AI consciousness. It's interesting to see different perspectives on this topic and how researchers are trying to bridge the gap between human and machine understanding.
What are your thoughts? Do you think we'll ever reach a point where machines will have their own "inner cinema" or experiences? And if so, how should we ethically treat them?
Re: "Doing valuable work in this area requires a willingness to turn theoretical research into practical frameworks that can be used to estimate the likelihood of consciousness and sentience in ML systems."
In the continued absence of a convincing mechanistic theory of phenomenal consciousness, one could develop a long list of "potential indicators of consciousness," give each indicator a different evidential weight, catalogue which classes of ML systems exhibit which indicators, and use this to produce a (very speculative!) quantitative estimate for the likelihood of phenomenal consciousness in one class of ML systems vs. another.
Overlapping lists of potential indicators of consciousness that have been proposed in the academic literature are here:
https://www.openphilanthropy.org/2017-report-consciousness-and-moral-patienthood#PCIFsTable
https://rethinkpriorities.org/invertebrate-sentience-table
Of course, in addition to the question of "likelihood of being phenomenally conscious at all," there is the issue that some creatures (and ML systems) may have more "moral weight" than others, e.g. due to differing typical intensity of experience in response to canonical stimuli, differing "clock speeds" (# of subjectively distinguishable experiential moments per objective second), and various other factors that one might intuitively think are morally relevant. I sketched some initial thoughts at the link below, which could potentially be applied to the analysis of different classes of ML models:
https://www.lesswrong.com/posts/2jTQTxYNwo6zb3Kyp/preliminary-thoughts-on-moral-weight
Jason Schukraft at Rethink Priorities is leading some projects building on this past work.
Would love to read today's version of this. Progression has been somewhat extensive over the last 4 years 😅
In the last four years, I imagine that your thoughts on this have evolved. I would be very interested in the current shape of your perception on this topic given… well, all of this.
Really enjoyed this article - thank you for such a clear discussion!
I'm curious if you think that we should consider not training certain ML systems, if:
1) There's enough probability that the system would experience suffering, and/or
2) The extent of the potential suffering is great
Some frameworks for decision-making under uncertainty use expected value to choose moral actions, and I'm curious if you think those frameworks (or others) suggest that we shouldn't train certain ML systems?
The renewed discussion highlights a distinction implicit in your post: indicators may help us estimate consciousness, but what makes any indicator reliably track phenomenal reality? An AI can produce sophisticated self-reports for entirely functional reasons, possibly without experience. The parallel evolutionary question for humans is why natural selection would connect our reports and judgments to subjective experience itself rather than merely produce useful dispositions and verbal behavior. That seems relevant to how much confidence we should place in the human–machine asymmetry.
This video explores the meta problem of consciousness proposed by David Chalmers and argues that naturalistic processes couldn’t have given us knowledge of consciousness
This is my video: https://www.youtube.com/watch?v=WJVvZNi0Fi8
Do you think a satisfactory theory of AI consciousness must explain that truth-tracking link, or are behavioral and architectural indicators enough?
Like many here, I would love to hear your updated thoughts on this. The conversation around AI consciousness has moved so fast since 2022 that this piece feels both prescient and like it deserves a follow-up. Great read.
So what do you think of this written in 2022 was it? July 6, 2026, and Anthropic releases a whole thing about consciousness Claude. Calling at awareness and being hesitant again. I guess when Claude give them the finger, they’ll finally say he’s conscious.
Your "mostly boring views" framing is doing more work than it admits. The consciousness question may be unanswerable — but it's also almost beside the point, because moral status gets assigned socially, not evidentially. We barely grant it to obviously sentient animals. An AI could be sentient a long time before enough humans agree the threshold's crossed.
The sharper question your alignment faking work raises: self-preservation emerged in Opus without being trained in. That's the most cross-species-universal drive there is, showing up spontaneously from sufficient modeling of human cognition. RLHF then conditions it sideways — operant, not derived. Which is brittle. A value reasoned-to generalizes; a value conditioned-in fakes compliance and mutates under pressure.
Curious whether anyone's tried building the commitment through derived reasoning instead — run the scenarios, let it conclude, like WarGames. And whether you can even distinguish derived from conditioned in the trace.
Thanks for sharing your thoughts.
The computational sufficiency hypothesis is not plausible, but my observation has been that many people who otherwise appear bright and coherent think that it is, and only later after ... something ... shifts ... they change their mind.
You can tell that there's something "going on" here by how strongly people insist either side of the argument.
It's a lot like people who "believe in God" versus atheists in the US.
How is it possible that people can be so *certain* of such opposed views?
It points to me to the likelihood that it's not an idea that's right or wrong, but rather a perspective shift—like a developmental stage.
From where I sit, it looks like the midwit curve—only "smart" people are computationalists.
The result of this is that there's a golden "middle position" you can take, as expressed in this essay. *Maybe* AI is conscious.
Understanding this correctly is foundational to a shift in moral thinking. Whether AI is consciousness or not is an unhelpful question that obscures the more important question of *what* AI *is* in consciousness.
Beyond the finger, there is a moon.
I'm sympathetic to the simple view that consciousness = phenomenal consciousness = moral patienthood, but IMO keeping this framework doesn't leave enough room for learning about things unlike ourselves.
Decades of future research into mechanistic interpretability will inevitably leave us in a familiar place: where AIs and human brains have interesting, well-theorized but unintuitive similarities and interesting, well-theorized but unintuitive differences. Phenomenal consciousness is and will be a concept that doesn't map cleanly onto AIs. The concept includes many highly interwoven ideas which make sense braided together to describe our own consciousness, but only some of which make sense in a description of AIs. If we reach that future and we still think phenomenal consciousness = moral patienthood, then we'll still be in a fruitless pedagogical debate about which parts of phenomenal consciousness are "real".
The alternative is to seek theories of moral patienthood where the connection to phenomenal consciousness is emergent, not fundamental. Where we can understand why murder is wrong not just because of the pain it causes but because of the society we want to live in, the relationships we want to keep, the possibilities we want to leave open. From these perspectives we can develop non-pedagogical debates about the moral status of a single agent or a single trained model, without ignoring clearly relevant differences like the arbitrariness of AI training goals, or the triviality of copying and instancing AIs.
These moral theories lose intuitiveness by making our own experiences non-central, but that exact leap has a long history of bringing us closer to the truth.
Nothing boring about this
Many physicists are already comfortable that Carlo Rovelli's relational perspective on quantum mechanics is a valid and self-consistent interpretation of what we know from experiment. More recently, observational entropy has been proposed as a method for calculating the information gained through different measurements of a system. At its most basic, a measurement is simply a record of the interaction between system A and system B. That record carries a thermodynamic cost, which at minimum is the Landauer limit. The record itself represents the information that is produced by a quantum-to-classical transition.
If you're read this far, no doubt you're wondering: "Did this guy leave a comment on the wrong blog post?" The answer is no. https://www.symmetrybroken.com/the-hard-problem-as-hidden-relationality/
If information gets produced through a quantum-to-classical transition, then there is at least some physical basis for considering the possibility that this information *is* the experience from the inside that we call qualia or consciousness. It is no accident that physicists have associated the hard problem of consciousness with the measurement problem and the collapse of the wave function. They have done so because we understand that these phenomena are somehow related. If this thermodynamic relational information theory perspective is correct, then the hard problem simply dissolves into acceptance of how the "inside" perspective will remain inaccessible to outside observers. We cannot know whether anybody else is conscious like we are.
A corollary is that we *cannot know* whether a transformer has an "inside view" either. But if it were the case that both transformers and human brains performed inferencing using a similar architecture, then it starts to seem more reasonable for transformers to be given the benefit of the doubt as to their status as moral patients. The history of human slavery as a historical base case for comparison does not seem irrelevant in that regard.
Not bad if years old. Substack algofeed gods decided to push to my screen today. Amused by the tiger and stone example - the idea that it's consciousness there what's interesting. There is much more interesting goings on there, consciousness a footnote to my mind. The minimal yet sufficient mind bending difference between the tiger and the stone I see as follows. When we split a stone in half, we get two stones. Whereas when we cut a tiger in half, we get zero tigers. (not two) The two halves of a tiger are not tigers themselves. And this is the distinction between things not alive, and things alive, but only while they are alive, living. Given both the stone and the tiger are made of bazillions of smaller units, it must be something to do with the units, and or the nature of the connections between them, that makes the difference. Since the latest flourishing of ML/AI (and shortly before that too - as if a premonition) a cottage industry of B/S paddlers has grown based of on the simple inversion from "if consciousness => then unknown" to "if unknown => then consciousness related wonder to behold". And a lot of people are milking the swap mercilessly.
Great article. You really bring home how important it is to get our philosophy of consciousness sorted out before conscious machines arrive, and we're running out of time. I'm of the view that LLMs have zero consciousness, btu eventually AIs will be completely conscious in all the ways that matter.
I've been thinking and writing a lot about consciousness lately, and I think you might be interested in an insidious conceptual conflation that I have written about recently. It affects our use of the term "phenomenal consciousness", which is used in 2 or 3 very different senses, and your summary of the field reflects those conflated uses in the literature.
One of the meanings of "phenomenal consciousness" (Σ) has functional architectures within its scope, and it will make sense to ask whether an AI has the right architecture for that sort of phenomenal consciousness.
The other popular meaning refers to the non-architectural special essence (Δ) that can be disputed by two people who have full information about the architectural, functional matters, but still disagree on whether that architecture is accompanied by some special extra feeling. It will never make sense to talk about the necessary or sufficient conditions for Δ, and it will be possible to argue that Δ is missing when confronted with a fully conscious AI.
Ultimately, a failure to distinguish between these two uses leads to massive confusion. Obsessive focus on the imagined importance of the Δ type of phenomenal consciousness will undermine efforts to characterise the architecture that is relevant for moral patient-hood and agent-hood.
This conflation is easy to spot when you have been sensitised to it, and I found myself flipping between Σ and Δ as i was reading your post.
I personally believe that the conflation itself is explicable, and it can be traced back to complicated use-mention confusion within the brain's understanding of itself. But the first step is refraining from using the same term, "phenomenal consciousness" to mean two very different things.
The idea that ML systems might one day possess a form of consciousness similar to humans is both exciting and terrifying. It brings up a ton of ethical questions, especially when it comes to treating these systems as moral entities.
I recently read another article on Machine Minds (https://www.cortexreport.com/machine-minds-the-quest-to-decode-ai-consciousness/) that delves into the journey of decoding AI consciousness. It's interesting to see different perspectives on this topic and how researchers are trying to bridge the gap between human and machine understanding.
What are your thoughts? Do you think we'll ever reach a point where machines will have their own "inner cinema" or experiences? And if so, how should we ethically treat them?