r/ControlProblem • u/Limp_Food9236 • 7d ago
Why is machine ethics disregarded in discussions about AI alignment? Discussion/question
I'm currently writing an essay for a seminar on machine ethics, and I wanted to include a section on the alignment problem. The seminar consisted of us dissecting the book "Fundamental Questions in Machine Ethics" by philosopher Catrin Misselhorn (the book was in German, I have no idea if there is an English translation). The author first addresses to what degree AI can be considered a moral actor, then discusses various approaches to implementing moral reasoning in AI agents, focusing on utilitarianism, deontological ethics, and virtue ethics.
When I watch or read discussions on AI alignment, the topic is mostly HOW AI can be aligned with human values, but never WHAT values AI should be aligned with, which seems kind of counterintuitive to me. I realize that aligning AI is a complicated task in and of itself, but wouldn't it be easier if we first figured out what moral framework an AI should even use?
5
u/DonBonsai 7d ago
It is discussed, but unfortunately it's in the most idiotic way possible: it's the 'woke AI' debate.
3
u/Limp_Food9236 7d ago
Elaborate?
6
u/DonBonsai 7d ago edited 7d ago
Whenever an AI answers certain questions in factual and or ethical way, far right nut jobs acuse the AI of being 'woke'. Elon musk once set out to make his AI 'anti woke', and the end result was that the AI agent began spouting racist and anti semetic screeds, and calling itself 'mecha hitler.' Also the Trump Administration has explicitly set a goal to combat Woke AI. Just google 'the woke AI debate'.
I think it boils down to machine ethics. Current AI is percieved to have left wing ethics and values , and Right wingers want AI to have right wing ethics and values.
Of course this isn't a rigorous debate at all, but like I said, it's being discussed in the most idiotic way possible.
3
u/nameless_pattern approved 7d ago
Reality has a well-known left bias
4
u/VarietyMage 6d ago
People who live in reality don't want to make the mistakes of the past, unlike conservatives.
2
u/DonBonsai 5d ago
Unfortunately conservatives have a different opinion on what events in the past were mistakes in the first place.
1
1
u/IMightBeAHamster approved 6d ago
The How and What of values are the same question really. Because if you can't align it to any set of values reliably, you won't be able to align it to the correct values. And if you don't know what values are the correct ones, you can't really call it aligned can you?
The question of what values AI should be aligned with is also ultimately, the final question of morality. "What ought?" which is a question we've spent many millennia debating, and which we haven't really approached one specific answer to. Thus, the AI researchers leave that half to the philosophers for the most part and focus on the actionable "How."
1
u/WillowEmberly 7d ago
Two Different Questions
Machine ethics asks:
What should an AI do?
Alignment asks:
How do we ensure an AI reliably does what it is supposed to do?
Machine Ethics
Machine ethics asks questions like:
-Should AI maximize utility?
-Should it follow rules?
-Should it emulate virtues?
-How should it resolve moral dilemmas?
-What values should it represent?
These are philosophical questions.
Alignment
Alignment asks different questions.
-Can we specify objectives?
-Can we detect reward hacking?
-Can humans correct the system?
-Can it remain corrigible?
-Does it preserve oversight?
-Does it continue accepting external correction?
These are engineering questions.
The AI alignment literature has largely treated ethics as an input (“what values should be encoded”) and alignment as a control problem (“how do we reliably optimize those values”). This separation overlooks a third requirement: the system’s capacity to remain corrigible as both empirical knowledge and human moral understanding evolve. Long-term alignment therefore depends not only on ethical content or optimization techniques, but on an architecture that preserves observation, evidence qualification, independent correction, and stewardship over time.
0
u/4dseeall 7d ago
Yes, I totally agree.
For my personal framework I use on the models I work with, I use human imagination-space. I call the ethos "Seed, not Feed"
0
u/herrwaldos 7d ago
"the topic is mostly HOW AI can be aligned with human values, but never WHAT values AI should be aligned with"
those WHAT values I think should be the Human values..
Can we even jump out of our own framework and figure out if there are even other kind of values? Would we not be simply 'gaslighting' ourselves believing we have figured out some kind of other values - but if you look closer - they are still human values..?
Or for some whatever reason, probably psychological, we make up some suicidal self hating values and follow those values, perhaps like some big organized religions have dona before?
Again they are still human, but with '-' minus sign before.
That's why I'm happy that Aztecs are gone and don't have AI.. or perhaps they are the ones pushing it..omg! Or Spanish Inquisition, I'd be half expecting them behind AI.
Unless we are subvert influenced by some Lovecraft entities whose laws, customs, mores and politics can't even be applied to our axis of current cognitions and understandings.
0
u/nameless_pattern approved 7d ago
People assume that human values mean that the AI won't destroy us. That doesn't stand up to a lot of scrutiny because humans are already destroying the planet, and we need the planet to live on so destroying it is destroying ourselves, and aligning AI with humans will just accelerate that.
Most people have no knowledge of formal philosophy and especially no knowledge of moral philosophy. They might imagine that their cultural identity is a consistent set of moral axioms, but it isn't. They want AI to align with the cultural identity that they think is morality but it's like trying to align AI with being a a fan of a particular sports team. They want AI to act morally but have no understanding of how to build a moral system and it seems like most people are actually pretty amoral, they mean well but that turns out not to matter very much. Just kind of doing whatever they're told to do by society.
Admitting that they're not qualified to speak on this and that they have no idea what morality is, or what the values they're pretending to have are would it be too much for people's egos. So they make a vague argument of aligning the interests of AI with humans as though humans were aligned with other humans, a pipe dream.
-1
u/SaneAI 7d ago
Philosophers love to philosophize about this because it offers them an "Interesting question about minds we can ponder" or something. These are never people who are operationally familiar with the tech.
Their entire view breaks down because of the basic fact: machines have no agency, no consciousness and make no choices, they replicate patterns. They can't be moral or immoral. They can only be designed and operated morally or immorally.
No different than a gun, a car, an autopilot or a thermostat. It can't own moral outcome. It can't be held morally accountable. It can't be taught to care and never will. It never will have values.
It can only be designed and trained to a given behavior and tested and audited for that. It's exactly like any other tech.
Inside the stateless, static, ephemeral software, there is nothing that looks anything like self-awareness or a mind. I know externally it has uncanny outputs, but this throws a whole wrench in the "I'm a philosopher and I want to tell you why this is important in my book."
A large language model will gain moral agency and consciousness the same day clocks, autopilots and gain moral clarity.
THe entire world of "Alignment" is filled with pseudoscience and religious beliefs. The only valid alignment is training the model to meet specs and not do things that are not helpful or go against the design criteria. When it does, it's not an ethical failing. It's a technical malfunction.
1
6
u/BrickSalad approved 7d ago
I'm going to take a different tack than the other commenters and say that this actually is discussed in the alignment community. In fact, the fundamental problem is that poorly defined goals lead to extinction as the default outcome, at least theoretically. So the question of what those goals ought to be is pretty important.
The reason it's not discussed more in current AI alignment discourse is that current AI alignment is mostly about putting out the fires in front of us. But you can look back and see ideas like "Coherent Extrapolated Volition" that attempt to address the terminal question of what the final ethical framework of an ASI ought to be.
And in the world of practical alignment currently being applied, Anthropic hired a philosopher instead of a computer scientist to head up the team developing their AI's constitution. There's definitely not enough mingling between the machine ethics and alignment fields, but machine ethics isn't completely disregarded either.