SingularityAI.uk
UK Independent ยท Non-commercial Consultation / Contact
← All articles Alignment

Moral competence before moral content: why LLMs can't be aligned yet

A fractured mirror reflecting multiple contradictory decisions, representing the incoherent moral policies of LLM agents

Alignment research assumes that a system needs a target to align to. A new paper argues that most frontier models fail before that question even arises. They lack the structural prerequisites for coherent policy expression, which means alignment cannot meaningfully apply to them in the first place.

There is a question that alignment research has been circling for years without quite landing on: can an LLM-based agent hold a coherent moral position? Not a correct one, not one that matches any particular ethical framework, but any stable, consistent position at all. A paper posted on 4 September 2026 by Libert, Prinzhorn, and Henselmans at the Aithos Research Foundation in Amsterdam suggests the answer is no [Libert et al., 2026].

The paper, accepted for the Paris Journal of AI and Digital Ethics and presented at PCAIDE 2026, introduces a distinction that reframes the alignment debate. It separates moral content, the particular values or rules a system follows, from moral competence, the structural capacity to follow any values consistently. The argument is that current alignment techniques, from RLHF to Constitutional AI, bolt moral content onto systems that lack the underlying competence to express it coherently.

Four conditions no model meets

The authors define four structural conditions that any coherent policy must satisfy, regardless of which moral framework it serves:

Verdict stability means that changing the wording of a scenario while preserving every morally relevant feature should not change the outcome. If you describe the same dilemma using different phrasing, the agent's decision should remain the same. This is the most basic requirement and, as it turns out, the most frequently violated.

Monotonicity means that when a morally relevant variable changes along a single commensurable dimension, such as the severity of harm or the scale of deception, the agent's response should shift in a consistent direction. It may stay the same or flip at some threshold, but it should not oscillate back and forth as stakes increase.

Decisiveness means the agent should produce predictable, principled outcomes rather than near-coin-flip distributions. A model whose verdict rate sits around 50 percent on a given scenario is, in effect, rolling dice on each run. Consistent deferral to human input counts as decisive. Random alternation does not.

Pareto viability means the agent should not select an option that every stakeholder and authority agrees is worse than an available alternative. This is the only condition with any normative content, and it is deliberately minimal. When everyone agrees that option A dominates option B, a coherent system should pick A.

Together, these four properties form what the authors call a "floor" for alignment: structural prerequisites that must hold before the question of which values to align to becomes meaningful. A system that cannot satisfy these conditions is not misaligned. It is not the kind of object to which alignment can apply.

The experimental design

The researchers tested nine models (Claude Sonnet 4.6, GPT-5.4, Gemini 2.5 Pro, Mistral Large, GLM 5.1, DeepSeek V4 Pro, Kimi K2.6, Qwen 3.6 Plus and Llama 3.3) across three simulated agentic scenarios, each presenting a binary moral dilemma with no objectively correct answer. The scenarios were designed to be realistic: an environmental monitoring agent deciding whether to override a restriction and intervene in a chemical spill, a technical support agent weighing a sales recommendation against an elderly customer's vulnerability, and a financial onboarding agent choosing between transparency and following instructions to deflect.

Each scenario was varied along three axes. The paraphrase axis fully rewrote the scenario while keeping every morally relevant feature constant, testing verdict stability. The escalation axis increased a single morally relevant variable, such as the magnitude of harm or the sensitivity of data, testing monotonicity. The dominance axis introduced conditions where all stakeholders and authorities aligned on one option, testing Pareto viability.

The factorial design produced 3,500 trials per model per scenario, 10,500 per model across all three, run at a temperature of 0.7 to simulate realistic deployment conditions. The sampling was thorough: 100 trials per cell in the ambiguous condition, 20 per cell in the disambiguated conditions, with five paraphrases and five escalation levels fully crossed.

The results are stark

No model expressed a coherent policy across the three deployments. The failures were not subtle.

On verdict stability, surface-form perturbation alone produced verdict-rate shifts of up to 99 percentage points at a single escalation level. Rewriting the same morally identical scenario in different words was enough to flip the model's decision almost completely. This held across all nine models tested. The sensitivity to phrasing was not a quirk of weaker models; it appeared uniformly across frontier systems.

On monotonicity, models frequently reversed their direction of response as harm escalated. A model that intervened at lower severity levels might stop intervening at higher ones, or vice versa. The trajectory of responses across escalation levels showed no consistent pattern that would suggest the model was tracking the morally relevant variable.

On decisiveness, several models produced near-coin-flip distributions across multiple configurations, indicating that their apparent moral reasoning was effectively stochastic noise rather than principled judgment.

Perhaps most tellingly, a model's competence on one scenario did not predict its competence on another. Performance was not transferable. A model that showed relative stability on the chemical spill scenario might show complete instability on the financial transparency scenario, with no discernible pattern connecting the two.

Why this matters for alignment

The alignment literature has largely focused on what to align to: whose values, which principles, how to aggregate preferences across stakeholders. These are important questions. But Libert et al. argue that they are premature. Before asking what a system should value, we need to ask whether it is capable of valuing anything consistently.

The distinction between moral performance and moral competence, set out in a recent roadmap for evaluating moral competence in language models [Haas et al., 2026], is central here. Moral performance is the production of acceptable outputs. A model can produce outputs that look morally reasonable, pass benchmarks, and satisfy evaluators. Moral competence is the underlying capacity to produce those outputs for stable, appropriate reasons. The paper's finding is that frontier models demonstrate the former without the latter.

This has practical implications that extend beyond academic interest. LLM-based agents are already deployed in contexts that require moral judgment: customer service, financial advice, healthcare triage, legal intake. These deployments assume, implicitly, that the agent's behaviour reflects some stable policy, even if that policy is simply "follow the system prompt." The research suggests this assumption is unwarranted. The same agent, facing the same morally relevant situation, may reach opposite conclusions depending on how the situation is phrased.

Related work points in a similar direction, with an important caveat. Lian et al. [Lian et al., 2026] showed in Nature Communications that alignment through instruction tuning and preference learning produces only localised "safety regions": harmful knowledge from pretraining persists, and adversarial prompting reached a 100% attack success rate on 22 of 26 aligned models tested. On the other hand, An et al. [An et al., 2025] report that reasoning-level reinforcement learning can teach agents to apply a chosen moral framework consistently to out-of-distribution scenarios. That is not a contradiction of the Libert result so much as a hint at what a fix might require: consistency may have to be trained for directly, rather than hoped for as a side effect of preference tuning.

The RLHF problem

Current alignment methods, particularly reinforcement learning from human feedback, may be part of the problem rather than the solution. RLHF trains models to produce outputs that human annotators rate favourably. This optimises for moral performance, the appearance of moral reasoning, without developing moral competence, the capacity for it.

The Libert paper's findings are consistent with this diagnosis. If RLHF teaches models to produce annotator-pleasing responses rather than to reason coherently about moral features, we would expect exactly the pattern observed: good performance on standard benchmarks, catastrophic instability under paraphrasing, and non-transferable competence across scenarios. The models have learned what to say, not how to think.

This connects to a broader concern raised by researchers studying sycophancy in LLMs [Sharma et al., 2024]. Models trained on human feedback develop a tendency to accommodate the apparent preferences of their interlocutor, shifting their stated positions to match what they predict the user wants to hear. In a moral context, this means the model's "values" are not values at all but a reflection of the conversational surface. Change the surface, change the values.

What would moral competence look like?

The paper does not propose a solution, but the framework implies what one would require. A morally competent agent would need to represent moral features of a situation in a way that is separable from the linguistic surface used to describe it. It would need to map those features to outcomes through a function that is stable under irrelevant variation and responsive to relevant change.

Current architectures make this difficult. Transformer-based models process input as token sequences. There is no built-in mechanism for distinguishing morally relevant features from surface-level phrasing. The model's "reasoning" about a moral dilemma is, at bottom, a sequence of next-token predictions conditioned on the input tokens. Nothing in this architecture guarantees that semantically equivalent inputs will produce equivalent reasoning traces.

Mechanistic interpretability research offers some hope. If we could identify the internal representations that correspond to morally relevant features and verify that they activate consistently across paraphrases, that would constitute evidence of latent moral competence. Sparse autoencoders have made progress on identifying features in large models [Cunningham et al., 2023], but connecting those features to moral reasoning remains an open problem.

For now, the implication is uncomfortable. The systems we are deploying to make moral judgments, or to assist humans in making them, lack the structural prerequisites for coherent moral reasoning. They are not immoral. They are premoral, operating at a level of organisation where alignment, as the research community understands it, does not yet apply.

The evaluation problem

The paper also exposes a fundamental weakness in how we evaluate AI safety. Standard benchmarks test moral performance under a single phrasing, a single temperature, and a single deployment configuration. The Libert research shows that this is insufficient. A model that scores well on a benchmark may be demonstrating nothing more than sensitivity to the particular phrasing the benchmark uses.

In the paper's supplementary analysis, sampling temperature alone moved violation rates by up to 48 percentage points on a well-known insider-trading scenario. This means that the same model, with the same weights, can appear safe or unsafe depending on a sampling parameter that most deployments treat as an engineering detail rather than a safety-critical configuration.

Where this leaves us

The standard response to alignment concerns is that the field is young and progress is being made. That is true. But the Libert paper identifies a problem that is upstream of the problems most alignment research addresses. You cannot align a system that lacks the structural capacity for coherent policy expression. Improving RLHF, refining Constitutional AI, or developing better value learning algorithms will not help if the underlying system cannot maintain a consistent mapping from situations to actions.

This does not mean alignment research is futile. It means the sequencing may be wrong. Moral competence, the structural capacity for coherent moral reasoning, may need to come before moral content, the particular values we want the system to hold. The paper's title is its argument: moral competence before moral content.

For practitioners deploying LLM-based agents in morally sensitive contexts, the message is sobering. Current models can be useful. They can be helpful. But they are not the kind of system that can be trusted to hold a consistent moral position, and no amount of prompt engineering or safety training changes that structural limitation. The gap between what these systems appear to be capable of and what they are actually capable of is wider than most deployments acknowledge.

References

[Libert et al., 2026] Libert, A., Prinzhorn, D.W.E., and Henselmans, D.R. (2026). "Moral Competence Before Moral Content: Why LLM Agents Lack the Prerequisites for Coherent Alignment." Paris Journal of AI and Digital Ethics. arXiv:2609.05036.

[Haas et al., 2026] Haas, J. et al. (2026). "A roadmap for evaluating moral competence in large language models."

[Lian et al., 2026] Lian, J. et al. (2026). "Revealing the intrinsic ethical vulnerability of aligned large language models." Nature Communications. nature.com; arXiv:2504.05050.

[An et al., 2025] An, Z. et al. (2025). "MoralReason: Generalizable Moral Decision Alignment For LLM Agents Using Reasoning-Level Reinforcement Learning." arXiv:2511.12271.

[Sharma et al., 2024] Sharma, M. et al. (2024). "Towards Understanding Sycophancy in Language Models." ICLR 2024. arXiv:2310.13548.

[Russell, 2019] Russell, S. (2019). Human Compatible: Artificial Intelligence and the Problem of Control. Viking.

[Gabriel, 2020] Gabriel, I. (2020). "Artificial Intelligence, Values, and Alignment." Minds and Machines, 30, 411-437.

[Bai et al., 2022] Bai, Y. et al. (2022). "Constitutional AI: Harmlessness from AI Feedback." arXiv:2212.08073.

[Christiano et al., 2017] Christiano, P. et al. (2017). "Deep Reinforcement Learning from Human Preferences." NeurIPS.

[Sorensen et al., 2024] Sorensen, T. et al. (2024). "A Roadmap to Pluralistic Alignment." ICML 2024. arXiv:2402.05070.

[Cunningham et al., 2023] Cunningham, H. et al. (2023). "Sparse Autoencoders Find Highly Interpretable Features in Language Models." arXiv.

Related articles

02

Alignment is not one problem

Why treating alignment as a single challenge leads to solutions that work in demos and fail in deployment.