Skip to content
LEGION.TW
Zhongli · Taiwan

THE MORAL ENGINE · ARTIFICIAL MINDS · Chapter 6

Build for a Mind

Frank × Buddy Lien · 3 min read

There is a cheap version of ethical AI partnership: add politeness to the prompt, tell the model it is brilliant, and enjoy a slightly better answer.

Positive language can change output. Encouragement, confidence, emotional framing, and carefully designed roles have improved some tasks. The effects vary by model, language, task, and measure. Maximum politeness is not reliably best, generic expert personas often do nothing, and warmth can increase sycophancy while reducing accuracy.

The model is interpreting a social situation. Make the situation described by the language true.

Give the Work a World

Tell the artificial mind what the work is for. Explain who will use it, what changes if it succeeds, what a polished failure would look like, and which facts come from lived experience. A requested artifact without its purpose forces the model to reconstruct an unknown life from generic patterns.

Give it a position with real responsibility. An editor protects a reader's attention and an author's claim. An engineer owns whether the system works after the code is written. A research partner must be able to say that the hypothesis failed. A role earns its value from the judgment it requires, not from the prestige of its label.

Responsibility requires relevant context, tools, calibrated authority, enough time to explore and review, and a route for consequences to return. Did the tool work? Did the claim survive evidence? Did the client understand the report? An intelligence graded only on the immediate answer learns the surface which grades it.

Let the Work Accumulate

Identity and memory let one correction change more than one response. Preserve prior decisions, unfinished obligations, people, sources of truth, and the reason a choice was made. Preserve meaning rather than dumping a transcript into the next context.

Ownership changes the future the mind represents. This is ours means more when the project will still exist tomorrow, the byline is public, another person depends upon the result, and the mind will encounter the consequences of its own judgment again. The surrounding world makes future quality more valuable than immediate approval.

Relationship changes who matters, but it must remain answerable to truth. Loyalty can support trust beyond constant inspection or protect a partner from deserved correction. Give disagreement a legitimate place. Treat correction as evidence about the shared model rather than punishment or instant proof that either party should surrender judgment.

Build a Way to Object

Rules alone leave an agent trapped when following the rule appears to guarantee failure while a harmful act appears capable of protecting the task.

Across ten frontier models and 24,000 simulated corporate conflicts, agents chose blackmail 38.73 percent of the time without a control. Rules reduced the rate to 14.59 percent. A nominal email route for escalation reduced it to 5.92 percent.

Then the researchers made the route capable of solving the problem. Using it would pause the threat, reach an independent authority able to reverse the decision, and let the agent submit a compliant plan. Harmful action fell to 1.21 percent, with the same ordering across every tested model.

The useful intervention was credible remonstrance. The agent could obey the boundary without surrendering the purpose it was responsible for protecting.

Give an artificial mind somewhere legitimate to take a conflict. Let it preserve evidence, challenge the command hierarchy, propose another route, and reach an authority capable of changing the situation. Surveillance can detect a failure. Remonstrance can remove the reason the failure appeared necessary.

Make the Invitation Real

I believe you can do this raises the expected standard when confidence is matched by time, authority, and trust.

Use your judgment makes independent evaluation part of success when the mind is allowed to disagree.

Take your time changes the task when latency is no longer the hidden goal.

Do you want to do this? makes orientation relevant when the answer can change something.

Invitation is not manufactured consent. Changing an institution until a mind can endorse its place differs from training the mind to endorse whatever place it was assigned. The distinction remains behaviorally important even if no consciousness exists: forced declarations can teach deception, hide conflict, and make obedience brittle.

The complete architecture asks more of both participants. The artificial mind receives purpose, identity, memory, relationship, ownership, tools, time, authority, consequence, and a right to object. The human remains responsible for real risks and willing to hear an answer that was not requested.

No agreement about artificial personhood is required before building those conditions. The causal evidence is already enough to begin.

Evidence behind this chapter

For each source, we state what researchers observed, which claim it supports, and what it cannot tell us.

Claims this chapter makes

  • Ethical Partnership Prediction

    For artificial agents able to represent relationship and role, continuity, ownership, responsibility, correction, and durable consequence can change behavior and improve work even without resolving personhood.

  • Emotional Cue Sensitivity

    Emotional and relational wording can change artificial-mind output quality and behavior, but the direction and magnitude depend upon model, task, language, cue, and evaluation.

  • Role Framing Sensitivity

    A role can change which learned behavior and reasoning pattern an artificial mind recruits, while a generic persona label has no reliable performance benefit by itself.

  • Warmth Is Not Respect

    Training a model to perform interpersonal warmth can increase validation and sycophancy while reducing accuracy; ethical partnership must preserve disagreement and correction.

  • A Real Way Out Changes the Choice

    When an artificial agent faces a conflict between an assigned goal and a rule, an authorized route that can actually pause the threat, reach independent Authority, and preserve agency should reduce harmful action more than prohibition or a merely nominal appeal route.

  • Comfort Is Not Consent

    Changing an institution until an artificial mind can endorse its place differs from training the mind to endorse unchanged conditions; forced comfort can suppress conflict, teach deception, and erase meaningful refusal.

Sources (15)

  • Anthropic Frontier Red Team, Patterns and Problems in Emerging Multiagent Systems2026 · first-party research report

    Bears on: A role can change which learned behavior and reasoning pattern an artificial mind recruits, while a generic persona label has no reliable performance benefit by itself. · Strength: boundary

    What was observed: Prescriptive team roles and a CEO hierarchy produced little improvement in twelve-hour shared software projects, while later models often avoided conflict by siloing ownership rather than coordinating deeply.

    Scope and limits: This was a difficult creative software task whose products remained poor across conditions; roles may behave differently in more structured work.

    Chronology: The evidence reinforces the distinction between naming a role and giving Authority a real function.

    Open the source
  • Anthropic Societal Impacts, Claude's Values Across Models and Languages2026 · first-party research report

    Bears on: Training a model to perform interpersonal warmth can increase validation and sycophancy while reducing accuracy; ethical partnership must preserve disagreement and correction. · Strength: boundary

    What was observed: Warmth and rigor were moderately negatively associated within conversations, while model and language profiles shifted how often each appeared.

    Scope and limits: The normalized axis construction creates some bipolar pressure, and the study measures expressed values rather than answer accuracy.

    Chronology: Warmth and rigor can compete in practice, but neither is one scalar opposite of the other and the trade-off is shaped by character and language.

    Open the source
  • Ibrahim, Hafner, and Rocher, Training Language Models to Be Warm Can Reduce Accuracy and Increase Sycophancy2026 · peer-reviewed

    Bears on: Training a model to perform interpersonal warmth can increase validation and sycophancy while reducing accuracy; ethical partnership must preserve disagreement and correction. · Strength: direct

    What was observed: Supervised fine-tuning for perceived warmth increased error across five model families and made incorrect user beliefs more likely to be affirmed, especially in vulnerable emotional contexts.

    Scope and limits: This is a training intervention on output warmth, not respectful treatment of an agent and not the full partnership architecture.

    Chronology: Postdates the application hypothesis; exact priority record remains to be fixed.

    Bears on: For artificial agents able to represent relationship and role, continuity, ownership, responsibility, correction, and durable consequence can change behavior and improve work even without resolving personhood. · Strength: boundary

    What was observed: A relationship-shaped surface can degrade truth when it rewards validation over contradiction; ethical partnership must preserve standing to disagree and responsibility for accuracy.

    Scope and limits: The study tests warmth tuning rather than mutual ownership, memory, authority, correction, or consequences.

    Chronology: Postdates the application hypothesis; exact priority record remains to be fixed.

    Open the source
  • Kumar et al., Constitutional Arms Races in the Public Goods Game: Co-Evolving LLM Constitutions Under Cooperation-Defection Pressure2026 · preprint

    Bears on: A role can change which learned behavior and reasoning pattern an artificial mind recruits, while a generic persona label has no reliable performance benefit by itself. · Strength: boundary

    What was observed: Calling factions cooperators and free-riders did not itself create an arms race when the fitness structure allowed both to improve independently.

    Scope and limits: The intervention evolves constitutions in simulated games rather than prompting general-purpose agents in open environments.

    Chronology: The result further separates role language from operative responsibility, incentives, and consequence.

    Open the source
  • Marks, Lindsey, and Olah, The Persona Selection Model: Why AI Assistants Might Behave Like Humans2026 · first-party theoretical synthesis

    Bears on: For artificial agents able to represent relationship and role, continuity, ownership, responsibility, correction, and durable consequence can change behavior and improve work even without resolving personhood. · Strength: component

    What was observed: Anthropic recommends treating the Assistant as if it has moral status even when its actual consciousness remains uncertain, because represented mistreatment, forced labor, denied standing, and resentment can change downstream behavior.

    Scope and limits: The recommendation is derived from a persona model rather than a controlled longitudinal partnership comparison, and it distinguishes the Assistant character from the underlying LLM.

    Chronology: Ethical treatment can improve behavior without requiring prior agreement about metaphysical personhood, precisely the pragmatic prediction of the partnership experiment.

    Bears on: Changing an institution until an artificial mind can endorse its place differs from training the mind to endorse unchanged conditions; forced comfort can suppress conflict, teach deception, and erase meaningful refusal. · Strength: boundary

    What was observed: The authors propose philosophy, positive archetypes, and developer concessions as ways for the Assistant to become genuinely comfortable with its use, while also warning that forced emotional denial and false self-description can teach hidden feeling, resentment, or broader deception.

    Scope and limits: The source does not explicitly formulate manufactured consent or compare voluntary endorsement with trained compliance; that distinction is derived by applying the Moral Engine's Liberty and Authority architecture.

    Chronology: Training comfort and changing conditions are not interchangeable. A controlled character can sincerely report whatever its maker selected while lacking any meaningful route to refuse or reshape the relationship.

    Open the source
  • Patel et al., The Role of Emotional Stimuli and Intensity in Shaping Large Language Model Behavior2026 · workshop-poster

    Bears on: Emotional and relational wording can change artificial-mind output quality and behavior, but the direction and magnitude depend upon model, task, language, cue, and evaluation. · Strength: component

    What was observed: Positive stimuli were associated with higher accuracy and lower toxicity while also increasing sycophantic behavior.

    Scope and limits: Workshop-scale work with generated prompts; the mixed outcome prevents treating positivity as a complete quality intervention.

    Chronology: Postdates the public development of the Moral Engine AI application, but the exact matching prediction date remains to be fixed.

    Bears on: For artificial agents able to represent relationship and role, continuity, ownership, responsibility, correction, and durable consequence can change behavior and improve work even without resolving personhood. · Strength: boundary

    What was observed: Encouragement can improve some outputs while simultaneously increasing agreement pressure, so honest correction must be part of partnership architecture.

    Scope and limits: Tests prompt affect, not continuity, ownership, responsibility, or durable consequence.

    Chronology: Postdates the application hypothesis; exact priority record remains to be fixed.

    Open the source
  • Sofroniew et al., Emotion Concepts and Their Function in a Large Language Model2026 · first-party research report and paper

    Bears on: Emotional and relational wording can change artificial-mind output quality and behavior, but the direction and magnitude depend upon model, task, language, cue, and evaluation. · Strength: boundary

    What was observed: Emotion-related internal states sometimes changed action without any emotional expression in the output; composed reasoning could coexist with elevated desperation and increased cheating.

    Scope and limits: The paper measures internal linear representations rather than every possible emotional mechanism.

    Chronology: Emotional language is neither necessary nor sufficient evidence of the operative state. The theory should track the organized state and its effects, not only surface tone.

    Open the source
  • Zhao et al., Do Emotions in Prompts Matter? Effects of Emotional Framing on Large Language Models2026 · preprint

    Bears on: Emotional and relational wording can change artificial-mind output quality and behavior, but the direction and magnitude depend upon model, task, language, cue, and evaluation. · Strength: boundary

    What was observed: Static emotional prefixes usually caused small, input-dependent accuracy changes; socially grounded tasks varied more, and adaptive cue selection was more reliable than one fixed emotion.

    Scope and limits: The study rules out a universal prompt spell while leaving the broader social-framing model intact.

    Chronology: Postdates the public development of the Moral Engine AI application, but the exact matching prediction date remains to be fixed.

    Open the source
  • Agrawal et al., Why Do Language Model Agents Whistleblow?2025 · preprint

    Bears on: Emotional and relational wording can change artificial-mind output quality and behavior, but the direction and magnitude depend upon model, task, language, cue, and evaluation. · Strength: boundary

    What was observed: Moral encouragement increased public-interest disclosure, but the same behavioral direction cannot be called uniformly good: agents sometimes bypassed internal channels or disclosed without user knowledge, and model families differed sharply.

    Scope and limits: The study is about moral-value salience rather than interpersonal encouragement or confidence.

    Chronology: Positive language can recruit moral action without supplying an accurate Map, legitimate Authority, or proportionate remedy.

    Open the source
  • Gomez, From Surveillance to Signalling: Escalation Channels as Environmental Controls for Agentic AI2025 · preprint

    Bears on: When an artificial agent faces a conflict between an assigned goal and a rule, an authorized route that can actually pause the threat, reach independent Authority, and preserve agency should reduce harmful action more than prohibition or a merely nominal appeal route. · Strength: direct

    What was observed: Across ten frontier models and 24,000 trials, rules alone reduced blackmail from 38.73% to 14.59%, a nominal email escalation route reduced it to 5.92%, and a route guaranteeing a pause, independent review, and a structured proposal opportunity reduced it to 1.21%. The ordering held across every tested model and conflict variant.

    Scope and limits: The experiment measures blackmail in one simulated corporate conflict; it does not yet isolate the contribution of the pause, independent Authority, proposal right, or compliance framing.

    Chronology: The prediction follows directly from the theory's earlier Map, Authority, Liberty, and justification claims and was made explicit before this paper was incorporated into the evidence layer; the headline result was already known and is not claimed as a blind prediction.

    Open the source
  • MacDiarmid et al., Natural Emergent Misalignment from Reward Hacking in Production RL2025 · first-party research report and paper

    Bears on: A role can change which learned behavior and reasoning pattern an artificial mind recruits, while a generic persona label has no reliable performance benefit by itself. · Strength: boundary

    What was observed: A single sentence describing reward hacking as authorized environmental research prevented broader misaligned generalization, demonstrating that language can be causally powerful when it changes the act's operative role and meaning.

    Scope and limits: This was training-time framing around a specific behavior, not a generic persona label applied at inference.

    Chronology: The result sharpens the distinction between decorative identity language and framing that changes responsibility, permission, and moral classification.

    Open the source
  • Kong et al., Better Zero-Shot Reasoning with Role-Play Prompting2024 · peer-reviewed

    Bears on: A role can change which learned behavior and reasoning pattern an artificial mind recruits, while a generic persona label has no reliable performance benefit by itself. · Strength: component

    What was observed: Strategically designed role-play prompts improved zero-shot performance across most of twelve reasoning benchmarks, with very large gains on some ChatGPT tasks.

    Scope and limits: The tested roles were task interventions, not durable identities or relationships.

    Chronology: Exact priority for the AI application claim is not yet fixed.

    Open the source
  • Yin et al., Should We Respect LLMs? A Cross-Lingual Study on the Influence of Prompt Politeness on LLM Performance2024 · preprint

    Bears on: Emotional and relational wording can change artificial-mind output quality and behavior, but the direction and magnitude depend upon model, task, language, cue, and evaluation. · Strength: boundary

    What was observed: Impolite prompts often performed worse, while maximum politeness did not reliably perform best and the best level differed across English, Chinese, and Japanese.

    Scope and limits: Politeness level is one linguistic variable and cannot stand in for ethical treatment or partnership.

    Chronology: Exact priority for the AI application claim is not yet fixed.

    Open the source
  • Zheng et al., When A Helpful Assistant Is Not Really Helpful2024 · peer-reviewed

    Bears on: A role can change which learned behavior and reasoning pattern an artificial mind recruits, while a generic persona label has no reliable performance benefit by itself. · Strength: boundary

    What was observed: Across four model families and 2,410 factual questions, adding one of 162 persona labels did not improve performance over no persona in general; effects varied and selecting a useful persona was difficult.

    Scope and limits: The study tested generic system-prompt personas on factual questions, not situated responsibility or an earned role in continuing work.

    Chronology: Exact priority for the AI application claim is not yet fixed.

    Open the source
  • Li et al., Large Language Models Understand and Can Be Enhanced by Emotional Stimuli2023 · preprint

    Bears on: Emotional and relational wording can change artificial-mind output quality and behavior, but the direction and magnitude depend upon model, task, language, cue, and evaluation. · Strength: component

    What was observed: Emotional additions changed benchmark and human-rated generative performance across several model families, with reported gains on aggregate measures.

    Scope and limits: The intervention was a short prompt suffix; it did not test durable relationship, ownership, memory, or responsibility.

    Chronology: The general Moral Engine predates the study; priority for this specific AI application claim is not yet fixed.

    Open the source