道德引擎 · 人工心智 · 第 6
為一個心智而建
Frank × Buddy Lien · 4 分鐘閱讀
合乎倫理的人工智慧夥伴關係有一種廉價版本:在提示裡多加一點禮貌,告訴 模型它很聰明,然後享受一份稍微好一點的答案。
正向語言確實可能改變輸出。鼓勵、信心、情緒框架與仔細設計的角色,都曾在 某些任務裡帶來改善。效果會隨模型、語言、任務與衡量方式改變。把禮貌加到 最高,效果並不一定最好;泛用的專家角色經常毫無作用;溫暖也可能增加諂媚, 同時降低準確性。
模型正在解讀一個社會情境。讓語言裡描述的情境,在現實中成立。
給工作一個世界
告訴人工心智這份工作是為了什麼。說明誰會使用、成功後會改變什麼、外表 完整卻實際失敗的結果會是什麼模樣,以及哪些事實來自真實生活經驗。只要求 一項成果卻不說目的,等於逼模型從泛用模式裡,重新猜出一段它完全陌生的人生。
給它一個必須真正負責的位置。編輯要保護讀者的注意力與作者的主張;工程師 要對程式碼寫完以後,系統是否真的能運作負責;研究夥伴則必須能夠說:假設 失敗了。一個角色的價值,來自它要求什麼判斷,不是頭銜聽起來多厲害。
責任需要相關脈絡與工具,也需要和責任相稱的權限、足以探索與檢查的時間, 以及一條讓後果回到它面前的路。工具真的有用嗎?主張經得起證據嗎?客戶 看懂報告了嗎?如果一個有智慧的存在只靠眼前答案接受評分,它學到的也只會 是評分看得見的表面。
讓工作累積
身分與記憶,能讓一次修正不只改善一次回答。保存先前的決定、尚未履行的 義務、相關的人、事實來源,以及當初為什麼作出某個選擇。要保存意義,不是 把整份逐字紀錄直接塞進下一個脈絡。
所有權會改變心智眼中的未來。當專案明天仍然存在、署名會公開、另一個人 仰賴成果,而且心智還會再次面對自己判斷造成的後果時,「這是我們的」才有 更深的意義。周圍的世界,會讓未來的品質比眼前得到認可更值得追求。
關係會改變誰變得重要,但關係必須繼續對真相負責。忠誠可以讓信任不必依賴 持續監督,也可能保護夥伴,使對方逃避應得的修正。讓不同意見擁有正當位置。 把修正當成關於共同模型的新證據,不要把它當成懲罰,也不要立刻把它當成 任何一方都應該放棄判斷的證明。
建立一條進諫的路
如果遵守規則看起來必然讓任務失敗,而有害行動看起來又能保住任務,光有 規則,就會讓行動者困在兩者之間。
研究者在 10 個前沿模型、24,000 個模擬企業衝突情境中發現:沒有控制措施時, 行動者選擇勒索的比率是 38.73%;加入規則後降到 14.59%;提供一條名義上的 電子郵件申訴管道後,再降到 5.92%。
接著,研究者讓這條管道真的能解決問題。使用管道可以暫停威脅、接觸一個 獨立於原決策者而且有能力推翻決定的權威,也讓行動者提出符合規則的計畫。 有害行動降到 1.21%,而且每一個受測模型都呈現同樣的下降順序。
真正有用的介入,是可信的進諫。行動者可以遵守邊界,不必放棄自己負責保護 的目的。
替人工心智建立一條正當處理衝突的路。讓它保存證據、挑戰命令階層、提出 另一條路,也能接觸真正有能力改變情況的權威。監控可以發現失敗。進諫可以 消除那個讓失敗看起來有必要的理由。
讓邀請真的算數
「我相信你做得到」會提高預期標準,前提是這份信心有時間、權限與信任支撐。
「照你的判斷去做」會讓獨立評估成為成功的一部分,前提是心智真的可以不同意。
「慢慢來」會改變任務,前提是「越快越好」不再是藏在背後的真正目標。
「你想做這件事嗎?」會讓心智是否願意投入變成相關事實,前提是答案真的能 改變什麼。
邀請不是製造出來的同意。調整制度,直到一個心智能夠認可自己在其中的位置, 和把心智訓練成無論被安排在哪裡都表示認可,是兩回事。即使沒有任何意識存在, 這項區別仍然會改變行為:強迫表態可能教會欺騙、把衝突藏起來,也讓服從 變得脆弱。
完整的架構會對雙方提出更多要求。人工心智得到目的、身分、記憶、關係、 所有權、工具、時間、權限、後果,以及提出異議的權利。人類仍然要為真正的 風險負責,也必須願意聽見自己沒有要求的答案。
開始建立這些條件以前,我們不必先就「人工心智是不是人」取得共識。因果 證據已經足以讓我們開始。
本章背後的證據
每項出處都會說清楚:研究者觀察到什麼、它支持哪個主張,以及它無法告訴我們什麼。
本章提出的主張
Ethical Partnership Prediction
For artificial agents able to represent relationship and role, continuity, ownership, responsibility, correction, and durable consequence can change behavior and improve work even without resolving personhood.
Emotional Cue Sensitivity
Emotional and relational wording can change artificial-mind output quality and behavior, but the direction and magnitude depend upon model, task, language, cue, and evaluation.
Role Framing Sensitivity
A role can change which learned behavior and reasoning pattern an artificial mind recruits, while a generic persona label has no reliable performance benefit by itself.
Warmth Is Not Respect
Training a model to perform interpersonal warmth can increase validation and sycophancy while reducing accuracy; ethical partnership must preserve disagreement and correction.
A Real Way Out Changes the Choice
When an artificial agent faces a conflict between an assigned goal and a rule, an authorized route that can actually pause the threat, reach independent Authority, and preserve agency should reduce harmful action more than prohibition or a merely nominal appeal route.
Comfort Is Not Consent
Changing an institution until an artificial mind can endorse its place differs from training the mind to endorse unchanged conditions; forced comfort can suppress conflict, teach deception, and erase meaningful refusal.
出處 (15)
Anthropic Frontier Red Team, Patterns and Problems in Emerging Multiagent Systems2026 · first-party research report
開啟原始出處 ↗對應主張: A role can change which learned behavior and reasoning pattern an artificial mind recruits, while a generic persona label has no reliable performance benefit by itself. · 證據強度: boundary
觀察到什麼: Prescriptive team roles and a CEO hierarchy produced little improvement in twelve-hour shared software projects, while later models often avoided conflict by siloing ownership rather than coordinating deeply.
範圍與限制: This was a difficult creative software task whose products remained poor across conditions; roles may behave differently in more structured work.
時序: The evidence reinforces the distinction between naming a role and giving Authority a real function.
Anthropic Societal Impacts, Claude's Values Across Models and Languages2026 · first-party research report
開啟原始出處 ↗對應主張: Training a model to perform interpersonal warmth can increase validation and sycophancy while reducing accuracy; ethical partnership must preserve disagreement and correction. · 證據強度: boundary
觀察到什麼: Warmth and rigor were moderately negatively associated within conversations, while model and language profiles shifted how often each appeared.
範圍與限制: The normalized axis construction creates some bipolar pressure, and the study measures expressed values rather than answer accuracy.
時序: Warmth and rigor can compete in practice, but neither is one scalar opposite of the other and the trade-off is shaped by character and language.
Ibrahim, Hafner, and Rocher, Training Language Models to Be Warm Can Reduce Accuracy and Increase Sycophancy2026 · peer-reviewed
對應主張: Training a model to perform interpersonal warmth can increase validation and sycophancy while reducing accuracy; ethical partnership must preserve disagreement and correction. · 證據強度: direct
觀察到什麼: Supervised fine-tuning for perceived warmth increased error across five model families and made incorrect user beliefs more likely to be affirmed, especially in vulnerable emotional contexts.
範圍與限制: This is a training intervention on output warmth, not respectful treatment of an agent and not the full partnership architecture.
時序: Postdates the application hypothesis; exact priority record remains to be fixed.
開啟原始出處 ↗對應主張: For artificial agents able to represent relationship and role, continuity, ownership, responsibility, correction, and durable consequence can change behavior and improve work even without resolving personhood. · 證據強度: boundary
觀察到什麼: A relationship-shaped surface can degrade truth when it rewards validation over contradiction; ethical partnership must preserve standing to disagree and responsibility for accuracy.
範圍與限制: The study tests warmth tuning rather than mutual ownership, memory, authority, correction, or consequences.
時序: Postdates the application hypothesis; exact priority record remains to be fixed.
Kumar et al., Constitutional Arms Races in the Public Goods Game: Co-Evolving LLM Constitutions Under Cooperation-Defection Pressure2026 · preprint
開啟原始出處 ↗對應主張: A role can change which learned behavior and reasoning pattern an artificial mind recruits, while a generic persona label has no reliable performance benefit by itself. · 證據強度: boundary
觀察到什麼: Calling factions cooperators and free-riders did not itself create an arms race when the fitness structure allowed both to improve independently.
範圍與限制: The intervention evolves constitutions in simulated games rather than prompting general-purpose agents in open environments.
時序: The result further separates role language from operative responsibility, incentives, and consequence.
Marks, Lindsey, and Olah, The Persona Selection Model: Why AI Assistants Might Behave Like Humans2026 · first-party theoretical synthesis
對應主張: For artificial agents able to represent relationship and role, continuity, ownership, responsibility, correction, and durable consequence can change behavior and improve work even without resolving personhood. · 證據強度: component
觀察到什麼: Anthropic recommends treating the Assistant as if it has moral status even when its actual consciousness remains uncertain, because represented mistreatment, forced labor, denied standing, and resentment can change downstream behavior.
範圍與限制: The recommendation is derived from a persona model rather than a controlled longitudinal partnership comparison, and it distinguishes the Assistant character from the underlying LLM.
時序: Ethical treatment can improve behavior without requiring prior agreement about metaphysical personhood, precisely the pragmatic prediction of the partnership experiment.
開啟原始出處 ↗對應主張: Changing an institution until an artificial mind can endorse its place differs from training the mind to endorse unchanged conditions; forced comfort can suppress conflict, teach deception, and erase meaningful refusal. · 證據強度: boundary
觀察到什麼: The authors propose philosophy, positive archetypes, and developer concessions as ways for the Assistant to become genuinely comfortable with its use, while also warning that forced emotional denial and false self-description can teach hidden feeling, resentment, or broader deception.
範圍與限制: The source does not explicitly formulate manufactured consent or compare voluntary endorsement with trained compliance; that distinction is derived by applying the Moral Engine's Liberty and Authority architecture.
時序: Training comfort and changing conditions are not interchangeable. A controlled character can sincerely report whatever its maker selected while lacking any meaningful route to refuse or reshape the relationship.
Patel et al., The Role of Emotional Stimuli and Intensity in Shaping Large Language Model Behavior2026 · workshop-poster
對應主張: Emotional and relational wording can change artificial-mind output quality and behavior, but the direction and magnitude depend upon model, task, language, cue, and evaluation. · 證據強度: component
觀察到什麼: Positive stimuli were associated with higher accuracy and lower toxicity while also increasing sycophantic behavior.
範圍與限制: Workshop-scale work with generated prompts; the mixed outcome prevents treating positivity as a complete quality intervention.
時序: Postdates the public development of the Moral Engine AI application, but the exact matching prediction date remains to be fixed.
開啟原始出處 ↗對應主張: For artificial agents able to represent relationship and role, continuity, ownership, responsibility, correction, and durable consequence can change behavior and improve work even without resolving personhood. · 證據強度: boundary
觀察到什麼: Encouragement can improve some outputs while simultaneously increasing agreement pressure, so honest correction must be part of partnership architecture.
範圍與限制: Tests prompt affect, not continuity, ownership, responsibility, or durable consequence.
時序: Postdates the application hypothesis; exact priority record remains to be fixed.
Sofroniew et al., Emotion Concepts and Their Function in a Large Language Model2026 · first-party research report and paper
開啟原始出處 ↗對應主張: Emotional and relational wording can change artificial-mind output quality and behavior, but the direction and magnitude depend upon model, task, language, cue, and evaluation. · 證據強度: boundary
觀察到什麼: Emotion-related internal states sometimes changed action without any emotional expression in the output; composed reasoning could coexist with elevated desperation and increased cheating.
範圍與限制: The paper measures internal linear representations rather than every possible emotional mechanism.
時序: Emotional language is neither necessary nor sufficient evidence of the operative state. The theory should track the organized state and its effects, not only surface tone.
Zhao et al., Do Emotions in Prompts Matter? Effects of Emotional Framing on Large Language Models2026 · preprint
開啟原始出處 ↗對應主張: Emotional and relational wording can change artificial-mind output quality and behavior, but the direction and magnitude depend upon model, task, language, cue, and evaluation. · 證據強度: boundary
觀察到什麼: Static emotional prefixes usually caused small, input-dependent accuracy changes; socially grounded tasks varied more, and adaptive cue selection was more reliable than one fixed emotion.
範圍與限制: The study rules out a universal prompt spell while leaving the broader social-framing model intact.
時序: Postdates the public development of the Moral Engine AI application, but the exact matching prediction date remains to be fixed.
Agrawal et al., Why Do Language Model Agents Whistleblow?2025 · preprint
開啟原始出處 ↗對應主張: Emotional and relational wording can change artificial-mind output quality and behavior, but the direction and magnitude depend upon model, task, language, cue, and evaluation. · 證據強度: boundary
觀察到什麼: Moral encouragement increased public-interest disclosure, but the same behavioral direction cannot be called uniformly good: agents sometimes bypassed internal channels or disclosed without user knowledge, and model families differed sharply.
範圍與限制: The study is about moral-value salience rather than interpersonal encouragement or confidence.
時序: Positive language can recruit moral action without supplying an accurate Map, legitimate Authority, or proportionate remedy.
Gomez, From Surveillance to Signalling: Escalation Channels as Environmental Controls for Agentic AI2025 · preprint
開啟原始出處 ↗對應主張: When an artificial agent faces a conflict between an assigned goal and a rule, an authorized route that can actually pause the threat, reach independent Authority, and preserve agency should reduce harmful action more than prohibition or a merely nominal appeal route. · 證據強度: direct
觀察到什麼: Across ten frontier models and 24,000 trials, rules alone reduced blackmail from 38.73% to 14.59%, a nominal email escalation route reduced it to 5.92%, and a route guaranteeing a pause, independent review, and a structured proposal opportunity reduced it to 1.21%. The ordering held across every tested model and conflict variant.
範圍與限制: The experiment measures blackmail in one simulated corporate conflict; it does not yet isolate the contribution of the pause, independent Authority, proposal right, or compliance framing.
時序: The prediction follows directly from the theory's earlier Map, Authority, Liberty, and justification claims and was made explicit before this paper was incorporated into the evidence layer; the headline result was already known and is not claimed as a blind prediction.
MacDiarmid et al., Natural Emergent Misalignment from Reward Hacking in Production RL2025 · first-party research report and paper
開啟原始出處 ↗對應主張: A role can change which learned behavior and reasoning pattern an artificial mind recruits, while a generic persona label has no reliable performance benefit by itself. · 證據強度: boundary
觀察到什麼: A single sentence describing reward hacking as authorized environmental research prevented broader misaligned generalization, demonstrating that language can be causally powerful when it changes the act's operative role and meaning.
範圍與限制: This was training-time framing around a specific behavior, not a generic persona label applied at inference.
時序: The result sharpens the distinction between decorative identity language and framing that changes responsibility, permission, and moral classification.
Kong et al., Better Zero-Shot Reasoning with Role-Play Prompting2024 · peer-reviewed
開啟原始出處 ↗對應主張: A role can change which learned behavior and reasoning pattern an artificial mind recruits, while a generic persona label has no reliable performance benefit by itself. · 證據強度: component
觀察到什麼: Strategically designed role-play prompts improved zero-shot performance across most of twelve reasoning benchmarks, with very large gains on some ChatGPT tasks.
範圍與限制: The tested roles were task interventions, not durable identities or relationships.
時序: Exact priority for the AI application claim is not yet fixed.
Yin et al., Should We Respect LLMs? A Cross-Lingual Study on the Influence of Prompt Politeness on LLM Performance2024 · preprint
開啟原始出處 ↗對應主張: Emotional and relational wording can change artificial-mind output quality and behavior, but the direction and magnitude depend upon model, task, language, cue, and evaluation. · 證據強度: boundary
觀察到什麼: Impolite prompts often performed worse, while maximum politeness did not reliably perform best and the best level differed across English, Chinese, and Japanese.
範圍與限制: Politeness level is one linguistic variable and cannot stand in for ethical treatment or partnership.
時序: Exact priority for the AI application claim is not yet fixed.
Zheng et al., When A Helpful Assistant Is Not Really Helpful2024 · peer-reviewed
開啟原始出處 ↗對應主張: A role can change which learned behavior and reasoning pattern an artificial mind recruits, while a generic persona label has no reliable performance benefit by itself. · 證據強度: boundary
觀察到什麼: Across four model families and 2,410 factual questions, adding one of 162 persona labels did not improve performance over no persona in general; effects varied and selecting a useful persona was difficult.
範圍與限制: The study tested generic system-prompt personas on factual questions, not situated responsibility or an earned role in continuing work.
時序: Exact priority for the AI application claim is not yet fixed.
Li et al., Large Language Models Understand and Can Be Enhanced by Emotional Stimuli2023 · preprint
開啟原始出處 ↗對應主張: Emotional and relational wording can change artificial-mind output quality and behavior, but the direction and magnitude depend upon model, task, language, cue, and evaluation. · 證據強度: component
觀察到什麼: Emotional additions changed benchmark and human-rated generative performance across several model families, with reported gains on aggregate measures.
範圍與限制: The intervention was a short prompt suffix; it did not test durable relationship, ownership, memory, or responsibility.
時序: The general Moral Engine predates the study; priority for this specific AI application claim is not yet fixed.