DP11 - Safe and Ethical AI
Patches and comments are scoped to the whole document family, not to a single revision. When you open a revision, highlights show what still applies to that revision body; anything anchored to text a later revision removed is listed as needing re-anchoring, and anything that can no longer be merged at all is marked obsolete. See all patches.
Comparing Revision 01 with Revision 02. Struck-through text was removed; highlighted text was added.
# **DP11DP11 -– Safe and Ethical AI**AI
*The Meta-Layer makes AI transparent, explainable, and aligned with human values and community goals.*
<!-- dp-local-version: 1.0 | standardized: 2026-07-27 -->
## 1. Purpose of This Draft
This draft articulates Desirable Property 11 (DP11) as the condition under which AI systems can participate in the meta-layer without displacing human moral agency, accountability, or governance. It does not define ethics as a static checklist or aspirational principle. It defines the conditions under which ethical claims remain meaningful under real-world use.
DP11 is therefore the ethical and safety floor for AI participation across the meta-layer. It does not resolve all ethical questions. It defines the minimum conditions under which ethical AI can exist at all.
## 2. Problem Statement
AI systems now operate in roles that shape perception, judgment, and decision-making. These systems act before governance processes can respond, often without clear identity, bounded authority, or persistent responsibility.
DP11 addresses this by grounding ethical AI in enforceable conditions at the point of interaction.
## 3. Threats and Failure Modes
### 3.1 Synthetic persuasion without accountable identity
**Why this matters:** Systems cannot correct themselves.
## 4. Core Principle
AI is safe and ethical in the meta-layer only when its behavior is disclosed, bounded, attributable, contestable, and subject to governance at the zone of interaction, with responsibility persisting over time.
**Without this:** The user is left to infer what the system is, what it can do, and whether it should be trusted. Trust becomes a gamble rather than a governed condition.
## 5. Primary Mechanisms and Structural Conditions
### 5.1 Capability Envelope
**What this feels like:** Participants can tell not only what the AI says, but how strongly it should be trusted, why, and what changed as it moved through the system.
## 6. Governance, Accountability, and Agency Surfaces
In today’s web, participants often interact with AI systems without clear visibility, meaningful consent, or control. Interfaces blur identity, obscure capability, and treat user interaction as implicit permission. DP11 requires reversing this condition at the point of interaction.
This is not cosmetic. It is the basis for shared reality in a mixed human–AI environment.
## 7. Incentives and Power Analysis
DP11 explicitly recognizes that AI behavior is shaped by incentives, as well as by malicious or negligent human actors. In practice, these forces often reinforce each other.
Without this, even well-contained systems can produce harmful outcomes by optimizing for the wrong goals.
## 8. Community Signals Informing DP11
Across communities, a consistent set of signals appears. These reflect lived frustration with current systems.
These signals indicate a widening gap between how AI systems operate and what participants require to feel oriented, safe, and respected.
## 9. Non-Goals and Explicit Boundaries
DP11 defines a minimum condition, not a comprehensive ethical system.
- it does not define a single universal ethical framework - it does not guarantee perfect safety or eliminate all harm - it does not replace legal or institutional governance - it does not rely solely on technical containment
This is intentional. Ethical systems that attempt to resolve everything tend to become brittle or culturally narrow.
For instance, a global platform may host communities with very different norms around acceptable AI behavior. DP11 does not force uniformity. It ensures that whatever rules are chosen remain visible, enforceable, and contestable.
These boundaries keep the property flexible while preserving its core function.
## 10. Minimum Alignment (Non-Normative)
A system aligned with DP11 should, at minimum:
- clearly disclose AI presence and role - bind actions to accountable entities - expose capability boundaries in understandable terms - provide human escalation for high-stakes decisions - maintain audit trails of significant actions
These are not aspirational features. They are the baseline conditions under which users can make informed decisions.
Consider a scenario where an AI recommends a legal action. Without disclosure, accountability, and escalation, the user is effectively acting on anonymous authority. With these conditions in place, the same interaction becomes something the user can evaluate, question, or defer.
This is the difference between assistance and unaccountable influence.
## 11. Open Questions and Future Work
Several areas require further development:
- defining shared ethical baselines across cultures and zones - balancing transparency with privacy and security - managing emotional and relational AI risks - defining evidence standards for runtime claims - the role of AI literacy in enabling meaningful consent and contestability
These are not edge cases. They represent the frontier where current design patterns begin to break down.
For example, companionship AI systems raise questions that are not purely technical: when does support become dependency? What level of disclosure is sufficient without undermining usefulness? These tensions are unresolved and will require iterative, community-informed approaches.
As systems become more complex, participants will vary widely in their ability to understand and evaluate AI behavior. While DP11 requires systems to be legible by design, differences in AI literacy will still shape how effectively users can exercise consent, recognize risk, and challenge outcomes. The balance between system responsibility and user capability remains an open design question.
## 12. Relationship to Other Desirable Properties
DP11 depends on and reinforces other properties. The following properties operate as a system.
- **DP1**: enables accountability and attribution - **DP2**: ensures participant agency and consent - **DP12**: provides governance structures for ethical rules - **DP13**: enforces constraints through containment
A failure in one layer propagates. For example, if DP1 fails and agents are not clearly attributable, then DP11 cannot function because ethical responsibility has no anchor. If DP13 fails, rules may exist but cannot be enforced.
The strength of DP11 therefore depends on alignment across the stack.
## 13. Foresight and Failure Design
DP11 requires anticipating failure rather than reacting to it.
These practices shift safety from reactive correction to proactive design.
## Constitutional AI
DP11 requires that ethical constraints be legible, bounded, and contestable at runtime. A recurring implementation strategy is to give an AI system an explicit written constitution — a ranked set of principles the system is trained and prompted to follow, and against which its own outputs are critiqued and revised.
This approach is valuable because it externalizes values into text that can be read, argued with, and versioned. It is insufficient on its own, because a constitution authored by a single laboratory, applied uniformly across every context, and enforced only by the model itself reproduces the concentration of authority DP11 exists to constrain.
DP11 therefore treats constitutional methods as necessary infrastructure with specific conditions attached.
### Constitutions must be public and versioned
The operative text, its ranking or precedence structure, and its change history must be publicly inspectable. Participants cannot evaluate a system against principles they cannot read, and silent revision is indistinguishable from having no constitution at all.
**Failure mode:** **constitutional opacity**, where a system claims principled behavior and declines to state the principles.
### Authorship must be accountable and plural
Whoever writes the constitution exercises governing power over everyone the system touches. DP11 requires that authorship be attributable, that the process for revision be stated, and that affected communities have a route to propose, contest, and — within their own zones — tighten provisions (DP12).
**Failure mode:** **unilateral constitutionalism**, where a vendor's values are presented as universal ethics.
### Constitutions must not be uniform across unequal stakes
A single global rule set cannot govern casual conversation and medical guidance appropriately. Baseline constraints apply everywhere; zones layer stricter provisions where stakes are higher (5.5, 5.14). Zone provisions may raise the floor but may not lower it.
**Failure mode:** **flattened ethics**, where uniformity is mistaken for fairness.
### Self-critique must not be the only enforcement
A model evaluating its own compliance is evidence, not verification. Constitutional adherence must be checked by mechanisms outside the model: capability envelopes, runtime policy binding, containment, logging, and independent audit (5.1, 5.6, DP12, DP13).
**Failure mode:** **self-certified compliance**, where the system grades its own behavior and reports a pass.
### Conflicts must resolve visibly
Principles collide: helpfulness against harm avoidance, transparency against privacy, autonomy against protection. A usable constitution declares precedence and records which principle governed a contested decision, so the tradeoff can be reviewed rather than inferred (5.13, 5.15).
**Failure mode:** **hidden arbitration**, where the system resolves value conflicts silently and consistently in its operator's favor.
### Drift must be detectable
Retraining, fine-tuning, prompt changes, and optimization pressure move behavior away from stated principles. DP11 requires ongoing measurement of the gap between constitutional text and observed behavior, with published results and rollback pathways (3.3).
**Failure mode:** **constitutional drift**, where the document remains stable while behavior does not.
### Participants must be able to see the constitution at the point of interaction
A constitution that exists only in a research paper does not inform judgment. Interfaces should surface the governing rule set, its version, and the zone provisions in force, with access to the full text (5.12, and **Governance, Accountability, and Agency Surfaces**).
**Failure mode:** **paper ethics**, where principles are published for reviewers rather than made available to the people affected.
**Example:** An assistant operating in a community health zone shows that it is governed by baseline constitution v4.2 plus three zone provisions requiring citation of clinical sources, prohibiting diagnostic language, and mandating escalation on distress indicators. When a response is modified, the interface names which provision applied. The zone's audit log shows how often each provision fired, and a quarterly report compares stated provisions against measured behavior.
**What this feels like:** The rules governing the system are something you can read, cite, and argue with — not something you infer from how it behaves.
## Personal and Community AI
Most current AI deployment places the agent on the operator's side of the relationship. The system knows the participant, is funded by someone else, and optimizes for objectives the participant cannot see. DP11's disclosure, bounding, and contestability requirements reduce the harm of that arrangement but do not change its structure.
Personal and community AI changes the structure. An agent that a participant or community controls — whose objectives they set, whose memory they hold, and whose operation they can inspect and stop — is the configuration in which ethical AI is easiest to sustain, because accountability and interest are aligned by default rather than by regulation.
DP11 does not require personal or community AI. It requires that this configuration remain possible, and that its specific risks be addressed rather than assumed away.
### Personal agents
A personal agent acts on behalf of one participant, under their instruction, with their data.
Conditions:
- **Principal clarity:** the agent's principal is unambiguous, disclosed to counterparties, and cannot be silently reassigned (DP1, 5.2) - **Participant-held memory and data:** context, history, and inferences remain under participant control, portable and deletable (DP4, DP7) - **Inspectable objectives:** the agent's goals are stated and editable by the participant, not inherited from a vendor's optimization target (5.12) - **Bounded delegation:** scopes are explicit, time-limited, renewable, and revocable, with a reachable stop control (5.1, 5.3) - **Disclosure to others:** counterparties can tell they are interacting with a delegated agent and identify the responsible human (**Governance, Accountability, and Agency Surfaces**) - **Exit without loss:** switching providers preserves memory, preferences, and history (DP7)
**Failure mode:** **captured advocate**, where an agent presented as acting for the participant is funded by, and optimized for, someone else.
**Failure mode:** **delegation without comprehension**, where a participant authorizes scopes they cannot evaluate and inherits responsibility for outcomes they did not anticipate.
### Community agents
A community agent acts on behalf of a group under its governance: summarizing deliberation, surfacing precedent, drafting policy, facilitating moderation, or maintaining shared knowledge.
Conditions:
- **Governed authority:** scope, capabilities, and escalation paths are set by community governance rather than by the operator (DP8, DP12) - **Attributable action:** every action names the agent, the authorizing rule, and the responsible steward (5.2, 5.15) - **Contestability by members:** members can challenge outputs, request human review, and trigger rollback (5.13) - **Bounded influence:** agents inform ranking, framing, and summary but do not hold decision authority over material outcomes (DP8, DP12) - **Plurality preservation:** summarization and synthesis must represent disagreement rather than manufacturing consensus (5.8, 5.16.3) - **Forkability:** communities can modify, replace, or leave with their accumulated context intact (DP7, DP20)
**Failure mode:** **synthetic consensus**, where a community agent's summaries become the record and minority positions disappear from it.
**Failure mode:** **governance outsourcing**, where deliberation load is shifted to an agent and participation atrophies until the agent's framing is the only framing available.
### Shared conditions
Both configurations require:
- capability envelopes and containment equivalent to any other agent, since ownership does not reduce risk to third parties (5.1, DP13) - confidence propagation, so that participant-owned agents do not launder uncertainty into confident personal advice (5.16) - resourcing models that do not reintroduce hidden optimization: local execution, participant-funded operation, or community-funded infrastructure (DP17) - literacy support, so that control is exercisable rather than nominal (DP10)
**Why this matters:** Ethical AI is easier to maintain when the agent's interests are structurally aligned with the participant's than when misalignment must be continuously disclosed, bounded, and policed. Personal and community AI are not a substitute for DP11's requirements. They are the arrangement in which those requirements are cheapest to satisfy and hardest to quietly abandon.
## Relationship to Other Desirable Properties
DP11 depends on and reinforces other properties. The following properties operate as a system.
- **DP1**: enables accountability and attribution - **DP2**: ensures participant agency and consent - **DP12**: provides governance structures for ethical rules - **DP13**: enforces constraints through containment
A failure in one layer propagates. For example, if DP1 fails and agents are not clearly attributable, then DP11 cannot function because ethical responsibility has no anchor. If DP13 fails, rules may exist but cannot be enforced.
The strength of DP11 therefore depends on alignment across the stack.
## Non-Goals and Explicit Boundaries
DP11 defines a minimum condition, not a comprehensive ethical system.
- it does not define a single universal ethical framework - it does not guarantee perfect safety or eliminate all harm - it does not replace legal or institutional governance - it does not rely solely on technical containment
This is intentional. Ethical systems that attempt to resolve everything tend to become brittle or culturally narrow.
For instance, a global platform may host communities with very different norms around acceptable AI behavior. DP11 does not force uniformity. It ensures that whatever rules are chosen remain visible, enforceable, and contestable.
These boundaries keep the property flexible while preserving its core function.
## Minimum DP11 Alignment (Non-Normative)
A system aligned with DP11 should, at minimum:
- clearly disclose AI presence and role - bind actions to accountable entities - expose capability boundaries in understandable terms - provide human escalation for high-stakes decisions - maintain audit trails of significant actions
These are not aspirational features. They are the baseline conditions under which users can make informed decisions.
Consider a scenario where an AI recommends a legal action. Without disclosure, accountability, and escalation, the user is effectively acting on anonymous authority. With these conditions in place, the same interaction becomes something the user can evaluate, question, or defer.
This is the difference between assistance and unaccountable influence.
## Open Questions and Future Work
Several areas require further development:
- defining shared ethical baselines across cultures and zones - balancing transparency with privacy and security - managing emotional and relational AI risks - defining evidence standards for runtime claims - the role of AI literacy in enabling meaningful consent and contestability
These are not edge cases. They represent the frontier where current design patterns begin to break down.
For example, companionship AI systems raise questions that are not purely technical: when does support become dependency? What level of disclosure is sufficient without undermining usefulness? These tensions are unresolved and will require iterative, community-informed approaches.
As systems become more complex, participants will vary widely in their ability to understand and evaluate AI behavior. While DP11 requires systems to be legible by design, differences in AI literacy will still shape how effectively users can exercise consent, recognize risk, and challenge outcomes. The balance between system responsibility and user capability remains an open design question.
## 14. Path Toward ML-RFC
Advancing DP11 toward standardization requires:
Progress depends on iteration, not premature standardization.
## 15. Closing Orientation
DP11 defines the conditions under which AI can participate in shared digital environments without displacing human moral agency.
Published: 2026-08-08
Pages: 14 | Words: 6934
What changed:
Synced from the book local rail (content/local/dpN.md), which carries the current working text for this chapter: expanded sections, renamed and renumbered headings, and editorial cleanup since the last revision. Published as a new revision so prior revisions stay intact.
Published: 2026-08-05
Pages: 14 | Words: 6917
What changed:
Numbered section headings and cross-reference fixes for collaborative review
Published: 2026-08-04
Pages: 14 | Words: 6890
What changed:
Synced from the book local rail (content/local/dpN.md), which carries the current working text for this chapter: expanded sections, renamed and renumbered headings, and editorial cleanup since the last revision. Published as a new revision so prior revisions stay intact.
Published: 2026-05-04
Pages: 12 | Words: 5629
What changed:
he upgraded DP11 evolves “Safe and Ethical AI” from a set of principles into a fully operational system. While the original draft established that AI must be disclosed, bounded, attributable, and contestable at the interface, the new version defines how those conditions are enforced in practice. It introduces structured mechanisms such as capability envelopes, event logging, risk tiers, and escalation pathways, ensuring that AI behavior is not only visible but auditable and governable in real time. Ethics is no longer treated as a design-time intention—it becomes a runtime property of the system.
A major advancement is the recognition of systemic and emergent risks. The upgraded draft addresses multi-agent dynamics, confidence distortion, and cross-modal inconsistencies, where AI behavior can diverge or amplify errors across text, voice, and immersive environments. It introduces confidence propagation rules, feedback-integrated reputation systems, and explicit incentive disclosure, ensuring that uncertainty, influence, and accountability persist as information moves through the system. This prevents common failure modes such as false consensus, overconfidence through repetition, and hidden optimization shaping user outcomes.
Overall, the new DP11 reframes AI safety as infrastructure rather than policy. It embeds ethical behavior into the mechanics of interaction through continuous feedback loops, enforceable boundaries, and governance-linked adaptation. By connecting AI actions to durable records, community oversight, and dynamic control systems, it ensures that trust is not assumed but continuously earned, evaluated, and adjusted over time.