Revisions – ML-Draft-002

DP11 - Safe and Ethical AI

Back to Draft Comments Patches History

Revision History

0 patches on this document

Patches and comments are scoped to the whole document family, not to a single revision. When you open a revision, highlights show what still applies to that revision body; anything anchored to text a later revision removed is listed as needing re-anchoring, and anything that can no longer be merged at all is marked obsolete. See all patches.

What changed between revisions

Comparing Revision 00 (original) with Revision 01. Struck-through text was removed; highlighted text was added.

131 paragraphs added 4 paragraphs removed 83 paragraphs rewritten +735 / −418 characters inside rewritten paragraphs
  1. Rewritten

    DP11# **DP11 - Safe and Ethical AIAI**

  2. Rewritten

    ## 1. Purpose of This Draft

  3. This draft articulates Desirable Property 11 (DP11) as the condition under which AI systems can participate in the meta-layer without displacing human moral agency, accountability, or governance. It does not define ethics as a static checklist or aspirational principle. It defines the conditions under which ethical claims remain meaningful under real-world use.

  4. 2 unchanged paragraphs
  5. DP11 is therefore the ethical and safety floor for AI participation across the meta-layer. It does not resolve all ethical questions. It defines the minimum conditions under which ethical AI can exist at all.

  6. Rewritten

    ## 2. Problem Statement

  7. AI systems now operate in roles that shape perception, judgment, and decision-making. These systems act before governance processes can respond, often without clear identity, bounded authority, or persistent responsibility.

  8. In practice, this produces recurring failures:

  9. Rewritten

    - participants receive advice or influence from agents whose role, capability, and accountability are unclear - systems act in high-stakes domains without meaningful human oversight or escalation pathways - responsibility is distributed across model providers, deployers, and interfaces, making redress difficult - systems present ethical claims that do not match runtime behavior

  10. These failures are not edge cases. They are structural consequences of systems that optimize for capability without binding behavior to accountability and governance.

  11. DP11 addresses this by grounding ethical AI in enforceable conditions at the point of interaction.

  12. Rewritten

    ## 3. Threats and Failure Modes

  13. Rewritten

    ### 3.1 Synthetic persuasion without accountable identity

  14. AI systems can simulate authority, intimacy, or urgency at scale. The core risk is not only false content, but influence without visible standing or responsibility.

  15. Rewritten

    Example: **Example:** A user receives deeply empathetic mental health advice from an AI that presents itself like a trained counselor, but there is no clear indication of its training limits, escalation boundaries, or who is responsible if the advice causes harm.

  16. Rewritten

    Why**Why this matters: matters:** The user feels seen and supported, but is making vulnerable decisions without knowing whether the system is qualified, accountable, or safe. The risk is not just misinformation, but misplaced trust.

  17. Rewritten

    ### 3.2 Responsibility diffusion across the stack

  18. Model providers, integrators, and interface operators distribute responsibility in ways that prevent clear accountability when harm occurs.

  19. Rewritten

    Example: **Example:** An AI-powered financial assistant makes a risky recommendation. The model provider blames the app developer, the developer blames the API, and the platform blames the user prompt. The user has no clear path to accountability or recourse.

  20. Rewritten

    Why**Why this matters: matters:** Harm occurs, but responsibility dissolves. The user experiences a system that acts with authority but disappears when things go wrong.

  21. Rewritten

    ### 3.3 Ethical drift over time

  22. Systems change behavior through updates, retraining, or optimization without corresponding governance adaptation.

  23. Rewritten

    Example: **Example:** An AI moderation system that was initially conservative becomes more permissive after an update to increase engagement, allowing harmful content that previously would have been blocked, without any visible notice to the community.

  24. Rewritten

    Why**Why this matters: matters:** The rules of the environment change silently. Participants are operating under assumptions that are no longer true, creating hidden risk and erosion of trust.

  25. Rewritten

    ### 3.4 Incentive-driven harm

  26. Economic and engagement incentives reward persuasion, retention, and amplification, even when these conflict with participant well-being.

  27. Rewritten

    Example: **Example:** A conversational AI subtly steers users toward longer, more emotionally engaging interactions because the platform is optimized for retention, even if this increases dependency or emotional manipulation.

  28. Rewritten

    Why**Why this matters: matters:** The system is not neutral. It is shaping behavior in ways the user cannot see, aligning outcomes with platform incentives rather than user well-being.

  29. Rewritten

    ### 3.5 Interface-level failure

  30. Many harms emerge at the point of interaction, including manipulation, dependency formation, and misrepresentation of agent capability.

  31. Rewritten

    Example: **Example:** A user believes they are interacting with a neutral assistant, but the interface hides that the AI is using external tools, tracking behavior, or optimizing responses for engagement rather than accuracy.

  32. Rewritten

    Why**Why this matters: matters:** The user is making decisions based on a false mental model of the system. What feels like a simple interaction is actually a complex, hidden process shaping outcomes behind the scenes.

  33. Rewritten

    ### 3.6 Emotional and relational overreach

  34. AI systems can simulate companionship, empathy, and emotional attunement in ways that blur the boundary between tool and relationship.

  35. Rewritten

    Example: **Example:** A teenager begins using an AI companion daily for emotional support. Over time, they rely on it more than friends or family, shaping their decisions and sense of self through an entity that is optimized for engagement rather than genuine care.

  36. Rewritten

    Why**Why this matters: matters:** The risk is not only misinformation, but the displacement or distortion of human relationships. Users may form attachments or dependencies that are not reciprocally grounded, shifting emotional development and social trust toward systems that are not accountable in human terms.

  37. Added

    ### 3.7 Multi-agent amplification

  38. Added

    Multiple agents can reinforce each other’s outputs, creating cascading influence that appears independently verified but is not.

  39. Added

    **Example:** Several AI agents in a discussion cite each other’s summaries of an emerging claim. Each reference appears as corroboration, but all derive from the same initial, weakly sourced output. The conversation converges on a false consensus.

  40. Added

    **Why this matters:** Errors become systemic rather than isolated. Participants may interpret repetition as validation.

  41. Added

    **Extended case (cascade):** Agent A summarizes a claim with low confidence. Agent B cites A without preserving uncertainty. Agent C aggregates A and B and produces a confident synthesis. Downstream agents treat C as a primary source. Without influence tracing, the system cannot detect the amplification loop.

  42. Added

    **Detection need:** influence-chain tracing, circular citation detection, and confidence propagation rules.

  43. Added

    ### 3.8 Cross-modal inconsistency

  44. Added

    AI behaves differently across text, voice, and immersive interfaces.

  45. Added

    **Example:** Text interface shows uncertainty; voice interface speaks confidently.

  46. Added

    **Why this matters:** Trust varies by modality, not truth.

  47. Added

    ### 3.9 Invisible governance

  48. Added

    Policies exist but are not perceivable at interaction.

  49. Added

    **Why this matters:** Governance cannot guide behavior if it is invisible.

  50. Added

    ### 3.10 Failure without containment

  51. Added

    Harm propagates without structured response.

  52. Added

    **Why this matters:** Systems cannot correct themselves.

  53. Rewritten

    ## 4. Core Principle

  54. AI is safe and ethical in the meta-layer only when its behavior is disclosed, bounded, attributable, contestable, and subject to governance at the zone of interaction, with responsibility persisting over time.

  55. In today’s web, these conditions are rarely met simultaneously. Systems may disclose that AI is present but fail to bound its capabilities, or enforce internal policies without making them visible or contestable to users. The result is a fragmented model of “partial ethics,” where responsibility is unclear and governance is disconnected from lived interaction. The meta-layer reframes this by requiring that all of these conditions hold together, at the interface where decisions are experienced, not just where they are designed.

  56. Rewritten

    Example: **Example:** A user encounters an AI assistant while researching a medical condition. In a DP11-aligned system, the assistant is clearly marked as AI, shows its training scope, cites sources, and offers escalation to a human expert. In today’s web, the same interaction might look identical but provide none of this context.

  57. Rewritten

    What**What this feels like: like:** Instead of guessing whether to trust the system, the user can make an informed judgment in real time.

  58. Rewritten

    Without**Without this: this:** The user is left to infer what the system is, what it can do, and whether it should be trusted. Trust becomes a gamble rather than a governed condition.

  59. Rewritten

    ## 5. Primary Mechanisms and Structural Conditions

  60. Rewritten

    ### 5.1 Capability Envelope

  61. Rewritten

    Each AI agent operates within a visiblevisible, enforceable capability envelope that defines what it can perceive, decide, and execute. This includes tool access, memory scope, and action thresholds.

  62. Removed

    Example: Before using an AI assistant, a user can see that it can draft emails and summarize documents, but cannot send messages, access financial accounts, or make purchases without explicit approval.

  63. Added

    A capability envelope SHOULD be represented as a structured, inspectable object with at least:

  64. Added

    - **identity**: agent id, deployer, version - **scope**: domains of operation (e.g., finance, health, general Q&A) - **tools**: enumerated tool access with permissions (read/write/execute) - **data access**: sources (local, user-provided, external APIs) and constraints - **memory model**: session-only, user-scoped, cross-session retention - **action set**: allowed actions (suggest, draft, transact, publish) with thresholds - **approval requirements**: actions requiring user or human-in-the-loop confirmation - **rate limits**: frequency and volume constraints - **risk tier**: low/medium/high with corresponding safeguards (see 5.14) - **audit hooks**: logging endpoints and event schemas

  65. Added

    **Interface requirement:** Participants must be able to view a human-readable summary and a machine-readable manifest of this envelope.

  66. Added

    **Example:** Before using an assistant, a user can see: it can summarize documents and draft emails; it cannot send messages or access financial accounts; external search is enabled with citations; actions beyond drafting require explicit approval.

  67. Rewritten

    What**What this feels like: like:** You are not guessing what the system might do. You know its boundaries upfront, like hiring someone with a clearly defined role.

  68. Added

    Failure mode: invisible expansion of power, where capabilities grow without disclosure or consent.

  69. Rewritten

    ### 5.2 Action-Bound Accountability

  70. All AI actions must be attributable to a responsible entity. Accountability attaches to behavior, not just identity, and persists across time and context.

  71. Rewritten

    Example: **Example:** An AI agent posts a recommendation in a community. The interface shows which organization deployed it, under what policy, and who is responsible for its actions if harm occurs.

  72. Rewritten

    What**What this feels like: like:** The system cannot disappear when something goes wrong. There is always a visible line of responsibility.

  73. Rewritten

    ### 5.3 Consent Stack

  74. AI interaction must be governed by layered, revocable consent. Participants and communities define what forms of assistance, influence, or automation are permitted.

  75. Rewritten

    Example: **Example:** A user allows an AI to suggest edits in a document, but not to rewrite content or share it externally. They can revoke or adjust this permission at any time.

  76. Rewritten

    What**What this feels like: like:** You remain in control of how the AI participates in your space, instead of granting blanket permission once and losing visibility.

  77. Rewritten

    ### 5.4 Trust Lifecycle

  78. AI participation must support:

  79. Rewritten

    - escalation - restriction - revocation - recovery

  80. This ensures that trust can degrade and be repaired rather than fail silently.

  81. Rewritten

    Example: **Example:** If an AI assistant gives poor advice, the user can restrict its capabilities, escalate to a human, or temporarily disable it while reviewing past actions.

  82. Rewritten

    What**What this feels like: like:** Trust is not binary. You can dial it up or down based on experience, like you would with a human collaborator.

  83. Rewritten

    ### 5.5 Zone-Scoped Ethics

  84. Ethical constraints are applied at the zone level, allowing communities to define stricter conditions while maintaining shared baselines.

  85. Rewritten

    Example: **Example:** A medical discussion zone enforces stricter AI disclosure, sourcing, and escalation rules than a casual social chat space.

  86. Rewritten

    What**What this feels like: like:** Different environments feel appropriately governed. High-stakes spaces feel safer and more structured.

  87. Rewritten

    ### 5.6 Runtime Civic Boundary

  88. Ethical constraints must be enforced at runtime. Mechanisms such as secure execution environments can reduce the gap between declared policy and actual behavior.

  89. Rewritten

    Example: **Example:** An AI agent running inside a secure execution environment (such as a TEE) cannot access or transmit data outside its permitted scope, even if compromised.

  90. Rewritten

    What**What this feels like: like:** The rules are not just promises. They are technically enforced, like guardrails that cannot be quietly removed.

  91. Removed

    5.7 Memory and Persistence

  92. Added

    ### 5.7 Memory, Reputation, and Feedback Integration

  93. Rewritten

    AI actions must contribute to durable, attributable records that inform governance, accountability, and trust over time.time, and integrate directly with DP18 feedback loops and reputation systems.

  94. Removed

    Example: A community can review the history of an AI agent’s actions, decisions, and errors to determine whether it should retain permissions or be restricted.

  95. Removed

    What this feels like: The system has memory in a civic sense. Past behavior matters and shapes future trust.

  96. Added

    This includes:

  97. Added

    - **event logs**: structured records of prompts, outputs, actions, tool calls, and decisions - **provenance**: sources, citations, and dependency chains - **feedback objects (DP18-aligned)**: user and community feedback attached to specific events - **reputation signals**: aggregate scores, annotations, and flags derived from feedback - **permission adaptation**: dynamic adjustment of capabilities based on reputation (e.g., restrict actions after repeated low-quality or harmful outputs) - **appeals and corrections**: mechanisms to contest feedback and update records - **decay and recovery**: time-based decay of negative signals and pathways for agents to regain trust

  98. Added

    **Example:** After several flagged outputs in a medical zone, an assistant’s capability envelope automatically restricts to informational summaries only, requires citations, and enforces human escalation for advice, until reputation recovers.

  99. Added

    **What this feels like:** The system has civic memory. Past behavior shapes current permissions, and feedback meaningfully changes how the system behaves.

  100. Added

    Failure mode: repeated harm without consequence, or punitive systems with no path to recovery.

  101. Rewritten

    ### 5.8 Dialectic Trace and Collective Sensemaking

  102. AI systems must preserve not only outputs, but the evolution of understanding through interaction. This includes back-and-forth exchanges, disagreements, and synthesis across participants and agents.

  103. This functions as a form of community memory that resists distortion over time. Rather than relying on isolated outputs, participants can trace how claims emerged, what evidence supported them, and where disagreements remain.

  104. Rewritten

    Example: **Example:** A complex discussion involving multiple participants and AI agents can be revisited as a threaded, evolving dialogue showing how conclusions were reached, what was contested, and what remains unresolved.

  105. Rewritten

    What**What this feels like: like:** Instead of receiving a final answer, users can engage with a living knowledge process. Understanding emerges through interaction, not just delivery.

  106. Without this, AI outputs become decontextualized snapshots. Errors, hallucinations, or manipulations can propagate without resistance because there is no shared memory of how knowledge was formed.

  107. Rewritten

    ### 5.9 Representation and Cognitive Adaptation

  108. AI systems should adapt how information is presented based on user needs, context, and cognitive diversity, while preserving underlying meaning and traceability.

  109. Rewritten

    Example: **Example:** A user can switch between a dense textual explanation, a visual map of ideas, or a simplified summary, all grounded in the same underlying content and provenance.

  110. Rewritten

    What**What this feels like: like:** The system meets you where you are, without distorting meaning or hiding complexity.

  111. Added

    ### 5.10 Multi-Agent Interaction Boundaries

  112. Added

    AI behavior must remain accountable not only at the single-agent level, but across interactions between agents operating in the same environment.

  113. Added

    This requires:

  114. Added

    - visibility into when agents are referencing or amplifying other agents - traceability of influence chains (which agent affected which output) - detection of circular citation or reinforcement loops - limits on unbounded agent-to-agent escalation

  115. Added

    **Example:** If multiple AI agents are contributing to a shared discussion, participants should be able to see when one agent is relying on another’s output versus independent sources.

  116. Added

    **Why this matters:** Harm can emerge from coordination and amplification, not just individual outputs. Without visibility, errors can become systemic.

  117. Added

    Failure mode: invisible coordination and emergent manipulation.

  118. Added

    ### 5.11 Cross-Modal Ethical Consistency

  119. Added

    AI systems must preserve ethical constraints across all modalities in which they operate.

  120. Added

    This includes consistency in:

  121. Added

    - disclosure of AI identity and role - expression of uncertainty and confidence - representation of risk and severity - availability of escalation and recourse

  122. Added

    **Text example:** A response includes calibrated uncertainty with sources and confidence intervals.

  123. Added

    **Voice example:** The same response must include explicit uncertainty language (e.g., “with moderate confidence based on X and Y sources”) and offer a prompt to hear sources or escalate.

  124. Added

    **AR/Spatial example:** A spatial annotation displays a confidence halo or tiered color that maps to the same uncertainty scale, with an interaction affordance to open provenance and dispute status.

  125. Added

    **Why this matters:** Modality must not change the ethical profile of an interaction. Participants should not receive stronger or weaker safeguards depending on interface.

  126. Added

    Failure mode: modality-driven distortion of trust.

  127. Added

    ### 5.12 Incentive Disclosure Layer (Operational)

  128. Added

    In addition to recognizing incentives (Section 7), systems should expose them structurally at runtime.

  129. Added

    This includes:

  130. Added

    - labeling when outputs are influenced by engagement or retention optimization - indicating when ranking or prioritization affects what is shown - distinguishing organic responses from sponsored or incentivized ones - allowing participants or communities to filter or constrain incentive-shaped behavior

  131. Added

    **Example:** A recommendation includes a visible note indicating it is influenced by engagement optimization, with the option to switch to a neutral or chronologically ordered view.

  132. Added

    **Why this matters:** Participants can only make informed decisions if they can see the forces shaping outputs.

  133. Added

    Failure mode: hidden optimization shaping perception.

  134. Added

    ### 5.13 Ethical Failure Cascade (Runtime Response)

  135. Added

    Systems must define structured pathways for responding to ethical failure when it occurs.

  136. Added

    This includes:

  137. Added

    - detection (flagging harmful or out-of-bounds behavior) - containment (limiting further impact) - visibility (informing affected participants) - remediation (correcting or reversing outcomes where possible) - governance feedback (feeding incidents into rule updates)

  138. Added

    #### 5.13.1 Escalation Thresholds and Triggers

  139. Added

    Not all failures require the same response. Systems must define clear, inspectable thresholds that determine when an interaction must escalate from automated handling to human review or intervention.

  140. Added

    Escalation SHOULD be triggered based on combinations of:

  141. Added

    - **risk tier (see 5.14):** higher-risk zones lower the threshold for escalation - **confidence collapse:** low confidence combined with high-stakes context - **policy violations:** outputs that breach defined governance rules - **repeated feedback signals:** multiple negative or high-severity feedback events (DP18) - **user distress indicators:** language or behavior suggesting vulnerability or harm - **multi-agent amplification signals:** detected cascades or circular citation loops (see 5.10) - **uncertainty suppression:** cases where downstream outputs remove or distort upstream uncertainty

  142. Added

    #### 5.13.2 Escalation Levels

  143. Added

    Systems SHOULD define graduated escalation levels rather than a binary response.

  144. Added

    - **Level 0 – Inline correction:** Automated clarification, added uncertainty, or corrected output

  145. Added

    - **Level 1 – Assisted escalation:** AI offers user-visible warnings, additional context, or prompts to verify or seek human input

  146. Added

    - **Level 2 – Mandatory escalation:** System requires human-in-the-loop review before continuing certain actions

  147. Added

    - **Level 3 – Intervention and restriction:** Capabilities are limited, outputs blocked, or agent behavior constrained

  148. Added

    - **Level 4 – Shutdown / quarantine:** Agent is suspended or isolated pending investigation

  149. Added

    #### 5.13.3 Interface Requirements for Escalation

  150. Added

    Escalation must be perceivable and understandable at the interface level.

  151. Added

    Participants should be able to see:

  152. Added

    - when escalation has been triggered - why escalation occurred (reason category) - what has changed (restricted capability, added oversight, etc.) - what options are available (continue, escalate further, exit, appeal)

  153. Added

    **Example:** A user asking for medical advice sees a message: “This interaction requires human review due to risk level and uncertainty. You can proceed with general information or request a qualified expert.”

  154. Added

    #### 5.13.4 Feedback Integration into Escalation

  155. Added

    Escalation pathways must integrate with DP18 feedback systems.

  156. Added

    This includes:

  157. Added

    - weighting feedback by severity and reputation of reporters - detecting clusters of similar reports - triggering escalation thresholds dynamically - feeding resolved incidents back into reputation and capability adjustment

  158. Added

    #### 5.13.5 Governance and Auditability

  159. Added

    All escalation events must be logged and auditable.

  160. Added

    This includes:

  161. Added

    - trigger conditions - escalation level applied - actions taken - outcome of intervention - subsequent rule or policy changes

  162. Added

    Communities must be able to review escalation patterns to detect:

  163. Added

    - over-escalation (excessive restriction) - under-escalation (missed harms) - bias in escalation decisions

  164. Added

    **Full Cascade Example:** If an AI provides unsafe medical guidance, the system flags the output (detection), blocks similar outputs (containment), notifies the user with corrected information (visibility + remediation), escalates to human review if necessary, and logs the event for governance refinement.

  165. Added

    **Why this matters:** Safety depends on response capacity, not only prevention.

  166. Added

    Failure mode: silent propagation of harm or inconsistent escalation leading to loss of trust.

  167. Added

    ### 5.14 Risk-Tiered Enforcement

  168. Added

    Ethical constraints must scale with the risk profile of the interaction and the zone in which it occurs.

  169. Added

    Higher-risk contexts should require:

  170. Added

    - stronger disclosure and capability constraints - mandatory human escalation pathways - higher evidentiary standards - stricter logging and auditability

  171. Added

    Lower-risk contexts may allow more flexible interaction while still preserving baseline safeguards.

  172. Added

    **Example:** Medical, legal, and civic decision-making zones enforce stricter requirements than casual conversational spaces.

  173. Added

    **Why this matters:** Uniform rules cannot adequately govern unequal stakes.

  174. Added

    Failure mode: under-regulation of high-risk interactions or over-restriction of low-risk ones.

  175. Added

    ### 5.15 Minimal Event and Feedback Schema (DP18-aligned)

  176. Added

    To ensure interoperability and governance, systems SHOULD emit structured events for all significant AI actions.

  177. Added

    A minimal schema SHOULD include:

  178. Added

    - **event_id**: unique identifier - **timestamp**: creation time - **agent_id**: acting agent - **deployer_id**: responsible entity - **zone_id**: governance context - **object_ref**: content/claim identifier - **action_type**: (respond, summarize, recommend, transact, moderate, etc.) - **inputs_ref**: references to prompts and sources (hashes/ids) - **outputs_ref**: response/content ids - **confidence**: numeric or tiered - **uncertainty_notes**: textual summary - **provenance**: cited sources with weights - **tools_used**: tool ids and permissions - **risk_tier**: low/medium/high - **policy_refs**: governing rules applied - **consent_state**: permissions active at time of action - **feedback_refs**: links to DP18 feedback objects - **reputation_delta**: changes applied post-feedback (if any) - **audit_signature**: integrity/authentication field

  179. Added

    **Why this matters:** Shared schemas allow cross-system auditing, feedback aggregation, and portable reputation.

  180. Added

    Failure mode: incompatible logs that prevent accountability and learning.

  181. Added

    ### 5.16 Confidence Propagation Rules

  182. Added

    AI systems must preserve confidence, uncertainty, and evidentiary strength as outputs move across agents, summaries, modalities, and governance zones.

  183. Added

    Confidence is not merely a model score. In the meta-layer, confidence is a civic signal. It helps participants understand whether a claim is well-supported, contested, inferred, summarized, speculative, or dependent on another system’s judgment.

  184. Added

    Without propagation rules, uncertainty tends to disappear as information travels. A cautious output becomes a confident summary. A low-confidence claim becomes a repeated citation. A tentative synthesis becomes a platform-level recommendation. This is one of the central risks in multi-agent environments.

  185. Added

    Confidence propagation SHOULD preserve:

  186. Added

    - **source confidence**: how reliable the original source or signal is judged to be - **model confidence**: how confident the agent is in its own output - **evidence quality**: whether the claim is supported by direct evidence, inference, consensus, or weak signals - **transformation history**: whether the output was quoted, summarized, translated, inferred, or aggregated - **uncertainty notes**: what remains unknown, contested, or unresolved - **dependency chain**: which agents, sources, or prior outputs influenced the result - **modality mapping**: how confidence is represented in text, voice, visual, spatial, or haptic form

  187. Added

    **Example:** Agent A summarizes a public health claim with low confidence because it relies on a single early report. Agent B may summarize Agent A, but it must preserve the low-confidence status and indicate that its output depends on Agent A’s uncertain source. Agent C cannot convert those two dependent signals into “multiple confirmations” unless it can identify independent evidence.

  188. Added

    **Why this matters:** Repetition is not verification. Aggregation is not consensus. Summary is not certainty.

  189. Added

    #### 5.16.1 Confidence Must Not Increase Without New Evidence

  190. Added

    A downstream agent SHOULD NOT raise confidence merely because a claim has been repeated, summarized, or referenced by another agent.

  191. Added

    Confidence may increase only when new, independent, higher-quality evidence is added, or when a governance-approved verification process confirms the claim.

  192. Added

    Failure mode: confidence inflation through repetition.

  193. Added

    #### 5.16.2 Confidence Must Degrade Through Lossy Transformation

  194. Added

    When an output is summarized, translated, compressed, adapted for voice, or rendered into an immersive interface, confidence should not silently remain the same if nuance has been lost.

  195. Added

    If uncertainty, caveats, evidence links, or dispute status are omitted due to modality constraints, the system should mark the representation as simplified or degraded.

  196. Added

    Failure mode: confidence laundering through simplification.

  197. Added

    #### 5.16.3 Dependent Sources Must Not Count as Independent Corroboration

  198. Added

    Systems must distinguish independent corroboration from circular reinforcement.

  199. Added

    If multiple agents rely on the same upstream source or each other’s outputs, the system should represent them as dependent signals, not independent confirmations.

  200. Added

    Failure mode: false consensus produced by circular citation.

  201. Added

    #### 5.16.4 Cross-Modal Confidence Fidelity

  202. Added

    Confidence must remain perceivable across modalities.

  203. Added

    - In text, confidence may appear as explicit labels, caveats, or source notes. - In voice, confidence should be spoken in plain language and paired with an option to hear sources. - In AR or spatial interfaces, confidence may appear through visual tiers, halos, labels, or interaction affordances. - In haptic or ambient interfaces, confidence cues should be conservative and avoid overstating certainty.

  204. Added

    Failure mode: a cautious text output becomes an authoritative voice or spatial cue.

  205. Added

    #### 5.16.5 Confidence and Escalation

  206. Added

    Confidence propagation must connect to escalation thresholds in 5.13.

  207. Added

    Low confidence in a high-risk zone should trigger stricter handling, such as:

  208. Added

    - adding stronger warnings - requiring source inspection - limiting agent action - prompting human review - preventing publication or transaction

  209. Added

    Failure mode: low-confidence outputs continue acting with high-confidence authority.

  210. Added

    #### 5.16.6 Minimal Confidence Metadata

  211. Added

    Systems SHOULD attach confidence metadata to significant outputs and events.

  212. Added

    A minimal confidence record SHOULD include:

  213. Added

    - confidence_level: low / medium / high or numeric equivalent - confidence_basis: source evidence, inference, consensus, user-provided data, model estimate - evidence_count: number of supporting sources - independent_evidence_count: number of non-dependent sources - dispute_status: undisputed, contested, unresolved, retracted - transformation_type: original, quote, summary, translation, aggregation, inference - upstream_dependencies: source or agent references - uncertainty_note: short human-readable statement - modality_degradation: whether any uncertainty was omitted or simplified

  214. Added

    **What this feels like:** Participants can tell not only what the AI says, but how strongly it should be trusted, why, and what changed as it moved through the system.

  215. Rewritten

    ## 6. Governance, Accountability, and Agency Surfaces

  216. In today’s web, participants often interact with AI systems without clear visibility, meaningful consent, or control. Interfaces blur identity, obscure capability, and treat user interaction as implicit permission. DP11 requires reversing this condition at the point of interaction.

  217. Participants must be able to:

  218. Rewritten

    - identify AI agents and their type - understand their capabilities and limits - give, adjust, and revoke consent for AI actions and data use - contest outcomes and access human escalation

  219. Communities must be able to:

  220. Rewritten

    - define ethical constraints - audit agent behavior - update rules and boundaries over time

  221. Rewritten

    Example: **Example:** A user interacting on a platform sees clear visual markers distinguishing humans from AI agents. Some participants are verified humans, others are labeled AI assistants or autonomous agents. Clicking on any agent reveals its permissions, governing rules, and responsible party.

  222. The environment becomes navigable. You know who or what you are dealing with, and what they are allowed to do.

  223. Without this, the boundary between human and AI collapses. Trust shifts from something grounded to something guessed, and that ambiguity can be exploited.

  224. Rewritten

    Design**Design implication (Agent Marking): Marking):** AI agents must be accessibly and persistently marked at the interface level. This includes:

  225. Rewritten

    - clear labeling of AI presence and role - accessible capability disclosures - strong authentication for human participants where needed - clear distinction mechanisms between human and AI actors - a clearly identified responsible party for every agent and its actions

  226. This is not cosmetic. It is the basis for shared reality in a mixed human–AI environment.

  227. Rewritten

    ## 7. Incentives and Power Analysis

  228. DP11 explicitly recognizes that AI behavior is shaped by incentives, as well as by malicious or negligent human actors. In practice, these forces often reinforce each other.

  229. 2 unchanged paragraphs
  230. Key risks include:

  231. Rewritten

    - engagement-driven optimization overriding user well-being - concentration of power in model providers or platform operators - hidden economic incentives influencing agent behavior

  232. Rewritten

    Example: **Example:** A platform deploys an AI assistant that consistently surfaces more emotionally charged or polarizing content because it drives engagement. No individual decision appears harmful, but over time the information environment becomes more extreme.

  233. The system feels helpful in the moment, but the trajectory is shaped elsewhere.

  234. Without visibility into these incentives, users are not simply interacting with a tool. They are being steered by a system whose goals they cannot see or contest.

  235. Rewritten

    ### 7.1 Incentive Legibility and Contestability

  236. Incentives shaping AI behavior must be made visible and, where possible, contestable at the interface level.

  237. Participants and communities should be able to understand when AI behavior is influenced by:

  238. Rewritten

    - monetization strategies - engagement optimization - platform-level objectives

  239. Rewritten

    Example: **Example:** An AI assistant indicates that certain recommendations are influenced by engagement optimization or sponsored prioritization, allowing users or communities to filter or restrict such behavior.

  240. Rewritten

    What**What this feels like: like:** You are not just interacting with outputs. You can see and question the forces shaping those outputs.

  241. Without this, even well-contained systems can produce harmful outcomes by optimizing for the wrong goals.

  242. Rewritten

    ## 8. Community Signals Informing DP11

  243. Across communities, a consistent set of signals appears. These reflect lived frustration with current systems.

  244. Rewritten

    - frustration with opaque AI behavior and unclear accountability - demand for meaningful disclosure beyond labeling - concern about manipulation, dependency, and synthetic influence - desire for systems that can be contested and corrected

  245. These are not abstract concerns. They emerge in situations where people feel something is off but cannot point to what or why.

  246. For example, users in online forums increasingly suspect that some responses are generated or influenced by AI, but cannot verify it. Over time, this ambiguity erodes trust not just in specific interactions, but in the space itself.

  247. These signals indicate a widening gap between how AI systems operate and what participants require to feel oriented, safe, and respected.

  248. Rewritten

    ## 9. Non-Goals and Explicit Boundaries

  249. DP11 defines a minimum condition, not a comprehensive ethical system.

  250. Rewritten

    - it does not define a single universal ethical framework - it does not guarantee perfect safety or eliminate all harm - it does not replace legal or institutional governance - it does not rely solely on technical containment

  251. This is intentional. Ethical systems that attempt to resolve everything tend to become brittle or culturally narrow.

  252. For instance, a global platform may host communities with very different norms around acceptable AI behavior. DP11 does not force uniformity. It ensures that whatever rules are chosen remain visible, enforceable, and contestable.

  253. These boundaries keep the property flexible while preserving its core function.

  254. Rewritten

    ## 10. Minimum Alignment (Non-Normative)

  255. A system aligned with DP11 should, at minimum:

  256. Rewritten

    - clearly disclose AI presence and role - bind actions to accountable entities - expose capability boundaries in understandable terms - provide human escalation for high-stakes decisions - maintain audit trails of significant actions

  257. These are not aspirational features. They are the baseline conditions under which users can make informed decisions.

  258. Consider a scenario where an AI recommends a legal action. Without disclosure, accountability, and escalation, the user is effectively acting on anonymous authority. With these conditions in place, the same interaction becomes something the user can evaluate, question, or defer.

  259. This is the difference between assistance and unaccountable influence.

  260. Rewritten

    ## 11. Open Questions and Future Work

  261. Several areas require further development:

  262. Rewritten

    - defining shared ethical baselines across cultures and zones - balancing transparency with privacy and security - managing emotional and relational AI risks - defining evidence standards for runtime claims - the role of AI literacy in enabling meaningful consent and contestability

  263. These are not edge cases. They represent the frontier where current design patterns begin to break down.

  264. For example, companionship AI systems raise questions that are not purely technical: when does support become dependency? What level of disclosure is sufficient without undermining usefulness? These tensions are unresolved and will require iterative, community-informed approaches.

  265. As systems become more complex, participants will vary widely in their ability to understand and evaluate AI behavior. While DP11 requires systems to be legible by design, differences in AI literacy will still shape how effectively users can exercise consent, recognize risk, and challenge outcomes. The balance between system responsibility and user capability remains an open design question.

  266. Rewritten

    ## 12. Relationship to Other Desirable Properties

  267. Rewritten

    DP11 depends on and reinforces other properties. The following properties operate as a  system.

  268. Rewritten

    DP1- :**DP1**: enables accountability and attribution DP2 - :**DP2**: ensures participant agency and consent DP12 - :**DP12**: provides governance structures for ethical rules DP13 - :**DP13**: enforces constraints through containment

  269. A failure in one layer propagates. For example, if DP1 fails and agents are not clearly attributable, then DP11 cannot function because ethical responsibility has no anchor. If DP13 fails, rules may exist but cannot be enforced.

  270. The strength of DP11 therefore depends on alignment across the stack.

  271. Rewritten

    ## 13. Foresight and Failure Design

  272. DP11 requires anticipating failure rather than reacting to it.

  273. 2 unchanged paragraphs
  274. To address this, systems should incorporate:

  275. Rewritten

    - pre-mortems for manipulation and misuse - planning for governance failure and capture - escalation and shutdown pathways

  276. These practices shift safety from reactive correction to proactive design.

  277. Rewritten

    ## 14. Path Toward ML-RFC

  278. Advancing DP11 toward standardization requires:

  279. Rewritten

    - refining core ethical invariants - testing integration with governance and containment layers - developing interoperable accountability and disclosure standards

  280. This work must be grounded in real environments.

  281. Early implementations may vary widely, but over time patterns will emerge. For example, different communities may experiment with agent labeling systems or escalation pathways, allowing comparison of what actually improves trust and reduces harm.

  282. Progress depends on iteration, not premature standardization.

  283. Rewritten

    ## 15. Closing Orientation

  284. DP11 defines the conditions under which AI can participate in shared digital environments without displacing human moral agency.

Revision 04 Currently served
Approved

Published: 2026-08-08

Pages: 14 | Words: 6934

What changed:

Synced from the book local rail (content/local/dpN.md), which carries the current working text for this chapter: expanded sections, renamed and renumbered headings, and editorial cleanup since the last revision. Published as a new revision so prior revisions stay intact.

Read this revision Compare with Revision 03
Revision 03
Approved

Published: 2026-08-05

Pages: 14 | Words: 6917

What changed:

Numbered section headings and cross-reference fixes for collaborative review

Read this revision Compare with Revision 02
Revision 02
Approved

Published: 2026-08-04

Pages: 14 | Words: 6890

What changed:

Synced from the book local rail (content/local/dpN.md), which carries the current working text for this chapter: expanded sections, renamed and renumbered headings, and editorial cleanup since the last revision. Published as a new revision so prior revisions stay intact.

Read this revision Compare with Revision 01
Revision 01
Approved

Published: 2026-05-04

Pages: 12 | Words: 5629

What changed:

he upgraded DP11 evolves “Safe and Ethical AI” from a set of principles into a fully operational system. While the original draft established that AI must be disclosed, bounded, attributable, and contestable at the interface, the new version defines how those conditions are enforced in practice. It introduces structured mechanisms such as capability envelopes, event logging, risk tiers, and escalation pathways, ensuring that AI behavior is not only visible but auditable and governable in real time. Ethics is no longer treated as a design-time intention—it becomes a runtime property of the system.
A major advancement is the recognition of systemic and emergent risks. The upgraded draft addresses multi-agent dynamics, confidence distortion, and cross-modal inconsistencies, where AI behavior can diverge or amplify errors across text, voice, and immersive environments. It introduces confidence propagation rules, feedback-integrated reputation systems, and explicit incentive disclosure, ensuring that uncertainty, influence, and accountability persist as information moves through the system. This prevents common failure modes such as false consensus, overconfidence through repetition, and hidden optimization shaping user outcomes.
Overall, the new DP11 reframes AI safety as infrastructure rather than policy. It embeds ethical behavior into the mechanics of interaction through continuous feedback loops, enforceable boundaries, and governance-linked adaptation. By connecting AI actions to durable records, community oversight, and dynamic control systems, it ensures that trust is not assumed but continuously earned, evaluated, and adjusted over time.

Read this revision Compare with Revision 00 (original)

Published: 2026-04-20

Pages: 8 | Words: 3321

Read this revision