DP21 - Multi-Modal Interactions & Experiences
Patches and comments are scoped to the whole document family, not to a single revision. When you open a revision, highlights show what still applies to that revision body; anything anchored to text a later revision removed is listed as needing re-anchoring, and anything that can no longer be merged at all is marked obsolete. See all patches.
Comparing Revision 00 (original) with Revision 01. Struck-through text was removed; highlighted text was added.
# DP21 – Multi-Modal Interactions and Experiences
# DP21 – Multi-modal
*Multi-Modal Interactions and Experiences — Meaning Before Presentation*
<!-- dp-local-version: 1.0 | standardized: 2026-07-27 -->
## 1. Purpose of This Draft
Failure mode: **fusion authoritarianism**, where the system produces a single authoritative view from contested inputs without preserving dissent or uncertainty.
## 7. Governance Requirements
## 7. Governance, Accountability, and Agency Surfaces
Multi-modal systems are governance systems because they shape perception, attention, and action.
Failure mode: **ungoverned modality expansion**, where new interface capabilities become active before the community has defined rules for them.
### 7.1 Accountability surfaces
Modality decisions are usually invisible to the people they affect. A caption is missing, a haptic cue is ambiguous, an overlay appears without a source, and there is no one to ask. DP21 requires named surfaces that make interface authority checkable:
- **A modality manifest per zone**, listing which modalities are active, which are optional, and which are prohibited - **A sensor permission record**, stating what is sensed, at what fidelity, for what purpose, retained for how long, and processed where - **An overlay registry** for spatial and immersive content, identifying publisher, governance zone, dispute status, and expiry - **An accessibility conformance statement** with known gaps, dates, and remediation owners, rather than an unqualified claim of compliance - **An AI mediation disclosure** covering summarization, translation, description, and transcription, including confidence and sampling practice (DP11–DP13) - **A degradation log**, recording when equivalence failed, what participants saw or heard instead, and who was notified
### 7.2 Participant agency surfaces
Multi-modal agency means participants shape how they are addressed and sensed, not only what they are shown:
- select preferred and prohibited modalities, and have those preferences persist and port (DP2, DP7) - refuse specific sensor classes without losing baseline participation - request an equivalent representation of any content encountered in an inaccessible form, with a response commitment - report modality failures through the modality they can actually use, including offline and assisted routes - see and contest AI-generated descriptions, captions, or summaries attributed to their own contributions - opt out of ambient, persuasive, or attention-shaping cues by default rather than by discovery - exit an immersive context immediately, with a predictable and always-available return path
**Example:** A participant using screen reader and haptic output sets a profile prohibiting ambient audio cues. Entering a public AR zone, they receive a structured spatial summary instead, with overlay sources listed and a route to challenge any unlabeled annotation.
**What this feels like:** The interface adapts to the person, and the person can see who decided what.
**Without this:** Accessibility becomes a feature request queue, and perception becomes something done to participants rather than with them.
Failure mode: **consent through exhaustion**, where refusing a modality is technically possible but practically requires abandoning participation.
## 8. Community Signals Informing DP21
DP21 reflects signals from disabled participants, educators, interpreters, low-bandwidth communities, immersive practitioners, and people whose primary interface is voice rather than screen.
Recurring signals include:
- "Accessibility arrives last and breaks first." Participants describe features shipped without equivalents, then repaired only after complaint. - "The audio version leaves out what matters." Alternative modalities are treated as summaries rather than as equivalents. - "I lost my place." Switching from phone to voice to headset loses context, state, and consent settings. - "Voice makes everything sound certain." Participants report that spoken interfaces strip uncertainty, provenance, and dissent. - "I don't know what is being sensed." Ambient microphones, cameras, and biometrics generate unease disproportionate to their stated purpose. - "AR content appeared over my street." Communities want authority over spatial annotation of places they inhabit. - "Captions are wrong in ways that change meaning." Automated transcription errors are experienced as misrepresentation, not inconvenience. - "Immersive spaces are hard to leave." Participants ask for predictable exits, especially for youth and first-time users. - "Low bandwidth means low status." Text-first and offline participants find themselves excluded from consequential discussion. - "Sign language is a language, not an accommodation." Requests for first-class support rather than fallback treatment (DP23).
These signals converge on one requirement: modality must not determine standing. DP21 encodes that requirement structurally rather than leaving it to product priorities.
Community signals inform DP21 without settling it. They are recorded so that later revisions can be tested against the experiences that motivated the property.
## 8.9. Evaluation Criteria
A DP21-aligned implementation should be evaluated against the following questions.
### 8.19.1 Continuity
- Does the same object retain identity across modalities? - Are trust signals, consent states, and governance zones preserved? - Can participants move across devices without losing essential context?
### 8.29.2 Accessibility
- Can participants with different abilities access the same meaning and agency? - Are accessibility tools receiving structured trust and governance information? - Are cognitive, sensory, motor, and linguistic needs considered?
### 8.39.3 Trust fidelity
- Do all modalities preserve provenance, uncertainty, confidence, and dispute status? - Are warnings and endorsements represented with appropriate severity? - Are simplified summaries faithful to the underlying record?
### 8.49.4 Privacy and sensor limits
- Are sensors governed explicitly? - Is data minimization enforced? - Can participants revoke modality-specific permissions? - Are local or privacy-preserving processing options used where possible?
### 8.59.5 Synchronization and freshness
- Are real-time states synchronized? - Is stale or delayed information labeled? - Can participants tell whether different modalities are showing the same system state?
### 8.69.6 Participant control
- Can participants choose preferred modalities? - Can they reduce stimulation or disable intrusive cues? - Can they control AI mediation and summarization?
### 8.79.7 Governance fit
- Are modality rules defined per zone? - Are immersive and ambient interfaces subject to community policy? - Are physical safety contexts handled responsibly?
## 9.10. Implementation Patterns
Implementation patterns translate DP21 from principle into practice. These are not rigid prescriptions, but recurring architectural and design approaches that help systems maintain semantic continuity, accessibility, and governance integrity across modalities. They are intended to guide builders in making consistent decisions when adapting the meta-layer to new devices, senses, and interaction environments.
### 9.110.1 Semantic-first architecture
Build semantic objects before rendering interfaces. Presentation layers should adapt governed meaning, not invent it.
### 9.210.2 Modality adapters
Create adapters for visual, voice, haptic, spatial, and assistive interfaces that consume the same underlying semantic objects.
### 9.310.3 Accessibility testing pipelines
Every core trust, consent, governance, and reputation signal should be tested across assistive modalities.
### 9.410.4 Context synchronization engines
State should travel across devices and modalities through privacy-preserving synchronization.
### 9.510.5 Sensor permission manifests
Multi-modal applications should publish clear manifests describing sensor access, purpose, retention, and sharing.
### 9.610.6 Degradation notices
When a modality cannot show full context, the interface should say what is missing and offer alternatives.
### 9.710.7 Voice uncertainty patterns
Voice systems should use standard language for uncertainty, provenance, dispute status, and request-for-detail prompts.
### 9.810.8 Spatial governance labels
AR and VR overlays should display or make available the governing zone, source, scope, and contestation path for spatial annotations.
### 9.910.9 Participant sensory profiles
Participants should be able to store preferences for stimulation level, modality, language, accessibility, and AI assistance.
### 9.1010.10 Multi-modal feedback receipts
Feedback submitted by voice, gesture, text, or spatial annotation should generate DP18-compatible receipts.
## 10.11. Relationship to Other Desirable Properties
### DP2 – Participant Agency and Empowerment
Multi-modal experiences make the meta-layer more present, memorable, and culturally accessible, especially through education, public installations, youth engagement, and immersive storytelling.
## 11. Open Questions for ML-RFC Development
## 12. Non-Goals and Explicit Boundaries
DP21 is often read as an argument for more interfaces. It is not. It is an argument that meaning, trust, and governance must survive the interface. The following boundaries are explicit.
**DP21 does not require every system to support every modality.** A text-only tool can be fully DP21-aligned. What DP21 requires is that supported modalities preserve meaning, that limits are declared, and that participants are not silently excluded.
**DP21 does not mandate immersive or spatial computing.** AR, VR, and ambient interfaces are permitted, not prescribed. Where they are used, they inherit heightened obligations for labeling, consent, exit, and physical safety.
**DP21 does not treat modality parity as visual equivalence.** Equivalence is about meaning, trust state, and available action, not about reproducing a layout in another form. A faithful audio equivalent may look nothing like the screen.
**DP21 does not authorize sensing as a convenience.** Cameras, microphones, biometrics, eye tracking, and location are governed under DP4 minimization. Improved personalization is not a sufficient justification for expanded sensing.
**DP21 does not permit ambient persuasion.** Cues that shape behavior below the threshold of notice are outside DP21 regardless of intent, because they cannot be consented to or contested.
**DP21 does not delegate meaning to AI.** Machine captioning, description, translation, and summarization are assistive and provisional. They may not become the authoritative record of what a participant expressed (DP11–DP14).
**DP21 does not replace human interpreters or accessibility professionals.** Automated equivalents extend reach; they do not discharge the obligation to provide human support where stakes are high.
**DP21 does not require perfect synchronization.** Cross-modal state will sometimes diverge. DP21 requires that divergence be detectable and disclosed, not that it be eliminated.
**DP21 does not make accessibility optional under deadline.** Degradation must be visible and bounded. Shipping an inaccessible path with a promise of later remediation is a DP21 failure, not a phase.
**DP21 does not govern content quality.** Whether an artifact is accurate or worthwhile belongs to DP14 and DP22. DP21 governs whether it can be perceived, trusted, and acted upon across forms.
**DP21 does not standardize hardware.** Device diversity is assumed, including low-cost, older, offline, and assistive devices. Requiring specific hardware to participate meaningfully is an exclusion, not a specification.
## 13. Minimum DP21 Alignment (Non-Normative)
A system claiming DP21 alignment SHOULD be able to demonstrate the following. These statements are non-normative and serve as an assessment aid rather than a conformance test.
### 13.1 Semantic core
Content and governance objects exist independently of presentation, with meaning, provenance, trust state, and available actions defined before rendering.
### 13.2 Declared modality support
There is a published statement of which modalities are supported, at what fidelity, and with what known gaps, per zone or surface.
### 13.3 Equivalence mapping
For each supported modality, there is a documented mapping showing how identity, trust signals, consent state, uncertainty, and dispute status are conveyed.
### 13.4 Accessibility floor
Text alternatives, captions, keyboard and switch access, adjustable timing, and screen reader compatibility exist as baseline conditions rather than optional enhancements, with conformance and gaps published.
### 13.5 Context continuity
A participant can move between devices and modalities with continuity of object identity, position, consent state, and governance zone, or receive an explicit notice that continuity was lost.
### 13.6 Sensor governance
Every sensor use has a declared purpose, retention limit, processing location, and refusal path, and refusal does not remove baseline participation.
### 13.7 AI mediation disclosure
Automated captioning, description, translation, and summarization are labeled, carry uncertainty where relevant, and are correctable by the participant or community affected.
### 13.8 Visible degradation
When a modality cannot faithfully represent an object, the system says so in that modality rather than presenting a degraded version as complete.
### 13.9 Participant profiles
Sensory, cognitive, and interaction preferences are storable, portable, and respected by default across surfaces (DP2, DP4, DP7).
### 13.10 Safety and exit
Immersive and physical-world interactions include situational awareness limits, prohibited-context rules, and an always-available exit that returns the participant to a predictable state.
## 14. Open Questions and Future Work
1. What minimum semantic object model enables modality-independent rendering? 2. How should trust signals map across visual, audio, haptic, spatial, and assistive forms? 3. What is the standard way to express uncertainty in voice interfaces? 4. What sensor permissions should be mandatory to disclose? 5. How should AR overlays represent governance zone, source, and dispute status? 6. What accessibility conformance standards should apply to meta-layer overlays? 7. How should multi-modal feedback objects be structured? 8. How should participant sensory preferences be stored and ported? 9. What are safe defaults for ambient interfaces? 10. How should systems indicate stale or desynchronized modality states? 11. How should AI summaries be audited for modality-specific distortion? 12. What physical-world safety constraints should apply to AR and haptic interactions? 13. How should multi-modal systems operate in low-bandwidth or crisis conditions? 14. What rights should participants have when modality-specific limitations affect governance participation?
## 12.15. Path Toward ML-RFC
DP21 is currently an ML-Draft and serves as exploratory scaffolding for how multi-modal interaction becomes part of meta-layer infrastructure.
DP21 will likely mature through several component RFCs rather than one monolithic standard.
## 13.16. Closing Orientation
DP21 ensures the meta-layer is not bound to a single interface, device, or sense.
Published: 2026-08-08
Pages: 13 | Words: 6151
What changed:
Synced from the book local rail (content/local/dpN.md), which carries the current working text for this chapter: expanded sections, renamed and renumbered headings, and editorial cleanup since the last revision. Published as a new revision so prior revisions stay intact.
Published: 2026-08-05
Pages: 13 | Words: 6137
What changed:
Numbered section headings and cross-reference fixes for collaborative review
Published: 2026-08-04
Pages: 13 | Words: 6127
What changed:
Synced from the book local rail (content/local/dpN.md), which carries the current working text for this chapter: expanded sections, renamed and renumbered headings, and editorial cleanup since the last revision. Published as a new revision so prior revisions stay intact.