
A robot copilot is allowed to read diagnostics and propose a docking task. The architecture diagram puts a command validator between the copilot and motion. Then someone gives both processes the same ROS 2 security credentials, or composes their nodes into one context. The diagram still shows a boundary. The deployed system may let the copilot publish directly to the controller.
ROS 2 security enclaves make that boundary enforceable at the middleware layer only when the credentials, process layout and permissions agree. Assign narrowly scoped identities to separate trust boundaries, generate permissions for the resolved communication graph, protect each identity’s private key, and test prohibited paths with the production deployment configuration.
The useful design question is specific: if the AI-facing process is compromised, which ROS interfaces can it actually use? A list of node names cannot answer it. An enclave-to-interface map, checked against the running system, can.
This article uses the DDS Security model and the official ROS 2 Jazzy security documentation as its baseline. Pin and qualify your ROS distribution, RMW implementation and DDS version together; do not assume these controls apply unchanged to every ROS middleware. The robot and interface names below are an illustrative design, not a tested hardware deployment.
The security identity belongs below the node graph
The ROS 2 security enclave design explains the essential mapping: nodes sharing a context share its middleware participant and security identity. Their permissions must be combined. An enclave selects security artifacts; it is not a hardware trusted execution environment or an operating-system sandbox.
Node namespaces and enclave paths are separate identifier spaces. A node called /r17/copilot does not acquire an independent identity merely because its name differs from /r17/controller. Conversely, an enclave path need not match the node namespace. Keep an explicit mapping between them.
That changes a common ROS 2 node-granularity decision. Composing perception stages that share one trust level may be reasonable. Composing a plugin-rich AI adapter with an actuator-facing command gateway deserves a security review even if it improves latency.
The practical test is about compromise, not naming. If arbitrary code runs in the copilot’s process, can it access the gateway’s objects, memory or key files? Separate contexts inside the same process do not provide a convincing answer to that threat. Where different authority must survive a process compromise, use separate processes and OS access controls; choose stronger host isolation when the threat model requires it.
Likewise, a different ROS_DOMAIN_ID is useful for graph organization but is not a credential. Do not treat selecting a domain number as proof that a participant belongs on the robot.
Draw the permission map before generating certificates
Consider a mobile robot with an AI operator interface, a deterministic command gate, a navigation process and a base controller. A recorder collects selected evidence. Start with these intended application capabilities:
| Enclave / process | Allowed application interfaces | Explicitly excluded |
|---|---|---|
/r17/copilot | Subscribe to /r17/diagnostics; request /r17/command_gate/propose | Navigation action calls, velocity publication, controller parameter writes |
/r17/command_gate | Reply to proposal requests; call /r17/navigate_to_pose | Raw motor interfaces, unrestricted parameter changes |
/r17/navigation | Serve the navigation action; publish /r17/cmd_vel; read approved localization and obstacle streams | AI tool execution, credential management |
/r17/base | Subscribe to velocity requests; publish base state | AI proposal handling, unrelated sensor access |
/r17/recorder | Subscribe to the selected diagnostics and state streams | Motion requests, parameter writes, bag playback onto the live graph |
This is a review artifact, not a complete SROS2 policy. Each row also needs the narrowly scoped infrastructure interfaces used by that implementation. The exact controller topic and command path depend on the robot.
Add three columns in the project version: executable or component-container identity, effective OS principal, and deployed credential location. Those columns reveal a frequent contradiction: separate enclave names backed by key files readable by every process.
The command gate remains responsible for the meaning and admissibility of a robot command. DDS permission to call a proposal service does not establish operator approval, a valid destination, acceptable speed or current robot mode. The gate must reject invalid requests even from an authenticated client.
For the same reason, do not forward an untrusted proposal unchanged through a privileged bridge. The bridge’s downstream credential proves that the bridge sent the message. It does not prove that the upstream requester was entitled to cause it. Record the authenticated request context, validate the permitted operation, and construct the downstream request from accepted fields. A caller-supplied operator_id is only a claim until the application authenticates it.
Review the interfaces hidden behind an action
A permission review that looks only at ordinary topic lists misses operational authority. The ROS 2 action design uses three services and two topics per action, with hidden names under /_action/.
For the illustrative /r17/navigate_to_pose action, review these directions:
| Suffix under the action name | Command-gate client | Navigation server |
|---|---|---|
/_action/send_goal | Request | Reply |
/_action/get_result | Request | Reply |
/_action/cancel_goal | Request | Reply |
/_action/feedback | Subscribe | Publish |
/_action/status | Subscribe | Publish |
Use the SROS2 action abstraction where appropriate, then inspect its generated DDS permissions for the pinned toolchain. A service expands into transport endpoints too; the table is expressed in ROS semantics, not literal DDS topic strings.
The interesting distinction is between observing motion and requesting it. A diagnostic copilot may need a filtered status mirror without needing an action client’s complete interface set. Granting broad action access just to display progress expands authority unnecessarily.
Cancellation needs its own application decision. An action permission does not express “may cancel only goals submitted by this operator.” If that restriction matters, enforce it in the gateway or action server with authenticated ownership context. A goal UUID is a correlation key, not an authorization secret.
Review parameters, lifecycle transitions and component loading with the same care. Permission to call set_parameters, change a lifecycle state, or load executable components can change the robot more profoundly than publishing an ordinary observation. Give operational tooling a dedicated, reviewed path rather than granting every node a convenient maintenance wildcard.
Generate permissions from resolved names, then inspect the result
The SROS2 policy format distinguishes topic publication/subscription, service request/reply and action call/execute. Profiles are consolidated within an enclave, with applicable explicit denials taking precedence over allowances. A profile is therefore useful policy input; it is not an independent identity inside a shared enclave.
Build the permission map from the actual launch configuration: final namespaces, remappings, node names, enabled components and conditional features. A development launch and a fleet launch can require different resolved resources even when they run identical binaries.
The official access-control tutorial shows that low-level permissions use DDS names, and demonstrates a topic remap failing under a restrictive policy. It also includes communication needed for parameters and discovery. Do not hand-convert a handful of visible topic names and assume the resulting policy covers the application.
Use this review sequence:
- Declare the approved application interfaces from the permission map.
- Resolve the launch variants that can reach production, including maintenance mode.
- Generate permissions using the matching SROS2 tooling; retain the source policy and generated output.
- Account for infrastructure traffic such as graph discovery, logging and parameter events. Disable unused application services where supported, and review each required grant.
- Compare the effective permissions with the previous release. Review every new resource pattern and every broader direction of access.
An observed graph can help discover missing interfaces, but it cannot decide which ones are legitimate. A recording from one successful mission may miss cancellation, recovery and startup behavior. It can also capture an accidental debug publisher and turn it into a permanent allowance if nobody reviews the result.
Treat wildcard additions as design changes. A rule intended to cover today’s namespace can cover tomorrow’s actuator interface too. Prefer explicit names for sensitive paths, and require a concrete explanation when a pattern is necessary. The question in review is not whether the generated XML looks tidy; it is what additional behavior the change authorizes.
Enforce startup and review domain governance separately
For a pre-provisioned deployment, the relevant startup settings follow the official Jazzy security setup:
1 | export ROS_SECURITY_KEYSTORE=/etc/robot/security |
The launcher must also select the intended enclave, for example through --ros-args --enclave /r17/copilot. These settings do not create credentials or define a least-privilege policy. Test the environment received by the actual service or container, not only an interactive shell.
Enforcement matters when security material is absent: a deployment must fail its secure startup requirement instead of silently continuing without security. Inspect inherited environment as well. An unexpected ROS_SECURITY_ENCLAVE_OVERRIDE can select a different enclave than the launch argument.
The DDS Security integration design separates the signed governance document, which describes domain protection, from the signed permissions document bound to a participant identity. Review both. A restrictive participant policy is not a substitute for checking how peers authenticate and how topic access checks and transport protection are configured.
For the protected robot domain, explicitly review unauthenticated-participant handling, join access control, topic read/write access checks, and the required discovery and data protection settings. Verify those choices on the deployed middleware. An encrypted packet capture demonstrates something about confidentiality; it does not prove that the wrong credential cannot publish an accepted command.
If secure communication fails, preserve the failing configuration and diagnose it. Changing Enforce or widening resource patterns until the graph appears healthy destroys the evidence needed to distinguish a policy defect from a transport defect.
Deploy one identity’s secrets, not the signing authority
The official deployment guidance separates the keystore’s public material, enclave artifacts and CA private material. Keep the signing authority’s private keys off deployed robots, and ship only the enclaves each target needs.
An enclave’s own private key still needs protection on the target. A read-only mount prevents modification through that mount; it does not prevent another process from reading and copying the key. Do not mount the complete fleet keystore into the AI container just because it is read-only. Restrict file access and credential mounts by process, and prevent alternative paths through host devices, privileged containers or container-management sockets.
This extends the Jetson container deployment contract: pin the identity and policy artifacts alongside the executable configuration. Provision distinct credentials where independent compromise containment or replacement is required. Copying one private key across the fleet makes those instances share the consequences of its theft.
For each release, retain a small security manifest:
1 | robot identity and deployment identifier |
Do not assume editing an XML file changes running participants. Define and test the artifact signing, rollout and participant-recreation procedure for your middleware. Replacing a key file also does not establish that every peer has stopped accepting a previously authenticated identity. The incident plan needs a tested containment path and evidence of when old authority actually stops working.
Test denials with evidence from both ends
Run these tests on an isolated graph or a secured test fixture before connecting physical actuation. Start with a positive control using the intended credential, exact topic type, compatible QoS and the same network path. Otherwise, “no message arrived” might mean a typo, failed discovery or incompatible QoS.
| Test | Evidence required before accepting the result |
|---|---|
| Copilot attempts velocity publication | Protected receiver accepts the authorized publisher but receives no copilot sample; security diagnostics identify the rejected endpoint or operation |
| Copilot tries the navigation action directly | Unauthorized goal never reaches application execution; the command-gate credential passes the corresponding controlled test |
| Participant starts without credentials | Protected peers refuse its application traffic; the secure service’s own startup check rejects missing artifacts |
| Remap an approved interface to a forbidden one | Effective policy denies the remapped resource without adding a wildcard to make the test pass |
| Exercise goal, feedback, result and cancellation | Every required action path works under the intended credential, including terminal-state handling |
| Run diagnostics and recording | Required observation succeeds without parameter-write, motion or playback authority |
| Remove a credential or test an expired grant | Expected startup/connection failure is visible; the robot follows its defined loss-of-command response |
| Copy a privileged key into the AI process in the fixture | Demonstrate the credential-theft consequence, then prove the real deployment prevents that file access |
The last test is deliberately different. Middleware sees a credential, not the moral identity of the executable holding it. If the copilot obtains the gateway’s private key and matching artifacts, an enclave policy alone cannot tell that the holder is the wrong program. That is why the OS principal and mount columns belong in the permission map.
Denied behavior may appear as endpoint-creation failure, failed matching, or another middleware-specific diagnostic. Define the expected signal for the pinned implementation. Preserve participant identity, policy digest, resolved resource name and receiver-side evidence; a CLI timeout alone is insufficient.
After security is enabled, also measure discovery, reconnection and workload timing with production-sized messages. Qualify the cost under the robot’s existing timing limits rather than assuming encryption is either free or prohibitively expensive.
Keep physical validity downstream of transport permission
An authorized navigation process can still publish an unreasonable velocity. An authorized sensor can still report a false observation. A compromised participant can abuse an interface it legitimately owns, and DDS access rules do not supply payload bounds or an application rate budget.
Keep semantic validation, resource limits and the independent robot safety path in place. Security withdrawal must lead to a defined controller response; an emergency stop must not depend on the AI process authenticating successfully.
For the next architecture review, choose one forbidden path: the copilot publishing a velocity request. Trace why it fails at the process boundary, credential store, middleware policy and receiver. Then repeat the test after changing composition, namespaces or deployment mounts. That evidence is far more useful than a diagram with a padlock beside the ROS graph.