본문으로 건너뛰기
이 페이지는 아직 사용자의 언어로 번역되지 않았거나 번역이 기술 검토를 기다리고 있습니다. 아래에는 영어 원문이 표시됩니다. 영어 페이지 열기

Device Agent Privacy and Consent

Status: PRIVATE PREVIEW

This page is for enterprise administrators and security teams who deploy the Control Zero device agent (the cz-agent binary, service cz-agentd) and need to know exactly what data leaves a device and what consent or privacy controls exist. It describes the device agent's current behavior and states where a control does not exist.

This page is about the device agent only. The separate shadow-AI discovery component has its own documentation.


The collection boundary​

The device agent emits seven groups of structured data. Each one is a closed, typed wire structure, and each one is enumerated here so the boundary is explicit rather than implied.

1. The local hook protocol​

When a governed host tool fires a hook, a shim forwards the call to the daemon over a local socket. Three types cross that boundary:

  • HookRequest — what the host chose to run: canonical_source (the host name), payload (the tool input as JSON, forwarded unchanged so the policy can be evaluated), cwd, git_remote (optional), shim_version, shim_signature, and shim_asserted_ancestry.
  • Provenance — what the daemon observed about the connecting peer: peer_pid, peer_uid, peer_sid, peer_exe_path, a content hash of the peer executable, ancestry_observed, shim_signature_valid, and a verdict string (host | not_host | unknown). A field the kernel cannot supply is None, and the peer identity fields are observed by the daemon, never taken from the request.
  • HookDecision — the finalised result returned to the shim: verdict (allow | deny | warn | hitl), origin (the deciding layer), the observed provenance, and an optional dlp outcome.

The findings in a decision's DLP outcome never carry matched bytes. masked_input redacts only spans matched by a mask rule, so text matched only by a detect rule stays in it unchanged:

  • DlpOutcome — action (detect | mask | block), findings, uncompilable_rule_ids, and masked_input. Each finding names the rule, its category and its action, not the matched text. masked_input is present only when a mask fired and the call proceeds: it is the tool input with the spans matched by mask rules redacted. Everything else, including text matched only by a detect rule, appears unchanged.
  • DlpFinding — a rule_id, the org's category, and the action; nothing else. masked_input is present only when a mask fired, changed the input, and the verdict lets the call proceed — the caller runs the tool on the redacted input instead of the original.

2. Discovery rows​

The discovery detectors return AiDiscovery rows. The row shape is:

  • id, discovery_type (ai_traffic | ai_process | ai_api_key | ai_tool), summary, details (a metadata-only string map), severity, first_seen, last_seen, and hostname.

The details map is where the privacy boundary is enforced. The module defines a hard list of forbidden detail keys — argv, args, argument, command, cmdline, cmd, env, environ, api_key, token, secret, password, value, credential — and the forbidden-key check flags any detail key that contains one of these strings, case-insensitively, naming the offending key. The process and traffic detectors' tests fail any row they emit that trips it. The key and host detectors have no such test.

What each detector actually reports, rather than what it saw:

  • Key scanning reports only file_path, line_number, key_type, and key_prefix (the first 8 characters of the matched text plus ...). The key value is never returned or transmitted.
  • Process scanning reports only process_name, pid, the owner's user id, the matched tool label (tool) and, when more than one instance matched, instance_count. Command-line arguments — which may contain API keys or secrets passed as flags — are deliberately not transmitted.
  • Traffic scanning reports only the endpoint (a known AI API hostname, or localhost:<port> for a known local inference port), a label, a port field that is always the literal 443, and a connection count. These are read from the active TCP connection table on Linux and macOS. On Windows it reports nothing. No payload data is captured.
  • Host detection reports the host name, the config path that proved presence, and an optional semantic-version string read from the host binary's --version output. Config file contents are never read into a row.

3. Inventory wire​

Inventory upload is not yet active. Nothing posts a batch to POST /api/device/inventory in production today. The cz-agent scan command reports discovery counts locally on the machine; it does not upload them.

The wire type that would carry an upload is defined as follows. Discovery rows are mapped into an InventoryBatch (scan_cycle_id, observed_at, rows) of InventoryRow objects, each with a kind from the closed vocabulary process | listener | egress | app | mcp_config | drift | weights | autostart | container | credential_reach | skill | plugin, an attribution (process | unknown), a metadata-only details map, and a severity. Egress rows carry dst_host and dst_port; credential-reach rows carry the same key metadata described above (path, line, type, prefix), never the credential itself.

4. Heartbeat​

The heartbeat is a fallback liveness signal carrying operational metadata, not discovery content. Its fields are agent_version, install_mode, posture, inventory_digest, spool_depth, spool_drops_total, per-writer last_ok_at timestamps, and daemon_absent. Every field except agent_version carries a hardcoded value in production today: install_mode is always user, posture is always an empty object, inventory_digest is always empty, spool_depth and spool_drops_total are always 0, the last_ok_at fields (stream, inventory, upload) are always absent, and daemon_absent is always false. The heartbeat therefore reports "never run" regardless of what the agent has done, and its freshness timestamps cannot distinguish "not run" from "ran and found nothing".

5. Policy fetch​

The daemon fetches its policy bundle with a machine-signed GET /api/policy. The machine-signed requests this agent sends (the map stream GET /api/device/map, the policy fetch and the heartbeat) all carry X-CZ-Machine-ID, X-CZ-Timestamp, and X-CZ-Signature headers, which is how the server identifies the enrolled machine.

6. Enrolment​

cz-agent up sends POST /api/device/login/start with hostname, os, install_mode (always user today) and shim_pubkey in the body, and the machine public key in the X-CZ-Pending-Key header. It then polls POST /api/device/login/poll with the device_code it was issued.

7. Map stream​

The daemon long-polls a machine-signed GET /api/device/map (every 30 seconds by default), sending its current policy_version and its enrolment session id as query parameters.


What is structurally absent​

The privacy property of this agent is not a runtime setting: the detectors are written to emit metadata only, and the process and traffic detectors' tests fail any row they emit that carries a forbidden key.

  • No file contents. No detector writes file contents into a row. The key detector reads config files to match patterns but records only path, line number, key type and an 8-character prefix. The host detector records the config path, not its contents. The forbidden-key check inspects detail key names, not values, and runs in the process and traffic detectors' tests.
  • No key values. Key scanning reports an 8-character prefix of the matched text, never the full secret. For a bare key match such as sk-..., that prefix includes the key's first characters.
  • No command lines. Process scanning reports process name, PID, the owner's user id, the matched tool label and an instance count only.
  • No traffic payload. Traffic scanning reports endpoint, label, port and count only.
  • No matched DLP bytes. DlpOutcome carries rule ids, categories, and actions, not the content the rules matched.
  • No executable contents. Provenance records a content hash of the peer executable, not the executable's contents.

The enforcement is in the detectors and the tests: the process and traffic detectors each have a privacy-boundary test that fails if any emitted row trips the forbidden-key check. The cz-agent scan command additionally counts and prints the number of would-be violations it saw, so a row that would have broken the boundary is observable at scan time rather than silently passed.


There is no consent state machine in the device agent: there is no opt-in or opt-out flag, no per-user consent record, and no telemetry toggle, so the only privacy-related states that actually exist are the structural boundary described above and the org-signature check below.

The daemon contains code for a trust manifest ({authorities, threshold, mode}), where required would demand customer signatures before the org signature and off ignores them. The production control plane does not send a manifest, and the daemon's policy fetch does not read customer signatures, so no customer-signature mode is in effect on any device today.

The org signature is always required: a bundle whose signature does not verify against the retained org public key is refused and the cache is left unchanged (fail closed).

The policy fetch also carries a freshness token (ETag, held as PolicyFetch.freshness) that folds in the server-side DLP gate, because a gate change rewrites the bundle without bumping policy_version. That token is a server-driven control, not an on-device consent choice.


What an administrator controls vs what an end user cannot change​

An administrator controls, through the control plane rather than on-device toggles:

  • The policy bundle itself — the content rules and DLP rules the daemon evaluates — which is org-signed and verified before use.
  • Customer co-signing of policy bundles (a trust manifest) is not administrator-configurable today: the control plane has no setting for it and does not send one to devices.
  • The DLP gate folded into the policy freshness token, which can disable or re-enable rule evaluation without a version bump.
  • The scan scope, indirectly, for key scanning only: cz-agent scan searches the current working directory plus $HOME (to a depth of 5, config and dotenv files only) for keys and checks host config paths under $HOME. Process and traffic scanning cover the whole machine's process list and TCP connection table regardless of directory.

An end user cannot change:

  • There is no per-user opt-out for discovery scanning, no consent dialog, and no preference file consulted before collection. The key scan runs over the current directory and $HOME, and the process and traffic scans over the whole machine, regardless of which user invokes it.
  • The device identity is held by the account that runs the daemon. Packaged installs (Linux systemd unit, macOS LaunchDaemon) run the daemon as root, and the identity file holding the signing key is set to owner-only (0600) permissions immediately after it is written, in root's state directory, so a non-administrator cannot read or replace it. If a user runs cz-agent up under their own account, the identity file is theirs. Removing it and enrolling again registers a new device rather than impersonating the old one.
  • The privacy boundary is not user-configurable. The detectors write only the metadata keys listed above, and no setting adds argv, environment values, or full key values. This is enforced by what the detectors emit and by the process and traffic detectors' tests that fail on a forbidden key, not by the type system: details is a string map, and cz-agent scan counts a row that trips the check rather than dropping it. Key rows carry the first 8 characters of the matched text, as described under Key scanning.