Skip to content

feat(telemetry): stamp operator resource attributes on every record - #34

Open
mufaddalq wants to merge 8 commits into
strands-agents:mainfrom
mufaddalq:feat/operator-resource-attributes
Open

mufaddalq wants to merge 8 commits into
strands-agents:mainfrom
mufaddalq:feat/operator-resource-attributes

Conversation

@mufaddalq

@mufaddalq mufaddalq commented Oct 8, 2026 •

Copy link
Copy Markdown

Description

A box record names its box only by a generated identifier, so an operator cannot attach a tenant, a user, or a conversation to the audit trail. Attribution then depends on headers that the untrusted workload chooses to send.

This change lets the operator set OTEL_RESOURCE_ATTRIBUTES in the environment of strands-box run. Every box record and every relayed agent payload then carries those attributes on its resource.

The box already read this variable in one place, by accident. Resource::builder() runs the SDK environment detector. Thus the box's own records got the attributes with no validation (a malformed entry was dropped, and a strands.box. key that the box does not stamp got through), and relayed payloads did not get them. This change replaces that read with a validated read that applies to the two lanes.

Load-bearing decisions:

  1. The source is the environment variable. It is the OpenTelemetry standard, it is per run, and it moves no frozen surface. A box run flag was rejected because it moves the CLI. A [telemetry] field was rejected because it is fixed per box and takes a name from the target labels.
  2. A bad value refuses the run before the workload starts. It does not skip the bad entry. The refusals are: more than 1024 bytes, not UTF-8, no =, an empty key, a repeated key, a bad percent escape, service.name, and each strands.box. key. The 1024-byte bound exists because the receiver stamps the list on each relayed resource entry. Without a bound, the operator value multiplies the stamp amplification that crates/telemetry/AGENTS.md records.
  3. On a relayed payload, only resource attributes under an operator key are replaced. Thus the workload cannot forge conversation.id. An attribute on a span, a log record, or a metric stays the agent's own data.

Interface. This adds one public method, TelemetryConfig::with_resource_attributes. The pub use set does not change. A run whose OTEL_RESOURCE_ATTRIBUTES has a bad entry now stops; before, the SDK dropped or passed the entry. There is no new box.toml key, no CLI flag, and no change to RECORD_VERSION, LIVE_VERSION, or PROTOCOL_VERSION. An operator whose shell already exports OTEL_RESOURCE_ATTRIBUTES=service.name=... for other tools gets a refusal and must change the variable for box runs.

Deferred: a way to "pull the audit stream per sandbox" (from the issue). A backend can filter on the new attribute, and each box already writes its own file.

Related Issues

Closes #28

Type of Change

New feature

Testing

The first commit message names the test for each claim. Each new test failed before its change. The forge guard and the logs and metrics routes were proved by mutation.

  • cargo test --no-fail-fast -p strands-box-telemetry -p strands-box --all-features on macOS arm64: 1009 passed, 6 ignored (the same as main), 0 skipped, 3 failed. The 3 failures are catalog timing tests in runtime_mcp_policy_staging. They also fail on main under load: interleaved runs of that suite gave 23/2 and 24/1 on this branch, and 25/0 and 23/2 on main.
  • Linux aarch64 (a privileged container, rust 1.98.1): strands-box-telemetry 94 passed; box_telemetry 28 passed, 1 ignored, 0 skipped.
  • cargo fmt --check is clean. cargo clippy --all-targets --all-features gives 0 warnings in crates/telemetry and crates/box. I did not run just clippy with -D warnings over the whole workspace.
  • Not run: Linux x86_64 (CI covers it), the containment suites in test-integ, and a measurement of the stamp memory peak at the 1024-byte bound.

Checklist

  • I have read the CONTRIBUTING document
  • I have reviewed and understand every line of code in this PR, including any generated by AI tools, and I can explain why it works
  • My change is focused and reasonably small; I have split unrelated work into separate PRs
  • I have added any necessary tests that prove my fix is effective or my feature works
  • I have updated the documentation accordingly
  • My changes generate no new warnings

By submitting this pull request, I confirm that you can use, modify, copy, and redistribute this contribution, under the terms of your choice.

🤖 Generated with Claude Code

muffadalq added 8 commits October 8, 2026 16:03
## Problem

A box record names its box only by a generated identifier. An operator cannot
attach a tenant, a user, or a conversation to a record. Attribution then
depends on headers that the workload chooses to send, or on a side table that
maps generated names to conversations (strands-agents#28).

Two more defects are in the same path. The box's own records read
OTEL_RESOURCE_ATTRIBUTES already, through the default detector of the
OpenTelemetry SDK. That read has no validation: the SDK drops a malformed
entry, and a strands.box. key that the box does not stamp gets through. A
relayed agent payload does not get the attributes, so the two lanes of one run
have different identities.

## Solution

The run process reads OTEL_RESOURCE_ATTRIBUTES from its own environment once,
and every box record and every relayed payload carries the attributes on its
resource.

- TelemetryConfig gets one method, with_resource_attributes. The pub use set
  of the telemetry crate does not change.
- validate parses the text in the OpenTelemetry syntax and decodes each
  percent escape. It refuses a value longer than 1024 bytes, an entry with no
  `=`, an empty key, a repeated key, a bad percent escape, service.name, and
  each strands.box. key. run refuses a value that is not UTF-8.
- The collector builds its resource without the environment detector. It
  keeps the telemetry.sdk.* keys and stamps the parsed list.
- On each of the three receiver routes, the receiver removes an agent
  resource attribute under an operator key, and then stamps the list. An
  attribute on a span, a log record, or a metric stays the agent's.

Interface: one new public method. A run whose OTEL_RESOURCE_ATTRIBUTES has a
bad entry now stops before the workload starts; before, the SDK dropped the
entry or passed it. There is no box.toml key, no CLI flag, and no change to
RECORD_VERSION, LIVE_VERSION, or PROTOCOL_VERSION. A box run flag was
rejected because it moves the frozen CLI. A [telemetry] field was rejected
because it is fixed per box and takes a name from the target labels.

## Tests

Each claim has a test beside it:

- Parse, decode, and refuse: config::tests
  operator_resource_attributes_parse_and_percent_decode,
  an_unset_or_empty_value_adds_no_attribute, a_blank_segment_is_ignored,
  a_malformed_operator_attribute_refuses,
  a_repeated_operator_attribute_refuses,
  a_reserved_operator_attribute_refuses, an_invalid_percent_escape_refuses,
  an_operator_attribute_value_over_the_bound_refuses.
- The box's own records carry the list, the box keys still win, and the
  telemetry.sdk.* keys stay: collector::tests
  operator_attributes_stamp_the_boxs_own_records,
  a_refused_operator_attribute_opens_no_collector.
- Relayed payloads on all three routes: receive::tests
  operator_attributes_stamp_a_relayed_agent_payload (no resource in the
  payload), an_agent_cannot_forge_an_operator_attribute,
  an_agent_attribute_the_operator_did_not_set_survives.
- The real binary: box_telemetry
  operator_resource_attributes_reach_every_record,
  a_reserved_operator_attribute_refuses_before_the_workload_starts,
  the_workload_does_not_see_operator_resource_attributes; and
  run::telemetry::tests::a_resource_attribute_value_that_is_not_utf8_refuses.

Each new test failed before its change. The forge guard was proved: with the
removal of operator keys deleted, an_agent_cannot_forge_an_operator_attribute
fails; with only the logs route or only the metrics route changed to pass no
list, both route tests fail on that route.

Gates on macOS arm64: cargo fmt --check is clean, and cargo clippy
--all-targets --all-features shows 0 warnings in crates/telemetry and
crates/box. cargo test --no-fail-fast -p strands-box-telemetry -p
strands-box --all-features: 1009 passed, 6 ignored (the same 6 as main), 0
skipped, 3 failed. The 3 failures are in runtime_mcp_policy_staging
(catalog timeouts and the 256-page cost ratio). They also fail on main under
load: interleaved runs of that suite gave 23/2 and 24/1 on this branch, and
25/0 and 23/2 on main.

Linux aarch64 (a privileged container on macOS, rust 1.98.1):
strands-box-telemetry 94 passed; box_telemetry 28 passed, 1 ignored, 0
skipped.

Not run: Linux x86_64, the containment suites (test-integ), and the stamp
memory peak at the 1024-byte bound, which is not measured.
@mufaddalq

Copy link
Copy Markdown
Author

Hi maintainers, could someone approve the pending workflow runs on this PR? As an outside contributor, CI and the PR-title check are at action_required, and Auto Strands Review is waiting on the manual-approval environment (run). Fork CI is green on my side, but I'd like the upstream checks and the review to run too. Thanks!

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[FEATURE] Operator-supplied sandbox and conversation identity on telemetry records

1 participant