One record that ties the steps together
Who would use it: a team running AI agents that needs to answer, after the fact, "what did the agent do, was it allowed, who signed off, and what did it cost?" for an audit, an incident review or a cost report.
It sits between the actions and evidence steps of the AgenTrust chain. Each fact is written as a small event in a format any storage system can accept. Events are tagged so they line up with the application's existing OpenTelemetry traces (OpenTelemetry is the common open standard for recording what software does). When a run's events are complete and saved, they can be turned into a signed TRACE receipt.
Agent applications can already record which model and tools they called. The oversight facts are the ones that stay scattered: in policy-engine logs, approval databases, cost modules and vendor dashboards. This gives them one shared format that leaves sensitive content out, and links them to the trace the application already produces.
Six event families
There are six kinds of event. Each one reports something that happened elsewhere, such as a decision made by a policy engine. This project only records decisions. Your policy engine makes them.
| Family | What it records |
|---|---|
| Policy decision | Whether a rule allowed, blocked or questioned an action (or errored), which rule it was, whether it was enforced, and when |
| Approval lifecycle | A request for human sign-off, through the final decision and what happened next, tied to a fingerprint of the exact action approved |
| Action execution | Each finished attempt to use a tool, call another agent, touch a file, make a web request or query a database |
| Data flow | Which kind of data moved from where to where, by label only, never the data itself |
| Usage | Tokens used and cost, per call and per run, with a note of where each cost figure came from |
| Evidence lifecycle | Progress points in a run, whether its record is complete, and whether it became a TRACE receipt |
Technical detail: schema format
Each family is a JSON Schema 2020-12 contract with valid and invalid fixtures. Approvals bind to an action digest. Actions cover resolved tool, MCP, A2A, file, HTTP and database attempts.
Privacy boundary: by default (the metadata-only profile), events may not contain prompts, model outputs, source code, tool inputs or results, passwords or access tokens. The IDs passed between agents only help line events up; they do not prove who an agent is or what it may do.
Keep the stack you already operate
The library attaches each event to whatever OpenTelemetry trace your application is already recording. It does not set up or replace any part of your monitoring pipeline, and where a standard OpenTelemetry field exists, that field wins over an AgenTrust one.
Technical detail: what the SDK does not install
The reference SDK uses the caller's current OpenTelemetry span. It does not install a provider, processor, exporter, collector, or global propagator.
Before you start: Python 3.11+, Git, and Bash on macOS, Linux, or Windows with WSL. Commands run from the repository root. No telemetry backend is needed for the first example.
Install the Python SDK
Install from the source checkout so the SDK and runnable examples use the same revision.
git clone https://github.com/agentrust-io/agentrust-telemetry cd agentrust-telemetry python3 -m venv .venv source .venv/bin/activate python -m pip install -e ".[otel]"
Run a complete event example
Run the included synthetic policy event. It defines every required field, validates the event, and prints it through a local log emitter.
python examples/manual_governance.py
Keep events linked when one agent hands work to another
When one agent passes work to another program, pass the run's IDs along with it so both sides' events line up in one story. Integration sketch: this block assumes your application already defines tracer and event. For a complete runnable version, use python examples/governed_workflow.py after installing ".[test]" as shown below.
Technical detail: direct calls versus queued work
For synchronous cross-process agent calls, propagate the caller's W3C context and the durable AgenTrust identifiers, then use the extracted context as the receiving span's remote parent. For asynchronous queue handoffs, start a new trace with links=[remote.link()] instead: that preserves causality without representing queued work as a synchronous child span.
from agentrust_telemetry import extract_context, inject_context carrier = {} inject_context(carrier, run_id="run-123", agent_id="planner") remote = extract_context(carrier) with tracer.start_as_current_span("worker", context=remote.otel_context): event.update(remote.event_fields(agent_id="worker"))
Treat IDs received from another agent as unverified. They do not prove who that agent is or what it is allowed to do.
Or use the TypeScript SDK
Use a separate terminal at the repository root, with Node.js and npm installed. Return to the root for the Python examples below.
The pre-alpha Node package validates the same fixtures and the same privacy profile as Python, and preserves nanosecond wire timestamps as decimal strings. It installs no OTel provider, exporter, or global propagator either.
cd packages/typescript npm ci npm run check
Expected first result: a printed policy event and an EmitResult with accepted=True and log_emitted=True. A missing span is expected without an active OpenTelemetry span. This is a made-up sample event; it does not enforce any policy.
Two runnable examples
Both ship in the repository. The second runs the whole path end to end: policy events, OpenTelemetry tracing, saving the complete record, and producing a TRACE receipt.
python examples/manual_governance.py
python -m pip install -e ".[test]"
python examples/governed_workflow.py
Map what your policy engine already emits
If you already use a policy engine, an adapter turns its decision log into these events for you. Adapters exist for Cedar and OPA (two widely used open-source policy engines) and for AGT (the Agent Governance Toolkit), in both Python and TypeScript.
Technical detail: what the event factory fills in
You should not have to hand-build envelopes. The event factory owns specification version, producer identity, event UUID, timestamp, correlation fields, and both schema and privacy validation, so adapters only translate a source record. TypeScript adapters take plain source records with no runtime dependency on the upstream package.
Cedar
Accepts the final Allow or Deny response plus determining policy IDs and caller-normalized diagnostic codes. Evaluation errors are recorded without rewriting the final decision, matching Cedar's skip-on-error semantics. Free-form error messages are rejected so content cannot leak through a diagnostic.
OPA
Accepts one decision-log event. Booleans map to allow or deny, an absent result maps to not-applicable, and structured results require an explicit mapper because OPA permits any JSON value. Input and result documents are never copied, and a trusted bundle digest stays required because a revision is not a content digest.
AGT bridge
Implements the batch sink shape used by Agent OS without making AGT a core dependency. The whole batch is normalized before the first event is emitted, so a malformed later event cannot cause partial delivery. Durable evidence callbacks must stay idempotent by event ID.
AGT's action-bound approval protocol maps as a linked sequence: the policy decision becomes a challenge, the approval request references that normalized policy event, and only a resolution becomes approved or rejected. Read the adapter reference.
Don't take the contract on trust
The repository ships sample events that must pass and sample events that must fail, plus a separate checker. The failing samples show what gets rejected: an event containing raw content, a timestamp in the wrong form, an all-zero trace ID, an action with no fingerprint, usage with no measurement, a total cost with no breakdown. Each one must be rejected.
python -m pip install -r conformance/requirements.txt python conformance/runner/validate.py python -m unittest discover -s tests -v
What the checker does not cover. It confirms events have the right shape and contain no prohibited content. It does not yet check how events are exported to OpenTelemetry (OTLP projection) or turned into TRACE receipts, so a passing run tells you about the event format only, not the whole pipeline.
The rules the schemas enforce
run_idnames one agent run and is kept for the record;trace_idis an optional link to the W3C trace used for day-to-day monitoring. One cannot stand in for the other.- Standard OpenTelemetry fields take precedence over AgenTrust extensions.
- By default, prompts, outputs, source code, tool inputs and results, passwords and access tokens are not allowed in any event.
- Monitoring data can be lost along the way. A record must never claim to be more complete than it is.
- An event reports a fact observed elsewhere. This project does not decide what is allowed.
What is not true yet
Stated here rather than in a footnote, because a governance evidence layer that oversells itself is worse than none.
The format is alpha. Version 0.1.0-alpha.6 has no stable code interface and no compatibility promise. It may change in ways that break existing users.
It checks the shape of what is reported, not whether it is true. A producer can report a wrong policy decision, identity, data label, token count or cost. A signed TRACE receipt can show the record was not changed, and a verified attestation can show where the software ran. Neither proves every reported fact is true.
Everyday monitoring is not audit evidence. OpenTelemetry may sample or drop data. The complete record has to be saved somewhere durable before any TRACE receipt is made, and the TRACE step here runs in software only: it needs a record explicitly marked complete plus settings from a trusted caller.
The in-memory evidence mode does not survive a restart. The callback mode defines how saves are confirmed and retried, but you own the storage, duplicate handling and recovery.
It has only been tested together locally with OpenTelemetry for Python, not across collectors, storage backends or other languages. Action events also record finished attempts only, not each step while an action is still running.
What it deliberately does not become
It does not collect, store or display data, and it is not a hosted service. It does not decide what is allowed or run approvals. It is not a framework for building agents, not a price list for models, and it does not repeat the automatic recording that OpenTelemetry GenAI and OpenInference already do.