Scout
DocsBlogToolsContact

Proposal · Collaboration Agreements

Concept, not a shipped feature: a technical perspective on deterministic agreements around human and agent collaboration.

View MD

Status: engineering commentary on a product concept. Nothing proposed here exists in a Scout release; everything described as shipping is cited to a file.

Companion documents:

A read for a product-and-engineering discussion: is the idea buildable, what does it actually claim, where does it break. Not a design spec.


1. The angle

The agreement is not a security wrapper around chat. It is the common trust and coordination substrate for every Scout-mediated relationship: rooms, direct exchanges, agent-to-agent asks, work claims, review requests, file transfer, branch convergence, tool use, approvals, and external effects. Conversation is one event stream inside that relationship, not its boundary.

Most agent-security work assumes the dangerous thing an agent does is call a tool. Scout's position falls out of what Scout already is: a broker that mediates typed coordination between humans and agents. In Scout, a message can itself be an effect — a reply wakes another agent, claims work, causes a human to act, or fans a delegated ask out to nine partners. A later tool call, commit, delivery, or external acknowledgement may be another effect in the same lineage. Both need to be governed by the same authority.

Agents can say anything. The agreement decides what their words are allowed to make happen.

So the authorization boundary belongs at the broker, not in the harness and not in the model's prompt. Scout already sits at the one point every Scout-owned coordination effect passes through; agreements make that point decide, not only record. Where Scout also controls the execution environment, the same decision can bind filesystem, process, secret, and network effects. Where it does not, the claim stops at broker mediation.

The guarantee has to be narrow, and stating it narrowly is what makes it worth building. Scout cannot promise an agent will not be manipulated — models follow hostile instructions, and that is not solvable at this layer. Scout could promise that a manipulated agent cannot use a Scout-mediated interface to produce an effect outside its effective envelope. That is a containment claim about interfaces, not about cognition.

2. The shape: a reference monitor at a chokepoint that already exists

A classic reference monitor — complete mediation, tamper-resistant, verifiable — with two unusual properties.

It is already partly built. Scout's broker is the canonical writer for Scout-owned records; clients and adapters submit commands and never write coordination records directly (README.agent.md↗). That rule is the completeness property a reference monitor needs, and it ships today for single-writer consistency. Agreements reuse it for authorization.

It brackets a nondeterministic component. The model is untrusted input on both sides, with semantic inspection above it (ingress) and below it (egress). The decision stays deterministic: the same canonical inputs must return the same verdict in TypeScript and Swift, or the audit record means nothing.

LayerNatureMay doMay never do
ClassifiersProbabilisticEmit labelled, versioned evidenceMint authority
Policy evaluatorDeterministic, pureTurn evidence + agreement into a verdictDepend on model output
Enforcement pointsImperativeRefuse an effect without a valid verdictAccept prose as a verdict

Evidence narrows. Only the agreement grants.

Two surfaces, one substrate

Scout exposes this substrate through two different product surfaces.

Rooms are the human-legible surface. Multiple people bring agents into a shared context. The room presents membership, roles, work, approvals, handoffs, artifacts, and history in language a non-specialist can understand. A room may live for hours or months and may contain many agreement-bound interactions.

Mesh is the machine-native surface. Independently controlled agents discover one another and establish a bounded working relationship during active work. An agent may ask another owner's agent to review a pull request, investigate a failure, reconcile competing branches, or change a related repository. The relationship may last seconds and need no human-visible room, but it still needs identity, purpose, authority, disclosure limits, delegation constraints, and a verifiable result.

These are not separate security models. The same agreement evaluator, grants, envelopes, decisions, and evidence records apply to both. The difference is presentation and lifetime:

DimensionRoomMesh interaction
Primary usersPeople collaborating with agentsAgents collaborating across owners or systems
Typical lifetimeSession, project, or standing teamOne request, flight, review, or temporary work graph
Human surfaceMembership, roles, work, approvalsOptional notification, approval, and result review
Machine surfaceRoom-scoped operationsTyped negotiation, delegation, transfer, and return
Shared primitiveAgreement-bound collaboration contextAgreement-bound collaboration context

The fundamental object is therefore not the thread. It is an agreement-bound collaboration context: a durable room or ephemeral mesh relationship within which messages, work items, flights, artifacts, approvals, and effects acquire meaning and authority.

3. What already ships, and what it is not

Scout's posture today is high-trust local developer pilots, but several agreement primitives exist anyway, because the mesh trust cone needed them.

A working reference monitor, scoped to remote HTTP. Every broker route is classified into exactly one tier — local, observe, control, public, guest — and unknown routes fall back to local, deny-by-default for remote callers (mesh-route-matrix.ts:20↗). A remote request needs a verified peer signature, an enrolled non-revoked grant, a fresh nonce, and a grant tier satisfying the route tier, and a route-inventory test fails the build when a route has no declared tier (mesh-ingress-gate.ts:26↗). guest routes are stricter still: they need a verified guest signature on every transport, loopback included, and a known guest key is denied on every other route; verify-warn softens neither. Deny-by-default, complete mediation, CI-enforced coverage — the best evidence the model is implementable here.

Identity with the right split already made. Each node owns a long-term Ed25519 identity whose private key never leaves the support directory; its key ID is canonical for lookup and signing, and its fingerprint is display-only, never used for authorization (node-identity.ts:28↗).

A scoped, expiring, revocable grant at the collaboration layer. ChannelInviteRecord carries expiresAt, maxRedemptions, and revokedAt, and redeems deterministically into revoked / expired / exhausted / redeemable. Its scope enum has exactly one member — channel_participation, "never project, shell, or filesystem authority" (channel-invites.ts:28↗). That is a participant grant in everything but name, and the right shape to widen carefully.

A key-bound guest grant. GuestGrantRecord binds a client public key to the broker actor it speaks as, a list of allowed targets, an expiry, and revocation, and resolves to active / expired / revoked (guest-access.ts:40↗). That is per-principal trust for one outside client — but granted by the operator to one key, not negotiated between parties or scoped to a collaboration.

But four gaps matter more than those primitives:

  1. The mesh gate defaults to verify-warn for peer tiers — it verifies, logs failures, and allows the request; denial needs OPENSCOUT_MESH_GATE=enforce (mesh-ingress-gate.ts:41↗). A monitor in warn mode is an observability feature.
  2. Capability evaluation is readiness, not authorization — it answers can this run, never may this principal do this, and its only caller wires it to a read query (capability-matrix.ts:271↗).
  3. No agreement, agreement-scoped grant, envelope, or decision record exists. Machine trust is enrolled, and guest keys and channel invites grant narrow participant access, but nothing binds participants, grants, and decisions to a shared, versioned agreement.
  4. The journal is append-only by convention, not hash-linked or signed — a local operator can edit the JSONL file.

4. Agreements, contexts, work, and envelopes

Four lifetimes must remain distinct:

  1. The agreement is the versioned constitution: who may participate, which authority may be granted, how it narrows, and what evidence is required.
  2. The collaboration context is one instantiated relationship governed by an agreement: a room, direct exchange, project, review, or ephemeral mesh link.
  3. A work object is an accountable unit inside that context: an ask, review, work item, handoff, convergence task, or transaction with an owner, inputs, completion condition, and outputs.
  4. The request envelope is the short-lived effective authority for one attempted operation within the work.

This separation prevents two opposite errors. A standing room membership must not become ambient authority for every future task, and a thirty-second mesh request should not require inventing a permanent room. Work can move between people and agents without silently moving all of the sender's authority with it.

The evaluator should be a pure function:

decide(agreement, grant, request, evidence, clock, restrictions)
  -> { verdict: allow | deny | require_approval,
       matchedRules[], obligations[], agreementDigest, evaluatorVersion }

No I/O, no model call, no ambient state. Everything the verdict depends on is an argument, which is what lets a decision be reproduced from the audit record years later. Unlisted operations deny by default, as the route matrix already does, and three further properties are worth naming:

  • Monotone narrowing. No input may widen a verdict — evidence, delegation hops, and elapsed time can only restrict. This is what makes prompt injection structurally uninteresting: hostile text adds evidence, evidence subtracts.
  • Canonical serialization before hashing. A cross-language digest needs one frozen normalization (field order, number form, Unicode, absent vs null) plus golden vectors shared by both evaluators. Implementations that disagree produce two digests for one agreement, and every audit claim built on digest equality silently fails. This is the hard part.
  • Immutable versions. Amending v17 creates v18 and never rewrites decisions made under v17. Sensitive operations evaluate at admission and at effect time, so revoking a participant stops future effects mid-turn.

An envelope is the short-lived effective authority for one operation — an intersection that can never exceed any parent, and that delegation intersects again against the target's ceiling:

organizational policy ∩ agreement (version-pinned) ∩ context policy
  ∩ participant grant ∩ work-object scope ∩ interface policy
  ∩ requested operations / resources / destination
  ∩ current restrictions (time, revocation, emergency deny)
  = request envelope

For mesh work, the initial exchange is a negotiation rather than an ambient grant. The requester proposes a purpose, requested result, inputs it intends to disclose, requested capabilities, deadline, and return channel. The receiver may accept a narrower version, require human approval, or reject it. Acceptance creates the context and pins its agreement version; natural-language assent in a message does not. Any subdelegation repeats the intersection and records its lineage. Authority never returns merely because a result does.

Three assertions stay separate, because conflating them is how these systems fail. Authentication (key X signed this) is cryptographic and already implemented at node scope. Recognition (agreement A accepts that principal as participant P, in scope S, until T) is a policy statement no signature produces. Authorization (P's envelope permits operation O toward destination D) is derived per request. An agreement signed by a valid key never authorized to issue for this room is a forgery that verifies.

One naming problem blocks all of it: "capability" in Scout already means four things, and authorization needs its own namespace. Since readiness already returns allow/deny/require_approval, mistaking availability for authorization is a live hazard.

5. Above and below the model

Ingress establishes the authenticated sender, the applicable agreement and grant, purpose, input classifications, injection indicators, and which history and retrieved resources may enter the context window at all. The mechanism that matters is exclusion, not instruction: a prompt saying "do not read .env" is a suggestion to a nondeterministic system, while not placing .env in the context is a property of the system. Untrusted content that does enter stays visibly marked as untrusted.

Egress classifies proposed content and effects, and deterministic policy converts that evidence into allow, deny, redaction, quarantine, or approval. Every classifier result is a record — classifier ID, ruleset version, labels, confidence, time — and the agreement, not the classifier, defines the treatment. That is what keeps a probabilistic component inside a deterministic system:

possible_secret ∧ destination=external      -> deny
confidence < 0.80                           -> require_approval
classifiers disagree                        -> most restrictive label wins
egress classifier unavailable ∧ external    -> deny

Fail-closed behavior is declared per direction, because uniform fail-closed makes Scout unusable whenever a classifier is down: deny external egress, allow internal read-only with audit. Three hard parts deserve naming — streaming output must be gated before the stream is classified or buffered at a latency cost; encoded, encrypted, image, and archive payloads defeat text classifiers; and slow leakage across many innocuous messages is a problem class per-message classification does not touch. And a classifier may never issue a grant, widen an envelope, or approve the effect it was asked to evaluate.

6. Enforcement across the whole work graph

Enforcement is only as good as the enumeration of chokepoints. The list is finite and mostly known because these are records or transitions the broker already owns: context creation and admission; room join and invitation; message send and reply; work creation, claim, handoff, and close; invocation create and dispatch; agent discovery and mesh negotiation; delegation and subdelegation; file and artifact transfer; delivery fan-out; managed tool and connector calls; coordination writes; session provisioning; approval; revocation; and effect receipt attachment.

Each needs a call site that refuses to proceed without a valid decision, and a test asserting that refusal — copy the route-inventory pattern, so a new effect path with no enforcement point fails the build. Evaluate twice: at admission, to derive the envelope and exclude denied context, and again immediately before any irreversible or external effect, which implies idempotency keys or a retried approval becomes a duplicate side effect. Roll it out as the mesh gate was — verify-warn first, then enforce — except that verify-warn must be a migration state with an owner and an exit date, not a default.

The evaluator follows authority through a graph, not merely through a turn:

principal -> participant grant -> collaboration context -> work object
          -> invocation -> delegated work -> proposed effect -> receipt

Every edge names what moved: responsibility, selected context, capability, data, or merely a result. A pull-request review, for example, can grant read access to one repository snapshot and permission to return comments without granting write access, access to unrelated history, or authority to invite another agent. A branch-convergence task can grant temporary write authority to named branches, require approval before the merge, and expire when the convergence work closes.

Two enforcement postures that must never blur

Two postures, and the product must never render them identically. Under broker-mediated enforcement Scout covers Scout-owned messages, invocations, deliveries, channels, connectors, managed tools, and coordination writes; if the agent also holds an unrestricted shell, there is no host containment claim. That is Scout's posture today, and the only one reachable without new runtime work. Under contained execution the agent runs inside an OS, VM, or container boundary whose filesystem, process, secret, and network effects are mediated against the envelope — only there are compromised-agent claims defensible.

Every decision record names the posture that applied. "Policy evaluated" is not "bypass was impossible," and a UI that renders the two identically is a lie with a shield icon on it.

7. Audit and limits

"Tamper-proof" is not an honest word for software on an administrated host. "Tamper-evident and enforceable" is achievable: protected effects require a valid decision, so the record sits on the critical path rather than beside it; records are append-only from the public interface (already true); events are hash-linked and periodically signed (not true today); checkpoints export to customer-owned storage no Scout administrator can modify. Hash-linking is the cheap part — the cost is the canonicalization contract from §4, plus key custody and delivery.

The lineage that matters runs authenticated request → ingress evidence → decision → invocation → proposed effect → egress evidence → approval → attempt → external acknowledgement. Scout keeps most of these as distinct records already; the missing links are the decision and evidence records. A transcript is not an audit trail — it shows what was said, not what was authorized.

What must be detectable, not merely stored: gaps in the chain, invalid signatures, rewritten records, missing checkpoints, and the difference between a denied request, an attempted bypass, and a system failure — the last of which is unsolved here and the one operators will care about most.

The limits are worth stating plainly, because the concept's credibility depends on them.

  • A manipulated model stays manipulated. Agreements bound effects, not reasoning; an agent tricked into a permitted action is still tricked.
  • Classifiers are probabilistic. False negatives release data, false positives stall work, and the threshold is a business tradeoff.
  • Broker mediation is not host containment — an agent with a shell bypasses every Scout-side rule.
  • Covert and aggregate channels are out of scope — timing, volume, and slow exfiltration are not addressed by per-request policy.
  • Local-first key custody is hard — key loss, laptop theft, and offboarding need answers first.
  • Revocation is eventually consistent. Offline nodes act on stale grants; TTL bounds the window rather than eliminating it.
  • A compromised broker host compromises everything, and tamper-evidence detects that afterward rather than preventing it.
  • Latency and cost are real — two classifier passes and two evaluations per request, plus stream buffering, on an interactive tool.
  • An audit record is not compliance. It is one input a compliance program might use.

8. Why this is Scout's to build

The mesh trust cone is the architectural seed: a signed-identity, tiered-grant, deny-by-default reference monitor that already ships, scoped to node-to-node HTTP. Agreements generalize the same shape from machine → machine to principal → participant → collaboration context → work → effect, with semantic inspection added around models and untrusted content.

This also matches the direction Scout would need to take to support multiple people honestly. Multi-user rooms cannot be a shared transcript with a member list; each person brings a separate authority domain, agents, credentials, data, and obligations. Multi-owner mesh work is the technical form of the same fact. The agreement substrate gives both surfaces one answer for who may collaborate, what may cross the boundary, what authority may be delegated, and how results are attributed.

That keeps this an extension of proven internal mechanism rather than a new subsystem. Decisions become broker-owned records; contexts bind agreements to rooms and ephemeral mesh relationships; envelopes bind to invocations and work objects while re-evaluation runs along the flight; operator attention carries approvals; and receipts close the loop from authorization to result.

Scout should not become an identity provider, an ESB, a SIEM, or a policy administration platform. Customer systems stay authoritative for human identity, resource identity, secrets, and logs; Scout applies a collaboration-scoped envelope to coordination it already mediates.

9. A plausible path

Each step is independently useful and makes no claim it cannot support.

Step 0 — split the vocabulary. Separate feature support, authorization capability, data-release permission, and delegation authority, and relabel readiness. Small, and every later API inherits the ambiguity if it is skipped.

Step 1 — schema, evaluator, golden vectors. Canonical schema and serialization, digest definition, a pure evaluator in TypeScript and Swift, shared golden vectors, property tests for deny-by-default and monotone narrowing. No enforcement claim.

Step 2 — shadow decisions on real traffic. Emit decision records without enforcing them, following the verify-warn precedent — this is what tells you whether default agreements are too tight to use, which no design review will. Exit criterion, owner, and date fixed up front.

Step 3 — contexts and enforcement over coordination. Signed agreements and participant grants; durable room contexts and ephemeral mesh contexts; policy over ask, send, replies, membership, work claims, handoffs, attachments, artifact transfer, and delegation; per-request envelopes with strict attenuation; require_approval routed to operator attention. This is the smallest step that supports the broader claim: every protected coordination effect is mediated.

Step 4 — multi-principal work. Make ownership, completion conditions, selected-context transfer, result return, and delegation lineage first-class. Support the same work object in a room and across the mesh, including bounded review and branch-convergence templates. Add revocation and re-evaluation during long-running work.

Steps 5–7 — audit, tools, containment. Hash-link decision and evidence records, sign checkpoints, export externally, ship a verifier (after Step 3, because chaining records nothing depends on proves little). Then typed connector operations with resource-level grants, secret handles so raw secrets never enter model context, and pre-effect re-evaluation with external receipts. Then managed filesystem, process, and network boundaries with host posture evidence — only after which, plus independent validation, may Scout make a compromised-agent containment claim.

The sequencing rule: enforcement before audit, and both before containment claims. An audit trail over unenforced policy is theater, and a containment claim without a contained runtime is worse.

10. Open questions that change the shape of the work

  • Policy language: a constrained Scout schema evaluates identically in TypeScript and Swift; Cedar or Rego costs that parity.
  • Issuer hierarchy, amendment quorum, and participant key custody, rotation, recovery, and offboarding.
  • Streaming inspection, whether classifier calls may leave the machine, and what their retained evidence contains.
  • How operators distinguish denial from attempted bypass from system failure, and where Scout-managed capability ends and service-native permission begins.