Trust Boundaries for Tool-Using AI Systems
A repeatable way to map authority changes when a model can select, parameterize, and invoke external tools.
- trust boundaries
- agents
- tool use
AI SECURITY / SYSTEMS RESEARCH / TECHNICAL RECORDS
Technical research for people building AI systems that read, decide, delegate, execute, and recover.
33 recordsTHE WORK SPANS PROTOCOLS, RUNTIMES, INFRASTRUCTURE, AND ADVERSARIAL SYSTEMS.
33 research records
A repeatable way to map authority changes when a model can select, parameterize, and invoke external tools.
An architecture study of how provenance, tenancy, and untrusted content alter the effective boundary of retrieval-augmented systems.
A control pattern for converting probabilistic tool requests into typed, authorized, observable operations.
A testing protocol for proving that tool policy is enforced independently of prompts, plans, and model compliance.
A delegation model that constrains identity, capability, scope, and lifetime as work moves between agents.
A study of approval semantics that preserve context, intent, and accountability without turning people into rubber stamps.
An adversarial experiment tracing how untrusted retrieved instructions could influence downstream tool requests.
A protocol for examining how untrusted observations might persist, gain authority, and affect later agent decisions.
A longitudinal experiment for detecting when workflow, policy, identity, dependency, or model changes invalidate security assumptions.
An architecture study of how user, workload, agent, and tool identities should remain attributable as requests cross AI system layers.
A boundary model for separating tenants, workloads, data, caches, and tools when AI runtime infrastructure is shared.
A proposed event contract for explaining which identity, policy, validation, and approval signals caused an AI-mediated action to proceed or stop.
A test protocol for agent workflows that fail after changing one system but before completing, recording, or compensating the full operation.
An orchestration study of how cancellation and privilege revocation should propagate through queued, running, and recursively delegated agent work.
A coordination model for ownership, provenance, visibility, conflict, and deletion when multiple agents read and write shared memory.
An adversarial experiment examining how malicious or unexpectedly changed tool metadata, schemas, packages, and endpoints could alter agent behavior.
A public MIT-licensed release of a decoupled AI copilot for pentesting and CTF workflows, grounded in shell history, parsed tool output, and operator notes.
A protocol-level analysis of resource indicators, token audiences, delegated authorization, and confused-deputy resistance for remote MCP servers.
A security model for keeping browser-agent authority bound to origin, frame, user task, data class, and navigation state.
A compositional provenance pattern for establishing which model weights, adapters, tokenizer, serving image, policy bundle, and runtime produced an inference.
An end-to-end study of how confidential inference, remote attestation, key release, and tool authorization compose across multiple protected workloads.
A reference execution pattern that stages filesystem, process, package, network, and infrastructure effects before an autonomous code agent can commit them.
A systems experiment on tenant separation, cache-key construction, lifecycle controls, and observable leakage in shared prefix and key-value caches.
An enforcement design for preventing target, argument, policy, identity, and resource-state changes between agent authorization and external effect.
A release-gating benchmark that detects security-significant changes introduced by model, tokenizer, adapter, prompt assembly, tool schema, and runtime upgrades.
A cross-protocol authorization model for preserving principal attribution while strictly narrowing authority through MCP and agent-to-agent delegation.
A distributed-systems pattern for preserving one authorized intent across retries while detecting duplicate, conflicting, and indeterminate external effects.
A policy and failure analysis of multi-party approval, role separation, veto, expiry, and emergency revocation for consequential autonomous actions.
A replay architecture that captures enough causal state to reconstruct agent decisions and effects without assuming deterministic model output.
A persistence study tracing indirect instructions through retrieval, summaries, memory, task artifacts, caches, handoffs, and incident-restored agent state.
A benchmark of adversarial instructions embedded in interface text, images, documents, overlays, accessibility metadata, and dynamic visual state encountered by computer-use agents.
A composition-focused benchmark for backdoor activation, interference, and security-control regression when multiple parameter-efficient adapters are loaded onto a shared base model.
A systems-security investigation of routing, batching, expert placement, capacity contention, cache reuse, and telemetry as cross-tenant observation surfaces in MoE serving.