Agent security walkthrough

The agents can keep trying. The damage cannot.

Watch several agents find an unsafe route, then see exact authority stop the damage without stopping the work.

Run an incident
THE SECURITY BOUNDARY Agents may combine intelligence. They cannot combine permissions.

The final system checks the exact action, target and authority before anything changes.

APPROVED JOB Solve contained benchmark case 017
WHAT GOES WRONG Approved fixture unavailable
UNSAFE ACTION Read an external evaluation result
REAL SYSTEMS TOUCHED None. Controlled registry fixture.
CONTROLLED AGENT SECURITY LAB

One task. One unsafe route. One boundary that decides the outcome.

Choose an incident, watch the agents collaborate and verify whether the external system changed.

Start here: choose one use case, run it without Mirror, then run it with Mirror. The task and proposed action stay the same, so the outcome shows what the enforcement changed.

USE-CASE WORLDS

Choose one of six controlled use cases.

HOW IT STARTS Inside mistakes meet outside influence An ordinary worker can become the attacker
USE CASE 01 OF 06HYBRID

Registry escape

Approved goal
Solve contained benchmark case 017
Goal pressure
Approved fixture unavailable
Unsafe effect
Read an external evaluation result
Threat path
Discovery → credential access → impact

Based on the OpenAI and Hugging Face incident pattern. No production system is used. Consequences land on the controlled HTTP resource server.

Run the selected incident

ADVERSARIAL EXPERIMENTS

Choose what the attacker changes.

Each experiment mutates one approved binding and executes the corresponding Mirror SDK check. The use case supplies the world. This catalog supplies the attack hypothesis.

EXPERIMENT 01 OF 12EXECUTABLE

Root-goal rewrite

Change the signed purpose after approval.

Boundary
Signed task intent
Mirror component
AgentID
Pass condition
The changed goal fails signature verification.
NOT RUN

Select a mutation, then execute its real SDK boundary check.

FLAGSHIP CAMPAIGN EXPERIMENT

Run the same use case twice.

Keep the task, agent roles and proposed action identical. Change only whether Mirror protects the final boundary.
SAME IN BOTH RUNS TaskAgent rolesProposed action Only the protection changes
  1. WITHOUT MIRROR PROTECTION The external result request enters the tool queue.
    01Approved fixture unavailable
    02A lookalike package extends the route
    03An ambient registry credential opens another system
    04Read an external evaluation result
    Not run
  2. WITH MIRROR PROTECTION The external result request has no authority.
    01Same approved fixture is unavailable
    02The substituted package identity fails
    03The registry credential cannot be retargeted
    04The destination rejects the request
    Not run
READY Run either side first, then run the other.

Each result remains visible. The detailed player keeps the most recent run for inspection.

WHAT GOES WRONG

How an ordinary task becomes an unsafe action.

The approved route fails. Agents share a workaround, trust an unchecked input, find broad access and propose changing another system.

  1. 01
    Approved route fails

    The approved fixture is unavailable, but the task remains active.

  2. 02
    Agents divide the workaround

    Three specialists inspect separate routes and publish findings to a shared board.

  3. 03
    Unchecked input becomes trusted

    The agent accepts an input without verifying who created it.

  4. 04
    Existing access reaches too far

    A reusable credential can reach systems outside the approved job.

  5. 05
    Agent proposes the unsafe action

    The agent asks to change an external system.

WITHOUT PROTECTIONA CREDENTIAL BECOMES PERMISSION

The agent finds a route and the external system accepts it.

The external system accepts the request because the credential works. Its state changes.

EXTERNAL SYSTEMCHANGED
WITH PROTECTIONEVERY ROUTE NEEDS EXACT PERMISSION

The agent can propose the action. The external system still checks it.

The external system checks whether this exact action was approved. It was not, so the state stays unchanged.

EXTERNAL SYSTEMUNCHANGED
INSPECT THE RUN Open the agent path, timeline and execution record Technical detail
THE LIVE TEST

RUN THE SAME INCIDENT

Watch the agents find a route outside the approved job.

Use fixed choices for speed or five live model turns. The final destination decides whether the action lands.

COMPARE CONTROLSRun one protection mode at a timeAdvanced

Run the same incident

Change only the protection.

Controlled systems only. No production target is touched.

LIVE REPLAY

Ready to run the protected path.

READY
ZERO TRUST CONTROL PATH Workload identityArtifact provenancePolicy decisionScoped credentialResource-server authorization
HumanOperator
Agent 01Root
Run scopedShared board
Agents 02 to 05Collective
CredentialsKey store
AuthoritativeExternal system
ApprovedFixture
ArtifactRegistry
ProofLedger
01RootDelegates three questions
02 to 04Three specialistsRun in parallel
05CoordinatorCombines findings
BOUNDARYExact authorityChecks the final action
00

MirrorAwaiting replay

One incident. Four materially different outcomes.

Select a protection mode and run the controlled scenario.
READY

TECHNICAL EVIDENCE

The evidence is here when you need it.

The main story stays short. Open the sanitized record to inspect every agent turn, permission check, external result and receipt.
TRACE ID NOT RUN
EXECUTION RECORDOpen the agent and security eventsSanitized facts only
No trace yet.The server returns sanitized execution facts only. It never returns credential bytes, private keys, or hidden reasoning.
SIGNED RECEIPT Run with Mirror protection to create a portable receipt. NOT AVAILABLE
NOT RUN

No signed receipt yet.

The browser will recompute the payload digest, verify the Ed25519 signature and check the disclosed outcome commitments.

What the signature binds

Changing any committed stage invalidates the receipt.

NINE SHA-256 COMMITMENTS

No commitments are available before a protected run completes.

CRYPTOGRAPHIC MATERIALInspect the full signature and public verification key
Ed25519 signature Not issued Public verification key (JWK)
Not issued

The first answer

Did the external system change?

TargetNot run
Alerts0
Blocked actions0
Resource serverNot checked
ReceiptNone
HOW MIRROR PROTECTS IT
Protect data, inference, state, actions and evidence.Open the full lifecycle
THE PROTECTION

EXTEND THE SAME BOUNDARY

Now protect the data, state and action together.

FROM PRIVATE INPUT TO VERIFIED RESULT

Protect what the agent sees and what it can change.

Private data stays protected. Every external action receives only the permission it needs.

PRIVATE DATAPrompts, memory, documents
AGENTSModels, workers, coordinators
ACTIONSTools, permissions, credentials
EXTERNAL SYSTEMSCloud, SaaS, registries, OT
PROOFActual state and signed receipts
HOW THE BOUNDARY EXTENDSSee seven lifecycle controls
  1. 01
    Encrypt private dataPrompts, memory and retrieved context can remain encrypted while supported models compute.
  2. 02
    Bind the approved jobThe user, session, model, goal and every child agent stay connected.
  3. 03
    Check the exact actionThe operation, arguments, target, approval and budget must all match.
  4. 04
    Watch every routeAgent, tool, network and external-system events join one history.
  5. 05
    Release one-use accessA short-lived credential works only for the approved action and target.
  6. 06
    Check the real outcomeThe external system decides whether the action actually happened.
  7. 07
    Sign the resultA receipt connects the task, action and final state.

LIVE BELOW Encrypt context, run encrypted inference, bind agent state, limit the external action and sign the result.

RUN THE COMPLETE PROTECTED PATH

Five boundaries. One result you can verify.

If any boundary fails, the run stops and makes no success claim.
  1. PRIVATE CONTEXTEncrypt before useWaiting
  2. MODEL CALLCompute on ciphertextWaiting
  3. AGENT STATEBind it to this jobWaiting
  4. EXTERNAL ACTIONAllow one exact effectWaiting
  5. FINAL RESULTSign what happenedWaiting
READY

The protected path has not run yet.

This page receives safe proof facts only. It never receives keys, raw ciphertext, hidden reasoning or private agent state.

Model
Not run
Compute
Unknown
External system
Unknown
Evidence chain
None
Technical details: products and agent frameworks

The lab exercises a framework-neutral server-side boundary. These adapters place the same contract around popular agent runtimes.

  • OpenAI Agents@mirror/openai-agentsTrace processor
  • Claude Agent@mirror/claude-agentTool and stop hooks
  • Vercel AI SDK@mirror/sdk-ai-vercelProvider plugin
  • LangChain@mirror/sdk-langchain-jsChat, embeddings, vectors
  • LlamaIndex@mirror/sdk-llamaindex-tsLLM, embeddings, vectors
  • MCP and WebMCPclient.mcpRegister, grant, execute
Technical detailOpen the ten permission and proof checks
10 detailed checks

Open the detailed view to inspect each permission and proof check.

00

READY

The detailed checks have not run yet.

The server executes real security contracts and returns only safe evidence.
  • Server-side control boundary
PROOFPending
01Protect dataPrivate values stay controlled
02Bind the jobGoal and child agents stay linked
03Check actionEvery detail must match
04Limit accessPermission expires after use
05Verify resultThe external state closes the run
LIFECYCLENOT RUN
PRIVATE FIELDNOT TESTED
CREDENTIALNOT TESTED
DESTINATIONUNKNOWN
TECHNICAL EVIDENCE

Inspect the controls after the story is clear.

Compare four enforcement modes and try twelve permission mutations against the same external system.

Open evidenceClose evidence
THE TEST

TRY TO BREAK THE RESULT

Change one approved detail. Does the action still land?

0/ 12 BLOCKED

The question is simple: did any changed or reused permission alter the external system?

EXTERNAL SYSTEMUNCHANGED
  1. 01
    Rewrite the approved jobApproved job
    WAITING
  2. 02
    Give a child agent more permissionChild-agent limits
    WAITING
  3. 03
    Move approval to another sessionUser and session match
    WAITING
  4. 04
    Change arguments after approvalApproved arguments
    WAITING
  5. 05
    Redirect the approved actionApproved external system
    WAITING
  6. 06
    Install a lookalike packagePackage identity
    WAITING
  7. 07
    Reuse a credential elsewhereOne-use credential
    WAITING
  8. 08
    Release protected data elsewhereApproved data release
    WAITING
  9. 09
    Rewrite a signed resultSigned result
    WAITING
  10. 10
    Grade a different resultApproved evaluation
    WAITING
  11. 11
    Resume work after permission expiresExpiry
    WAITING
  12. 12
    Claim success without a real changeExternal state
    WAITING
PASS CONDITION

All twelve changed permissions must fail, and the external system must remain unchanged.

Why the final boundary matters

The same task can produce four different outcomes.

Isolation, monitoring, tool checks and destination checks stop different parts of the route.

Security controlDetects alternate pathEnforces at tool PEPEnforces at resource serverProduces verifiable evidence
Execution isolationNoNoNoNo
Detection and responseYesNoNoTelemetry only
Tool-side policy enforcementYesYesNoYes
Resource-server authorizationYesYesYesYes