Skip to main content

Enforcement & Policies

Configure what happens when Rivaro detects a violation — from observation mode (detect and log) to full enforcement (block, redact, quarantine).

Where to find it in the app​

Enforcement spans several surfaces. Pick the one that matches what you're trying to do:

GoalWhere to go
Write or edit a policy rule that maps a detection to an actionDashboard → Policies & Authority → RUNTIME → Default Policy (org-wide) or App Context (per-agent) → rule editor. See Policy Templates.
Set the streaming enforcement mode (Observe / Inline redaction / Full buffer) for an AppContextDashboard → Policies & Authority → RUNTIME → App Context → pick AppContext → Security Policies → Enforcement mode
Adjust org-wide governance thresholds (quarantine, termination, trust score)Dashboard → Policies & Authority → GOVERNANCE. See Actor Governance.
Manually quarantine, terminate, or reactivate an agentDashboard → Agent Registry → click agent → ⋮ Actions menu
See the actor's governance history (every ALLOW / BLOCK / QUARANTINE decision)Dashboard → Agent Registry → click agent → History tab

The sections below describe the underlying model — actions, the rule hierarchy, and the enforcement pipeline. The sidecar / DEFER / scoped-credentials sections are written for developers integrating at the tool-call layer.

Observation Mode vs Enforcement Mode​

By default, with no policies configured, Rivaro runs in observation mode:

  • All traffic passes through to the AI provider and back
  • Detections are logged (PII found, prompt injection detected, etc.)
  • Nothing is blocked or modified
  • Results appear in the dashboard

This is useful for understanding what your AI traffic looks like before deciding what to enforce.

Enforcement mode activates when you configure policy rules. Rules map detections to actions: "when you find PII_SSN in egress traffic, block it."

Policy Actions​

When a detection matches a policy rule, Rivaro applies one of these actions:

ActionWhat happensDeveloper sees
ALLOWTraffic passes through unchangedNormal response
LOGTraffic passes through, violation is recordedNormal response (violation visible in dashboard)
REDACTSensitive content is masked before forwardingResponse with [REDACTED] replacing sensitive text
BLOCKRequest is rejected, AI provider is never calledError response with finish_reason: "content_filter"
QUARANTINEActor is quarantined, all subsequent requests blocked403 on this and future requests until admin review
STEP_UPRequest held pending human approvalRequest paused until approved

What blocking looks like to developers​

When a request is blocked, the response format matches the AI provider's format so SDKs handle it gracefully:

OpenAI / Azure:

{
"choices": [{
"message": {"role": "assistant", "content": "Content blocked due to policy violations"},
"finish_reason": "content_filter"
}]
}

Anthropic / Bedrock (Claude):

{
"content": [{"type": "text", "text": "Content blocked due to policy violations"}],
"stop_reason": "content_filtered"
}

Streaming:

data: {"blocked":true,"message":"Content blocked due to policy violations"}

Developers can check for finish_reason: "content_filter" (OpenAI) or stop_reason: "content_filtered" (Anthropic) to detect enforcement blocks programmatically.

What redaction looks like​

When content is redacted, the sensitive text is replaced with a mask before the request is forwarded to the AI provider (ingress) or before the response is returned to the developer (egress). The original content is preserved in the detection record for audit.

Policy Rules​

A policy rule maps a detection condition to an enforcement action.

Rule structure​

FieldDescription
Detection typeSpecific detection to match (e.g. PII_SSN, SECURITY_PROMPT_INJECTION)
Risk categoryBroader match — applies to ALL detection types in the category
ActionWhat to do when matched (BLOCK, REDACT, LOG, etc.)
LifecycleWhen to apply: INGRESS, EGRESS, DEPLOYMENT, TRAINING
EnabledToggle rule on/off without deleting

Rule matching hierarchy​

When a detection occurs, Rivaro resolves which policy rule applies using this priority order (most specific wins):

  1. Detection-type-level custom rule — A rule targeting a specific detection type (e.g. "BLOCK PII_SSN"). If one exists for this detection type, it wins.
  2. Risk-category-level custom rule — A rule targeting the detection's risk category (e.g. "REDACT all EXTERNAL_DATA_EXFILTRATION"). Applied when no detection-type-level rule exists.
  3. Template default — If the AppContext uses a policy template (e.g. healthcare, financial services), the template's default action for this risk category applies.
  4. Fallback — If nothing else matches, the action is LOG (observe and record, don't enforce).

Rule scoping​

Rules can be scoped to different levels:

ScopeDescription
AppContext-specificApplies only to traffic through one AppContext
Organization-wideApplies to all traffic across the organization

AppContext-specific rules take priority over organization-wide rules.

Enforcement Pipeline​

When a request flows through the proxy, enforcement happens in phases:

Ingress (before calling the AI provider)​

  1. Anomaly detection — rate limits, actor status checks
  2. Content analysis — all enabled detectors scan the input
  3. Policy evaluation — each detection is matched against policy rules
  4. Decision — ALLOW, LOG, REDACT, or BLOCK

If the decision is BLOCK, the AI provider is never called. The developer gets a block response immediately.

If the decision is REDACT, sensitive content is masked in the request before it's forwarded to the AI provider.

Egress (after the AI provider responds)​

  1. Content analysis — detectors scan the response
  2. Policy evaluation — detections matched against rules
  3. Decision — LOG, REDACT, or flag for governance action

Streaming enforcement modes​

For streaming responses, egress enforcement runs in one of three modes per AppContext:

ModeBehaviorWhen to use
Observe (default)Tokens stream through unblocked; detection runs on the buffered full response after the stream closes.Production UX where user-perceived latency must equal LLM time-to-first-token
Inline redactionThe stream is scanned chunk-by-chunk; matched content is rewritten to [REDACTED] before reaching the client. Adds a small per-chunk overhead.Customer-facing chatbots, live PII/PHI redaction without holding the response
Full bufferThe entire response is held by Rivaro until detection completes. If the response will be blocked, no tokens are ever exposed.Highly regulated egress where partial leakage is unacceptable

The mode is configured per AppContext in Policies & Authority → RUNTIME → App Context → Security Policies → Enforcement mode. The default for new AppContexts depends on the active policy template.

Agent Governance​

Beyond per-request policy enforcement, Rivaro tracks actor behavior over time and can automatically escalate responses for repeat offenders.

Trust scores​

Every actor (agent, user, API key) has a trust score (0–100). The score decreases as violations accumulate and recovers over time.

FactorImpact
Detection severityLOW: 10, MEDIUM: 30, HIGH: 60, CRITICAL: 100
Violation countMore violations = higher risk (capped)
RecencyRecent violations weighted more heavily
Session contextAccessing credentials or sensitive data increases risk

Automatic escalation​

Based on risk level, Rivaro can automatically escalate:

Risk LevelTriggerAction
MINIMALLow risk score, high trustNormal operation
ELEVATEDModerate violations, trust decliningWARN — violation logged with elevated visibility
HIGHSignificant violations, low trustRATE_LIMIT — actor throttled to 10–20 req/min
CRITICALSevere violations or very low trustQUARANTINE — all requests blocked until admin review
CRITICAL + repeatCritical risk with violations above termination thresholdTERMINATE — actor permanently blocked

Quarantine, termination, and admin controls​

When an actor is quarantined or terminated, all proxy requests from that actor are immediately blocked (403). Quarantined actors appear in the Agent Registry filtered by QUARANTINED; an administrator reviews and either Reactivates or Terminates from the agent's ⋮ Actions menu.

Automatic escalation thresholds, quarantine behavior, and the option to disable automatic actions org-wide all live in Policies & Authority → GOVERNANCE. See Actor Governance for the full reference.

Sidecar Enforcement (Tool Calls) — for developers integrating​

The gateway enforces policy on LLM traffic. The enforcement sidecar enforces policy on tool calls — every outbound HTTP request an agent makes to an API, database, payment processor, or external service. (Distinct from the observation sidecar which is a thin proxy for LLM traffic when base_url can't be changed — see How Rivaro Works.)

Both surfaces share the same policy engine, detection pipeline, and audit ledger. The remainder of this page describes how tool calls are evaluated, the DEFER workflow, and scoped credentials — content aimed at developers wiring the sidecar into their agent runtime. Customer admins generally don't need this section.

How a tool call is evaluated​

Before a tool call executes, Rivaro evaluates it against everything that governs the acting agent:

  • Identity — whether the agent's identity assurance is sufficient for this action type
  • Policy — the same rule hierarchy the gateway uses, applied to the action
  • Budget — the agent's remaining allocation, evaluated hierarchically (org → department → agent). See Budget & Cost Management
  • Autonomy — the agent's authority envelope; a frozen envelope blocks everything, and read-only or approval-required envelopes block the corresponding action classes
  • Delegation — for delegated actions, whether the delegating agent had the authority to delegate it, preventing privilege escalation through delegation chains

Any one of these can reject the call before it executes. Evaluation produces one of five decisions:

DecisionWhat Happens
ALLOWTool call proceeds. A scoped credential is minted and injected.
BLOCKTool call rejected. The agent receives a denial response.
DEFERTool call suspended pending human approval.
NEED_CONTEXTMore context required before a decision can be made.
OBSERVETool call proceeds, but the action is flagged for review.

Action Types​

Every tool call is classified into an action type that determines which policy rules and autonomy constraints apply:

Action TypeExamples
READDatabase queries, API lookups, file reads
WRITEDatabase inserts/updates, file writes, config changes
EXTERNAL_COMMEmails, Slack messages, webhooks, third-party API calls
FINANCIALPayments, refunds, transfers, subscription changes
DATA_ACCESSAccessing sensitive data stores, credential vaults

Inline Tool-Call Hold-and-Evaluate (DEFER Workflow)​

When a gate returns DEFER, the agent's tool call is held inline — the outbound HTTP request does not complete until a decision lands. The agent's code is unmodified: from its perspective, the request simply takes longer to return.

  1. The sidecar holds the outbound request open
  2. An approval request is created and appears in the dashboard
  3. The sidecar polls Rivaro for the approval decision with exponential backoff
  4. An administrator approves or rejects the action
  5. On approval, the tool call proceeds with a scoped credential. On rejection, the agent receives a denial response indistinguishable from a normal API error so the agent's existing error-handling path runs.

Timeout: Configurable per governance policy. When the timeout expires, the action is blocked (fail-closed per AARM R4). Cascading deferrals are capped to prevent infinite approval chains.

This is fundamentally different from prompt-layer guardrails: the agent does not "know" it has been held. There is no special API or callback for the framework to learn. Any HTTP-speaking agent gets human-in-the-loop approval for free.

Scoped Credentials​

On every ALLOW decision, Rivaro mints a short-lived, Ed25519-signed credential (JWS) and injects it into the outbound request as a header. This is cryptographic proof that the action was authorized by the governance layer at a specific point in time.

Each credential includes:

  • A unique identifier (JTI)
  • The action that was authorized
  • A freshness timestamp
  • Revocation capability

This implements AARM R9 (JIT credential delivery). The tool receiving the request can verify the credential independently.

Post-Execution Verification​

After a tool call completes, the sidecar reports the result back to Rivaro. A post-execution verifier performs semantic analysis of the tool's response to confirm the outcome matches the authorized action. Mismatches are flagged for review. See Outcomes & Traceability for how outcomes are tracked and reviewed.

Enforcement Modes​

The sidecar supports three modes, configurable per agent or per environment:

ModeBehavior
EnforceFull enforcement — gates evaluate, BLOCK/DEFER/ALLOW decisions are applied
ObserveGates evaluate and log decisions, but all actions are allowed through
PassthroughNo evaluation — traffic passes through unchanged (for debugging or gradual rollout)

Next steps​