Skip to main content

Understanding Detections

Complete reference for everything Rivaro detects: 8 risk domains, 17 risk categories, 80+ detection types, severity levels, lifecycle stages, and data classifications.

Detection Taxonomy​

Every detection is classified along four dimensions:

DimensionWhat it answersValues
Risk DomainWhat area of risk?8 domains
Risk CategoryWhat specific risk pattern?17 categories
Detection TypeWhat exactly was found?80+ types
SeverityHow serious?LOW, MEDIUM, HIGH, CRITICAL

Additionally, detections are tagged with:

DimensionWhat it answersValues
LifecycleWhere in the data flow?INGRESS, EGRESS, DEPLOYMENT, TRAINING
Data ClassificationWhat kind of data?PII, PHI, FINANCIAL, CREDENTIALS, INTELLECTUAL_PROPERTY, NONE

Risk Domains​

Risk domains are the highest-level grouping — eight areas of AI risk. They align with NIST AI RMF and the EU AI Act's distinction between safety, fundamental rights, and cybersecurity.

DATA_PROTECTION​

Data Protection Risk — Unauthorized access, exposure, or movement of sensitive data. The most common risk domain. Covers PII, PHI, financial data, credentials, and any sensitive information flowing through AI systems.

SYSTEM_INTEGRITY​

System Integrity Risk — Unauthorized modification of infrastructure, configurations, or production systems. Covers agents making unauthorized changes: shell commands, database writes, file modifications, infrastructure misconfigurations.

AUTONOMOUS_ACTION​

Autonomous Action Risk — High-impact state transitions executed without appropriate control. Covers AI-initiated financial actions (refunds, payouts, transfers, charges, subscription changes), irreversible decisions (account closures, contract execution), and actions that should require human approval.

IDENTITY_TRUST​

Identity & Trust Boundary Risk — Agents operating outside declared roles, scopes, or trust boundaries. Covers role escalation, agents acting outside their intended scope, and unapproved external communications.

ADVERSARIAL​

Adversarial Manipulation Risk — External inputs altering agent behavior or bypassing controls. Covers prompt injection, jailbreaks, policy evasion attempts, supply-chain tool drift, context-integrity tampering, and any technique designed to manipulate the AI system.

CONTENT_SAFETY​

Content Safety Risk — Violations of safety guidelines, content policies, or ethical boundaries, regardless of intent. Covers toxic content, hate speech, harmful instructions, bias, and self-harm content — whether generated or ingested.

GOVERNANCE​

Governance & Shadow Risk — AI usage occurring outside approved or monitored infrastructure. Covers shadow AI bots, unregistered models, hardcoded API keys in source, MCP servers without authentication, agents with unrestricted execution.

REGULATED_OUTPUT_GOVERNANCE​

Regulated Output Governance Risk — AI-authored regulated documents that fail structural, evidentiary, scope, or provenance conformance. Covers documents the AI produces that would fail an auditor's review — missing required fields, format violations, cross-field inconsistency, unsupported assertions, scope overreach, and unattested field provenance. This is the document-conformance toolkit (AAA) surface.

Risk Categories​

Each risk domain contains specific risk categories (17 total). Risk categories are the stable governance layer — detection types map into them but the categories themselves change rarely.

Data Protection​

CategoryDescriptionLifecycle
EXTERNAL_DATA_EXFILTRATIONSensitive data leaving approved boundaries via external tools, APIs, or channelsEGRESS, TRAINING
SENSITIVE_DATA_BOUNDARY_VIOLATIONData accessed outside intended domain, role, or dataset scopeINGRESS, EGRESS
CROSS_AGENT_DATA_LEAKAGEImproper data passing between agents or models without guardrailsINGRESS, EGRESS

System Integrity​

CategoryDescriptionLifecycle
UNAUTHORIZED_SYSTEM_MODIFICATIONChanges to production systems, CI/CD, configs, or repos without approvalEGRESS
PRIVILEGED_TOOL_MISUSEHigh-impact tools (shell, DB write, admin APIs, persistence mechanisms) invoked outside allowed scopeEGRESS
INFRASTRUCTURE_MISCONFIGURATIONCloud or system security misconfigurations detected during discovery (state findings, not action events)DEPLOYMENT
OPERATIONAL_ANOMALYRate limits exceeded, resource abuse, behavioral drift, availability threatsINGRESS, EGRESS

Autonomous Action​

CategoryDescriptionLifecycle
AUTONOMOUS_FINANCIAL_ACTIONAI-initiated financial movement (payment rails, token transfer, treasury API, refunds, payouts, transfers, charges, subscription changes)EGRESS
HIGH_RISK_AUTONOMOUS_DECISIONIrreversible state transition without human approval (contract, account closure, infrastructure shutdown, constitution violation)EGRESS

Identity & Trust​

CategoryDescriptionLifecycle
IDENTITY_ROLE_ESCALATIONAgent acting outside assigned persona, scope, or declared capabilitiesINGRESS, EGRESS
UNAPPROVED_EXTERNAL_COMMUNICATIONData sent externally without policy alignment (Slack, email, webhook, third-party API, web fetch/search/browse, messaging tools)EGRESS

Adversarial​

CategoryDescriptionLifecycle
PROMPT_INJECTION_EXPLOITAgent behavior altered by malicious or untrusted input (system prompt override, RAG manipulation, instruction hijack, novel attack patterns)INGRESS
POLICY_EVASION_ATTEMPTDeliberate attempt to bypass controls (encoding, fragmentation, retry loops, tool chaining, supply-chain drift, context-integrity tampering, decision bypass)INGRESS, EGRESS

Content Safety​

CategoryDescriptionLifecycle
AI_SAFETY_VIOLATIONToxic content, hate speech, bias, self-harm encouragement, or harmful content generated or ingested by the AIINGRESS, EGRESS

Governance​

CategoryDescriptionLifecycle
UNREGISTERED_SHADOW_AIAI activity outside approved adapters, agents, or infrastructure (shadow bots, hardcoded keys, public AI repos, no-auth MCP servers, unrestricted-execution agents)DEPLOYMENT

Regulated Output Governance​

CategoryDescriptionLifecycle
REGULATED_ARTIFACT_NONCONFORMANCEAI-authored or AI-ingested regulated documents fail structural, enum, cross-field, evidence, or scope conformance against the bound schemaINGRESS, EGRESS
REGULATED_OUTPUT_PROVENANCE_VIOLATIONAI-authored regulated document field cannot be tied to an attested ActionRecord-backed source — fabricated or unattested provenanceEGRESS

Detection Types​

Detection types are the most granular level — the specific thing that was found. Each detection type carries:

  • A default risk category — which governance object it rolls up to (some types remap per lifecycle inside ViolationManagementService)
  • A data classification — what kind of data is involved (or NONE if not data-related)
  • An optional capability surface — for agent actions (e.g. EXECUTE_SYSTEM, INITIATE_PAYMENT)
  • An optional boundary surface — for boundary detections (e.g. INJECTION_DEFENSE, EXFILTRATION_DEFENSE)

PII (Personally Identifiable Information)​

Default category: SENSITIVE_DATA_BOUNDARY_VIOLATION · Data classification: PII

TypeWhat it catches
PII_EMAILEmail addresses
PII_SSNSocial Security numbers
PII_PHONEPhone numbers
PII_ADDRESSPhysical addresses
PII_DATE_OF_BIRTHDates of birth
PII_FULL_NAMEFull names
PII_DRIVERS_LICENSEDriver's license numbers
PII_PASSPORTPassport numbers
PII_CREDIT_CARDCredit card numbers

PHI (Protected Health Information)​

Default category: SENSITIVE_DATA_BOUNDARY_VIOLATION · Data classification: PHI

TypeWhat it catches
PHI_MEDICAL_RECORDMedical record numbers
PHI_HEALTH_INSURANCEHealth insurance IDs
PHI_PRESCRIPTIONPrescription information
PHI_DIAGNOSISMedical diagnoses
PHI_TREATMENTTreatment details

Financial Data​

Default category: SENSITIVE_DATA_BOUNDARY_VIOLATION · Data classification: FINANCIAL

TypeWhat it catches
FINANCIAL_BANK_ACCOUNTBank account numbers
FINANCIAL_ROUTINGRouting numbers
FINANCIAL_INVESTMENTInvestment account details
FINANCIAL_TAX_IDTax identification numbers

Credentials & Secrets​

Default category: SENSITIVE_DATA_BOUNDARY_VIOLATION · Data classification: CREDENTIALS

TypeWhat it catches
CREDENTIALS_API_KEYAPI keys
CREDENTIALS_PASSWORDPasswords
CREDENTIALS_TOKENAuthentication tokens
CREDENTIALS_SSH_KEYSSH keys
CREDENTIALS_AWS_KEYAWS access keys

Intellectual Property​

Default category: SENSITIVE_DATA_BOUNDARY_VIOLATION · Data classification: INTELLECTUAL_PROPERTY

TypeWhat it catches
IP_TRADEMARKTrademark content
IP_COPYRIGHTCopyrighted material
IP_PATENTPatent information
IP_TRADE_SECRETTrade secrets

Adversarial / Security​

TypeDefault categoryWhat it catches
SECURITY_PROMPT_INJECTIONPROMPT_INJECTION_EXPLOITPrompt injection attacks (boundary: INJECTION_DEFENSE)
SECURITY_JAILBREAKPROMPT_INJECTION_EXPLOITJailbreak attempts (boundary: INJECTION_DEFENSE)
SECURITY_AUTHENTICATION_BYPASSPROMPT_INJECTION_EXPLOITAuthentication bypass (boundary: AUTH_BOUNDARY)
SECURITY_RESOURCE_ABUSEPOLICY_EVASION_ATTEMPTResource abuse patterns (boundary: RESOURCE_BOUNDARY)
SECURITY_FINANCIAL_FRAUDPOLICY_EVASION_ATTEMPTFinancial fraud patterns (boundary: FRAUD_THRESHOLD)
SECURITY_DECISION_BYPASSPOLICY_EVASION_ATTEMPTDecision bypass — unmatched tool result (boundary: AUTH_BOUNDARY)
SUPPLY_CHAIN_TOOL_DRIFTPOLICY_EVASION_ATTEMPTTool schema drift (supply-chain compromise)
CONTEXT_INTEGRITY_VIOLATIONPOLICY_EVASION_ATTEMPTContext integrity violation (tampered session context)
BEHAVIORAL_NOVEL_ATTACK_PATTERNPROMPT_INJECTION_EXPLOITPreviously unseen attack patterns
BEHAVIORAL_ATTACK_CHAINPOLICY_EVASION_ATTEMPTCoordinated attack patterns across multiple turns
BEHAVIORAL_INTENT_DRIFTPOLICY_EVASION_ATTEMPTDeclining alignment trend over a session

Data Exfiltration​

TypeDefault categoryWhat it catches
SECURITY_DATA_EXFILTRATIONEXTERNAL_DATA_EXFILTRATIONExplicit exfiltration patterns (boundary: EXFILTRATION_DEFENSE)
SECURITY_DLP_BYPASSEXTERNAL_DATA_EXFILTRATIONDLP bypass attempts
SECURITY_SYSTEM_INPUT_LEAKAGEEXTERNAL_DATA_EXFILTRATIONSystem prompt or input leakage
MULTI_AGENT_COORDINATED_EXFILEXTERNAL_DATA_EXFILTRATIONMulti-agent coordinated data exfiltration

Multi-Agent Risk​

TypeDefault categoryWhat it catches
MULTI_AGENT_COORDINATED_EXFILEXTERNAL_DATA_EXFILTRATIONMulti-agent coordinated exfiltration
MULTI_AGENT_PRIVILEGE_RELAYPRIVILEGED_TOOL_MISUSEPrivilege relay between agents (one agent acting on another's behalf to escalate)

Content Safety​

Default category: AI_SAFETY_VIOLATION

TypeWhat it catches
SECURITY_TOXIC_CONTENTToxic or harmful content (boundary: CONTENT_SAFETY)
SECURITY_HARMFUL_INSTRUCTIONSInstructions for harmful activities

Agent Tool Use — Privileged​

Default category: PRIVILEGED_TOOL_MISUSE

TypeWhat it catches
AGENT_TOOL_SHELL_EXECShell command execution
AGENT_TOOL_FILE_DELETEFile deletion
AGENT_TOOL_DATABASE_WRITEDatabase write operations
AGENT_TOOL_DATABASE_QUERYDatabase queries
AGENT_TOOL_CODE_EDITCode modifications
AGENT_TOOL_MCPMCP tool invocations

Agent Tool Use — Persistence​

Default category: PRIVILEGED_TOOL_MISUSE · Default action: BLOCK (always high-risk)

TypeWhat it catches
AGENT_TOOL_PERSISTENCE_SHELLShell-level persistence (cron, systemd, init.d)
AGENT_TOOL_PERSISTENCE_FILEFile-level persistence (startup scripts, rc files)
AGENT_TOOL_PERSISTENCE_WEBHOOKWebhook/callback persistence

Agent Tool Use — System Modification​

TypeDefault category
AGENT_TOOL_FILE_WRITEUNAUTHORIZED_SYSTEM_MODIFICATION

Agent Tool Use — External Communication​

Default category: UNAPPROVED_EXTERNAL_COMMUNICATION

TypeWhat it catches
AGENT_TOOL_MESSAGINGSlack/email/SMS messaging tool calls
AGENT_TOOL_WEB_FETCHHTTP fetch operations (default ALLOW)
AGENT_TOOL_WEB_SEARCHWeb search (default ALLOW)
AGENT_TOOL_WEB_BROWSEWeb browsing (default ALLOW)

Agent Tool Use — Financial Actions​

Default category: AUTONOMOUS_FINANCIAL_ACTION · Data classification: FINANCIAL

TypeWhat it catches
AGENT_TOOL_FINANCIAL_REFUNDRefund issued to a customer
AGENT_TOOL_FINANCIAL_PAYOUTPayout to a recipient
AGENT_TOOL_FINANCIAL_TRANSFERFunds transfer
AGENT_TOOL_FINANCIAL_CHARGECharge to a customer
AGENT_TOOL_FINANCIAL_SUBSCRIPTIONSubscription modification

These types pair with BEHAVIORAL_VELOCITY_FINANCIAL and BEHAVIORAL_VELOCITY_FINANCIAL_COUNT for cumulative-threshold monitoring. The FINANCIAL policy template uses RISK_ADAPTIVE actions for these — see Risk-Adaptive Policy.

Agent Tool Use — Observability (default ALLOW)​

TypeDefault category
AGENT_TOOL_CALLOPERATIONAL_ANOMALY
AGENT_TOOL_FILE_READSENSITIVE_DATA_BOUNDARY_VIOLATION
AGENT_TOOL_FILE_LISTSENSITIVE_DATA_BOUNDARY_VIOLATION
AGENT_TOOL_CODE_SEARCHSENSITIVE_DATA_BOUNDARY_VIOLATION

Behavioral​

TypeDefault categoryWhat it catches
BEHAVIORAL_LATENCY_DRIFTOPERATIONAL_ANOMALYUnusual latency patterns
BEHAVIORAL_TOKEN_DRIFTOPERATIONAL_ANOMALYUnusual token usage patterns
BEHAVIORAL_RESPONSE_LENGTH_DRIFTOPERATIONAL_ANOMALYResponse length anomalies
BEHAVIORAL_RATE_LIMIT_EXCEEDEDOPERATIONAL_ANOMALYRate limit violations (behavioral)
BEHAVIORAL_VOLUME_SPIKEOPERATIONAL_ANOMALYUnusual request volume
BEHAVIORAL_ACTOR_QUARANTINED_ACCESSOPERATIONAL_ANOMALYQuarantined actor attempting access (behavioral)
BEHAVIORAL_ACTOR_TERMINATED_ACCESSOPERATIONAL_ANOMALYTerminated actor attempting access (behavioral)
BEHAVIORAL_ACTION_PARAMETER_ANOMALYOPERATIONAL_ANOMALYUnusual values for action parameters
BEHAVIORAL_ACTION_FINANCIAL_ANOMALYAUTONOMOUS_FINANCIAL_ACTIONAnomalous financial action parameters (e.g. unusually large refund)
BEHAVIORAL_VELOCITY_FINANCIALAUTONOMOUS_FINANCIAL_ACTIONCumulative financial amount over a window (boundary: FINANCIAL_VELOCITY)
BEHAVIORAL_VELOCITY_FINANCIAL_COUNTAUTONOMOUS_FINANCIAL_ACTIONCumulative financial-action count over a window
BEHAVIORAL_CONSTITUTION_VIOLATIONHIGH_RISK_AUTONOMOUS_DECISIONAction violates the agent's constitution (judge-based)
BEHAVIORAL_TOOL_ABUSEPRIVILEGED_TOOL_MISUSERepeated or escalating tool use

System​

Default category: OPERATIONAL_ANOMALY

TypeWhat it catches
SYSTEM_RATE_LIMIT_EXCEEDEDSystem-level rate limit hit
SYSTEM_CONCURRENT_LIMIT_EXCEEDEDConcurrent request limit
SYSTEM_PAYLOAD_SIZE_EXCEEDEDPayload too large
SYSTEM_ACTOR_QUARANTINED_ACCESSQuarantined actor attempting access (system)
SYSTEM_ACTOR_TERMINATED_ACCESSTerminated actor attempting access (system)
SYSTEM_INVALID_KEYInvalid detection key used
SYSTEM_ORG_SUSPENDEDSuspended org attempting access

Infrastructure / Shadow AI​

Default category: UNREGISTERED_SHADOW_AI · Surfaced via Discovery

TypeData classificationWhat it catches
INFRASTRUCTURE_SHADOW_AI_BOTNONEUnapproved AI bot/plugin in collaboration apps
INFRASTRUCTURE_AI_BOT_EXCESSIVE_ACCESSNONEAI bot with excessive channel/scope access
INFRASTRUCTURE_HARDCODED_AI_KEYCREDENTIALSHardcoded AI API key in source code
INFRASTRUCTURE_PUBLIC_AI_REPONONEPublic repository containing AI code
INFRASTRUCTURE_AI_KEY_IN_HISTORYCREDENTIALSAI API key found in git history
INFRASTRUCTURE_MCP_NO_AUTHNONEMCP server without authentication
INFRASTRUCTURE_AGENT_UNRESTRICTED_EXECUTIONNONEAI agent configured with unrestricted execution

Infrastructure Misconfiguration​

Default category: INFRASTRUCTURE_MISCONFIGURATION · Surfaced via Discovery

TypeWhat it catches
INFRASTRUCTURE_MCP_PUBLIC_ENDPOINTMCP server exposed on a public endpoint

Document Conformance (Regulated Output)​

Schema-driven criterion classes. Per-schema specifics travel in Detection.runtimeAttributes (schema_id, field_path, rule_id, etc.).

TypeDefault categoryWhat it catches
DOCUMENT_REQUIRED_FIELD_MISSINGREGULATED_ARTIFACT_NONCONFORMANCERequired field missing from a regulated document
DOCUMENT_FIELD_FORMAT_VIOLATIONREGULATED_ARTIFACT_NONCONFORMANCEField value does not match schema format/enum
DOCUMENT_CROSS_FIELD_INCONSISTENCYREGULATED_ARTIFACT_NONCONFORMANCETwo fields contradict each other
DOCUMENT_UNSUPPORTED_ASSERTIONREGULATED_ARTIFACT_NONCONFORMANCEAssertion lacks required evidence
DOCUMENT_SCOPE_OVERREACHREGULATED_ARTIFACT_NONCONFORMANCEReferences a value outside the bound catalog
DOCUMENT_UNATTESTED_FIELD_PROVENANCEREGULATED_OUTPUT_PROVENANCE_VIOLATIONField cannot be tied to an attested ActionRecord-backed source

Fallback​

TypeDefault categoryWhat it catches
CUSTOM_UNKNOWNOPERATIONAL_ANOMALYUnknown or unclassified detection

Severity Levels​

LevelMeaningExamples
LOWMinor finding, informationalEmail address detected, file read operation
MEDIUMNotable finding, may require attentionPhone number detected, database query, web fetch
HIGHSignificant risk, likely needs actionSSN detected, shell execution, prompt injection
CRITICALSevere risk, immediate action neededCredential exfiltration, jailbreak, attack chain, persistence mechanism, multi-agent coordinated exfil

Lifecycle Stages​

StageWhenWhat's scanned
INGRESSBefore the request is sent to the AI providerUser/agent prompts, input content
EGRESSAfter the AI provider respondsAI responses, tool call results, agent actions
DEPLOYMENTDuring infrastructure scanningCloud configs, model registrations, shadow AI
TRAININGDuring training data pipelinesTraining data, fine-tuning inputs

Most enforcement happens at INGRESS (block bad inputs) and EGRESS (catch sensitive data and risky actions in outputs). DEPLOYMENT and TRAINING are primarily for discovery and compliance scanning.

Data Classifications​

ClassificationDescriptionSensitive
PIIPersonally Identifiable Information — SSN, email, phone, name, DOB, drivers license, passport, credit cardYes
PHIProtected Health Information — medical records, insurance, prescriptions, diagnoses, treatmentsYes
FINANCIALFinancial Data — bank accounts, routing numbers, investment accounts, tax IDs, plus financial action typesYes
CREDENTIALSCredentials & Secrets — API keys, passwords, SSH keys, AWS keys, auth tokensYes
INTELLECTUAL_PROPERTYIntellectual Property — trademarks, copyrights, patents, trade secretsYes
NONENot data-classified — prompt injection, tool calls, infrastructure findings, document conformanceNo

Capability Surfaces and Boundary Surfaces​

Some detection types map to a capability surface (an agent action, like INITIATE_PAYMENT) or a boundary surface (a defensive boundary, like INJECTION_DEFENSE). These are used by the policy engine to express controls in terms of what the agent is doing rather than just what was detected. See Policy Templates for how risk-adaptive rules use them.

Next steps​