Discovery & Shadow AI
Automatically map your AI footprint — cloud agents, source code, MCP servers, collaboration apps, and direct browser-based AI usage — before you can govern it.
Where to find it in the app
Dashboard → Administration → AI Estate → DISCOVERY.
The DISCOVERY stage of the AI Estate page has three sub-tabs:
| Sub-tab | What's here |
|---|---|
| Attack Surface Monitors | Cloud, code, and network discovery channels. The Discovered Assets view inside this tab is the asset inventory — filter by status (PENDING_APPROVAL, APPROVED, BLOCKED, ACTIVE, etc.), category, or risk level. |
| Browser AI | Shadow AI detection — direct browser-based ChatGPT / Claude / Perplexity usage caught by the Rivaro browser extension. |
| Tool Integrations | MCP servers discovered in your environment — bridges to the MCP configuration UI in Tool Integrations. |
Click any discovered asset to open its detail panel with findings, source channels, and the approve / deny / promote actions. See Asset Management for the full approval workflow.
Overview
Discovery runs continuously across your infrastructure, finding AI assets you may not know exist. Every discovered asset enters an approval workflow before it can be used by governed agents. Shadow AI detection catches direct AI usage (ChatGPT, Claude, etc.) happening outside your proxy.
Discovery works through channels — configured integrations with your infrastructure. Each channel type uses a different collection method and targets a different part of your environment.
Discovery channels
Configure discovery channels under Administration → AI Estate → DISCOVERY → Attack Surface Monitors → Channels. Rivaro supports the following channel types:
| Channel Type | Display Name | Mode | What it finds |
|---|---|---|---|
CLOUD_AI_SERVICES | Cloud Accounts | Scheduled | AWS, GCP, and Azure accounts — Bedrock Agents, AgentCore, SageMaker, Vertex AI, Azure ML endpoints, models, and AI-specific risks |
COLLABORATION_PLATFORM | Collaboration Apps | Scheduled | Slack, Teams, Google Workspace — unauthorized AI bots, plugins, and integrations |
SOURCE_CODE | Code Repositories | Scheduled | GitHub, GitLab, Bitbucket — AI dependencies, hardcoded API keys, agent code |
NETWORK_ENDPOINT | Network Scan | Agent callback | Running MCP servers, AI agent runtimes, and AI/MCP API endpoints on internal networks — requires a deployed network scanner agent |
AGENT_DATA | Scanner Agent | Agent callback | Pre-collected findings from a deployed scanner agent |
OPENCLAW_CONFIG | OpenClaw Config | Agent push | Agents, models, channels, and tools imported from local OpenClaw configuration |
MANUAL_ENTRY | Manual Entry | Manual | Admin-created assets — auto-approved on creation |
Channel configuration
Each channel has these common fields:
| Field | Description |
|---|---|
name | Display name for this channel |
channelType | One of the types above |
active | Whether the channel runs on its schedule |
pollingIntervalSeconds | How often to run (scheduled channels) |
configuration | Channel-specific settings (non-sensitive) |
lastRunAt | Timestamp of most recent scan |
lastRunStatus | SUCCESS, FAILED, or RUNNING |
lastRunAssetCount | Assets found in last run |
lastRunRiskCount | Risk findings in last run |
Sensitive credentials (API keys, tokens, secrets) are stored separately in an encrypted credential store — never in the channel configuration JSON.
Network scanner agent
For NETWORK_ENDPOINT channels, the channel configuration page lets you generate a downloadable Python agent. The agent is bound to a per-channel detection key, scans your internal network, and posts results back to Rivaro. Deploy it anywhere with network access to your internal AI infrastructure.
What Gets Discovered
Each discovered asset is classified by type and category:
| Category | Examples |
|---|---|
| AI_SERVICE | OpenAI, Anthropic, Vertex AI endpoints in use |
| AI_AGENT | Bedrock Agents, AgentCore agents, LangChain/CrewAI/AutoGen runtimes |
| AI_MODEL | Deployed models, fine-tuned versions, model registries |
| MCP_SERVER | Discovered MCP servers (gateway and tool endpoints) |
| DATA_STORAGE | Vector databases, embedding stores, training data repositories |
| ML_PIPELINE | Training pipelines, fine-tuning jobs, MLflow experiments |
| SOURCE_CODE | Repositories with AI dependencies or hardcoded keys |
| IDENTITY_ACCESS | Service accounts and roles with AI service permissions |
| COLLABORATION_BOT | AI bots and plugins in collaboration apps |
Asset risk findings
Each discovered asset can have associated findings — specific security or compliance issues detected during scanning:
| Finding field | Description |
|---|---|
detectionType | e.g. CREDENTIAL_EXPOSURE, MISCONFIGURATION, INFRASTRUCTURE_MCP_PUBLIC_ENDPOINT |
severity | CRITICAL, HIGH, MEDIUM, LOW |
status | ACTIVE, RESOLVED, IGNORED |
description | Human-readable description of the finding |
detectedContent | What was found (masked in UI) |
status | OPEN → IN_PROGRESS → RESOLVED |
Asset Approval Workflow
Rivaro defaults to zero-trust / default-deny: every new asset starts as PENDING_APPROVAL. No agent can use an unapproved asset.
Approval lifecycle
| Status | Meaning |
|---|---|
PENDING_APPROVAL | Discovered, awaiting security team review |
APPROVED | Reviewed and explicitly approved for use |
BLOCKED | Reviewed and denied — agents cannot access |
ACTIVE | Approved and currently in use by governed agents |
PROMOTED | Graduated to a governed entity (agent, data source, model) |
REMOVED | Asset no longer detected in environment |
ARCHIVED | Deprecated, kept for audit history |
The approval request includes a riskScore (0–100) calculated from the asset's findings. Reviewers can add notes before approving or denying.
Promoting an asset
Approved assets can be promoted — graduated into a fully governed entity with an AppContext, detection key, and full enforcement. This is how shadow infrastructure becomes official, monitored infrastructure.
| Promoted entity type | What it becomes |
|---|---|
| AGENT | A registered agent identity with trust score tracking |
| DATA_SOURCE | A governed data source with access controls |
| MODEL | An approved model with allowed-model list enforcement |
| INTEGRATION | A governed integration with policy enforcement |
| SERVICE | An approved AI service endpoint |
Multi-Source Correlation
The same asset may be discovered by multiple channels. Rivaro deduplicates using an externalId fingerprint — the same fingerprint from two channels links to one asset, with confidence increasing with each additional source.
| Observation type | Confidence | How it's detected |
|---|---|---|
| DISCOVERED | SUSPECTED → INFERRED | Found by a scanner/channel scan |
| RUNTIME_USAGE | CONFIRMED | Seen in live agent traffic through the proxy |
| CODE_REFERENCE | INFERRED | Found in source code as an import or API call |
| IAM_POLICY | INFERRED | Service account has permission to access it |
Shadow AI Detection
Shadow AI is direct use of AI services (ChatGPT, Claude, Perplexity, etc.) that bypasses your proxy — typically via a browser. The Rivaro Shadow AI browser extension monitors this activity and applies your policies in real time.
How it works
- Install the Chrome extension and configure it with your organization's detection key (managed under Administration → API Credentials).
- The extension monitors supported AI domains:
chatgpt.com,claude.ai,bard.google.com,bing.com/chat,poe.com,perplexity.ai, and more. - When a user types a prompt and submits it, the extension captures the content and sends it to Rivaro's detection engine.
- Rivaro runs the same detection pipeline as the proxy — PII, PHI, credentials, prompt injection, etc.
- The response action is applied directly in the browser.
Browser AI activity is visible at Administration → AI Estate → DISCOVERY → Browser AI.
Shadow AI policy actions
| Action | What the user sees |
|---|---|
| BLOCK | Modal appears, submission is prevented |
| REDACT | Modal shows sanitized version; user can copy and resubmit |
| LOG | Submission proceeds, violation is logged in the dashboard |
| ALLOW | No action, submission proceeds normally |
Shadow AI analytics
The Browser AI sub-tab tracks:
- Session trends — daily session counts and week-over-week change
- Violations by severity — CRITICAL / HIGH / MEDIUM / LOW breakdown
- Compliance rate — percentage of sessions with no violations
- Risk users — top users by risk score and violation count
- Cost exposure — estimated API cost of shadow usage, productivity hours
- Compliance by framework — HIPAA, GDPR, and other framework-level metrics
Zero Trust inventory
Shadow AI detection surfaces an unverified asset inventory including:
- Agent runtimes — LangChain, AutoGen, CrewAI instances running without governance
- MCP servers — unauthenticated or public MCP endpoints
- AI bots — Slack/Teams bots with excessive AI access
- Public endpoints — ML infrastructure exposed to the internet
Next steps
- Asset Management — Manage the approved asset inventory
- Agent Management — Register and govern agents
- Compliance Reporting — Use discovery data in compliance reports