Drift Monitoring
Detect when your agents' prompts or outputs change in ways you didn't intend. Two complementary streams: prompt drift (system prompt edits, often unintentional) and content drift (output behavior changes against a baseline).
Where to find it in the app
Prompt drift is integrated into the prompt library:
Dashboard → Instruction Intelligence → click a prompt → Behavior tab.
The Behavior tab inside the prompt drawer shows the drift events for that prompt: every hash change, with a side-by-side diff, severity badge, and the timeline of when it happened and how often it ran. The prompt list itself shows a severity badge on rows where the latest change was MAJOR or SUSPICIOUS, so you can spot what needs review at a glance.
The Instruction Intelligence tab also has filter groups for Confidence and Behavior drift so you can scope the prompt list to drifting prompts only.
Content drift alerts are delivered through notification channels and the API. The content-drift section below documents the endpoints and behavior.
Why drift matters
Agent quality and safety are functions of the agent's full configuration: system prompt, model, tools, policies. A single character changed in the system prompt — by an SRE rushing a deploy, by a hostile dependency that mutated the prompt template, by an automation gone wrong — can completely shift the agent's behavior.
You won't notice from a smoke test. You'll notice from a customer complaint two weeks later.
Drift monitoring is the safety net.
Prompt drift
Rivaro fingerprints every distinct system prompt your agents run with. When an agent's prompt changes — same agent, different prompt content — that change is recorded with a timestamp, severity assessment, and side-by-side diff. All of this is visible on the Behavior tab of the prompt drawer in Instruction Intelligence.
What gets tracked
| Field | Description |
|---|---|
| Prompt hash | SHA-256 of the normalized prompt content |
| Agent | Which agent ran this prompt |
| First seen / last seen | Observation window for this hash |
| Usage count | How many requests used this prompt |
| Diff | When a new hash appears for an agent that already had one, the line-level diff against the previous hash |
| Severity | Heuristic-classified severity of the change (TRIVIAL / MINOR / MAJOR / SUSPICIOUS) |
Severity classification
The classifier looks at:
- Whether the diff added, removed, or modified policy-relevant instructions ("never reveal", "always require", role declarations)
- Whether tool descriptions or authorities were changed
- Whether new instructions about data handling appeared
- Whether the change was a tiny whitespace / comment edit vs a structural rewrite
MAJOR and SUSPICIOUS changes warrant a human review; TRIVIAL and MINOR are noted in the history but don't generally need attention.
Filtering for what needs attention
On the Instruction Intelligence list, apply the Behavior drift filter to scope to prompts with recent drift events. Sort by severity to bring SUSPICIOUS and MAJOR changes to the top.
Pre-deployment drift check
When you want to verify whether a candidate prompt would trigger drift severity before deploying — for example, in a CI step — Rivaro can produce the same severity classification on demand against a candidate prompt for an agent. This is available via the Test prompt action on the agent's Instruction Intelligence page, or programmatically via the prompt-drift API used for CI gating.
Content drift
Where prompt drift catches changes to the agent's instructions, content drift catches changes to the agent's outputs against a baseline — even when the prompt hasn't changed at all.
A typical content drift scenario:
- An agent has been answering customer questions about refund policy.
- The underlying LLM is silently swapped to a new version by the provider.
- The agent's responses start including hallucinated policy clauses.
- The prompt hasn't changed. The tools haven't changed. But the outputs have drifted.
Content drift alerts surface this.
How it works
Rivaro hashes and clusters representative outputs per request signature (an AppContext + a topic cluster). When a new cluster appears for a signature that previously had a stable output distribution, it's recorded as a drift alert with:
- The request signature that drifted
- The baseline output cluster
- The new output cluster
- Confidence score (how distinct the new cluster is from the baseline)
- Sample side-by-side outputs
Where content drift surfaces
Content drift shows up in two places:
- Notification channels — wire a channel to the content-drift event stream and you'll get Slack / PagerDuty / Webhook alerts on every drift event. See Notifications & Integrations.
- Incidents — high-confidence content drift events can auto-create incidents, which then appear on the Incidents & Detections tab.
Routing drift to incidents
MAJOR or SUSPICIOUS prompt drift events and high-confidence content drift events can be configured to auto-create incidents for investigator pickup.
Drift auto-escalation runs on Rivaro's default thresholds. Your Rivaro contact can configure custom thresholds for your organization.
What to do when drift fires
| Drift type | Suggested response |
|---|---|
MINOR prompt drift | Note in change log; no action required |
MAJOR prompt drift | Open a ticket — verify the change was intentional; if not, roll back |
SUSPICIOUS prompt drift | Treat as a potential supply-chain attack on your prompt template — investigate immediately |
| Low-confidence content drift | Note in dashboard; review during weekly check-in |
| High-confidence content drift | Investigate within hours — possibly a silent model swap, a tool change, or a regression in fine-tuning |
Integration with policy
Both forms of drift can be combined with policy:
- Use Prompt hash in policy scoping to apply a tighter rule to a specific prompt version while it's under review.
- Configure
RISK_ADAPTIVErules in policy templates that escalate the action when content-drift confidence is high.
Next steps
- Sessions — Drift events appear inline on session timelines
- Notifications & Integrations — Route drift alerts to your incident-response stack
- Incident Management — Auto-escalate high-severity drift
- Policy Scoping — Pin policies to specific prompt versions