Skip to main content

Drift Monitoring

Detect when your agents' prompts or outputs change in ways you didn't intend. Two complementary streams: prompt drift (system prompt edits, often unintentional) and content drift (output behavior changes against a baseline).

Where to find it in the app​

Prompt drift is integrated into the prompt library:

Dashboard → Instruction Intelligence → click a prompt → Behavior tab.

The Behavior tab inside the prompt drawer shows the drift events for that prompt: every hash change, with a side-by-side diff, severity badge, and the timeline of when it happened and how often it ran. The prompt list itself shows a severity badge on rows where the latest change was MAJOR or SUSPICIOUS, so you can spot what needs review at a glance.

The Instruction Intelligence tab also has filter groups for Confidence and Behavior drift so you can scope the prompt list to drifting prompts only.

How content drift is delivered

Content drift alerts are delivered through notification channels and the API. The content-drift section below documents the endpoints and behavior.

Why drift matters​

Agent quality and safety are functions of the agent's full configuration: system prompt, model, tools, policies. A single character changed in the system prompt — by an SRE rushing a deploy, by a hostile dependency that mutated the prompt template, by an automation gone wrong — can completely shift the agent's behavior.

You won't notice from a smoke test. You'll notice from a customer complaint two weeks later.

Drift monitoring is the safety net.

Prompt drift​

Rivaro fingerprints every distinct system prompt your agents run with. When an agent's prompt changes — same agent, different prompt content — that change is recorded with a timestamp, severity assessment, and side-by-side diff. All of this is visible on the Behavior tab of the prompt drawer in Instruction Intelligence.

What gets tracked​

FieldDescription
Prompt hashSHA-256 of the normalized prompt content
AgentWhich agent ran this prompt
First seen / last seenObservation window for this hash
Usage countHow many requests used this prompt
DiffWhen a new hash appears for an agent that already had one, the line-level diff against the previous hash
SeverityHeuristic-classified severity of the change (TRIVIAL / MINOR / MAJOR / SUSPICIOUS)

Severity classification​

The classifier looks at:

  • Whether the diff added, removed, or modified policy-relevant instructions ("never reveal", "always require", role declarations)
  • Whether tool descriptions or authorities were changed
  • Whether new instructions about data handling appeared
  • Whether the change was a tiny whitespace / comment edit vs a structural rewrite

MAJOR and SUSPICIOUS changes warrant a human review; TRIVIAL and MINOR are noted in the history but don't generally need attention.

Filtering for what needs attention​

On the Instruction Intelligence list, apply the Behavior drift filter to scope to prompts with recent drift events. Sort by severity to bring SUSPICIOUS and MAJOR changes to the top.

Pre-deployment drift check​

When you want to verify whether a candidate prompt would trigger drift severity before deploying — for example, in a CI step — Rivaro can produce the same severity classification on demand against a candidate prompt for an agent. This is available via the Test prompt action on the agent's Instruction Intelligence page, or programmatically via the prompt-drift API used for CI gating.

Content drift​

Where prompt drift catches changes to the agent's instructions, content drift catches changes to the agent's outputs against a baseline — even when the prompt hasn't changed at all.

A typical content drift scenario:

  • An agent has been answering customer questions about refund policy.
  • The underlying LLM is silently swapped to a new version by the provider.
  • The agent's responses start including hallucinated policy clauses.
  • The prompt hasn't changed. The tools haven't changed. But the outputs have drifted.

Content drift alerts surface this.

How it works​

Rivaro hashes and clusters representative outputs per request signature (an AppContext + a topic cluster). When a new cluster appears for a signature that previously had a stable output distribution, it's recorded as a drift alert with:

  • The request signature that drifted
  • The baseline output cluster
  • The new output cluster
  • Confidence score (how distinct the new cluster is from the baseline)
  • Sample side-by-side outputs

Where content drift surfaces​

Content drift shows up in two places:

  1. Notification channels — wire a channel to the content-drift event stream and you'll get Slack / PagerDuty / Webhook alerts on every drift event. See Notifications & Integrations.
  2. Incidents — high-confidence content drift events can auto-create incidents, which then appear on the Incidents & Detections tab.

Routing drift to incidents​

MAJOR or SUSPICIOUS prompt drift events and high-confidence content drift events can be configured to auto-create incidents for investigator pickup.

Custom escalation thresholds

Drift auto-escalation runs on Rivaro's default thresholds. Your Rivaro contact can configure custom thresholds for your organization.

What to do when drift fires​

Drift typeSuggested response
MINOR prompt driftNote in change log; no action required
MAJOR prompt driftOpen a ticket — verify the change was intentional; if not, roll back
SUSPICIOUS prompt driftTreat as a potential supply-chain attack on your prompt template — investigate immediately
Low-confidence content driftNote in dashboard; review during weekly check-in
High-confidence content driftInvestigate within hours — possibly a silent model swap, a tool change, or a regression in fine-tuning

Integration with policy​

Both forms of drift can be combined with policy:

  • Use Prompt hash in policy scoping to apply a tighter rule to a specific prompt version while it's under review.
  • Configure RISK_ADAPTIVE rules in policy templates that escalate the action when content-drift confidence is high.

Next steps​