Skip to main content

Training-Data Connectors

Scan the data your models train on — for PII, PHI, credentials, regulated content, and license-incompatible material — before it ever reaches the training pipeline. This is the TRAINING lifecycle stage of the Rivaro detection engine.

LLM gateway covers what flows through the agent at runtime. Training-data connectors cover what the model learned from in the first place.

Where to find it in the app​

Dashboard → Administration → AI Estate → TRAINING DATA → Connectors.

(Training-data governance is enabled per organization — your Rivaro contact can turn it on for you.)

The TRAINING DATA stage of the AI Estate page is where you add, configure, and monitor connectors. From this page:

  • Click Add connector and pick a type (S3, GCS, Azure Blob, Database) to open the configuration form.
  • Each connector row shows its last poll time, error counts, and finding totals.
  • Click a connector to see its findings stream and edit its configuration.
  • Test connection on the connector page validates credentials and lists a sample of accessible files / rows without writing any findings.

To apply policy rules to what the connector finds:

Dashboard → Policies & Authority → TRAINING DATA stage → pick the connector → Data Privacy tab.

This is where you write TRAINING-stage rules that flag or block specific data classes from being ingested.

What gets scanned​

Each connector polls a data source on a schedule, samples files or rows, and runs the configured detectors against the content. Detections are tagged with lifecycle: TRAINING and feed:

  • Discovery findings under the connector as a source
  • Compliance evidence for ISO 42001 Clause 8.6 (Data for AI Systems), HIPAA training-data audits, GDPR data-quality requirements
  • Policy enforcement — TRAINING-stage policy rules can flag or block ingestion of specific data classes

Supported connectors​

ConnectorSourcePolling pattern
S3 Training DataAmazon S3 bucketSample N files per poll, scan first M lines
GCS Training DataGoogle Cloud Storage bucketSame sampling pattern
Azure Blob Training DataAzure Blob Storage containerSame sampling pattern
Database Training DataAny JDBC-reachable database (MySQL, Postgres, Snowflake, etc.)Configurable SQL sample query

All four are first-party connectors built into Rivaro. If you need a source that isn't listed, talk to your Rivaro contact.

How polling works​

For each connector instance, Rivaro:

  1. Authenticates using the credentials you configured (IAM, service account, DB user, etc.).
  2. Samples up to a configurable number of files (default 50) or rows from the source.
  3. For each file/row up to a configurable size cap (default 10 MB) — larger files are skipped with a warning.
  4. Reads the first sample of textual content (default 100 lines for files, full row for DB).
  5. Runs the connector's enabled detectors against the content at TRAINING lifecycle.
  6. Persists detections as findings with the connector ID as the source.
  7. On the next poll, picks up where it left off (offset/cursor tracked per connector instance).

This is sampling, not full scanning, by design — you can scan a multi-TB bucket without burning a week of compute. If you need full coverage, lower the file-size limit, raise the per-poll cap, and run more frequently — or run an explicit Test connection scan against a specific path.

The polling interval, file caps, and sample sizes are all configurable on the connector's detail page.

Adding an S3 connector​

Click Add connector, pick S3 Training Data, and fill in:

FieldDescription
Connector nameDisplay name (e.g. "ML Training Bucket")
Bucket nameThe S3 bucket containing training data
RegionAWS region (us-east-1, us-west-2, eu-west-1, etc.)
Access key ID / Secret access keyProgrammatic credentials with s3:GetObject + s3:ListBucket on the bucket. Stored encrypted.
Prefix (optional)Restrict scanning to a key prefix (e.g. datasets/2026/)
Enabled detectorsPick which detector families to run (PII, PHI, CREDENTIALS, PCI_DATA, etc.)
Polling intervalHow often to sample (default 1 hour)

After save, click Test connection to verify access. Once that passes, the connector starts polling automatically.

For GCS and Azure Blob the fields are analogous — bucket / container name, region or storage account, service account JSON or shared key.

Adding a database connector​

Click Add connector, pick Database Training Data, and fill in:

FieldDescription
JDBC URLE.g. jdbc:postgresql://db.example.com:5432/customers
Username / PasswordRead-only credentials
Sample queryA bounded SQL query returning the columns to scan (e.g. SELECT comment_text FROM support_tickets WHERE created_at > NOW() - INTERVAL '7 days' LIMIT 100)
Enabled detectorsDetector families to run
Polling intervalHow often to re-run the sample query

The sample query should return a representative sample, not the entire table. Rivaro is scanning, not ingesting — keep the query bounded.

Policy on TRAINING-stage detections​

TRAINING-stage detections feed into the policy engine like any other detection, but with lifecycle: TRAINING. To apply a TRAINING-stage rule:

  1. Open Policies & Authority → TRAINING DATA stage.
  2. Pick the connector.
  3. Switch to the Data Privacy tab.
  4. Click Add rule, set the detection type (e.g. PII_SSN), and choose an action (BLOCK, LOG, REDACT, etc.).

The action — BLOCK here — flags the file/row in the connector's finding stream so your data pipeline can quarantine it before ingestion. See Policy Templates for the full rule structure.

Push-mode webhook ingestion​

For sources you'd rather push from than have us poll, every connector has a webhook receiver. The connector's detail page shows the webhook URL and token under Webhook ingestion. POST any JSON payload to that URL; Rivaro extracts the textual content via the connector's mapping and runs the detectors. This is the integration pattern for Slack, GitHub, and Jira — they post events, Rivaro scans them as training-stage content.

Best practices​

  • Start with sampling, not full coverage. A few hundred files a day catches almost all systematic issues. Save full scans for compliance audits.
  • Use prefix scoping aggressively. If your bucket has 50 TB of logs, scoping to datasets/ saves compute and noise.
  • Set enabled detectors per dataset, not globally. A medical-records bucket needs PHI; a marketing-leads bucket needs PII; a logs bucket needs credentials. Don't make every connector run every detector.
  • Test before going live. Always Test connection before relying on a connector in production.

Next steps​