Training-Data Connectors
Scan the data your models train on — for PII, PHI, credentials, regulated content, and license-incompatible material — before it ever reaches the training pipeline. This is the TRAINING lifecycle stage of the Rivaro detection engine.
LLM gateway covers what flows through the agent at runtime. Training-data connectors cover what the model learned from in the first place.
Where to find it in the app
Dashboard → Administration → AI Estate → TRAINING DATA → Connectors.
(Training-data governance is enabled per organization — your Rivaro contact can turn it on for you.)
The TRAINING DATA stage of the AI Estate page is where you add, configure, and monitor connectors. From this page:
- Click Add connector and pick a type (S3, GCS, Azure Blob, Database) to open the configuration form.
- Each connector row shows its last poll time, error counts, and finding totals.
- Click a connector to see its findings stream and edit its configuration.
- Test connection on the connector page validates credentials and lists a sample of accessible files / rows without writing any findings.
To apply policy rules to what the connector finds:
Dashboard → Policies & Authority → TRAINING DATA stage → pick the connector → Data Privacy tab.
This is where you write TRAINING-stage rules that flag or block specific data classes from being ingested.
What gets scanned
Each connector polls a data source on a schedule, samples files or rows, and runs the configured detectors against the content. Detections are tagged with lifecycle: TRAINING and feed:
- Discovery findings under the connector as a source
- Compliance evidence for ISO 42001 Clause 8.6 (Data for AI Systems), HIPAA training-data audits, GDPR data-quality requirements
- Policy enforcement — TRAINING-stage policy rules can flag or block ingestion of specific data classes
Supported connectors
| Connector | Source | Polling pattern |
|---|---|---|
| S3 Training Data | Amazon S3 bucket | Sample N files per poll, scan first M lines |
| GCS Training Data | Google Cloud Storage bucket | Same sampling pattern |
| Azure Blob Training Data | Azure Blob Storage container | Same sampling pattern |
| Database Training Data | Any JDBC-reachable database (MySQL, Postgres, Snowflake, etc.) | Configurable SQL sample query |
All four are first-party connectors built into Rivaro. If you need a source that isn't listed, talk to your Rivaro contact.
How polling works
For each connector instance, Rivaro:
- Authenticates using the credentials you configured (IAM, service account, DB user, etc.).
- Samples up to a configurable number of files (default 50) or rows from the source.
- For each file/row up to a configurable size cap (default 10 MB) — larger files are skipped with a warning.
- Reads the first sample of textual content (default 100 lines for files, full row for DB).
- Runs the connector's enabled detectors against the content at TRAINING lifecycle.
- Persists detections as findings with the connector ID as the source.
- On the next poll, picks up where it left off (offset/cursor tracked per connector instance).
This is sampling, not full scanning, by design — you can scan a multi-TB bucket without burning a week of compute. If you need full coverage, lower the file-size limit, raise the per-poll cap, and run more frequently — or run an explicit Test connection scan against a specific path.
The polling interval, file caps, and sample sizes are all configurable on the connector's detail page.
Adding an S3 connector
Click Add connector, pick S3 Training Data, and fill in:
| Field | Description |
|---|---|
| Connector name | Display name (e.g. "ML Training Bucket") |
| Bucket name | The S3 bucket containing training data |
| Region | AWS region (us-east-1, us-west-2, eu-west-1, etc.) |
| Access key ID / Secret access key | Programmatic credentials with s3:GetObject + s3:ListBucket on the bucket. Stored encrypted. |
| Prefix (optional) | Restrict scanning to a key prefix (e.g. datasets/2026/) |
| Enabled detectors | Pick which detector families to run (PII, PHI, CREDENTIALS, PCI_DATA, etc.) |
| Polling interval | How often to sample (default 1 hour) |
After save, click Test connection to verify access. Once that passes, the connector starts polling automatically.
For GCS and Azure Blob the fields are analogous — bucket / container name, region or storage account, service account JSON or shared key.
Adding a database connector
Click Add connector, pick Database Training Data, and fill in:
| Field | Description |
|---|---|
| JDBC URL | E.g. jdbc:postgresql://db.example.com:5432/customers |
| Username / Password | Read-only credentials |
| Sample query | A bounded SQL query returning the columns to scan (e.g. SELECT comment_text FROM support_tickets WHERE created_at > NOW() - INTERVAL '7 days' LIMIT 100) |
| Enabled detectors | Detector families to run |
| Polling interval | How often to re-run the sample query |
The sample query should return a representative sample, not the entire table. Rivaro is scanning, not ingesting — keep the query bounded.
Policy on TRAINING-stage detections
TRAINING-stage detections feed into the policy engine like any other detection, but with lifecycle: TRAINING. To apply a TRAINING-stage rule:
- Open Policies & Authority → TRAINING DATA stage.
- Pick the connector.
- Switch to the Data Privacy tab.
- Click Add rule, set the detection type (e.g.
PII_SSN), and choose an action (BLOCK,LOG,REDACT, etc.).
The action — BLOCK here — flags the file/row in the connector's finding stream so your data pipeline can quarantine it before ingestion. See Policy Templates for the full rule structure.
Push-mode webhook ingestion
For sources you'd rather push from than have us poll, every connector has a webhook receiver. The connector's detail page shows the webhook URL and token under Webhook ingestion. POST any JSON payload to that URL; Rivaro extracts the textual content via the connector's mapping and runs the detectors. This is the integration pattern for Slack, GitHub, and Jira — they post events, Rivaro scans them as training-stage content.
Best practices
- Start with sampling, not full coverage. A few hundred files a day catches almost all systematic issues. Save full scans for compliance audits.
- Use prefix scoping aggressively. If your bucket has 50 TB of logs, scoping to
datasets/saves compute and noise. - Set enabled detectors per dataset, not globally. A medical-records bucket needs PHI; a marketing-leads bucket needs PII; a logs bucket needs credentials. Don't make every connector run every detector.
- Test before going live. Always Test connection before relying on a connector in production.
Next steps
- Policy Templates — TRAINING-stage rules
- Discovery & Shadow AI — Connector findings appear in discovery
- Compliance Reporting — TRAINING evidence in framework reports (ISO 42001 Clause 8.6)
- Understanding Detections — Which detectors run on training data