Skip to main content

AWS SageMaker Integration

Route AWS SageMaker endpoint traffic through Rivaro for runtime enforcement. Supports custom-deployed models, JumpStart-hosted models, and any inference endpoint exposed under runtime.sagemaker.{region}.amazonaws.com.

How Rivaro proxies SageMaker​

SageMaker's runtime API uses path /endpoints/{endpointName}/invocations. Rivaro registers this pattern via ProxyMappingRegistry, so requests to:

https://your-org.rivaro.ai/endpoints/{endpointName}/invocations

are proxied to:

https://runtime.sagemaker.{region}.amazonaws.com/endpoints/{endpointName}/invocations

with the AWS SigV4 headers forwarded unchanged.

Authentication​

Like Bedrock, SageMaker uses AWS SigV4 for authentication — there is no API key. Your AWS SDK signs the request, Rivaro proxies the signed request to SageMaker, AWS validates the signature against the endpoint's resource policy.

The Rivaro detection key is provided via the X-Detection-Key header (separate from AWS authentication).

SDK Configuration​

Python (boto3)​

import boto3
from botocore.config import Config

sagemaker = boto3.client(
'sagemaker-runtime',
region_name='us-east-1',
endpoint_url='https://your-org.rivaro.ai',
config=Config(inject_host_prefix=False)
)

def add_detection_key(params, **kwargs):
params['headers']['X-Detection-Key'] = 'detect_live_your_key_here'

sagemaker.meta.events.register('before-sign.sagemaker-runtime.*', add_detection_key)

response = sagemaker.invoke_endpoint(
EndpointName='my-llama-endpoint',
Body=b'{"inputs": "Hello"}',
ContentType='application/json'
)

Streaming (invoke_endpoint_with_response_stream)​

Streaming SageMaker endpoints are supported. The response stream is forwarded unbuffered for token streaming, with enforcement mode applying egress detection on close.

AppContext Configuration​

When you create an AppContext for SageMaker traffic, the configuration map supports:

KeyDescription
sagemakerRegionAWS region of the SageMaker endpoint (e.g. us-east-1)
sagemakerEndpointNameThe endpoint name (my-llama-endpoint)
responseFormattext-generation, hf-tgi, vllm, or custom — determines how Rivaro parses the response for egress detection

For responseFormat: "custom", Rivaro still applies pattern + ML detection against the raw response bytes but can't apply format-aware policies (like "redact assistant turn only").

Discovery: auto-classification​

Assets in the discovery pipeline are auto-matched to the SageMaker adapter when they have:

  • Cloud provider aws + service sagemaker in their metadata
  • An endpoint URL matching runtime.sagemaker.*.amazonaws.com
  • An ARN containing sagemaker

Once matched, the asset can be promoted to a governed AppContext in one click — see Asset Management.

Supported endpoints​

EndpointMethodDescription
/endpoints/{endpointName}/invocationsPOSTStandard SageMaker invocation
/endpoints/{endpointName}/invocations-with-response-streamPOSTStreaming invocation

Required headers​

HeaderRequiredDescription
X-Detection-KeyYesYour Rivaro detection key
AuthorizationYesAWS SigV4 signature (set by AWS SDK)
X-Amz-DateYesRequest timestamp (set by AWS SDK)
X-Amz-Security-TokenConditionalIf using temporary credentials
Content-TypeYesTypically application/json

Next steps​