All three routes share the same AWS credential chain and region resolution — no separate configuration is needed. Requests to the Mantle endpoint are authenticated with
AWS_BEARER_TOKEN_BEDROCK when set, or SigV4-signed via the standard boto3 credential chain otherwise.
Prerequisites
- AWS credentials — any source supported by the boto3 credential chain:
- IAM instance role (EC2, ECS, Lambda — zero config)
AWS_ACCESS_KEY_ID+AWS_SECRET_ACCESS_KEYenvironment variablesAWS_PROFILEfor SSO or named profilesaws configurefor local development
- boto3 — install with
cd ~/.mibyan/mibyan-agent && python -c "import pm; pm.sync_venv(['bedrock'], explicit=True)" - IAM permissions — at minimum:
bedrock:InvokeModelandbedrock:InvokeModelWithResponseStream(for inference)bedrock:ListFoundationModelsandbedrock:ListInferenceProfiles(for model discovery)bedrock:GetInferenceProfile(only ifmodel.defaultis an application inference profile ARN — used to size the context window from the wrapped model)
Quick Start
Configuration
After runningmibyan model, your ~/.mibyan/config.yaml will contain:
Region
Set the AWS region in any of these ways (highest priority first):bedrock.regioninconfig.yamlAWS_REGIONenvironment variableAWS_DEFAULT_REGIONenvironment variable- Default:
us-east-1
Guardrails
To apply Amazon Bedrock Guardrails to all model invocations:guardrailConfig) and on the Claude route (InvokeModel headers via the Anthropic Bedrock SDK, so prompt caching and thinking are kept). A blocked request surfaces as a content-filter refusal rather than as model text. stream_processing_mode only applies to Converse.
AWS does not apply Guardrails to the Mantle Responses endpoint used by openai.gpt-5.x models (AWS docs); use a Converse-served model when a guardrail is required.
Model Discovery
Mibyan auto-discovers available models via the Bedrock control plane. You can customize discovery:Prompt caching (cachePoint)
Mibyan automatically applies prompt caching on the Bedrock Converse API path by insertingcachePoint markers after the system prompt, tool definitions, and the latest message. Because sending a cachePoint block to a model that doesn’t support it raises a ValidationException, markers are only added for models on a known-good allowlist (Anthropic Claude and Amazon Nova model IDs); unknown models default to no cache markers. Claude models normally use the AnthropicBedrock SDK path, which has its own prompt caching — the Converse cachePoint path covers Nova and the bearer-token Claude fallback. No configuration needed; cache reads/writes show up in usage accounting.
Context-window probing
For models whose context window isn’t in Mibyan’ static table, Mibyan can probe the real limit by sending oversized requests at fixed tiers (~1.3M and ~2.2M tokens) and parsing themaximum reported in Bedrock’s length-validation error. Probed values feed the same metadata cache as the static table; stale cached entries that under-report a model’s window (e.g. entries seeded before a model’s 1M window went GA) are dropped automatically in favor of the larger known value.
Application inference profiles. An ARN such as arn:aws:bedrock:us-west-2:123456789012:application-inference-profile/abcdef123456 names no model, so neither the probe nor the static table can size it. Mibyan calls bedrock:GetInferenceProfile in the ARN’s region and sizes the window from the model the profile wraps (1M for a profile wrapping Claude Sonnet 4.6). Without that permission the 128,000-token default applies and a WARNING names the profile; set model.context_length explicitly to override either way.
Available Models
Bedrock models use inference profile IDs for on-demand invocation. Themibyan model picker shows these automatically, with recommended models at the top:
Cross-Region InferenceModels prefixed with
us. use cross-region inference profiles, which provide better capacity and automatic failover across AWS regions. Models prefixed with global. route across all available regions worldwide. OpenAI openai.* model IDs are served by Bedrock Mantle in the configured region and don’t use inference-profile prefixes.Switching Models Mid-Session
Use the/model command during a conversation:
Diagnostics
- Whether AWS credentials are available (env vars, IAM role, SSO)
- Whether
boto3is installed - Whether the Bedrock API is reachable (ListFoundationModels)
- Number of available models in your region
Gateway (Messaging Platforms)
Bedrock works with all Mibyan gateway platforms (Telegram, Discord, Slack, Feishu, etc.). Configure Bedrock as your provider, then start the gateway normally:config.yaml and uses the same Bedrock provider configuration.
Troubleshooting
”No API key found” / “No AWS credentials”
Mibyan checks for credentials in this order:AWS_BEARER_TOKEN_BEDROCKAWS_ACCESS_KEY_ID+AWS_SECRET_ACCESS_KEYAWS_PROFILE- EC2 instance metadata (IMDS)
- ECS container credentials
- Lambda execution role
aws configure or attach an IAM role to your compute instance.
”Invocation of model ID … with on-demand throughput isn’t supported”
Use an inference profile ID (prefixed withus. or global.) instead of the bare foundation model ID. For example:
- ❌
anthropic.claude-sonnet-4-6 - ✅
us.anthropic.claude-sonnet-4-6

