azure-foundry provider supports Microsoft Foundry (formerly Azure AI Foundry) and Azure OpenAI. A single Foundry resource can host models with two different wire formats:
- OpenAI-style —
POST /v1/chat/completionson endpoints likehttps://<resource>.openai.azure.com/openai/v1. Used for GPT-4.x, GPT-5.x, Llama, Mistral, and most open-weight models. - Anthropic-style —
POST /v1/messageson endpoints likehttps://<resource>.services.ai.azure.com/anthropic. Used when Microsoft Foundry serves Claude models via the Anthropic Messages API format.
Prerequisites
- A Microsoft Foundry or Azure OpenAI resource with at least one deployment
- The deployment’s endpoint URL
- Either an API key (from the Azure Portal under “Keys and Endpoint”) or the Azure AI User RBAC role on the Foundry resource if you plan to use Microsoft Entra ID (the keyless path Microsoft recommends). Some tenants may show the role as Foundry User during Microsoft’s rename rollout.
Quick Start
- Sniff the URL path — URLs ending in
/anthropicare recognised as Microsoft Foundry Claude routes. - Probe
GET <base>/models— if the endpoint returns an OpenAI-shaped model list, Mibyan switches tochat_completionsand prefills a picker with the returned deployment IDs. - Probe Anthropic Messages shape — fallback for endpoints that do not expose
/modelsbut do accept the Anthropic Messages format. - Fall back to manual entry — private/gated endpoints that reject every probe still work; you pick the API mode and type a deployment name by hand.
models.dev, provider metadata, and hardcoded family fallbacks) and stored in config.yaml so the model can size its own context window correctly.
Microsoft Entra ID (keyless, RBAC) — recommended
Microsoft recommends keyless authentication with Microsoft Entra ID for production Foundry workloads. Mibyan supports Entra ID for both API surfaces:- OpenAI-style (
api_mode: chat_completions/codex_responses) — GPT-4/5, Llama, Mistral, DeepSeek, etc. - Anthropic-style (
api_mode: anthropic_messages) — Claude models on Microsoft Foundry.
Azure AI User grants both surfaces; some tenants may display Foundry User) and Microsoft documents the same inference scope (https://ai.azure.com/.default) for both. Under the hood:
- OpenAI-style uses the OpenAI Python SDK’s native callable
api_key=contract — the SDK mints a fresh JWT per request automatically. - Anthropic-style uses an
httpx.Clientwith a request event hook installed byagent.azure_identity_adapter.build_bearer_http_client, because the Anthropic SDK does not accept callableauth_tokennatively. The hook rewritesAuthorization: Bearer <fresh-jwt>per outbound request. Same Microsoft RBAC, same Foundry scope — the SDK contract is the only difference.
Why use Entra ID?
- No long-lived API keys to rotate or revoke.
- RBAC-driven access — grant or remove
Azure AI Useron the Foundry resource, no config rewrite needed. - Access and audit logs are segmented by assignee instead of all callers sharing one static key.
- Single auth surface for Azure VMs, AKS pods, App Service, Functions, Container Apps, and Foundry Agent Service via managed identity.
- Workload identity and service-principal flows for CI/CD pipelines.
One-time setup (Azure side)
- In the Azure Portal, open your Foundry resource → Access control (IAM) → Add → Add role assignment.
- Pick the Azure AI User role (or Foundry User if your tenant has the renamed role).
- Assign it to:
- Your user account for local development with
az login. - A managed identity or workload identity for Azure-hosted compute (recommended for production).
- A Foundry Agent Service hosted agent’s agent identity when Mibyan runs inside a hosted agent.
- A service principal for CI/CD pipelines when workload identity is not available.
- Your user account for local development with
- Wait ~5 minutes for the role to propagate.
One-time setup (Mibyan side)
azure-identity is installed automatically on first use via Mibyan’ lazy-install path. To pre-install:
Configuration written to config.yaml
config.yaml:
scope— the OAuth resource scope. Defaults to Microsoft’s documented inference scope (https://ai.azure.com/.default). Override only if your resource was provisioned against a non-standard audience.
azure-identity directly from the standard AZURE_* environment variables — see the credential resolution order below. Set those in ~/.mibyan/.env or your deployment environment, exactly as Microsoft’s SDK reference describes.
No secrets land in ~/.mibyan/.env for Entra mode — azure-identity caches tokens in-process (and where available, in your OS keychain / ~/.IdentityService).
Credential resolution order
azure-identity’s DefaultAzureCredential walks this chain on each token request, stopping at the first credential that returns a token:
- Environment credential —
AZURE_TENANT_ID+AZURE_CLIENT_ID+AZURE_CLIENT_SECRET(orAZURE_CLIENT_CERTIFICATE_PATH/AZURE_FEDERATED_TOKEN_FILE). - Workload Identity —
AZURE_FEDERATED_TOKEN_FILE(AKS federated tokens / OIDC). - Managed Identity — IMDS endpoint (
169.254.169.254) for virtual machines;IDENTITY_ENDPOINTfor App Service / Functions / Container Apps. Foundry Agent Service hosted agents use the hosted agent’s agent identity. - Visual Studio Code — Azure account extension.
- Azure CLI —
az loginsession. - Azure Developer CLI —
azd auth login. - Azure PowerShell —
Connect-AzAccount. - Broker (Windows / WSL only) — Web Account Manager.
gateway.multiplex_profiles: true): every source in that chain resolves from the process — the launch profile’s AZURE_*, its az login session, the host’s managed identity. A served profile that sets no AZURE_* of its own is therefore refused instead of borrowing the launch identity (the same rule the Vertex adapter applies to Application Default Credentials). Give each profile its own AZURE_TENANT_ID + AZURE_CLIENT_ID + AZURE_CLIENT_SECRET (or AZURE_FEDERATED_TOKEN_FILE) in its .env; AZURE_CLIENT_ID alone opts that profile into the host’s user-assigned managed identity. Single-profile runs (mibyan, mibyan -p beta) keep the full chain.
Deployment patterns
Local development:- Enable system-assigned identity on the compute resource.
- Grant the identity
Azure AI User(orFoundry User) on the Foundry resource. - Set
model.auth_mode: entra_idin config.yaml — no env vars needed.
- Set
AZURE_CLIENT_IDto the user-assigned identity’s client ID soDefaultAzureCredentialpicks the right one.
- Create the hosted agent and grant that agent’s identity
Azure AI User(orFoundry User) on the Foundry resource. Mibyan usesManagedIdentityCredentialfrom inside the hosted agent; role assignment belongs on the agent identity, not just the parent project or your user.
- Annotate the pod’s service account with the workload identity client ID.
- The pod’s federated token file is auto-detected via
AZURE_FEDERATED_TOKEN_FILE. model.auth_mode: entra_idworks without further config changes.
- Set
AZURE_TENANT_ID,AZURE_CLIENT_ID,AZURE_CLIENT_SECRETin the runner env.
Sovereign clouds (Government, China)
ExportAZURE_AUTHORITY_HOST (e.g. https://login.microsoftonline.us for Azure Government, https://login.partner.microsoftonline.cn for Azure China). azure-identity reads it directly.
Health checks
mibyan doctor runs a 10 s probe against DefaultAzureCredential when model.auth_mode: entra_id, reporting which inner credential won (env vars present, managed identity endpoint reachable, etc.).
mibyan auth shows a structured status block:
provider: auto — session titles, context compression, smart approval) reuse the main session’s Entra token provider instead of re-authenticating; a --api-key string on the CLI still overrides it for one-off testing.
Limitations
- Anthropic-style endpoints use an httpx event hook. The Anthropic Python SDK does not accept a callable
auth_tokennatively (≤ 0.86.0). Mibyan installs a request event hook on a customhttpx.Clientthat mints a fresh JWT per outbound request and rewritesAuthorization: Bearer <jwt>. This is functionally equivalent to the OpenAI SDK’s nativeCallable[[], str]contract but adds one indirection layer. If the Anthropic SDK adds first-class callable-auth support in a future release, Mibyan will switch to it transparently. - Batch jobs and
multiprocessing.Pool. The Entra token provider is a closure that cannot be pickled across process boundaries.batch_runner.pyautomatically drops the callable from the worker config and lets each worker process rebuild its own provider fromconfig.yaml— no user action required, but each worker pays one chain walk at startup. - No bearer JWT persistence in
auth.json. Mibyan does not duplicateazure-identity’s internal token cache; cold starts walk the credential chain on first inference.
Configuration (written to config.yaml)
After running the wizard you’ll see something like this:
~/.mibyan/.env:
OpenAI-style endpoints (GPT, Llama, etc.)
Azure OpenAI’s v1 GA endpoint accepts the standardopenai Python client with minimal changes:
- GPT-5.x, codex, and o-series auto-route to the Responses API. Microsoft Foundry deploys GPT-5 / codex / o1 / o3 / o4 models as Responses-API-only — calling
/chat/completionsagainst them returns400 "The requested operation is unsupported.". Mibyan detects these model families by name and upgradesapi_modetocodex_responsestransparently, even whenconfig.yamlstill readsapi_mode: chat_completions. GPT-4, GPT-4o, Llama, Mistral, and other deployments stay on/chat/completions. api_mode: responsesis accepted as a spelling ofcodex_responses. The alias works onmodel.api_mode, onfallback_providersentries and on per-taskauxiliary.<task>.api_mode(e.g. anauxiliary.visionroute to a GPT-5.x deployment), and selects the same Responses adapter.max_completion_tokensis used automatically. Azure OpenAI (like direct OpenAI) requiresmax_completion_tokensfor gpt-4o, o-series, and gpt-5.x models. Mibyan sends the right parameter based on the endpoint.- Pre-v1 endpoints that require
api-version. If you have a legacy base URL likehttps://<resource>.openai.azure.com/openai?api-version=2025-04-01-preview, Mibyan extracts the query string and forwards it viadefault_queryon every request (the OpenAI SDK otherwise drops it when joining paths).
Anthropic-style endpoints (Claude via Microsoft Foundry)
For Claude deployments, use the Anthropic-style route:/v1is stripped from the base URL. The Anthropic SDK appends/v1/messagesto every request URL — Mibyan removes any trailing/v1before handing the URL to the SDK to avoid double-/v1paths.api-versionis sent viadefault_query, not appended to the URL. Azure Anthropic requires anapi-versionquery string. Baking it into the base URL produces malformed paths like/anthropic?api-version=.../v1/messagesand returns 404. Mibyan passesapi-version=2025-04-15via the Anthropic SDK’sdefault_queryinstead.- Bearer auth is used instead of
x-api-key. Azure’s Anthropic-compatible route requiresAuthorization: Bearer <key>rather than Anthropic’s nativex-api-keyheader. Mibyan detectsazure.comin the base URL and routes the API key through the SDK’sauth_tokenfield so the right header reaches the upstream. - 1M context window beta header is kept. Azure still gates the 1M-token Claude context (Opus 4.6/4.7, Sonnet 4.6) behind the
anthropic-beta: context-1m-2025-08-07header. Mibyan keeps that beta header on Azure paths (it’s stripped from native Anthropic OAuth requests because some subscriptions reject it, but Azure requires it). - OAuth token refresh is disabled. Azure deployments use static API keys. The
~/.claude/.credentials.jsonOAuth token refresh loop that applies to Anthropic Console is explicitly skipped for Azure endpoints to prevent the Claude Code OAuth token from overwriting your Azure key mid-session. mibyan doctorprobes the same route. The/anthropicroute has noGET /models, so the connectivity check sends a one-tokenPOST /v1/messageswith the same Bearer auth andapi-versionquery the runtime uses; a 200 (or a 400 from the Messages API) reports the endpoint as healthy, 401/403 as an auth problem.
Alternative: provider: anthropic + Azure base URL
If you already have provider: anthropic configured and just want to point it at Microsoft Foundry for Claude, you can skip the azure-foundry provider entirely:
AZURE_ANTHROPIC_KEY set in ~/.mibyan/.env. Mibyan detects azure.com in the base URL and short-circuits around the Claude Code OAuth token chain so the Azure key is used directly with x-api-key auth.
key_env is the canonical snake_case field name; api_key_env (and the camelCase keyEnv / apiKeyEnv) are accepted as aliases. If both key_env and AZURE_ANTHROPIC_KEY/ANTHROPIC_API_KEY are set, the key_env-named env var wins.
Model discovery
Azure does not expose a pure-API-key endpoint to list your deployed model deployments. Deployment enumeration requires Azure Resource Manager authentication (az cognitiveservices account deployment list) with an Azure AD principal, not the inference API key.
What Mibyan can do:
- Azure OpenAI v1 endpoints (
<resource>.openai.azure.com/openai/v1) exposeGET /modelswith the resource’s available model catalog. Mibyan uses this list to prefill the setup wizard’s model picker and the in-session/model azure-foundrypicker (CLI, TUI, Desktop, gateway), so you can switch deployments without re-runningmibyan setup. - Microsoft Foundry
/anthropicroutes: detected via URL path, model name entered manually (no/modelsthere — the/modelpicker shows only the current selection and anyproviders.azure-foundry.modelsyou declare). - Private / firewalled endpoints: manual entry with a friendly “couldn’t probe” message.
- Entra ID (
model.auth_mode: entra_id, noAZURE_FOUNDRY_API_KEY): the/modelpicker lists the provider as soon asmodel.base_url(orAZURE_FOUNDRY_BASE_URL) is set — no token is minted just to show the row.
/models, declare them in config.yaml; they are listed first, ahead of the live catalog:
model.base_url while Azure Foundry is the active provider; set AZURE_FOUNDRY_BASE_URL as well if you want the row to stay populated after switching to another provider.
Environment variables
The Azure SDK reads the
AZURE_* env vars directly. Mibyan never inspects them other than to report which sources are present in mibyan doctor output.
Troubleshooting
401 Unauthorized on gpt-5.x deployments. Azure serves gpt-5.x on/chat/completions, not /responses. Mibyan handles this automatically when the URL contains openai.azure.com, but if you see a 401 with an Invalid API key body, check that api_mode in your config.yaml is chat_completions.
404 on /v1/messages?api-version=.../v1/messages.
This is the malformed-URL bug from pre-fix Azure Anthropic setups. Upgrade Mibyan — the api-version parameter is now passed via default_query rather than baked into the base URL, so the SDK can’t corrupt it during URL joining.
Wizard says “Auto-detection incomplete.”
The endpoint rejected both the /models probe and the Anthropic Messages probe. This is normal for private endpoints behind a firewall or with an IP allow-list. Fall back to manual API mode selection and type your deployment name — everything still works, Mibyan just can’t prefill the picker.
Wrong transport picked.
Run mibyan model again and the wizard will re-probe. If the probe still picks the wrong mode, you can edit config.yaml directly:
auth_mode: entra_id.
- Run
az loginto refresh your developer session (the cached token may have expired). - Verify the
Azure AI User(orFoundry User) role assignment took effect:az role assignment list --assignee <user-or-identity-id>should list it on your Foundry resource. Role propagation can take up to 5 minutes. - For user-assigned managed identities, double-check
AZURE_CLIENT_IDmatches the identity attached to the compute resource. - Run
mibyan doctor— the Azure Entra probe reports whether token acquisition succeeded and includes a remediation hint.
mibyan doctor after deploying to the target environment. Common causes include an unreachable token service or stale local login state — prefer workload identity in CI, set AZURE_TENANT_ID+AZURE_CLIENT_ID+AZURE_CLIENT_SECRET when using a service principal, or run az login for local development.
401 on Anthropic-style endpoint with Entra ID.
Verify the same Azure AI User (or Foundry User) role is assigned on the Foundry resource (it covers both /openai/v1 and /anthropic paths). If the OpenAI-style probe works during the wizard but claude-* requests fail at runtime, the most common cause is a stale model.entra.scope left over from an earlier wizard run — delete the entra.scope line from config.yaml so the runtime falls back to the default https://ai.azure.com/.default scope.
Related
- Environment variables
- Configuration
- AWS Bedrock — the other major cloud provider integration
- Microsoft: Configure Entra ID for Foundry — upstream documentation for the keyless path

