AIAgent class. run_agent.py is now a thin facade: the loop itself lives in agent/conversation_loop.py, each turn phase in agent/turn_*.py (iteration prep, API call, API error, overflow, truncation, recovery), constructor wiring in agent/agent_init.py, and everything from prompt assembly to tool dispatch to provider failover in focused agent/*.py modules mixed into AIAgent.
Core Responsibilities
AIAgent is responsible for:
- Assembling the effective system prompt and tool schemas via
prompt_builder.py - Selecting the correct provider/API mode (chat_completions, codex_responses, anthropic_messages)
- Making interruptible model calls with cancellation support
- Executing tool calls (sequentially or concurrently via thread pool)
- Maintaining conversation history in OpenAI message format
- Handling compression, retries, and fallback model switching
- Tracking iteration budgets across parent and child agents
- Flushing persistent memory before context is lost
Two Entry Points
chat() is a thin wrapper around run_conversation() that extracts the final_response field from the result dict.
API Modes
Mibyan supports three API execution modes, resolved from provider selection, explicit args, and base URL heuristics:
The mode determines how messages are formatted, how tool calls are structured, how responses are parsed, and how caching/streaming works. All three converge on the same internal message format (OpenAI-style
role/content/tool_calls dicts) before and after API calls.
Mode resolution order:
- Explicit
api_modeconstructor arg (highest priority) - Provider-specific detection (e.g.,
anthropicprovider →anthropic_messages) - Base URL heuristics (e.g.,
api.anthropic.com→anthropic_messages) - Default:
chat_completions
Turn Lifecycle
Each iteration of the agent loop follows this sequence:Message Format
All messages use OpenAI-compatible format internally:assistant_msg["reasoning"] and optionally displayed via the reasoning_callback.
Message Alternation Rules
The agent loop enforces strict message role alternation:- After the system message:
User → Assistant → User → Assistant → ... - During tool calling:
Assistant (with tool_calls) → Tool → Tool → ... → Assistant - Never two assistant messages in a row
- Never two user messages in a row
- Only
toolrole can have consecutive entries (parallel tool results)
Interruptible API Calls
API requests are wrapped in_interruptible_api_call() which runs the actual HTTP call in a background thread while monitoring an interrupt event:
/stop command, or signal):
- The API thread is abandoned (response discarded)
- The agent can process the new input or shut down cleanly
- No partial response is injected into conversation history
Tool Execution
Sequential vs Concurrent
When the model returns tool calls:- Single tool call → executed directly in the main thread
- Multiple tool calls → executed concurrently via
ThreadPoolExecutor- Exception: tools marked as interactive (e.g.,
clarify) force sequential execution - Results are reinserted in the original tool call order regardless of completion order
- Exception: tools marked as interactive (e.g.,
Execution Flow
Agent-Level Tools
Some tools are intercepted byagent/tool_executor.py (called from agent/conversation_loop.py) before reaching handle_function_call():
These tools modify agent state directly and return synthetic tool results without going through the registry.
Callback Surfaces
AIAgent supports platform-specific callbacks that enable real-time progress in the CLI, gateway, and ACP integrations:
Budget and Fallback Behavior
Iteration Budget
The agent tracks iterations viaIterationBudget:
- Default: 500 iterations (configurable via
agent.max_turns) - Each agent gets its own budget. Subagents get independent budgets capped at
delegation.max_iterations(default 50) — total iterations across parent + subagents can exceed the parent’s cap - At 100%, the agent stops and returns a summary of work done
Fallback Model
When the primary model fails (429 rate limit, 5xx server error, 401/403 auth error):- Check
fallback_providerslist in config - Try each fallback in order
- On success, continue the conversation with the new provider
- On 401/403, attempt credential refresh before failing over
- A Codex Responses turn that stalls on reasoning-only output (three consecutive continuations with no visible text or tool call) also fails over to the next fallback with reason
incomplete_response; if the stall consumed the iteration budget, the fallback gets exactly one bounded grace call
auxiliary.* config section.
Compression and Persistence
When Compression Triggers
- Preflight (before API call): If conversation exceeds 50% of model’s context window
- Gateway auto-compression: If conversation exceeds 85% (more aggressive, runs between turns)
What Happens During Compression
- Memory is flushed to disk first (preventing data loss)
- Middle conversation turns are summarized into a compact summary
- The last N messages are preserved intact (
compression.protect_last_n, default: 20) - Tool call/result message pairs are kept together (never split)
- A new session lineage ID is generated (compression creates a “child” session)
Session Persistence
After each turn:- Messages are saved to the session store (SQLite via
mibyan_state.py) - Memory changes are flushed to
MEMORY.md/USER.md - The session can be resumed later via
/resumeormibyan chat --resume

