ContextCompressor with an alternative strategy for managing conversation context. For example, a Lossless Context Management (LCM) engine that builds a knowledge DAG instead of lossy summarization.
How it works
The agent’s context management is built on theContextEngine ABC (agent/context_engine.py). The built-in ContextCompressor is the default implementation. Plugin engines must implement the same interface.
Only one context engine can be active at a time. Selection is config-driven:
context.engine to the plugin’s name.
Directory structure
Each context engine lives inplugins/context_engine/<name>/:
The ContextEngine ABC
Your engine must implement these required methods:Class attributes your engine must maintain
The agent reads these directly for display and logging:Optional methods
These have sensible defaults in the ABC. Override as needed:Per-turn context selection and observation
compress() answers “context is too long → make it shorter”. Two optional,
no-op-default hooks cover the orthogonal selection / observation axis, so an
engine no longer has to force should_compress() to True and abuse
compress() as a per-turn callback:
- No-op by default, fail-open. Both default to
return None. A missing hook, an exception, or an invalid return value leaves the request untouched — so a failing engine is never worse than not installing one. The host also identity-checks for the inherited ABC default and skips it entirely, so non-implementing engines (including the built-in compressor) pay no per-request work at all. select_context()is request-only. The returned list replaces the messages for a single provider call; persisted history is never written. ReturningNone,[], a non-list, or a list containing non-dicts all fall open to the unmodified request.- Ordering / cache stability. The hook runs before prompt cache-control and every request sanitizer, so (a) a replacement still passes the same validation as any request, and (b) the no-op default leaves the request byte-identical — prompt-cache behaviour is unchanged for non-implementing engines. An engine that replaces the list changes only its own cache prefix. Evaluated per provider request (re-runs on retries).
on_turn_complete()is post-turn observation only; treatmessagesas read-only. Coverage is best-effort: it fires from the standard turn-finalization seam. Some abnormal early-return paths in the loop (e.g. a content-policy block or a provider terminal failure) persist and return without routing through finalization, so they do not currently emit this hook — treat it as a best-effort observation for completed turns, not a guaranteed callback for every early exit. Unifying all terminal paths behind one finalization seam is a separate follow-up.
Stable message identity: message_uid
Every message the host has persisted carries message_uid: a 32-hex id
(uuid4().hex) minted once, at the row’s first insert, and stored in
messages.message_uid. It is the key to use when an engine keeps its own
per-message state (a verbatim store, a summary DAG, per-message embeddings)
and needs to recognise a message it has already seen. The physical row id
(_row_id) is not that key: it is re-issued by every copy and only present on
some restore paths.
What the host guarantees:
- Present on every engine surface once a row exists: the
compress()input list,on_turn_complete()clones,post_llm_call’sconversation_history,on_session_end()messages, and every restored history (CLI, TUI, ACP, gateway, compression’s durable-snapshot adoption). It is restored unconditionally, unlike_row_id. The one exception is a row written before the column existed: on a large store the upgrade mints those over the next few opens, so a legacy row can briefly arrive without one. Treat a missing uid as “no identity yet”, never as an error. - Kept across every host copy of the same logical message: in-place
compaction generations and their concurrent-tail clones, rotation-child
handoff copies and foreign-tail clones,
replace_messagesre-issues, rewind, export/import, and/branch/ Desktop branch copies (the child session’s copied rows keep the parent’s uids). - Kept across content rewrites of the same row: the persist override, the
sanitizer’s row-addressed rewrite, the interrupted-stream fill. Treat
(message_uid, content)as a version of the message; never fail closed on a content change under a known uid. - Merges keep the first constituent’s uid and record the rest. Whenever
the host folds one durable message into another (alternation repair’s
consecutive-user and consecutive-assistant merges, the compressor restating
an in-flight task onto its summary carrier, the real user anchor folded into
a trailing scaffolding turn, micro-compaction’s adjacent-user merge), the
composite keeps the uid of the constituent whose text comes first and
records the others in
_absorbed_message_uids(text order). The list is persisted on the survivor’s row (messages.absorbed_message_uids) and restored with it, so after a restart an engine still sees thatA\n\nBis the host’s fold ofAandBrather than a new message. Merges made while restoring a history (repair_alternation=True) record the witness the same way. A row the host discards rather than folds (a provisional verification candidate superseded by the final answer, or an assistant turn whose text sits beside multimodal content, which is never joined) is not recorded: the witness names text that lives on in the composite, nothing else. - Engine-authored rows keep the uid the engine sets. If your
compress()output pre-stampsmessage_uidon a summary carrier (or any row it emits), the host writes that value through the commit, both in place and on rotation, and every later restore and copy returns it; rows without one are minted at insert. An engine can therefore recognise its own rows by uid. - A uid names one logical message, not one row. Copies of a message share it by design, and a host path that re-appends an edited or merged message next to a still-active earlier row (a persist override on a restored list, a merged survivor flushed as a new row) leaves two active rows with one uid. Never assume per-row uniqueness in the active set; key your own state on the uid and treat the later row as the current version.
- Tool calls get per-occurrence ids too. Provider tool-call ids repeat
(Mibyan mints deterministic
call_<12hex>ids for identical calls, and models reuse ids), so an assistant message carries_tool_call_uids, a{tool_call_id: uid}map for itstool_calls, and each tool-result message carries the matching_tool_call_uid. Calls that repeat a provider id inside one response share one uid (every result carries it, so no call looks unanswered). When a fold leaves one message holding calls from two turns that share an id, that id maps to a list of uids, one per occurrence intool_callsorder. The provider-facingidinsidetool_callsis untouched. Both are minted at the assistant row’s first insert, paired onto the result when it is flushed (same batch, or from the live list when the result lands in a later flush) and on restore (from the preceding assistant row), persisted (messages.tool_call_uids,messages.tool_call_uid), and kept across the same copy and rewrite paths asmessage_uid. A result whose call was never persisted with a uid has none. - Never on the wire.
message_uid,_absorbed_message_uids,_tool_call_uidsand_tool_call_uidare inPERSISTENCE_ONLY_MESSAGE_FIELDS: stripped from every outgoing provider copy and ignored by the token estimator. Engines that never read them are unaffected. - Absent only before the row exists. The current turn’s user message has
no uid during a preflight
compress()that runs before the turn-start flush; the same dict object receives it at that flush. Stores upgraded from an older schema backfill a uid onto every existing row once (schema v31), and an insert trigger mints one for any row an older build writes into a v31 store afterwards (that build’s live dicts still lack it until restored).
When to use these hooks — and when NOT to
- Implement
select_context()only when your engine must replace the per-request context — retrieval-augmented selection, topic/branch routing, role switching. It is the only verb that can swap which messages enter a request: thepre_llm_callplugin hook is inject-only by documented design (it appends to the user message and never rewrites the list, to preserve the prompt-cache prefix). If you don’t need replacement, don’t implement it. - If your plugin only needs post-turn observation / ingestion (indexing,
memory sync, analytics), implement a memory provider (
sync_turn()— see Memory Provider Plugins) instead of a context engine. A context engine takes ownership of the session’s compaction policy; a memory provider observes turns without owning anything.on_turn_complete()exists as the observation mirror for engines that already needselect_context()— so the same component can learn from the turn it just routed — not as a general-purpose turn callback. - Prompt-cache impact of a real
select_context(). A non-no-op selection naturally changes the prompt-cache prefix for the turns where it changes the selection — that request’s prefix no longer matches the provider’s cached prefix, so those turns re-write cache instead of reading it. Engines should return stable selections when nothing has changed (same object or an equal list), and only reshape the context when the routing decision actually differs; a selection that shuffles per turn silently forfeits cache reuse every turn.
Engine tools
Context engines can expose tools the agent calls directly. Return schemas fromget_tool_schemas() and handle calls in handle_tool_call():
Registration
Via directory (recommended)
Place your engine inplugins/context_engine/<name>/ (bundled) or ~/.mibyan/plugins/<name>/ (user-installed; $mibyan_HOME/plugins/<name>/). The __init__.py must export a ContextEngine subclass or a register(ctx) that calls ctx.register_context_engine(...). Setting context.engine: <name> is the activation — a user-installed engine does not need a plugins.enabled entry. Bundled names win on collision.
Via general plugin system
A general plugin can also register a context engine:AIAgent (parent, subagents, gateway
sessions) needs its own engine so a child’s update_model() cannot mutate the parent’s budget.
Mibyan therefore calls engine.clone_for_agent() on the registered instance at each agent init.
The default is copy.deepcopy(self); override it when the engine holds state that cannot be
deep-copied (locks, SQLite or HTTP connections) and return a fresh engine sharing the durable
backend while copying only the mutable budget fields. If the clone raises, the agent falls back to
the built-in compressor and logs Context engine 'X' could not be safely copied for this agent.
Lifecycle
on_session_reset() is called on /new or /reset to clear per-session state without a full shutdown.
Configuration
Users select your engine viamibyan plugins → Provider Plugins → Context Engine, or by editing config.yaml:
compression config block (compression.threshold, compression.protect_last_n, etc.) is specific to the built-in ContextCompressor, with one explicit exception: compression.model_thresholds (per-model threshold overrides) is part of the context-engine contract. The host assigns the resolved map to engine.model_thresholds before the initial update_model() call, and the base-class update_model() applies it (longest substring match, falling back to the engine’s configured threshold). Engines that override update_model() own their own compaction policy and may honor or ignore the map — from agent.context_compressor import resolve_model_threshold to reuse the same resolution logic. For everything else, your engine should define its own config format if needed, reading from config.yaml during initialization.
Testing
tests/agent/test_context_engine.py for the full ABC contract test suite.
Thread safety
Whencompression.context_timeout_seconds > 0 (the default), Mibyan runs the
whole compression pass — including your engine’s compress() and boundary
callbacks, and any memory provider’s on_pre_compress /
on_session_switch — on a pooled daemon thread with a host-side timeout.
Your engine must therefore assume:
- Calls may arrive on an arbitrary pooled thread. Do not rely on thread
affinity or
threading.localstate shared with the conversation thread. - The message list you receive is a private deep snapshot; mutating it in place is allowed (legacy contract), but the mutation only becomes visible if the pass commits. After a host timeout your still-running work is discarded — never publish to external/durable state outside the commit.
- Passes for different sessions can run concurrently on pool siblings; a single engine/provider instance shared across sessions must be thread-safe.
See also
- Context Compression and Caching — how the built-in compressor works
- Memory Provider Plugins — analogous single-select plugin system for memory
- Plugins — general plugin system overview

