How It Works
Two files make up the agent’s memory:
Both are stored in
~/.mibyan/memories/ and are injected into the system prompt as a frozen snapshot at session start. The agent manages its own memory via the memory tool — it can add, replace, or remove entries.
Character limits keep memory focused. Memory does not auto-compact: when a
write would exceed the limit, the
memory tool returns an error instead of
silently dropping entries. The agent then makes room itself — consolidating or
removing entries in the same turn before retrying (see What Happens When Memory
is Full). Note that replace is also bound
by the limit: swapping an entry for a longer one can still overflow, so the new
content must be shortened (or another entry removed) to fit.How Memory Appears in the System Prompt
At the start of every session, memory entries are loaded from disk and rendered into the system prompt as a frozen block:- A header showing which store (MEMORY or USER PROFILE)
- Usage percentage and character counts so the agent knows capacity
- Individual entries separated by
§(section sign) delimiters - Entries can be multiline
Memory Needs Session Boundaries
The whole memory system is built around the moment a session ends:MEMORY.md and USER.md carry the essentials into the next session, and session_search fills the gaps once the old context is gone. Inside a single session none of that machinery has a reason to run — everything important is still in the live context, so the agent rarely consults session_search and mostly compacts memory entries instead of curating them.
This matters on messaging platforms (Telegram, Discord, etc.), where a chat is deliberately one continuous session that survives restarts, gateway crashes, and machine reboots. Shutting the machine down overnight does not end the session — the next message picks it up exactly where it left off. If you never reset, a chat can run for weeks as a single session: convenient, but it grows expensive (compaction runs repeatedly over an ever-longer history) and the learning loop of forget → recall from memory → search past sessions almost never gets to fire. Fresh memory entries also stay invisible to the running session because of the frozen snapshot above.
Practice: run /new at natural boundaries — a finished task, a change of topic, the start of a day. Each boundary is when memory pays off: the agent re-reads the updated MEMORY.md/USER.md snapshot, starts from a cheap short context, and reaches for session_search when it actually needs history. On the CLI this mostly takes care of itself (every invocation is a new session); on gateways the boundary is yours to create.
Troubleshooting: “I told it to remember, and the next session it forgot”
The most common report looks like this: you tell the agent where something lives (an Obsidian vault, a project directory, a server), it answers “Done, I’ll remember that”, and a fresh session has no idea what you mean. Work through these in order — the first one explains the large majority of cases.-
Check whether the write actually happened. Memory only persists when the model calls the
memorytool; a sentence like “I’ve added that to my memory” is just text. Open the file and look for the entry:If the fact is not there, the model claimed a save it never made. Small local models (roughly under 30B parameters) and models with weak tool-calling do this often — they produce the confirmation without the tool call. Ask explicitly (“use thememorytool to save the vault path/srv/vault”) and confirm the entry landed in the file. If it keeps happening, the fix is a stronger model for setup, not more instructions; once the entries exist, a smaller model reads them fine because they arrive in the system prompt. -
Check the write wasn’t staged. With
write_approval: true, writes outside the interactive CLI are held for review and never reach the file until approved — run/memory pendingand/memory approve all. See Controlling memory writes. -
Check you are reading the same memory you wrote. Memory is per profile:
mibyan -p work(orwork chat/work gateway start) reads~/.mibyan/profiles/work/memories/, not~/.mibyan/memories/. A CLI session in the default profile and a Telegram bot on another profile do not share notes.mibyan profile listshows what exists. -
Check memory is enabled.
memory.memory_enabled: false(ormemoryunderagent.disabled_toolsets) removes the tool entirely — the model cannot save anything, whatever it says. See Configuration. -
Remember the snapshot is frozen at session start. A fact saved in the current session is visible to the next session, not to another session that was already running. Start a new session (
/new, or a fresh CLI invocation) after the write.
.env (those are credentials and settings, not memory) and facts mentioned in passing without asking for them to be saved. For a location the agent needs on every run of a recurring task, a skill is often the better home than a memory entry — it loads only when relevant and does not compete for the 2,200-character budget.
Memory Tool Actions
The agent uses thememory tool with these actions:
- add — Add a new memory entry
- replace — Replace an existing entry with updated content (uses substring matching via
old_text) - remove — Remove an entry that’s no longer relevant (uses substring matching via
old_text)
read action — memory content is automatically injected into the system prompt at session start. The agent sees its memories as part of its conversation context.
Substring Matching
Thereplace and remove actions use short unique substring matching — you don’t need the full entry text. The old_text parameter just needs to be a unique substring that identifies exactly one entry:
replace overwrites the whole matched entry with content — old_text only locates the entry, it is not cut out and replaced. The new content must be the complete new entry, including every part of the old one you want to keep. (A whole-entry old_text equal to the entry itself is matched exactly and wins over substring matches.)
Two Targets Explained
memory — Agent’s Personal Notes
For information the agent needs to remember about the environment, workflows, and lessons learned:
- Environment facts (OS, tools, project structure)
- Project conventions and configuration
- Tool quirks and workarounds discovered
- Completed task diary entries
- Skills and techniques that worked
user — User Profile
For information about the user’s identity, preferences, and communication style:
- Name, role, timezone
- Communication preferences (concise vs detailed, format preferences)
- Pet peeves and things to avoid
- Workflow habits
- Technical skill level
What to Save vs Skip
Save These (Proactively)
The agent saves automatically — you don’t need to ask. It saves when it learns:- User preferences: “I prefer TypeScript over JavaScript” → save to
user - Environment facts: “This server runs Debian 12 with PostgreSQL 16” → save to
memory - Corrections: “Don’t use
sudofor Docker commands, user is in docker group” → save tomemory - Conventions: “Project uses tabs, 120-char line width, Google-style docstrings” → save to
memory - Completed work: “Migrated database from MySQL to PostgreSQL on 2026-01-15” → save to
memory - Explicit requests: “Remember that my API key rotation happens monthly” → save to
memory
Skip These
- Trivial/obvious info: “User asked about Python” — too vague to be useful
- Easily re-discovered facts: “Python 3.12 supports f-string nesting” — can web search this
- Raw data dumps: Large code blocks, log files, data tables — too big for memory
- Session-specific ephemera: Temporary file paths, one-off debugging context
- Information already in context files: SOUL.md and AGENTS.md content
Capacity Management
Memory has strict character limits to keep system prompts bounded:What Happens When Memory is Full
When you try to add an entry that would exceed the limit, the tool returns an error:- Read the current entries (shown in the error response)
- Identify entries that can be removed or consolidated
- Use
replaceto merge related entries into shorter versions - Then
addthe new entry
Practical Examples of Good Memory Entries
Compact, information-dense entries work best:Duplicate Prevention
The memory system automatically rejects exact duplicate entries. If you try to add content that already exists, it returns success with a “no duplicate added” message.Security Scanning
Memory entries are scanned for injection and exfiltration patterns before being accepted, since they’re injected into the system prompt. Content matching threat patterns (prompt injection, credential exfiltration, SSH backdoors) or containing invisible Unicode characters is blocked.Session Search
Beyond MEMORY.md and USER.md, the agent can search its past conversations using thesession_search tool:
- All CLI and messaging sessions are stored in SQLite (
~/.mibyan/state.db) with FTS5 full-text search - Search queries return actual messages from the DB — no LLM summarization, no truncation
- The agent can find things it discussed weeks ago, even if they’re not in its active memory
- The agent can also scroll forward/backward inside any session it finds
session_search vs memory
Memory is for critical facts that should always be in context. Session search is for “did we discuss X last week?” queries where the agent needs to recall specifics from past conversations.
Learning Journey (/journey)
The learning journey is a timeline view of everything Mibyan has learned — saved skills and memory entries plotted over time (oldest at top, newest at bottom), with a playable “constellation” scrubber that replays the build-up. The same graph data drives three surfaces:
- Classic CLI / standalone —
mibyan journey(aliases:mibyan learning,mibyan memory-graph) renders the timeline in the terminal. Flags:--playanimates the build-up (--fpsto tune it),--width/--heightoverride the render size,--no-colordisables color, and--jsondumps the raw graph payload. - TUI —
/journey(aliases:/learning,/memory-graph) opens the timeline as an overlay. - Desktop app —
/journeyopens the Star Map / memory-graph panel, an interactive visual of the same nodes.
/learn result or a foreground skill_manage create), created by the background review, or used at least once. Bundled skills and hand-written skills that have never been used stay out of the timeline.
Beyond viewing, the journey is also where you prune and correct what Mibyan has learned:
The same
list / delete <id> / edit <id> subcommands work from the in-chat /journey command on the CLI, and the desktop panel offers edit/delete on nodes directly.
Configuration
memory_enabled and user_profile_enabled to false turns the
built-in stores off completely: the memory tool is dropped from the schema and
its guidance block is dropped from the system prompt, so the model is never told
about a tool it cannot use. An external provider set via memory.provider
(Hindsight, Mem0, Honcho, …) is unaffected and keeps its own tools — use this
when you want a third-party memory backend instead of the built-in files.
Listing memory under agent.disabled_toolsets is the heavier switch: it hides
external provider tools too.
With only memory_enabled: false (user profile still on), the tool stays —
it backs the profile store — but the system prompt swaps the full memory
guidance for a narrower profile-only block. The tool schema advertises only the
user target, and direct or staged writes to disabled MEMORY.md are rejected.
The inverse configuration advertises only memory and rejects USER.md writes.
Controlling memory writes (write_approval)
By default the agent saves memory freely — including from the background
self-improvement review that runs after a turn. If you’d rather approve saves
first, set memory.write_approval: true. It’s a simple on/off gate applied to
both foreground turns and the background review:
To turn memory off entirely (not just gate it), set bothReview staged writes from the CLI or any messaging platform:memory_enabled: falseanduser_profile_enabled: false. When both built-in stores are disabled, the built-inmemorytool is automatically hidden.
write_approval: true, and every save — especially the unprompted background
ones — waits for your yes/no before it ever enters your profile.
A staged replace or remove (the background review stages these even with the
gate off) records the full entry it targets, and /memory pending shows it.
Approval applies to exactly that entry: if it changed after the write was staged,
the write is refused and stays pending for you to reject. A replace/remove
staged before this pinning existed has no verifiable target and is refused too:
reject it and recreate the change. /memory approve lists the full text of
every entry it overwrote or removed.
Background review notifications (display.memory_notifications)
After a turn, the background self-improvement review may quietly save a memory
or update a skill. This is Mibyan’ consent-aware learning loop: repeated
corrections and durable workflow lessons become compact memory entries or
procedural skills, while write_approval can stage those writes for review
before they affect future sessions. By default it surfaces a short
💾 Memory updated line in chat so you know it happened. Control how chatty
that is:
This only governs the gateway chat notification. The review itself, and
writes to your memory/skill stores, are unaffected by this setting. Set it
per-platform via display.platforms.<platform>.memory_notifications.
Successful skill batches name each applied operation in both on and verbose
mode, including supporting-file writes/removals and skill deletion. Staged writes
awaiting approval and rolled-back batches are not reported as completed changes.
Batch summaries use the applied results rather than assuming requested writes ran.
Running the review on a cheaper model (auxiliary.background_review)
The review runs on your main chat model by default, replaying the
conversation — which is already warm in the prompt cache, so it’s cheap cache
reads. On an expensive main model you can run the review on a cheaper model
instead:
auto (or set it to your main model) and nothing changes — the
review keeps running on the main model with the full warm-cache replay.
Same-model review reasoning
A review using the same model as the parent always inherits the parent’s reasoning effort. Settingauxiliary.background_review.reasoning_effort does not override it, whether the route is auto or explicitly selects the parent provider/model.
Reasoning settings, the system prompt, the full conversation snapshot, and tool definitions stay byte-identical to the parent at fork birth so the review can reuse its prompt-cache prefix. Changing only the review’s thinking level would break that parity. There is no independent-effort switch for same-model reviews.
To reduce review work without changing the main conversation’s effort, adjust memory.nudge_interval / skills.creation_nudge_interval, disable automatic reviews as described below, or route reviews to a different model. A different-model route uses a digest and does not share the parent’s warm prefix; on that route auxiliary.background_review.reasoning_effort IS honored (unset = the routed provider’s default). A one-time warning is printed when the key is set but the review stays on the main model. These frequency and routing controls do not decouple same-model reasoning.
Disabling automatic reviews (enabled)
The review fork can burn a meaningful share of total tokens on busy hosts.
Operators can disable it without zeroing nudge intervals:
enabled: false, automatic post-turn forks do not spawn; manual
/refine still works.
Capping review cost (max_input_tokens)
The review loop replays the conversation on every provider request it makes,
so a single review can multiply input tokens across its tool iterations.
max_input_tokens caps the SUM of replayed input tokens for one review; the
loop stops before crossing it. <= 0 means unlimited.
auxiliary:; a top-level background_review: block is not read.
Fork usage is persisted in session_model_usage with task='background_review'
and a completion line is written to agent.log
(Background review complete: thread=bg-review calls=… in=… out=… result=…).
Allowing a narrowly scoped extra review tool (extra_tools)
Background review can use memory, skill-management, and read-only file tools
by default. If a profile provides another tool that is safe for unattended
review, opt it in by name:
Local models: reviews wait for an idle GPU (defer)
On a cloud provider the review finishes in seconds and runs alongside
whatever you do next. When the review’s runtime is the managed local
llama-server (Settings → Local models), the same fork occupies the GPU your
next prompt needs — for minutes on a large model — and sending a new prompt
cancels it, discarding the learning. So on the managed local runtime, reviews
are deferred by default: queued at turn end and executed once the machine
has been quiet for a short settle window. Nothing about the review itself
changes — same model, same full-transcript replay, same writes — only the
execution moment moves.
Queued reviews coalesce per session (a newer turn’s snapshot replaces the
older one — the review replays the whole conversation, so nothing is lost),
a review preempted by a new prompt is re-queued instead of discarded, and a
review that has waited longer than
defer_max_age_s runs even if the machine
never goes idle. Explicit /refine always runs immediately. The queue is
in-memory: reviews still pending when the app exits are dropped, same as an
in-flight fork would have been.
Controlling skill writes (skills.write_approval)
Skills use the same on/off gate, but the review UX differs because a
SKILL.md is far too large to read in a chat bubble:
write_approval: true, skill writes (create / edit / patch / write_file /
delete) always stage regardless of origin. You review the one-line gist
inline, but the full diff stays out-of-band:
/skills diff on the CLI / dashboard / the staged file under
~/.mibyan/pending/skills/<id>.json when you want to read the whole change.
Full details in Gating agent skill writes.
External Memory Providers
For deeper, persistent memory that goes beyond MEMORY.md and USER.md, Mibyan ships with 7 external memory provider plugins — Honcho, OpenViking, Mem0, Holographic, RetainDB, ByteRover, and Supermemory — and more, such as Hindsight, are available from the plugin catalog viamibyan plugins install <name>.
External providers run alongside built-in memory (never replacing it) and add capabilities like knowledge graphs, semantic search, automatic fact extraction, and cross-session user modeling.

