Skill metadata
Reference: full SKILL.md
The following is the complete skill definition that Mibyan loads when this skill is triggered. This is what the agent sees as instructions when the skill is active.
Karpathy’s LLM Wiki
Build and maintain a persistent, compounding knowledge base as interlinked markdown files. Based on Andrej Karpathy’s LLM Wiki pattern. Unlike traditional RAG (which rediscovers knowledge from scratch per query), the wiki compiles knowledge once and keeps it current. Cross-references are already there. Contradictions have already been flagged. Synthesis reflects everything ingested. Division of labor: The human curates sources and directs analysis. The agent summarizes, cross-references, files, and maintains consistency.When This Skill Activates
Use this skill when the user:- Asks to create, build, or start a wiki or knowledge base
- Asks to ingest, add, or process a source into their wiki
- Asks a question and an existing wiki is present at the configured path
- Asks to lint, audit, or health-check their wiki
- References their wiki, knowledge base, or “notes” in a research context
Wiki Location
Location: Set viaWIKI_PATH environment variable (e.g. in ${mibyan_HOME:-~/.mibyan}/.env).
If unset, defaults to ~/wiki.
Architecture: Three Layers
SCHEMA.md defines structure, conventions, and tag taxonomy.
Resuming an Existing Wiki (CRITICAL — do this every session)
When the user has an existing wiki, always orient yourself before doing anything: ① ReadSCHEMA.md — understand the domain, conventions, and tag taxonomy.
② Read index.md — learn what pages exist and their summaries.
③ Scan recent log.md — read the last 20-30 entries to understand recent activity.
- Creating duplicate pages for entities that already exist
- Missing cross-references to existing content
- Contradicting the schema’s conventions
- Repeating work already logged
search_files for the topic
at hand before creating anything new.
Initializing a New Wiki
When the user asks to create or start a wiki:- Determine the wiki path (from
$WIKI_PATHenv var, or ask the user; default~/wiki) - Create the directory structure above
- Ask the user what domain the wiki covers — be specific
- Write
SCHEMA.mdcustomized to the domain (see template below) - Write initial
index.mdwith sectioned header - Write initial
log.mdwith creation entry - Confirm the wiki is ready and suggest first sources to ingest
SCHEMA.md Template
Adapt to the user’s domain. The schema constrains agent behavior and ensures consistency:confidence and contested are optional but recommended for opinion-heavy or fast-moving
topics. Lint surfaces contested: true and confidence: low pages for review so weak claims
don’t silently harden into accepted wiki fact.
raw/ Frontmatter
Raw sources ALSO get a small frontmatter block so re-ingests can detect drift:sha256: lets a future re-ingest of the same URL skip processing when content is unchanged,
and flag drift when it has changed. Compute over the body only (everything after the closing
---), not the frontmatter itself.
Tag Taxonomy
[Define 10-20 top-level tags for the domain. Add new tags here BEFORE using them.] Example for AI/ML:- Models: model, architecture, benchmark, training
- People/Orgs: person, company, lab, open-source
- Techniques: optimization, fine-tuning, inference, alignment, data
- Meta: comparison, timeline, controversy, prediction
Page Thresholds
- Create a page when an entity/concept appears in 2+ sources OR is central to one source
- Add to existing page when a source mentions something already covered
- DON’T create a page for passing mentions, minor details, or things outside the domain
- Split a page when it exceeds ~200 lines — break into sub-topics with cross-links
- Archive a page when its content is fully superseded — move to
_archive/, remove from index
Entity Pages
One page per notable entity. Include:- Overview / what it is
- Key facts and dates
- Relationships to other entities ([[wikilinks]])
- Source references
Concept Pages
One page per concept or topic. Include:- Definition / explanation
- Current state of knowledge
- Open questions or debates
- Related concepts ([[wikilinks]])
Comparison Pages
Side-by-side analyses. Include:- What is being compared and why
- Dimensions of comparison (table format preferred)
- Verdict or synthesis
- Sources
Update Policy
When new information conflicts with existing content:- Check the dates — newer sources generally supersede older ones
- If genuinely contradictory, note both positions with dates and sources
- Mark the contradiction in frontmatter:
contradictions: [page-name] - Flag for user review in the lint report
_meta/topic-map.md that groups pages by theme for faster navigation.
log.md Template
Core Operations
1. Ingest
When the user provides a source (URL, file, paste), integrate it into the wiki: ① Capture the raw source:- URL → use
web_extractto get markdown, save toraw/articles/ - PDF → use
web_extract(handles PDFs), save toraw/papers/ - Pasted text → save to appropriate
raw/subdirectory - Name the file descriptively:
raw/articles/karpathy-llm-wiki-2026.md - Add raw frontmatter (
source_url,ingested,sha256of the body). On re-ingest of the same URL: recompute the sha256, compare to the stored value — skip if identical, flag drift and update if different. This is cheap enough to do on every re-ingest and catches silent source changes.
search_files to find
existing pages for mentioned entities/concepts. This is the difference between
a growing wiki and a pile of duplicates.
④ Write or update wiki pages:
- New entities/concepts: Create pages only if they meet the Page Thresholds in SCHEMA.md (2+ source mentions, or central to one source)
- Existing pages: Add new information, update facts, bump
updateddate. When new info contradicts existing content, follow the Update Policy. - Cross-reference: Every new or updated page must link to at least 2 other
pages via
[[wikilinks]]. Check that existing pages link back. - Tags: Only use tags from the taxonomy in SCHEMA.md
- Provenance: On pages synthesizing 3+ sources, append
^[raw/articles/source.md]markers to paragraphs whose claims trace to a specific source. - Confidence: For opinion-heavy, fast-moving, or single-source claims, set
confidence: mediumorlowin frontmatter. Don’t markhighunless the claim is well-supported across multiple sources.
- Add new pages to
index.mdunder the correct section, alphabetically - Update the “Total pages” count and “Last updated” date in index header
- Append to
log.md:## [YYYY-MM-DD] ingest | Source Title - List every file created or updated in the log entry
2. Query
When the user asks a question about the wiki’s domain: ① Readindex.md to identify relevant pages.
② For wikis with 100+ pages, also search_files across all .md files
for key terms — the index alone may miss relevant content.
③ Read the relevant pages using read_file.
④ Synthesize an answer from the compiled knowledge. Cite the wiki pages
you drew from: “Based on [[page-a]] and [[page-b]]…”
⑤ File valuable answers back — if the answer is a substantial comparison,
deep dive, or novel synthesis, create a page in queries/ or comparisons/.
Don’t file trivial lookups — only answers that would be painful to re-derive.
⑥ Update log.md with the query and whether it was filed.
3. Lint
When the user asks to lint, health-check, or audit the wiki: ① Orphan pages: Find pages with no inbound[[wikilinks]] from other pages.
[[links]] that point to pages that don’t exist.
③ Index completeness: Every wiki page should appear in index.md. Compare
the filesystem against index entries.
④ Frontmatter validation: Every wiki page must have all required fields
(title, created, updated, type, tags, sources). Tags must be in the taxonomy.
⑤ Stale content: Pages whose updated date is >90 days older than the most
recent source that mentions the same entities.
⑥ Contradictions: Pages on the same topic with conflicting claims. Look for
pages that share tags/entities but state different facts. Surface all pages
with contested: true or contradictions: frontmatter for user review.
⑦ Quality signals: List pages with confidence: low and any page that cites
only a single source but has no confidence field set — these are candidates
for either finding corroboration or demoting to confidence: medium.
⑧ Source drift: For each file in raw/ with a sha256: frontmatter, recompute
the hash and flag mismatches. Mismatches indicate the raw file was edited
(shouldn’t happen — raw/ is immutable) or ingested from a URL that has since
changed. Not a hard error, but worth reporting.
⑨ Page size: Flag pages over 200 lines — candidates for splitting.
⑩ Tag audit: List all tags in use, flag any not in the SCHEMA.md taxonomy.
⑪ Log rotation: If log.md exceeds 500 entries, rotate it.
⑫ Report findings with specific file paths and suggested actions, grouped by
severity (broken links > orphans > source drift > contested pages > stale content > style issues).
⑬ Append to log.md: ## [YYYY-MM-DD] lint | N issues found
Working with the Wiki
Searching
Bulk Ingest
When ingesting multiple sources at once, batch the updates:- Read all sources first
- Identify all entities and concepts across all sources
- Check existing pages for all of them (one search pass, not N)
- Create/update pages in one pass (avoids redundant updates)
- Update index.md once at the end
- Write a single log entry covering the batch
Archiving
When content is fully superseded or the domain scope changes:- Create
_archive/directory if it doesn’t exist - Move the page to
_archive/with its original path (e.g.,_archive/entities/old-page.md) - Remove from
index.md - Update any pages that linked to it — replace wikilink with plain text + “(archived)”
- Log the archive action
Obsidian Integration
The wiki directory works as an Obsidian vault out of the box:[[wikilinks]]render as clickable links- Graph View visualizes the knowledge network
- YAML frontmatter powers Dataview queries
- The
raw/assets/folder holds images referenced via![[image.png]]
- Set Obsidian’s attachment folder to
raw/assets/ - Enable “Wikilinks” in Obsidian settings (usually on by default)
- Install Dataview plugin for queries like
TABLE tags FROM "entities" WHERE contains(tags, "company")
OBSIDIAN_VAULT_PATH to the
same directory as the wiki path.
Obsidian Headless (servers and headless machines)
On machines without a display, useobsidian-headless instead of the desktop app.
It syncs vaults via Obsidian Sync without a GUI — perfect for agents running on
servers that write to the wiki while Obsidian desktop reads it on another device.
Setup:
~/wiki on a server while you browse the same
vault in Obsidian on your laptop/phone — changes appear within seconds.
Pitfalls
- Never modify files in
raw/— sources are immutable. Corrections go in wiki pages. - Always orient first — read SCHEMA + index + recent log before any operation in a new session. Skipping this causes duplicates and missed cross-references.
- Always update index.md and log.md — skipping this makes the wiki degrade. These are the navigational backbone.
- Don’t create pages for passing mentions — follow the Page Thresholds in SCHEMA.md. A name appearing once in a footnote doesn’t warrant an entity page.
- Don’t create pages without cross-references — isolated pages are invisible. Every page must link to at least 2 other pages.
- Frontmatter is required — it enables search, filtering, and staleness detection.
- Tags must come from the taxonomy — freeform tags decay into noise. Add new tags to SCHEMA.md first, then use them.
- Keep pages scannable — a wiki page should be readable in 30 seconds. Split pages over 200 lines. Move detailed analysis to dedicated deep-dive pages.
- Ask before mass-updating — if an ingest would touch 10+ existing pages, confirm the scope with the user first.
- Rotate the log — when log.md exceeds 500 entries, rename it
log-YYYY.mdand start fresh. The agent should check log size during lint. - Handle contradictions explicitly — don’t silently overwrite. Note both claims with dates, mark in frontmatter, flag for user review.

