Skip to main content
Hybrid local search over notes, docs, and transcripts.

Skill metadata

Reference: full SKILL.md

The following is the complete skill definition that Mibyan loads when this skill is triggered. This is what the agent sees as instructions when the skill is active.

QMD — Query Markup Documents

Local, on-device search engine for personal knowledge bases. Indexes markdown notes, meeting transcripts, documentation, and any text-based files, then provides hybrid search combining keyword matching, semantic understanding, and LLM-powered reranking — all running locally with no cloud dependencies. Created by Tobi Lütke. MIT licensed.

When to Use

  • User asks to search their notes, docs, knowledge base, or meeting transcripts
  • User wants to find something across a large collection of markdown/text files
  • User wants semantic search (“find notes about X concept”) not just keyword grep
  • User has already set up qmd collections and wants to query them
  • User asks to set up a local knowledge base or document search system
  • Keywords: “search my notes”, “find in my docs”, “knowledge base”, “qmd”

Prerequisites

Node.js >= 22 (required)

SQLite with Extension Support (macOS only)

macOS system SQLite lacks extension loading. Install via Homebrew:

Install qmd

First run auto-downloads 3 local GGUF models (~2GB total):

Verify Installation

Quick Reference

Setup Workflow

1. Add Collections

Point qmd at directories containing your documents:

2. Add Context Descriptions

Context metadata helps the search engine understand what each collection contains. This significantly improves retrieval quality:

3. Generate Embeddings

This processes all documents in all collections and generates vector embeddings. Re-run after adding new documents or collections.

4. Verify

Search Patterns

Fast Keyword Search (BM25)

Best for: exact terms, code identifiers, names, known phrases. No models loaded — near-instant results.
Best for: natural language questions, conceptual queries. Loads embedding model (~3s first query).

Hybrid Search with Reranking (Best Quality)

Best for: important queries where quality matters most. Uses all 3 models — query expansion, parallel BM25+vector, reranking.

Structured Multi-Mode Queries

Combine different search types in a single query for precision:

Query Syntax (lex/BM25 mode)

HyDE (Hypothetical Document Embeddings)

For complex topics, write what you expect the answer to look like:

Scoping to Collections

Output Formats

qmd exposes an MCP server that provides search tools directly to Mibyan via the native MCP client. This is the preferred integration — once configured, the agent gets qmd tools automatically without needing to load this skill.

Option A: Stdio Mode (Simple)

Add to ~/.mibyan/config.yaml:
This registers tools: mcp_qmd_search, mcp_qmd_vsearch, mcp_qmd_deep_search, mcp_qmd_get, mcp_qmd_status. Tradeoff: Models load on first search call (~19s cold start), then stay warm for the session. Acceptable for occasional use. Start the qmd daemon separately — it keeps models warm in memory:
Then configure Mibyan to connect via HTTP:
Tradeoff: Uses ~2GB RAM while running, but every query is fast (~2-3s). Best for users who search frequently.

Keeping the Daemon Running

macOS (launchd)

Linux (systemd user service)

MCP Tools Reference

Once connected, these tools are available as mcp_qmd_*: The MCP tools accept structured JSON queries for multi-mode search:

CLI Usage (Without MCP)

When MCP is not configured, use qmd directly via terminal:
For setup and management tasks, always use terminal:

How the Search Pipeline Works

Understanding the internals helps choose the right search mode:
  1. Query Expansion — A fine-tuned 1.7B model generates 2 alternative queries. The original gets 2x weight in fusion.
  2. Parallel Retrieval — BM25 (SQLite FTS5) and vector search run simultaneously across all query variants.
  3. RRF Fusion — Reciprocal Rank Fusion (k=60) merges results. Top-rank bonus: #1 gets +0.05, #2-3 get +0.02.
  4. LLM Reranking — qwen3-reranker scores top 30 candidates (0.0-1.0).
  5. Position-Aware Blending — Ranks 1-3: 75% retrieval / 25% reranker. Ranks 4-10: 60/40. Ranks 11+: 40/60 (trusts reranker more for long tail).
Smart Chunking: Documents are split at natural break points (headings, code blocks, blank lines) targeting ~900 tokens with 15% overlap. Code blocks are never split mid-block.

Best Practices

  1. Always add context descriptions — qmd context add dramatically improves retrieval accuracy. Describe what each collection contains.
  2. Re-embed after adding documents — qmd embed must be re-run when new files are added to collections.
  3. Use qmd search for speed — when you need fast keyword lookup (code identifiers, exact names), BM25 is instant and needs no models.
  4. Use qmd query for quality — when the question is conceptual or the user needs the best possible results, use hybrid search.
  5. Prefer MCP integration — once configured, the agent gets native tools without needing to load this skill each time.
  6. Daemon mode for frequent users — if the user searches their knowledge base regularly, recommend the HTTP daemon setup.
  7. First query in structured search gets 2x weight — put the most important/certain query first when combining lex and vec.

Troubleshooting

”Models downloading on first run”

Normal — qmd auto-downloads ~2GB of GGUF models on first use. This is a one-time operation.

Cold start latency (~19s)

This happens when models aren’t loaded in memory. Solutions:
  • Use HTTP daemon mode (qmd mcp --http --daemon) to keep warm
  • Use qmd search (BM25 only) when models aren’t needed
  • MCP stdio mode loads models on first search, stays warm for session

macOS: “unable to load extension”

Install Homebrew SQLite: brew install sqlite Then ensure it’s on PATH before system SQLite.

”No collections found”

Run qmd collection add <path> --name <name> to add directories, then qmd embed to index them.

Embedding model override (CJK/multilingual)

Set QMD_EMBED_MODEL environment variable for non-English content:

Data Storage

  • Index & vectors: ~/.cache/qmd/index.sqlite
  • Models: Auto-downloaded to local cache on first run
  • No cloud dependencies — everything runs locally

References