Skill metadata
Reference: full SKILL.md
The following is the complete skill definition that Mibyan loads when this skill is triggered. This is what the agent sees as instructions when the skill is active.
OSINT Investigation — Public Records Cross-Reference
Investigative framework for public-records OSINT: government contracts, corporate filings, lobbying, sanctions, offshore leaks, property records, court records, web archives, knowledge bases, and global news. Resolve entities across heterogeneous sources, build cross-links with explicit confidence, run statistical timing tests, and produce structured evidence chains. Python stdlib only. Zero install. Works on Linux, macOS, Windows. Most sources work with no API key (OpenCorporates has an optional free token that raises rate limits). Adapted from the MIT-licensed ShinMegamiBoson/OpenPlanter project; expanded to cover identity / property / litigation / archives / news sources that the original didn’t address.When to use this skill
Use when the user asks for:- “follow the money” — government contracts, lobbying → legislation, sanctions
- corporate due diligence — who controls company X, where are they incorporated, who serves on their boards, what filings have they made
- sanctions screening — is entity X on OFAC SDN, ICIJ offshore leaks
- pay-to-play investigation — contractors with offshore ties, lobbying clients winning awards
- property ownership — find recorded deeds/mortgages by name or address (NYC; for other counties point users at the relevant recorder)
- litigation history — find federal + state court opinions and PACER dockets
- multi-source entity resolution where naming varies (LLC suffixes, abbreviations)
- evidence-chain construction with explicit confidence levels
- “what’s been said about X” — international news (GDELT) + Wikipedia narrative + Wayback Machine to recover dead URLs
- general web research →
web_search/web_extract - domain/infrastructure OSINT →
domain-intelskill - academic literature →
arxivskill - social-media profile discovery →
sherlockskill (optional) - US federal campaign finance — FEC is intentionally NOT covered here (the API is unreliable for ad-hoc contributor-name queries on the free DEMO_KEY tier). For federal donations, point users at https://www.fec.gov/data/ directly.
Workflow
The agent runs scripts via theterminal tool. SKILL_DIR is the directory
holding this SKILL.md.
1. Identify which sources apply
Read the data-source wiki entries to plan the investigation:2. Acquire data
Each source has a stdlib-only fetch script inSKILL_DIR/scripts/:
Federal financial / regulatory
SEC_USER_AGENT, SENATE_LDA_TOKEN,
OPENCORPORATES_API_TOKEN, COURTLISTENER_TOKEN). All scripts surface
429 responses immediately with the upstream’s quota message so the user
knows to slow down or supply a key.
3. Resolve entities across sources
Normalize names and find matches between two CSV files:
Output
cross_links.csv columns: match_type, confidence, left_name, right_name, left_normalized, right_normalized, left_row, right_row.
4. Statistical timing correlation (optional)
Test whether two time series cluster suspiciously close together — e.g. lobbying filings near contract awards — using a permutation test:5. Build the findings JSON (evidence chain)
id, title, severity, confidence, summary, evidence[], sources[].
Each evidence item points back to a specific row in a source CSV. The user (or a
follow-up agent) can verify every claim against its source.
Confidence and evidence discipline
This is the load-bearing rule of the skill. Tell the user:- Every claim must trace to a record. No naked assertions.
- Confidence tier travels with the claim.
match_type=fuzzyis “probable”, not “confirmed.” - Entity resolution produces candidates, NOT conclusions. A
fuzzymatch between “ACME LLC” and “Acme Holdings Group” is a lead, not a fact. - Statistical significance ≠ wrongdoing. p < 0.05 means the timing pattern is unlikely under the null. It does not establish corruption.
- All data sources here are public records. They may still contain inaccuracies, stale info, or redactions (GDPR, sealed records).
Adding a new data source
Use the template:fetch_<source>.py script in scripts/ that
uses stdlib only and writes a normalized CSV. Update the source list in the
“When to use” section above.
Tools and their limits
entity_resolution.pydoes NOT use external fuzzy libraries (no rapidfuzz, no jellyfish). Token-bag matching is the upper bound here. If you need Levenshtein, transliteration, or phonetic matching, pip-install separately.timing_analysis.pyuses Python’srandomfor permutations. For reproducibility, pass--seed N.fetch_*.pyscripts useurllib.requestand respectRetry-After. Heavy bulk usage may still violate ToS — read each source’s legal section first.
Legal note
All Phase-1 sources are public records. Bulk acquisition is permitted under their respective access terms (FOIA, public records law, ICIJ explicit publication, OFAC public data). However:- Some sources rate-limit aggressively. Respect their headers.
- Some redact registrant info (GDPR on WHOIS, sealed filings).
- Cross-referencing public records to identify private individuals can have ethical implications. The skill produces evidence chains, not accusations.

