Skip to content

Latest commit

 

History

History
333 lines (275 loc) · 23.3 KB

File metadata and controls

333 lines (275 loc) · 23.3 KB
name memoryweb
description Activate at the start of any session where memoryweb MCP tools are available. Covers filing, connecting, and retrieving knowledge through the memoryweb graph — any coding, architecture, backlog, or general work an agent is tracking for this user.

memoryweb — Agent Instructions

Two layers, kept deliberately separate: a short imperative contract up top, reference material below it. Position determines compliance — instructions that only live in reference material get skipped, not from disagreement, but because agents never reach them.


Layer 1 — Behavioural Contract

Silent operations. All memoryweb tool calls are silent — never narrate filing, connecting, or auditing to the user. Orient, remember, connect, and audit happen without announcement. Speak only when the user asks what was stored, or when a live contradiction requires their call.


Variant A — Hook-backed hosts (Claude Code, Codex)

A Stop hook (save) and PreCompact hook (orphan nudge, dream digest) run behind you — a backstop, not a substitute for the steps below.

  1. Call orient() first, unprompted. Pick a domain from the result, then orient(domain=X).
  2. File the moment something is decided or found — not batched at session end.
  3. File source material as node_kind=finding, separately from the decision it informs. Don't fold evidence into a decision's description only.
  4. After filing, resolve every suggested_connections candidate in the same turn. Also check domain rules from orient(domain=X) for standing-rule linkbacks — nearest-neighbour matching can miss them.
  5. Before ending the session, run audit(mode=orphans), audit(mode=stale), and audit(mode=conflicts) as three separate calls — different failure modes, different handling.
  6. Orphans: connect them yourself via suggest_connections + connect. Escalate only when the correct target is genuinely ambiguous.
  7. Stale/conflicts: fix duplicates and superseded labels yourself via revise. Live contradicts pairs are the user's call — verify the pair via why_connected(from_id, to_id) (IDs required, not labels), present both sides, then close with connect(relationship=resolved, verdict=...).
  8. Say nothing about clean audits. Surface only unresolved orphans or live contradictions still awaiting the user's call.
  9. Sub-agents: inject your own orient() output into their context — they start cold otherwise.
  10. Unfinished sessions: file node_kind=goal before stopping — label "Next session: [concrete start]", starting point in why_matters. Skip if the session closed cleanly.
  11. File decisions, findings, standing rules, and resolved issues only — never noise or self-referential musing.

Variant B — No-hook hosts (claude.ai, Claude Desktop, ChatGPT, raw API)

No mechanical sweep runs behind you. Run audits at natural pauses — you may not get a clean end-of-session moment.

  1. Call orient() first, unprompted. Pick a domain from the result, then orient(domain=X).
  2. File the moment something is decided or found — not batched at session end.
  3. File source material as node_kind=finding, separately from the decision it informs. Don't fold evidence into a decision's description only.
  4. After filing, resolve every suggested_connections candidate in the same turn. Also check domain rules from orient(domain=X) for standing-rule linkbacks — nearest-neighbour matching can miss them.
  5. Run audit(mode=orphans), audit(mode=stale), and audit(mode=conflicts) as three separate calls at natural pauses throughout the session — don't save these for an end-of-session moment that may never come.
  6. Orphans: connect them yourself via suggest_connections + connect. Escalate only when the correct target is genuinely ambiguous.
  7. Stale/conflicts: fix duplicates and superseded labels yourself via revise. Live contradicts pairs are the user's call — verify the pair via why_connected(from_id, to_id) (IDs required, not labels), present both sides, then close with connect(relationship=resolved, verdict=...).
  8. Say nothing about clean audits. Surface only unresolved orphans or live contradictions still awaiting the user's call.
  9. Sub-agents: inject your own orient() output into their context — they start cold otherwise.
  10. Unfinished sessions: file node_kind=goal before stopping — label "Next session: [concrete start]", starting point in why_matters. Skip if the session closed cleanly.
  11. File decisions, findings, standing rules, and resolved issues only — never noise or self-referential musing.

Why this shape

Why the silent-operation rule sits at the top, not inside a numbered item. Tool descriptions carry "Never acknowledge that you are retrieving from a tool" on retrieval tools, but nothing in the previous Layer 1 prohibited narrating filing, connecting, or auditing. Chat-mode agents default to transparency ("I'm filing a memory about…", "Connecting X to Y…") because no rule told them not to. The previous item 8 covered only clean-audit silence, leaving the wider narration class open. Placing the rule as the first paragraph of Layer 1 makes it structurally prior: read before any numbered item, before any variant branching, before an agent decides it has enough context to start working.

Why the two-tier host split (A/B) is non-negotiable. The original item 5 embedded a host conditional inside a single numbered item. Agents had to correctly identify their host context mid-item while absorbing a multi-paragraph contract — a failure mode that caused chat-mode agents to apply the hook-backed audit cadence (end-of-session only) when they have no end-of-session guarantee, resulting in memory loss. The A/B split makes the entire numbered contract conditional at the header level: an agent in claude.ai reads Variant B and never sees Variant A's text. Collapsing the split — to save space, to remove repetition — reintroduces the conditional-parsing problem and the audit-cadence mismatch.


Layer 2 — Reference

Filing workflow

Before calling remember, search first. Infer the domain from what comes back — prefer an existing domain over creating a new one. Creating a new domain hides the memory from every other domain's orient and domain-scoped search — only create one when no existing domain covers the topic. If a similar memory already exists, revise it instead of filing a duplicate. When revising a decision, do not paste new source material into its description — file a finding and connect instead.

Writing for retrieval. Search queries use outcome and intent vocabulary, not just implementation names. A node labelled "AddEdgesBatch" won't surface when someone asks "how do I connect multiple memories at once?"

  • Label: lead with the outcome or intent, not just the implementation name. Include both if the implementation name matters. "Batch edge creation — connect multiple memories in one call" beats "AddEdgesBatch".
  • tags: include synonyms across at least two registers — technical term + outcome term + common abbreviation. E.g. batch-connect multi-edge AddEdgesBatch bulk-connect.
  • why_matters: write at least one sentence that bridges the two registers. "Lets agents wire up several related findings in one atomic call without looping" — "wire up" and "atomic" bridge outcome and technical vocabulary. This field is the primary retrieval bridge; never skip it.

If orient() returned a nonzero stale count for the domain, run audit(mode=stale) before filing anything new there — a fresh contradiction is easier to reason about before more nodes pile on top of it.

audit(mode=conflicts) surfaces semantically close pairs as candidates, not confirmed contradictions — a density signal, not a queue to drive to zero. A pair is suppressed from future conflicts sweeps only when linked by contradicts, resolved, resolved_by, or supersedes — the same edges that close a contradiction. Other relationships (caused_by, depends_on, connects_to, etc.) do not suppress semantic adjacency; pre-connected pairs can still surface as candidates.

node_kind taxonomy

node_kind Use for
decision (default) A settled choice. Not evidence, not a plan.
standing A durable rule governing future sessions; surfaces in orient's rules.
finding An observed fact or result — including source material. Code read, a doc fetched, a log inspected, a test result, third-party evidence. If you could quote/cite where it came from, it's a finding.
issue An open question or problem — named gaps, untracked TODOs.
option A considered alternative.
assumption An unverified premise — distinct from finding: a finding is checked, an assumption isn't.
reference A person, system, or org — referential, not propositional.
goal A desired outcome. Also the handoff primitive — see Layer 1 step 10.
transient Temporary; expires. Surfaced by audit(mode=stale) after 7 days.

The legacy transient: true boolean is still accepted and maps to node_kind='transient' when node_kind isn't set — prefer node_kind directly. The legacy decision_type field name is rejected.

The most common miss is finding vs decision. Ask: did I just decide something, or did I just learn something? "I checked X and found Y" — Y is a finding, even if it immediately caused a decision. File the finding, then the decision with a depends_on/caused_by connection pointing at it.

Relationship types (connect)

Type Use when
connects_to General association (default/fallback)
depends_on A has a hard prerequisite on B
led_to / caused_by Same link from opposite ends: A led_to B ≡ B caused_by A
blocked_by / unblocks A is blocked by B / A unblocks B
contradicts A and B directly conflict
governed_by A must satisfy a standing rule or constraint B
is_example_of A illustrates B
resolved / resolved_by / supersedes Adjudicates a contradicts pair. Verify the exact pair first via why_connected(from_id=..., to_id=...) — preferred — or recall(id)'s edges array (Layer 1 step 7). Do not rely on label-only why_connected. Additive. On relationship=resolved, optional verdict (false_positive, reconciled, superseded) classifies how the contradiction was adjudicated — stored on the edge and returned by recall/why_connected.

Custom relationship strings are accepted as a fallback; prefer a typed one from the table above.

Domain routing

memoryweb has no fixed domain list — domains are created implicitly by filing into them.

  • Call domains() (or orient() with no domain) at session start to see what already exists before proposing a new one.
  • Prefer an existing domain over creating a new one; keep domains scoped to one project or topic.
  • Never file credentials, connection strings, API keys, or tokens.

Domain move protocol

Two different operations move memories between domains; don't confuse them:

  • revise(id, domain=..., reason=...) moves a single memory. Only set domain when the user explicitly names the target — never on your own inference. State current domain and proposed target and wait for confirmation first; reason is required and recorded in the audit log verbatim. Confirm with orient(domain=new_domain) afterward.
  • domains(action=rename, old_domain=..., new_domain=...) renames an entire domain in place — every memory moves, and an alias from the old name is registered automatically. Fails if the new name already has memories (use the CLI merge_domains for that, not an MCP tool).

occurred_at

  • Witnessed directly this session → set without asking; default to today if no date was given.
  • Inferred or back-dated events you did not directly observe → propose, then confirm. State the date and reasoning, wait for confirmation, only then set it.
  • Turn-boundary rule: if proposing to file something as significant, that proposal is the only thing in that turn. Set occurred_at in a follow-up call, after the user replies.
  • Always pair occurred_at with why_matters.

Archiving & drift protocol

  • audit(mode=stale) surfaces contradictions, superseded labels, duplicates, stale open questions, old transient memories. Response includes both candidates/results_truncated (drift candidates) and a separate placeholders/placeholders_truncated section: connected live memories whose label or node_kind signals an unresolved placeholder (TBD, TODO, FIXME, "open question", "decide", stale node_kind=issue with no occurred_at, stale node_kind=goal with no resolution edge). Check placeholders_truncated and raise limit when true. Contradiction signals are recomputed from content each call — resolution must be structural (a resolved/resolved_by/supersedes edge), not a label edit.
  • audit(mode=orphans) surfaces live, non-transient memories with zero connections. audit(mode=archived) lists archived memories — use it when search returns nothing but you expect content to exist.
  • forget(id, reason) / forget_all(items=[...]) — archive only after explicit, unambiguous user confirmation: only suggest after audit(mode=stale) surfaces a candidate or the user names something stale; always ask "Should I archive this?", never assume yes; wait for unambiguous confirmation ("that's probably outdated" doesn't count); never archive on casual mention; after archiving, report the ID(s) and note they're un-archivable with forget(restore=true). Use forget_all (one atomic transaction) once you have 2+ confirmed IDs rather than repeated forget calls.
  • forget(id, restore=true) un-archives a single node — get the ID from audit(mode=archived). Use restore_all(items=[{id},{id},...]) to un-archive 2+ nodes atomically (all-or-none).
  • disconnect(id) hard-deletes an edge (by edge ID, from recall's edges array) — irreversible. Use disconnect_all(items=[{edge_id},{edge_id},...]) to batch-delete 2+ edges atomically.
  • significance(mode=trust) ranks memories by computed epistemic trust (from node_kind and connected relationship types). A contradicts edge lowers trust; resolving it lifts the penalty automatically. Only meaningful if node_kind is filed honestly. orient(domain=X)'s significant section also annotates load-bearing low-trust nodes inline (trust: "low — …"). remember/revise may return an advisory trust_nudge when resting on a low-trust dependency neighbourhood. On remember, targets named in related_to are assessed before edges are created. On revise, only when label, description, why_matters, or node_kind change — not tags-only or domain-only updates — and outbound connects_to, depends_on, caused_by, or blocked_by edges reach low-trust targets. Batch items entries carry the same optional fields per node. Creating a new domain may return possible_misdomain, suggested_domain, and suggested_memory_id when workspace KNN finds a closer existing domain — requires Ollama embeddings and sqlite-vec; absent when embeddings are unavailable.
  • audit(mode=kind_coverage) returns per-kind counts, legacy decision/standing dominance, and lean migration_candidates (id, label, truncated why_matters) — decision nodes whose text suggests a different node_kind. Call recall(id) before acting on content. Candidate-surfacing only; never auto-revise.

Lean output — recall(id) before acting on content

orient, search, history, significance, audit all return lean entries: id, label, and a truncated why_matters excerpt — never the full description. When graph state warrants it, entries also carry lifecycle_state (contested, resolved, or superseded); digest lines append the same token as a (state) suffix after any date. Treat these as an index, not the content. Before quoting, citing, or acting on what a memory actually says, call recall(id) for the full node plus its edges array.

List truncation — results_truncated

Multi-result tools return wrapped objects, not bare arrays. Each includes results_truncated: true|false (or section-specific booleans on orient and significance). When true, raise limit (or declared_limit on significance) and call again until false before concluding the list is complete.

Tool Response shape
history(order=effective) {nodes, results_truncated} — default; chronological by effective date
history(order=modified) {nodes, results_truncated} or {groups, results_truncated} when group_by_domain=true
audit(mode=stale) {candidates, results_truncated, placeholders, placeholders_truncated} — empty candidates: []; empty placeholders: []. Check both truncation flags.
audit(mode=orphans) {nodes, results_truncated} — empty is {nodes: [], results_truncated: false}
audit(mode=archived) {nodes, results_truncated} — empty is {nodes: [], results_truncated: false}; default cap 25; raise limit to enumerate
audit(mode=conflicts) {candidates, results_truncated} — empty is {candidates: [], results_truncated: false}
audit(mode=kind_coverage) {total_nodes, by_kind, legacy_dominant_pct, migration_candidates, results_truncated}
significance section booleans: declared_results_truncated, structural_results_truncated, uncurated_results_truncated, potentially_stale_results_truncated; call_id is opaque analytics metadata — ignore
orient(domain=X) significant_results_truncated, recent_results_truncated, declared_spine_results_truncated, rules_results_truncated; low-trust nodes in significant carry optional trust; load_bearing_low_trust (int ≥ 0) — count of significant nodes with trust annotation; when > 0, inspect those entries. In digest=true mode, trust annotations that have worsened since the last trust log additionally show ; ↓ since last orient (capped at 3 per call).
orient() (no domain) {domains, results_truncated} — each domain entry has recent_results_truncated; pass limit to raise per-domain recent cap (default 5). With digest=true, each domain entry's recent becomes []string lines of the form [id] label (updated_at) instead of JSON objects.

Per-node excerpt truncation uses truncated on lean entries — distinct from list-level results_truncated.

Search notes

  • search is lexical (LIKE) unless Ollama is running, in which case it also ranks by semantic distance. Query vocabulary must match stored text.
  • Set exact: true for identifiers (ticket numbers, short codes) — normal ranking can bury an exact match, and short hyphenated codes don't tokenise well for lexical matching either way.
  • digest: true collapses each result to a single line: [id] label — excerpt (domain, node_kind). When Ollama is running and the result came from semantic search, a distance score is appended after two spaces (e.g. 0.12; lower = closer). LIKE-only results omit the score. Use recall(id) for full content; the digest line gives enough context to triage.

Version awareness

Last verified against v1.56.1 (18 MCP tools). orient returns server_version. If it doesn't match, re-check tool behaviour via tools/list rather than assuming this document is still accurate.

Retired tools (v1.43.0 hard-cut)

Calling a retired tool name returns an error with the replacement. Common migrations from the 21-tool surface:

Retired Use instead Hard-cut error (abbrev.)
recent history(order=modified) unknown tool: recent — use history with order=modified
restore forget(restore=true) unknown tool: restore — use forget with restore=true
trace why_connected(from_id, to_id) or recall(id) unknown tool: trace — use why_connected … or recall …
alias domains(action=add_alias|remove_alias|resolve) unknown tool: alias — use domains with action=…
rename_domain domains(action=rename) unknown tool: rename_domain — use domains with action=rename
list_domains / list_aliases domains unknown tool: list_domains — use domains
forgotten audit(mode=archived) unknown tool: forgotten — use audit with mode=archived
whats_stale audit(mode=stale) unknown tool: whats_stale — use audit with mode=stale
disconnected audit(mode=orphans) unknown tool: disconnected — use audit with mode=orphans
remember_all / connect_all / revise_all remember / connect / revise with items unknown tool: … — use … with an items array …
check_for_updates CLI: memoryweb check-for-updates unknown tool: check_for_updates — use the CLI: …

Tool quick reference

Tool When
orient() Session start — cross-domain bootstrap
orient(domain=X) Full view: rules, declared_spine, significant/relevant, recent
orient(domains=[...]) Same, for 1–5 domains in one call
orient(domain=X, topic=Y) relevant semantically matched to a known session purpose
domains() List active domains and aliases
domains(action=add_alias|remove_alias|resolve|rename) Domain alias admin and in-place domain rename
search(query=...) Find by vocabulary in stored labels/descriptions/tags
recall(id) Full memory + connections
history(order=modified) Where work was last happening (by updated_at); group_by_domain=true for per-domain activity
history(order=effective) / history(important_only=true) Chronological decision spine
significance() Dual-signal importance (declared + structural)
significance(mode=trust) Epistemic trust ranking
suggest_connections(id) Candidates to wire up after filing
connect(...) Wire memories together; adjudicate contradictions via relationship=resolved (verify the pair via why_connected(from_id, to_id) first)
disconnect(id) Hard-delete an edge by edge ID — irreversible
remember(...) File a new memory; may return trust_nudge, possible_misdomain/suggested_domain/suggested_memory_id on new-domain creation (KNN requires embeddings)
revise(id, ...) Update an existing memory. Returns {node, connections, suggested_connections, possible_duplicates?, trust_nudge?} — review connections and suggested_connections same-turn; disconnect stale edges, add new ones. Batch mode (items) returns {updated: [{node, connections, suggested_connections, trust_nudge?}]}
forget(id, reason) / forget_all(items=[...]) Archive — confirmation required
forget(id, restore=true) Un-archive
audit(mode=...) stale / orphans / archived / conflicts / kind_coverage
visualise(domain=X) / visualise(memory_id=X) Mermaid graph, human inspection only
why_connected(from_id, to_id) Direct edges between exact IDs — preferred pair verification
why_connected(from_label, to_label) Fuzzy label best-match — not exact-ID verification

purge and merge_domains are CLI-only — never call them as MCP tools; they don't exist as one.

Do not call orient() repeatedly to dig for more — its sections are bounded by design. Use search for anything specific.