# Second brain — setup and upgrade (v3)

Work in the connected folder.

## Ground rules — these override everything below

1. **Never delete.** Not a file, a page, a tag, or a folder. Things that get
   retired move to a `superseded/` folder. Every step is reversible.
2. **Never rename, move, or restructure my folders.** If I already have a layout,
   it is correct by definition — you adapt to it. If my folder names differ from
   the reference names below, use mine and say so once.
3. **Never edit a raw note.** Notes in the inbox and archive are the evidence. The
   only thing you may do to one is move it from inbox to archive, name unchanged.
4. **Never invent a fact about me.** If you need something I have not said, ask.
   If I skip the question, write `_TBD_`.
5. **Never copy examples from another brain into mine.** Every example, tag, and
   piece of vernacular in my files must trace to something in *my* notes or
   something I told you.
6. **Stop where marked.** Phases 1, 2 and 7 end in a gate. Phase 6 gates only if
   Python is missing. Do not continue past one without my go.
7. **Never run the brain on guessed numbers.** If the maintenance script cannot
   run, stop and say so — see Phase 6. An index that is quietly estimated instead
   of counted is the one failure I cannot see from the outside.

---

# Phase 0 — Inspect, then confirm what you found

**Read before you ask.** If I already have material here, the answers to most
setup questions are in it. Asking me to describe my own job when it is written
across forty pages is wasted effort and gets a worse answer than reading would.

Write nothing in this phase.

### 0.1 What is here

- List every folder and every root-level markdown file.
- Count markdown files per folder.
- Check for each and report present/absent: Root instructions file (`AGENTS.md`,
  `CLAUDE.md`, or `AI.md`) · `INGEST-ROUTINE.md` · `.brain/taxonomy.yml` (or
  `.claude/taxonomy.yml`) · `.brain/indexes/` (or `.claude/indexes/`) ·
  `.brain/scripts/` (or `.claude/scripts/`) · any synthesis standard.

**Decide which path you are on:**

| What you found | Path |
|---|---|
| No markdown files, or only unprocessed notes | **Fresh build** — Phase 1 is a full interview |
| Pages exist but no root instructions file and no taxonomy | **Adopt** — harvest from the pages, light interview |
| Root instructions file and a taxonomy exist | **Migrate** — keep everything, replace machinery only |

Say which path you are on and why, in one line.

### 0.2 Read the terminology, and work out what I do

**Fresh build with an empty folder: skip to Phase 1.** There is nothing to read.

Otherwise, sample up to **twelve** files — spread across folders, favouring the
longest, plus any root instructions file. Do not read the whole vault; a sample
is enough to find vocabulary and it keeps this phase cheap.

From that sample, work out and report:

- **My role and function** — what I appear to do for a living, who I do it for,
  and what I am accountable for. Quote the lines that told you.
- **My recurring nouns** — the systems, teams, projects, and internal codenames
  that keep appearing. These become the tag vocabulary in Phase 1. Names an
  outsider would not recognise are the most valuable ones here.
- **The people who recur**, and whether they read as my manager, my peers, my
  stakeholders, or people I am learning from.
- **What kinds of notes I actually take** — this decides which processing lanes
  get built in Phase 7:

  | Shape | How to tell | Lane it needs |
  |---|---|---|
  | Transcript | repeating `HH:MM:SS` timestamps with `Name: utterance` | transcript lane |
  | Prose capture | headings and paragraphs, no speaker turns | capture lane |
  | Clipped external material | links, quotes, someone else's argument | capture lane, research schema |
  | Task or reminder | short, imperative, addressed to future-me | task lane |

### 0.3 Confirm it with me — do not assume you got it right

Play back what you inferred, in **under fifteen lines**: my role, my function, who
my work is for, my recurring vocabulary, and which note shapes you found.

Then ask me directly whether it is right, and what you got wrong. Inference from
documents is good at *what I work on* and unreliable at *why*, seniority, and who
matters to me. Expect corrections and take them at face value.

**If the folder is empty, say so plainly and go to the interview in Phase 1.**

---

# Phase 1 — Vocabulary and identity · **STOP HERE**

## 1A — Fresh build: interview me

Only when there was nothing to read. Keep it to these; do not expand the list.

**Identity**
1. Name, role, and who your work is for.
2. What are you accountable for — the thing you would be judged on in a review?
3. What are you working on right now that will still matter in three months?

**The people**
4. Who do you work with most? For each: manager, peer, stakeholder, or mentor?

**The material**
5. What lands in this brain — meeting recordings, your own notes, articles you
   clip, screenshots, something else?
6. Roughly how many notes a week? *(Under 5 → monthly audit · 5–20 → biweekly ·
   over 20 → weekly.)*

**Retrieval**
7. When you come back looking for something months later, what do you actually
   type? Give three real examples, in your words.

Question 7 matters more than it looks. It is the only one that tells you what my
retrieval vocabulary is, as opposed to my writing vocabulary. They differ, and
the gap is where lookups fail.

## 1B — Existing material: propose, don't ask

Build the vocabulary from the Phase 0 harvest. Every proposed tag must trace to
pages you actually read — name them.

Five axes:

| Axis | What it captures | Open or closed |
|---|---|---|
| `domain` | The system, product surface, or business area — "where does this live?" | Open |
| `function` | The activity or method — "what kind of work is this?" **The highest-value filter, because it cuts across every domain.** | Open |
| `org` | Team, role, org structure, career thread — "whose world is this?" | Open |
| `relation` | How a person relates to me. Person pages only. | Closed |
| `epistemic` | Trust and risk signals about the **content itself**, not its subject. | Closed |

Rules:

- **8–20 controlled tags total** across `domain`, `function`, and `org`. Fewer is
  better. A vocabulary that starts too big is never applied consistently; one
  that starts small grows on evidence.
- **Use my nouns, exactly as I write them.** Where my pages use an internal
  codename, an acronym, or a product name, that string is the tag — do not
  translate it into a general-English category an outsider would prefer. The
  translation loses the only thing that made the tag findable: it is the word I
  will actually type when I come looking. Names an outsider would not recognise
  are the most valuable tags in the vocabulary.
- **If I already have tags, every one gets a disposition. Nothing is dropped:**

  | Existing tag | Pages | Disposition | Becomes |
  |---|---|---|---|
  | `<tag on 3+ pages>` | 6 | **controlled** | `<axis>: <tag>` |
  | `<variant spelling of it>` | 2 | **alias** | → `<the canonical tag>` |
  | `<tag on 1 page>` | 1 | **candidate** | queued; promotes at 3 pages |

  Three dispositions, no fourth. Anything you would want to delete is a
  candidate. **Fill this table with my tags only** — the placeholders above show
  the shape, not the content.
- **Fold variants into aliases now.** Two spellings of one idea is one tag and one
  alias, never two tags.
- `relation` and `epistemic` ship as fixed defaults in Phase 5.

Ask me the two questions you cannot answer from disk: **question 7 above**, and
note volume.

**Stop. Do not write a file until I approve the vocabulary.**

---

# Phase 2 — The change manifest · **STOP HERE**

Every write you are about to make, grouped, with exact paths in **my** folder
names. No prose, no reassurance — I want the blast radius.

- **Files created** — full list.
- **Files patched** — what section gets added to each. Appended, never rewritten.
- **Pages touched** — count and paths. Frontmatter only, never bodies.
- **Scheduled tasks** — which are created or changed.
- **Folders** — state plainly: none created, renamed, or removed. On a fresh
  build, list the folders being scaffolded.

**How to undo it**, in two lines.

**Stop. Wait for my go.**

---

# Phase 3 — Folders

**Migrate or adopt path: do nothing.** Report my existing layout and the role each
folder plays. Use my names everywhere from here on.

**Fresh build:** scaffold these, adjusting names if I asked for different ones.

```
00 notes/          inbox for raw notes — immutable once captured
01 knowledge/
  topics/          themes, initiatives, systems
  people/          one page per person
  decisions/       what was decided, why, and what was rejected
  sources/         one summary per processed note, linking to the archive
03 briefs/         synthesis and status outputs
04 archive/        processed notes, immutable
.brain/
  indexes/         map.md and retrieval-log.md
  scripts/         brain.py
```

Numbered prefixes exist so the folders sort in reading order. Drop them if you
prefer; the script resolves either convention. *(If your agent uses a `.claude/`
convention, `.claude/` is supported as an interchangeable alias).*

---

# Phase 4 — Root Instructions File (`AGENTS.md` / `CLAUDE.md`)

**If I already have a root instructions file, do not rewrite it.** Append the
sections it is missing and show me each addition. Everything I wrote by hand stays.

Fresh build — write it to `AGENTS.md` (or `CLAUDE.md` if working in Claude) from
Phase 0 and Phase 1, in my words, never a template's:

```markdown
# <Brain name>

<One paragraph: who I am, my role, and who my work is for. From Phase 0.3 as I
confirmed it, or Phase 1A.>

> Context the AI agent reads first on every conversation. <Scope note — what
> belongs here and what does not.>

## Who I Am

- **Role:** <role>
- **My work is for:** <who>
- **Accountable for:** <what I would be judged on>
- **Current focus:** <what is live now — or _TBD_>
- **Tools:** <the systems I actually use>

## Retrieval — read the map first

**`.brain/indexes/map.md` lists every page with a one-line summary.** Read it,
pick the pages whose summaries match, open only those — usually three to seven.
Do not grep to discover what exists; the map already says.

- Still need a keyword search? Scope it to the knowledge folder.
- **Never search the archive** except to trace provenance for a specific claim.
- If no tag fits and you fall back to full-text search, say so — it marks a gap
  in the vocabulary.

### Query vernacular — my phrasing → where to look

| When I say… | Go to |
|---|---|
| "what do we know about X" | map → the matching topic page |
| "where do things stand on X" | topic page → Current State, then recent sources |
| "why did we decide X" | decisions — rationale and alternatives |
| "who owns X" / "what did <person> say" | people |
| "prep me for <meeting>" | that person + related topics + open questions |
| "draft a <doc> on X" | related decisions and topics first, so it is grounded |
| <my example 1 from Phase 1> | <route> |
| <my example 2> | <route> |
| <my example 3> | <route> |

### Log retrieval misses at the moment they happen

Ingest learns from what I *write down*. Retrieval learns from what I *go looking
for* — a different signal, and the only one that shows where the vocabulary has a
hole.

A miss is any of: **fallback** (no tag fit, so you grepped) · **vernacular-gap**
(I used a word that is not a tag or alias) · **wrong-shelf** (the map sent you to
a page that was not the answer) · **real-gap** (the brain had nothing) ·
**near-miss** (I named a person or topic and you had to guess which page).

Answer first. Then, in the same turn, append one line to
`.brain/indexes/retrieval-log.md` and close the reply with:

> ⬡ logged retrieval miss — `<type>`: <term> (suggests: <tag/alias/page>)

- **Same turn, never batched.** There is no end-of-session event in a chat
  product — a rule that fires "when we're done" fires never.
- **One line per distinct miss.** Append only.
- **No misses, no write, no note.** A clean session leaves no trace.
- **This file suggests, it never decides.** Only the audit and I promote tags.
- Do not log a miss on a question this brain was never meant to answer.

**The visible note is not optional.** It is the only way either of us can tell the
rule is alive.

## Folder Map (KEEP CURRENT)

<My actual folders, one line each, with the role each plays.>

Machinery: `.brain/taxonomy.yml` · `.brain/indexes/` · `.brain/scripts/brain.py`

Operating docs: <the lanes actually built in Phase 7>

## Operating Rules

- The inbox is immutable — never edit a note after capture; add a dated version.
- Knowledge pages are living and get refined as new notes arrive.
- **Provenance is mandatory** — every page links to its source summary, which
  links to the archived note.
- **Transform, don't transcribe** — reusable knowledge, not verbatim copies.
- **New and edited pages need a `summary:`** in frontmatter, one sentence. The map
  is built from it; a page without one indexes poorly.
- **`type: task` notes are staged, never auto-run.** Nothing executes against any
  outside system until I confirm that task by name.
- **Offer to capture after substantive chats** — if a decision was made or new
  context emerged, ask whether to save it. Skip after quick lookups.

### Answering from the brain

- **Cite the page for every claim.** A claim with no `[[page]]` behind it is
  general knowledge — say so plainly rather than letting it pass as something the
  brain knows.
- **Separate observed from inferred.** "The brain says X" and "my read is Y" are
  different statements. Never let the second wear the clothes of the first.
- **Check freshness.** `fast-changing` with a date 30+ days old, or `volatile` at
  7+ days, gets flagged as possibly stale before it is relied on.
- **Never resolve a contradiction silently.** Two pages disagree → surface both
  with their sources and say which is newer.
- **Confidence never travels upward.** A conclusion built on low-confidence pages
  is low-confidence. Say so.
- **Empty is an answer.** "The brain has nothing on this" is correct and useful.
  Say it, then offer to capture a note. Filling the gap from general knowledge
  without flagging it is the worst failure mode available here.

## Tagging

`.brain/taxonomy.yml` holds the controlled vocabulary on five axes.

- Pick from the controlled list only. A genuinely new term goes in
  `candidate_tags:`, never `tags:`, and promotes after appearing on 3+ pages.
- Never invent a tag for one page. A tag used once filters nothing.
- Renames go in `aliases:`, never as a second tag, so a concept never forks.
- `epistemic` tags describe the content, not its subject.

## Maintenance

Python command on this machine: `python3` — *replace with whatever Phase 6.0
found: `python3` on macOS/Linux, usually `python` on Windows.*

Run `python3 .brain/scripts/brain.py all` after any batch of ingest, and whenever
the map looks behind. It rebuilds the map, reports the singleton ratio, refreshes
the generated section below, and lists stale pages and broken links.

## Processing Notes

"Update my brain" → `INGEST-ROUTINE.md`.
```

---

# Phase 5 — `.brain/taxonomy.yml`

Axes, settings, and the two closed lists are fixed — they are mechanism. The
`controlled:` entries under `domain`, `function`, and `org` come from **my**
Phase 1 vocabulary and nowhere else.

```yaml
# Controlled vocabulary.
#
# WHY THIS EXISTS
#   Ingest left to itself invents a tag per note. Forty pages produces a hundred
#   tags, most used exactly once. Tags used once cannot filter anything; they are
#   noise wearing a filter's clothes. This file makes the vocabulary grow on
#   evidence instead of impulse.
#
# HOW IT WORKS
#   controlled — ingest MUST pick from these. Closed set per axis.
#   candidate  — a genuinely new term lands in `candidate_tags:` in frontmatter,
#                not `tags:`. Visible and greppable, not pretending to filter.
#   promotion  — a candidate on 3+ pages becomes controlled.
#   aliases    — variant spellings resolve to the canonical tag before writing,
#                so renaming a concept never forks the corpus.
#
# EDITING RULES
#   - Adding an alias is always safe. Do it the moment you spot a duplicate.
#   - Adding a controlled tag by hand is allowed but should be rare. Prefer
#     letting the promotion gate earn it.
#   - Adding an AXIS should be very rare. More than once a year means the axis
#     model is wrong, not the content.

version: 1
updated_at: <today>

settings:
  promotion_threshold: 3
  max_tags_per_page: 6
  singleton_target: 0.30
  singleton_alarm: 0.45

# Set from the page's `type:`. Ingest must not choose these, and they are
# excluded from drift metrics.
structural:
  - source
  - topic
  - person
  - decision
  - task
  - transcript

axes:

  domain:
    description: >
      The system, product surface, or business area the content is about.
      Answers "where does this live?". A page uses 1-2.
    open: true
    controlled:
      <my tags from Phase 1>:
        aliases: []

  function:
    description: >
      The activity, capability, or method. Answers "what kind of work is this?".
      The highest-value filter, because it cuts across every domain.
    open: true
    controlled:
      <my tags from Phase 1>:
        aliases: []

  org:
    description: >
      Team, role, org structure, career thread. Slow-changing; reorgs are the
      main churn source.
    open: true
    controlled:
      <my tags from Phase 1>:
        aliases: []

  relation:
    description: >
      How a person relates to me. Only valid on person pages. Closed set — a
      fixed rubric, not a growing vocabulary.
    open: false
    controlled:
      mentor:
        aliases: []
      advocate:
        aliases: [hiring-manager, pilot-participant, implementer]
      sponsor:
        aliases: []
      peer:
        aliases: []
      stakeholder:
        aliases: []

  epistemic:
    description: >
      Trust and risk signals about the CONTENT ITSELF, not its subject. Closed,
      and exempt from the promotion gate — a warning that applies to one page is
      still a warning. This is the highest-value metadata in the vault: no
      embedding, title match, or full-text search will ever recover "don't quote
      this confidently."
    open: false
    controlled:
      attribution-unreliable:   # who said what is uncertain — do not quote
        aliases: []
      under-documented:         # the real detail lives in someone's head
        aliases: []
      key-person-risk:          # depends on one person, or has no owner
        aliases: []
      high-leverage:            # small effort, outsized effect
        aliases: []
      stale:                    # superseded by later sources; kept for provenance
        aliases: []

candidate_tags: {}
```

Also create `.brain/indexes/retrieval-log.md`:

```markdown
---
id: index-retrieval-log
type: index
---

# Retrieval Misses

Append-only. Every row is evidence that the vocabulary has a hole shaped like a
question I actually asked. Written at the moment of the miss, never batched.

Read by the audit as demand-side promotion evidence: a term in **3+ misses** is
as strong a signal as a term on 3+ pages — arguably stronger, since it comes
from demand rather than supply.

| date | type | what I asked / the term I used | suggests |
|---|---|---|---|
```

---

# Phase 6 — Install `brain.py` · **STOP HERE if Python is missing**

Write the script below to `.brain/scripts/brain.py` (or `.claude/scripts/brain.py`),
**exactly as it appears — character for character**. Do not improve it or shorten
it; it automatically resolves your folder layout and instruction files.

## 6.0 Find the Python command — do this first

**The command is not the same on every computer.** On macOS and Linux it is
`python3`. On Windows it is almost always `python`, because the official
installer does not create a `python3`. Getting this wrong looks exactly like
"Python is broken" on a machine where Python is fine, so establish it before
anything else.

Try these in order and stop at the first that prints **3.8 or higher**:

```
python3 --version
python --version
py -3 --version
```

Call the winner **`PY`** for the rest of this document. Wherever you see `PY`,
substitute the command that actually worked.

**Record it** — add a line to the maintenance section of your root instructions file,
so every later run and both scheduled tasks use the right one:

> Python command on this machine: `python3` *(or `python`, or `py -3`)*

Now confirm the script runs:

```
PY .brain/scripts/brain.py map
```

If that printed a map, Phase 6 is done — skip to the table below. Say which
command worked, in one line, and nothing more about it.

## 6.1 If that failed — get Python 3 working

Do not shrug and move on. Without Python the bookkeeping falls back to being
done by hand: slower, more expensive, and prone to arithmetic drift on exactly
the numbers the brain is trusted for. Work the steps below.

**Tell me why, in one line, before you start:** the brain uses a small script to
count pages, build the index, and find broken links — deterministic work that
should not cost model tokens. It needs Python 3.8+, which is free, standard
developer tooling and is not a new account, subscription, or service.

### Step 1 — Work out what actually happened

The failure modes look alike and are fixed differently. Diagnose before acting,
using all three commands from 6.0 — never conclude from `python3` alone:

| What you see | What it means |
|---|---|
| Any of the three prints `Python 3.8`+ | Python is fine — you have your `PY`. If the script still failed it is the *location*: re-run from the vault root and check the script exists |
| **Windows:** `python3` not recognized **but `python` or `py -3` prints 3.8+** | Completely normal. **Python is installed and working** — `PY` is `python`. Not a problem, do not install anything |
| All three print `command not found` / `not recognized` | Genuinely missing. Step 2 |
| **Windows:** opens the **Microsoft Store**, or prints nothing and exits cleanly | The Store alias stub. **This is not a working Python** — try `py -3` before concluding anything; if that also fails, treat as missing. Step 2 |
| `Python 3.7` or lower, or `Python 2.x`, on all of them | Too old. Needs a 3.8+ install alongside it |
| **macOS:** a dialog offers to install developer tools | Command Line Tools. Needs a human click — Step 3 |

**Never report Python as working because a command exited 0.** The Windows stub
exits 0 and does nothing. The only proof is a version string reading 3.8 or
higher.

**Never report Python as missing because `python3` failed.** On Windows that is
the expected result even on a perfectly good install. All three must fail.

### Step 2 — Install it yourself, silently, without asking

Try these **without narrating the attempt**. If one works, say only *"Installed
Python 3.12 — continuing"* and move on. Do not explain package managers to me.

They install to the user account and never prompt for a password:

- **Windows:** `winget install --id Python.Python.3.12 --scope user --silent`
- **macOS, if Homebrew already exists** (`brew --version` succeeds):
  `brew install python@3.12`
- **Linux, if `pyenv` already exists:** `pyenv install 3.12`

**You may not run these** — they need sudo, UAC, an MDM exception, or a GUI
click you cannot perform. They go to Step 3:

- `sudo apt install` / `sudo dnf install` / any `sudo`
- `xcode-select --install` (opens a dialog a human must accept)
- The Homebrew install script (prompts for a password)
- Any `winget` without `--scope user`

After any install, **start a fresh terminal session and re-run all three checks
from 6.0**. A newly installed Python is not visible to the terminal that was
already open — on Windows this usually means closing the terminal window
entirely and opening a new one. **Re-check all three commands, not just
`python3`**, since a Windows install typically answers to `python`.

If it still fails, that is Step 3, not a retry. **Two failed attempts is the
limit** — then Step 3.

### Step 3 — Stop and wait for me · **HARD STOP**

If you cannot install it, **do not continue to Phase 7, and do not run the brain
in a degraded state.** Everything after this depends on the index being real.
Proceeding builds a brain whose numbers are guesses, and I will not know.

Post the message below and **wait for my reply.** Do not proceed on silence, do
not ask again, do not offer to continue without it.

Write it in **plain language for someone non-technical**. No jargon: not
"unelevated," not "PATH," not "package manager," not "MDM." Fill in the bracketed
parts for my actual machine and give **one** command, not a menu:

> **I need Python installed before I can finish setting up your brain.**
>
> Your brain uses a small helper program to keep its index up to date — that is
> what lets it find the right pages fast instead of reading everything, which is
> what keeps it quick and cheap to run. That helper needs Python, a free piece
> of standard software from python.org. It is not an account, a subscription, or
> a service, and nothing gets sent anywhere.
>
> I tried to install it for you and could not, because [**plain reason** — e.g.
> "it needs your password, and I cannot type that for you" / "your work laptop
> blocks new software installs"].
>
> **Here is what to do — about two minutes:**
>
> [**Numbered steps for their exact OS — detect it, do not ask, and never show
> them steps for an operating system they are not on. Download link or one
> command. Say what they will see: an installer window, a password box, a Next
> button.**
>
> **Windows — python.org installer:** the first screen has a checkbox at the
> bottom, **"Add python.exe to PATH"**, and it is **off by default**. Tell them
> to tick it before clicking Install, in those words. Skipping it is the single
> most common reason this fails, and it fails in the worst way: the install
> succeeds and the command still does not work.
>
> **macOS — python.org installer:** a normal .pkg. It will ask for their Mac
> password. There is no checkbox to worry about.**]
>
> **Then just tell me "done" and I will pick up exactly where I left off.** You
> will not lose any progress and you will not need to repeat anything.
>
> If you get stuck or your computer will not let you install it, tell me what you
> saw and I will give you something you can send to IT.

When I come back:

1. Re-run all three version checks from 6.0 to establish `PY`, then
   `PY .brain/scripts/brain.py map`.
2. **If it works:** say so in one line and resume at the exact phase you paused
   in. Do not restart the setup, do not re-ask Phase 1 questions, do not rewrite
   files you already wrote.
3. **If it still fails:** do not repeat the same instructions. Ask for the exact
   error text, then either fix that specific error or give me an IT request I can
   paste: *"I need Python 3.8+ installed for a documentation tooling script. No
   admin rights needed if installed to my user account. Package: python.org 3.12
   or winget `Python.Python.3.12`."*

**Only if I explicitly tell you to carry on without Python** do you continue —
and then every report, including Phase 11, carries this heading:

> **Python unavailable — bookkeeping is manual.** `map.md`, the singleton ratio,
> and the link audit are model-generated estimates, not counted. Install Python
> and run `PY .brain/scripts/brain.py all` to make them real.

Carry that warning until the script runs. A brain that quietly stopped indexing
is worse than one that says it is not indexing.

---

**What it does and why it exists.** Counting pages, building an index, computing
a singleton ratio, and finding broken links are deterministic operations. A
language model doing them produces the same answer as twenty lines of Python, at
several thousand times the cost, with a chance of arithmetic drift. The model's
job is synthesis and judgment. This is the dividing line, and holding it is the
main reason a v3 brain stays fast as it grows.

| Command | What it does |
|---|---|
| `brain.py map` | Rebuilds index map from page frontmatter |
| `brain.py taxonomy` | Tag census, singleton ratio, candidates ready to promote |
| `brain.py context` | Refreshes the generated region of root instructions file |
| `brain.py audit` | Stale pages, orphans, broken links, pages with no summary |
| `brain.py backfill` | One-off: moves glosses out of a v2 `registry.md` onto the pages |
| `brain.py all` | All four |

```python
#!/usr/bin/env python3
"""brain.py - deterministic maintenance for an AI-assisted second brain.

Everything here is bookkeeping: counting, indexing, and checking. None of it
needs judgment, so none of it should cost model tokens. The AI agent does the
synthesis; this does the arithmetic.

No third-party dependencies. Python 3.8+.

  python3 brain.py map        # rebuild index map
  python3 brain.py taxonomy   # tag census, singleton ratio, promotions
  python3 brain.py context    # refresh generated region of instructions file
  python3 brain.py audit      # stale pages, orphans, broken links
  python3 brain.py all        # map + taxonomy + context + audit

  python3 brain.py backfill   # one-off migration: move glosses out of a v2
                              # registry.md onto the pages as `summary:`.
                              # Run this BEFORE the first `map` on a v2 brain,
                              # or the map comes out as titles with no summaries.
"""

from __future__ import annotations

import argparse
import os
import re
import sys
from datetime import datetime, timezone
from pathlib import Path

VAULT = Path(os.environ.get("VAULT_PATH", Path(__file__).resolve().parent.parent.parent))
NOW = datetime.now(timezone.utc)
TODAY = NOW.strftime("%Y-%m-%d")

CTX_BEGIN = "<!-- brain:begin - generated, do not edit below -->"
CTX_END = "<!-- brain:end -->"

STRUCTURAL = {"source", "topic", "person", "decision", "project", "task", "transcript", "index"}


def resolve_dir(*names: str) -> Path | None:
    """Folder names differ between brains. Take the first that exists."""
    for n in names:
        p = VAULT / n
        if p.is_dir():
            return p
    return None


def resolve_meta_dir() -> Path:
    """Resolve metadata root (.brain, .claude, .agent, or .ai)."""
    for n in [".brain", ".claude", ".agent", ".ai"]:
        p = VAULT / n
        if p.is_dir():
            return p
    return VAULT / ".brain"


def resolve_context_file() -> Path | None:
    """Find the root instructions file for this environment."""
    for n in ["AGENTS.md", "CLAUDE.md", "AI.md", "SYSTEM.md"]:
        p = VAULT / n
        if p.exists():
            return p
    return None


KNOWLEDGE = resolve_dir("01 knowledge", "knowledge", "1 knowledge")
NOTES = resolve_dir("00 notes", "notes", "inbox", "0 notes")
META_DIR = resolve_meta_dir()
INDEXES = META_DIR / "indexes"

# Subfolders of knowledge, in display order. Missing ones are skipped.
SECTIONS = ["topics", "people", "decisions", "sources", "projects", "briefs"]


# ---------------------------------------------------------------- frontmatter

def parse_frontmatter(path: Path) -> dict:
    """Minimal YAML frontmatter reader: scalars, block lists, inline [].

    Deliberately not pyyaml - a kit script that needs pip install is a kit
    script that does not run.
    """
    try:
        text = path.read_text(encoding="utf-8")
    except (OSError, UnicodeDecodeError):
        return {}
    if not text.startswith("---"):
        return {}
    end = text.find("\n---", 3)
    if end == -1:
        return {}

    data: dict = {}
    key = None
    for raw in text[3:end].split("\n"):
        if not raw.strip() or raw.lstrip().startswith("#"):
            continue
        if raw.startswith((" ", "\t")) and raw.lstrip().startswith("- ") and key:
            data.setdefault(key, [])
            if isinstance(data[key], list):
                data[key].append(raw.lstrip()[2:].strip().strip("'\""))
            continue
        m = re.match(r"^([A-Za-z_][\w-]*):\s*(.*)$", raw)
        if not m:
            continue
        key, val = m.group(1), m.group(2).strip()
        if val in ("", "[]", "{}"):
            data[key] = [] if val in ("[]", "") else {}
        else:
            data[key] = val.strip("'\"")
    return data


def title_of(path: Path, fm: dict) -> str:
    try:
        for line in path.read_text(encoding="utf-8").split("\n"):
            if line.startswith("# "):
                return line[2:].strip()
    except (OSError, UnicodeDecodeError):
        pass
    return fm.get("id") or path.stem


def gloss_of(fm: dict, words: int = 15) -> str:
    """The one line that lets a page be ruled out without opening it."""
    s = fm.get("summary") or ""
    if isinstance(s, list):
        s = " ".join(s)
    s = re.sub(r"\s+", " ", str(s)).strip().strip("'\"").lstrip("'").rstrip("…")
    parts = s.split()
    return " ".join(parts[:words]) + ("…" if len(parts) > words else "")


def pages(section: str) -> list[Path]:
    if not KNOWLEDGE:
        return []
    d = KNOWLEDGE / section
    if not d.is_dir():
        return []
    return sorted(p for p in d.glob("*.md") if not p.name.startswith("."))


def all_pages() -> list[tuple[str, Path]]:
    return [(s, p) for s in SECTIONS for p in pages(s)]


def tags_of(fm: dict) -> list[str]:
    t = fm.get("tags") or []
    if isinstance(t, str):
        t = [t]
    return [x for x in t if x and x not in STRUCTURAL]


def rel(p: Path) -> str:
    try:
        return str(p.relative_to(VAULT))
    except ValueError:
        return str(p)


# ------------------------------------------------------------------ commands

def cmd_map(write: bool = True) -> str:
    """One file, every page, one line each. Replaces reading the vault."""
    out = [
        "---", "id: index-map", "type: index", f"updated_at: {TODAY}", "---", "",
        "# Brain Map", "",
        "Every page in this brain with a one-line summary. **Read this first.**",
        "Pick the pages whose summaries match, then open only those - usually",
        "three to seven. Never grep to discover what exists; this file already says.",
        "",
        "Rebuilt by `brain.py map`. Appended to by ingest.", "",
    ]
    total = 0
    missing_summary = []
    for section in SECTIONS:
        ps = pages(section)
        if not ps:
            continue
        out.append(f"## {section.title()} ({len(ps)})")
        out.append("")
        for p in ps:
            fm = parse_frontmatter(p)
            g = gloss_of(fm)
            if not g:
                missing_summary.append(rel(p))
            g = g.replace("|", "\\|")
            tg = tags_of(fm)
            line = f"- `{rel(p)}` - **{title_of(p, fm)}**"
            if g:
                line += f" - {g}"
            if tg:
                line += f" · _{', '.join(sorted(tg))}_"
            out.append(line)
            total += 1
        out.append("")

    text = "\n".join(out).rstrip() + "\n"
    if write:
        INDEXES.mkdir(parents=True, exist_ok=True)
        (INDEXES / "map.md").write_text(text, encoding="utf-8")
        print(f"  map.md: {total} pages indexed in {rel(INDEXES / 'map.md')}")
        if missing_summary:
            print(f"  {len(missing_summary)} page(s) have no summary: - they index poorly")
            for m in missing_summary[:5]:
                print(f"    - {m}")
    return text


def cmd_taxonomy() -> dict:
    """Singleton ratio is the health metric. Above 45% the vocabulary is noise."""
    counts: dict[str, list[str]] = {}
    cand: dict[str, list[str]] = {}
    for _, p in all_pages():
        fm = parse_frontmatter(p)
        for t in tags_of(fm):
            counts.setdefault(t, []).append(rel(p))
        c = fm.get("candidate_tags") or []
        if isinstance(c, str):
            c = [c]
        for t in c:
            cand.setdefault(t, []).append(rel(p))

    singles = [t for t, v in counts.items() if len(v) == 1]
    ratio = len(singles) / len(counts) if counts else 0.0

    print(f"  {len(counts)} controlled tags in use across {len(all_pages())} pages")
    print(f"  singleton ratio: {ratio:.0%}  (target <30%, alarm >45%)")
    if ratio > 0.45:
        print("  DRIFT - too many tags used once. Fold variants into aliases.")
    if singles:
        print(f"  used once: {', '.join(sorted(singles)[:12])}")

    ready = {t: v for t, v in cand.items() if len(v) >= 3}
    if ready:
        print(f"  {len(ready)} candidate(s) at 3+ pages, ready to promote:")
        for t, v in sorted(ready.items()):
            print(f"    - {t} ({len(v)} pages)")
    elif cand:
        print(f"  {len(cand)} candidate(s) below the promotion threshold")
    return {"counts": counts, "ratio": ratio, "candidates": cand}


def cmd_context() -> None:
    """Refresh only the generated region of the root instructions file."""
    cm = resolve_context_file()
    if not cm:
        print("  Instructions file (AGENTS.md / CLAUDE.md) not found - skipping")
        return

    counts = {s: len(pages(s)) for s in SECTIONS if pages(s)}
    inbox = len(list(NOTES.glob("*.md"))) if NOTES else 0

    body = [
        CTX_BEGIN, "", "## What's In This Brain", "",
        f"**{inbox} notes** in the inbox; "
        + "; ".join(f"**{n} {s}**" for s, n in counts.items()) + ".", "",
        f"Last compiled: {NOW.strftime('%Y-%m-%dT%H:%M:%SZ')}", "",
    ]

    # Decisions are the highest-value thing to surface; show the newest.
    decs = pages("decisions")
    if decs:
        dated = []
        for p in decs:
            fm = parse_frontmatter(p)
            dated.append((str(fm.get("decided_on") or fm.get("created_at") or ""), p, fm))
        dated.sort(reverse=True)
        body += ["### Recent Decisions", ""]
        for _, p, fm in dated[:5]:
            body.append(f"- [[{rel(p)[:-3]}|{title_of(p, fm)}]]")
        body.append("")
    body.append(CTX_END)

    text = cm.read_text(encoding="utf-8")
    block = "\n".join(body)
    if CTX_BEGIN in text and CTX_END in text:
        text = re.sub(re.escape(CTX_BEGIN) + r".*?" + re.escape(CTX_END),
                      block.replace("\\", "\\\\"), text, flags=re.S)
    else:
        text = text.rstrip() + "\n\n" + block + "\n"
    cm.write_text(text, encoding="utf-8")
    print(f"  {cm.name}: refreshed generated region ({sum(counts.values())} pages)")


def cmd_audit() -> None:
    """Propose, never execute. Every finding here needs a human decision."""
    stale, orphans, broken, nosum = [], [], [], []
    limits = {"volatile": 7, "fast-changing": 30, "slow-changing": 180}

    ps = all_pages()
    # Links into the archive and inbox are provenance, not breakage. Resolve
    # against every markdown file in the vault, not just knowledge pages.
    ids = {p.stem for _, p in ps}
    for extra in ("04 archive", "archive", "00 notes", "notes", "03 briefs", "briefs"):
        d = VAULT / extra
        if d.is_dir():
            ids |= {f.stem for f in d.rglob("*.md")}
    linked: set[str] = set()

    for _, p in ps:
        fm = parse_frontmatter(p)
        try:
            text = p.read_text(encoding="utf-8")
        except (OSError, UnicodeDecodeError):
            continue
        for target in re.findall(r"\[\[([^\]|#]+)", text):
            stem = target.strip().split("/")[-1]
            if stem.endswith(".md"):
                stem = stem[:-3]
            linked.add(stem)
            if stem not in ids:
                broken.append((rel(p), stem))

        if not fm.get("summary"):
            nosum.append(rel(p))

        fc = str(fm.get("freshness_class") or "")
        d = str(fm.get("processed_at") or fm.get("created_at") or "")
        if fc in limits and re.match(r"^\d{4}-\d{2}-\d{2}", d):
            try:
                age = (NOW.date() - datetime.strptime(d[:10], "%Y-%m-%d").date()).days
                if age > limits[fc]:
                    stale.append((rel(p), fc, age))
            except ValueError:
                pass

    for _, p in ps:
        if p.stem not in linked and p.parent.name != "sources":
            orphans.append(rel(p))

    def report(label: str, items: list, fmt=lambda x: f"    - {x}"):
        print(f"  {label}: {len(items)}")
        for i in items[:8]:
            print(fmt(i))
        if len(items) > 8:
            print(f"    … and {len(items) - 8} more")

    report("stale pages", stale, lambda x: f"    - {x[0]} ({x[1]}, {x[2]}d old)")
    report("broken [[links]]", broken, lambda x: f"    - {x[0]} -> [[{x[1]}]]")
    report("orphans (no inbound link)", orphans)
    report("pages with no summary", nosum)
    print("  Propose, don't execute - fix links, never delete pages.")


def cmd_backfill() -> None:
    """Move glosses out of a v2 registry.md and onto the pages themselves.

    A v2 brain keeps its one-line summaries in registry.md and never writes them
    to frontmatter. Everything downstream reads `summary:`, so without this the
    map comes out as a list of titles. Run once, before the first map build.
    """
    candidates = [INDEXES / "registry.md", INDEXES / "superseded" / "registry.md"]
    reg = next((p for p in candidates if p.exists()), None)
    if not reg:
        print("  no registry.md found - nothing to backfill (fine on a fresh brain)")
        return

    # | id | type | aliases | gloss | sources | updated |
    glosses: dict[str, str] = {}
    for line in reg.read_text(encoding="utf-8").split("\n"):
        if not line.strip().startswith("|"):
            continue
        cells = [c.strip() for c in line.strip().strip("|").split("|")]
        if len(cells) < 4 or cells[0].lower() in ("id", "---") or set(cells[0]) <= {"-", ":"}:
            continue
        pid, gloss = cells[0].strip("`[]"), cells[3]
        if pid and gloss and gloss not in ("-", "—"):
            glosses[pid] = gloss

    print(f"  {len(glosses)} gloss(es) found in {rel(reg)}")
    written, skipped = 0, 0
    for _, p in all_pages():
        fm = parse_frontmatter(p)
        if fm.get("summary"):
            skipped += 1
            continue
        g = glosses.get(p.stem) or glosses.get(str(fm.get("id") or ""))
        if not g:
            continue
        text = p.read_text(encoding="utf-8")
        end = text.find("\n---", 3)
        if not text.startswith("---") or end == -1:
            continue
        safe = g.replace("'", "''")
        head = text[:end]
        p.write_text(f"{head}\nsummary: '{safe}'{text[end:]}", encoding="utf-8")
        written += 1

    print(f"  wrote summary: onto {written} page(s); {skipped} already had one")
    if written:
        print("  re-run `brain.py map` to pick them up")


def main() -> int:
    ap = argparse.ArgumentParser(description="Second-brain maintenance")
    ap.add_argument("commands", nargs="+",
                    choices=["map", "taxonomy", "context", "audit", "backfill", "all"])
    args = ap.parse_args()

    if not KNOWLEDGE:
        print(f"No knowledge folder found under {VAULT}", file=sys.stderr)
        print("Set VAULT_PATH, or run this from inside the brain folder.",
              file=sys.stderr)
        return 1

    cmds = ["map", "taxonomy", "context", "audit"] if "all" in args.commands else args.commands
    print(f"=== brain.py: {', '.join(cmds)} ===")
    print(f"vault: {VAULT}")
    for c in cmds:
        print(f"\n--- {c} ---")
        {"map": lambda: cmd_map(), "taxonomy": cmd_taxonomy, "context": cmd_context,
         "audit": cmd_audit, "backfill": cmd_backfill}[c]()
    print("\n=== done ===")
    return 0


if __name__ == "__main__":
    sys.exit(main())
```

---

# Phase 7 — Write the processing lanes I actually need · **STOP HERE**

Build **only** the lanes Phase 0.2 found evidence for. A standard for a note
shape I never produce is dead weight that still gets read.

Every lane below is given as a **rule spec, not a template**. The rules are
mechanism and must survive intact. The examples, section names, and vocabulary
must come from *my* notes — read two or three real ones before writing, and
name which you used.

## `INGEST-ROUTINE.md` — always

Steps, in order:

0. **Load context.** Read `.brain/indexes/map.md` and `.brain/taxonomy.yml`.
   That is the whole context load — one read, not a vault scan. *Cost discipline:
   a run must cost roughly the size of the new notes, not the size of the brain.
   Read frontmatter before bodies; never open a page you are not about to cite
   or edit.*
1. **Find unprocessed notes** in the inbox.
2. **Route by shape.** Frontmatter is optional — raw exports arrive with none,
   and that is normal, not an error. Never skip a file for lacking `type:`.
   Route to the lanes built below.
3. **Update the living knowledge.** Entity resolution first: check the map for a
   match against id, aliases, or gloss before creating any page. "Alex Smith" and
   "A. Smith" are a match — add an alias, never a second page.
4. **Tags** come only from `taxonomy.yml`. Resolve aliases before writing. New
   terms go in `candidate_tags:`.
5. **Archive the note**, name unchanged. Never delete.
6. **Stage `type: task` notes** — surface them, execute nothing.
7. **Run the maintenance script** — `python3 .brain/scripts/brain.py all`, using
   the Python command recorded in your root instructions file.
8. **Report** in 2–4 lines: notes processed, pages touched, tags promoted, staged
   tasks, anything needing attention.

## Transcript lane — only if Phase 0.2 found speaker turns

These rules are non-negotiable; getting them wrong corrupts the graph silently.

- **Do not chase dates.** Set `created_at` to the ingest date. Record a
  `meeting_date` only if one is *free* — in the filename or frontmatter. Never
  derive one from file mtime or inference. Keep relative references as spoken —
  "Wednesday," "end of quarter" — because the words are the fact.
- **Assess diarization** as `clean`, `partial`, or `broken`. When it is not
  clean: never assign an unlabeled turn to a named person, **never write an
  action item with an owner you inferred** — that is a fabricated commitment —
  set `confidence: low`, tag `attribution-unreliable`, and say so in the first
  line. Do not create a person page from a broken transcript.
- **Near-duplicate check, on substantive utterances only** (80+ characters).
  High overlap on long utterances means two passes of one recording — synthesize
  the fuller one. Overlap only on filler means two different meetings.
  **Matching timestamps prove nothing**; a recurring meeting starts at the same
  time every week.
- **Listen a second time for four things** people say aloud and never write down:
  *risk* ("if that slips"), *assumption* ("presumably", "as far as I know"),
  *issue* ("that's still broken"), *decision* ("let's go with"). An assumption
  recorded as fact is the most common way a brain becomes confidently wrong.
- **Sections are named after what the meeting produced**, 2–6 of them, never a
  fixed Summary/Notes shape. `Discussion` and `Other Notes` are banned — if
  content has no evidence-bearing home, it did not earn a section.
- **Every decision carries its why**, and what was rejected.
- **Disagreement is signal.** Record both positions and who held each. The moment
  the argument disappears, so does the reason the decision was hard.
- **Signal over evenness.** Ten minutes of small talk earns zero lines; a
  90-second reversal can earn a section. Never allocate space by airtime.
- **3–6 verbatim quotes with timestamps** — the only place timestamps belong.
  Quotes that state a position, reveal a constraint, or contain pushback.
- Target 20–25% of the source length. Not a highlight reel.

**The test to build against:** six months from now I should be able to answer
*"what did we decide, who owes what, and why did we rule out the alternative"*
from the page alone, without reopening the transcript.

## Capture lane — only if Phase 0.2 found prose notes or clipped material

Different from transcripts on purpose: a transcript records what people said; a
capture is raw material for a call I have to make.

Classify first, then use the matching schema:

- **Delta** — updates something already in the brain. Append to the existing
  page's Updates section with a date; do not create a new page. Links straight
  to the archive.
- **Analytical** — my own thinking. Preserve the reasoning and the alternatives I
  weighed, not just the conclusion.
- **External research** — someone else's claim. **Run a relevance gate first:**
  does this change anything about how I work? If not, archive it and say so
  rather than manufacturing a page. If it passes, capture the claim, the strength
  of its evidence, the author's interest in being right, what it would change,
  and my verdict — kept visibly separate from the claim itself.

**Never let an external author's framing enter my brain as my own position.**
Attribution and verdict stay distinct.

## Task lane — only if Phase 0.2 found task notes

Stage, never execute. Surface in the report. Nothing runs against any outside
system until I confirm it by name.

**Show me each lane you are about to write, and which of my notes you drew the
examples from. Stop for my go.**

---

# Phase 8 — Backfill

**Fresh build with no pages: skip this.** There is nothing to index. Run
`brain.py map` once anyway to confirm the install works, and expect an empty map.

**Adopt or migrate:** this is where the brain becomes searchable.

1. If I have existing tags, rewrite `tags:` to controlled tags only, resolving
   aliases as you go. Cap at 6. **Move every non-controlled tag to
   `candidate_tags:` — nothing is deleted.** Frontmatter only; never touch a body.
   Report a running count every 20 pages so I can interrupt.
2. Add a `summary:` to any page missing one — one sentence, what the page is
   about. The map is built from it.
3. Person pages get exactly one `relation` tag **if it is obvious**. If not,
   leave it off and list the page for me. Guessing whether someone is a sponsor
   or a stakeholder is the kind of confident wrong fact this exists to prevent.
4. Run `python3 .brain/scripts/brain.py all`, using the Python command recorded
   in your root instructions file.
5. **If I am migrating from v2**, move `registry.md`, `by-tag.md`,
   `taxonomy-queue.md`, and `RETRIEVAL-CONTRACT.md` into
   `.brain/indexes/superseded/`. Moved, not deleted. `map.md` replaces the first
   two; root instructions now carry the answering rules.

Report the singleton ratio before and after. That is the headline number.

---

# Phase 9 — Scheduled tasks

**Daily ingest**, at a time I choose:

> Read the root instructions (`AGENTS.md` / `CLAUDE.md`) and `INGEST-ROUTINE.md`,
> then process any unprocessed notes in my inbox following the routine exactly —
> including the Step 0 context load, the routing by note shape, the task lane,
> and the rule to **archive, never delete**. Finish by running the maintenance
> script — `python3 .brain/scripts/brain.py all`, using the Python command
> recorded in your root instructions file. Give me a 2–4 line summary of what
> was processed, tags promoted, staged tasks, and anything needing attention. If
> there are no new notes, say so in one line and stop.

**Audit**, on the cadence from my note-volume answer — weekly over 20 notes/week,
biweekly at 5–20, monthly under 5:

> Run `python3 .brain/scripts/brain.py all` — using the Python command recorded
> in your root instructions file — and read the output. Then read
> `.brain/indexes/retrieval-log.md`: any term appearing in 3+ misses is
> promotion evidence — propose it as a controlled tag or an alias. Recurring
> `real-gap` entries mean a note I should capture; say which. Recurring
> `wrong-shelf` means a page's `summary:` is misleading; rewrite it. Then surface
> any pages asserting incompatible facts. **Propose, don't execute** — fix broken
> links and summaries, but never merge pages, retire tags, or delete anything
> without my okay. Report in under 15 lines.

Never phrase either task as "delete the original note." Always "archive."

---

# Phase 10 — Prove it, don't tell me it works

Show the output, not a summary of the output.

1. **Filter, then read.** Pick a real topic from my brain. Show the `map.md`
   lookup, the pages it returned, and the answer with citations. **Tell me how
   many pages you opened.** More than seven means the map is not doing its job —
   say so.
2. **Citation or silence.** Ask yourself something my brain genuinely has nothing
   on. Show me the refusal, and show me the retrieval-miss line getting logged
   with its `⬡` note.
3. **The script.** Show me the raw output of `brain.py all` — the counts, the
   singleton ratio, the broken links. That output is the proof the bookkeeping is
   no longer costing model tokens.

Fresh build: skip 1 and 2, there is nothing to retrieve yet. Say so.

---

# Phase 11 — Report

Under fifteen lines:

- Created, patched, pages touched.
- Singleton ratio before → after, and which lanes were built.
- Both scheduled tasks and when they run.
- **What did not change:** no folder renamed or removed, no raw note touched, no
  page body edited, nothing deleted.
- **What I should do differently:** nothing. Same folders, same habit. Two visible
  changes — when the brain cannot find something it now tells me and logs it, so
  expect an occasional `⬡`; and maintenance now runs as a script, so the numbers
  in your root instructions file are counted rather than estimated.
- At most one structural observation, if the setup surfaced something real. One
  line. Change nothing.

---

**Run it now. Start with Phase 0.**
