Structure Beats Magic
The vocabulary, defined

Glossary

The non-concept vocabulary of Structure Beats Magic: field terms, distinctions and borrowed ideas that deserve a shared definition and a link — but aren't ownable coined concepts (those live in the concept library, Nick-Milo-style). A volatile, promotable layer: a term that keeps recurring or earns its own article gets promoted to a concept.

Knowledge management

Zettelkasten

Niklas Luhmann's method of atomic, densely-linked notes (one idea per note, wired together). The historical root of the modern second brain; SBM's take is "let the AI do the labor, keep the judgment." → Content Intelligence; article Zettelkasten 2.0.

Personal Knowledge System

The whole apparatus by which one person captures, structures, and reuses what they know — vault, rules, skills, publishing. Broader than PKM (note-taking): a system, not a note collection. → Sovereign Personal System · Personal Data Warehouse.

Personal Knowledge Management (PKM)

The established field of managing one's own notes and knowledge (Obsidian/Notion/Roam culture). SBM positions as a superset: KM re-architected for AI, one level above the note-taking-tool cohort. → Structure Beats Magic.

Atomic note / atomic unit

The smallest self-contained one-idea unit that recomposes into articles, slides, courses, flashcards. The building block of the second brain. → Play With The Blocks · Content To Training; cf. Zettelkasten, Decompose & Reassemble.

Corpus

A whole set of documents made machine-readable together, so a model can range over all of it at once instead of one file at a time. → article Your Bookshelf Is Already a Knowledge Base.

Personal data

EXIF

Metadata embedded in a photo file (GPS, timestamp, camera) that makes an unstructured photo pile queryable without tagging anything by hand. → article Your Photos Are Already a Map.

Geocoding / reverse geocoding

Turning GPS coordinates into named places (city, venue) and back; how location data becomes human-readable, with fallbacks when a lookup fails.

Enrichment

A pass that reads a record's context (e.g. a day's date) and writes derived facts into it from source data — new data created, not typed. → Derived Insight.

Declared vs derived profile

A self-reported profile (aspirational, stale) versus one inferred from actual observed behaviour (current, honest). The gap is why derived data is more valuable. → Derived Insight · A Modeled Self.

Tooling & facilitation

Obsidian

A markdown-based note app treated as an interface over a plain-file vault — the app is rented, the files are owned. → Rent The Ai Own The Structure · Document Vault.

Mermaid

Text-based diagram syntax (flowcharts, timelines, Gantt) that markdown tools render into pictures — "text in, picture out". A living diagram: generated from the source, always current. → Living Diagrams.

Excalidraw

A sketch/whiteboard tool integrable with an Obsidian vault, turning it into a drawing surface — where drawing it together happens on a personal scale.

Positioning

Lowest Hanging Fruit

The audience or opportunity that's easiest to reach and most ready to say yes — the people already looking for what you offer, targeted first before the harder market. A positioning discipline: name your persona precisely, then start where the pull already exists. → Curated Sources; book The Lowest Hanging Fruit (Orsolya Toth).

Data & modelling

Data Modeling

Designing the shape of information: the entities that matter, how they relate, the rules that hold. Not SQL or tables — the conceptual picture where humans and machines agree on meaning. Jaco's core craft. → Visual Thinking · Drawing It Together; MDDE owns the coined variants Business Friendly Data Modeling · Collaborative Data Modeling.

Decompose & Reassemble

Breaking a body of knowledge into atomic, reusable units and recomposing them into new wholes (article, course, talk) without rebuilding. The Lego move applied to knowledge. → candidate concept; overlaps Play With The Blocks · Content To Training.

Frontmatter

The YAML metadata block at the top of a markdown file (type, domain, status, links) that turns a loose note into a queryable data record. The typed fields are the model. → The Vault Is The Data Model (article) · Document Vault.

Markdown (.md)

Plain-text format, greppable and versionable for decades; the open, portable source-of-truth unit for notes, articles and skills. You own the files, not a vendor's database. → Rent The Ai Own The Structure.

YAML

Human-readable key-value markup used for frontmatter and config; itself validatable as valid/invalid, so structure can be checked, not hoped for.

Ontology

A machine-checkable agreement about what exists, how it relates, and what must hold. Not a taxonomy (hierarchy) and not a schema (storage) — the constraints are what earn its keep. → An Ontology Is A Contract.

Knowledge graph

Entities and typed relationships as data, traversable by machines. A model, distinct from the graph database that may or may not store it. → A Graph Is A Model Not A Database.

Semantic layer

Where business meaning, metrics and rules are defined once so every consumer agrees. Answers what it means, where a catalogue answers where it lives. → Metadata Says Where Ontology Says What.

Entity resolution

Deciding that several records denote the same real thing. Every join, count and graph edge is wrong until it is settled. → Resolve The Entity First.

Data lineage

Where data came from and how it was transformed — the past. Distinct from observability, which asks whether it is healthy now. → Uptime Is Not Correctness.

Data product

Data plus its contract: schema, owner, SLA, access policy, docs. Without the contract it is a dataset someone shipped. → A Dataset Is Not A Data Product.

Master data vs reference data

Master data are the core business entities (customer, product); reference data are the value lists that qualify them (country codes, currencies) — the latter usually adopted from standards, not invented. → Resolve The Entity First.

Bi-temporal

Modelling both when something was true in the world and when the system recorded it — what lets a store answer historical questions correctly. → Facts Expire.

Medallion (bronze/silver/gold)

A storage-refinement convention, frequently mis-sold as an architecture: it addresses where data sits, not integration, meaning or governance. → Meaning Belongs In The Data.

DCAT

W3C-standaard om datasets, distributies en catalogi te beschrijven. Het punt is interoperabiliteit: je vindt hem al, dus verzin geen eigen metadataschema. → Meaning Before Mechanism.

PROV-O

W3C-vocabulaire voor herkomst — wie of wat een gegeven heeft voortgebracht, waaruit en wanneer. Zet lineage in de metadata zelf in plaats van in een apart hulpmiddel.

SBVR

OMG-standaard voor bedrijfsvocabulaire en -regels, expliciet ontworpen voor natuurlijke taal. De uitzondering op grafen die taal niet kennen. → An Ontology Is A Contract.

Normal form (1NF/2NF/3NF)

Opeenvolgende eisen die redundantie en update-anomalieën uit een relationeel model halen: atomaire waarden, geen partiële en geen transitieve afhankelijkheden. Ideaal voor OLTP; in analytics is denormaliseren vaak een bewuste keuze.

Semantic medallion

Bronze/silver/gold heruitgelegd als semantiek-gradatie — geen, lokale, globale betekenis — in plaats van als schoonmaak-gradatie. → Meaning Belongs In The Data.

AI & tooling

DuckDB

An in-process analytical (OLAP) SQL database used as the local, rebuildable view over the file corpus — the query layer beside the vault, never the source of truth. → A Brain That Publishes Itself.

MCP (Model Context Protocol)

A connection/server that lets the AI query a live source directly rather than being handed pasted data. One of four tools (with prompts, skills, plugins) to reach for deliberately. → article Prompts, Skills, Plugins, MCP.

Skill (Claude Code)

A procedure written once as a markdown file, invoked as /name, that the AI runs the same way every time — the unit that compounds. → Runbooks; article Skills Are the Unit That Compounds.

Prompt

A one-off instruction typed in the moment; the right tool only when you'll do the thing once. Past that, promote it to a skill. → Stop Prompting Start Directing.

RAG (retrieval-augmented generation)

Handing a model the right retrieved passages to reason over instead of relying on its memory — structure feeding the reasoning. → The Reasoning Layer.

Idempotent

A process safe to re-run: it skips what's already done and never clobbers existing work. What makes an enrichment or import trustworthy to run daily. → The Validation Loop.

Importer vs connector

An importer lands a source's export into a local table; a connector lets the AI query the live source directly. Two ways to bring a source under structure. → Connect And Dispatch.

Embedding

A vector that places a piece of text in meaning-space, so similar ideas sit close together. It is a similarity instrument, not a store of truth — closeness is not correctness. → Retrieval Is Not Memory.

Vector database

A store optimised for nearest-neighbour search over embeddings. Often unnecessary: pgvector or DuckDB serve most workloads, and a vector store is retrieval, not memory. → Retrieval Is Not Memory · A Graph Is A Model Not A Database.

Chunking

Cutting a document into retrievable pieces. The first production bottleneck and the largest quality lever — whatever the cut destroys, no later stage recovers. → Chunking Is A Design Decision.

Reranking

A second pass that reorders retrieved candidates by actual relevance, usually with a cross-encoder. Without it, proximity beats relevance. → Grounding Is An Outcome.

Hybrid search

Combining dense (vector) and sparse (keyword/BM25) retrieval, merged by rank fusion. Recovers exact matches that embeddings miss. → Grounding Is An Outcome.

Top-K

How many retrieved chunks get passed to the model. Too few starves the answer, too many spends the attention budget on padding. → The Attention Budget.

BM25

The classic keyword-ranking function; the sparse half of hybrid search. Still beats embeddings on names, codes and exact phrases.

Grounding

The outcome of an answer being verifiably tied to a trusted source — as opposed to RAG, which is only the technique for attempting it. → Grounding Is An Outcome.

GraphRAG

Retrieval over a graph of entities and relationships rather than over flat chunks; answers multi-hop questions that similarity search cannot reach. → A Graph Is A Model Not A Database.

Corrective RAG (CRAG)

Judging retrieval quality before generating, and re-retrieving or rewriting the query when it is poor. Cheaper than repairing a wrong answer. → The Verifier Is The Bottleneck.

Self-RAG

The model decides whether retrieval is needed at all, what to fetch, and whether the result is usable. The no-retrieval path is the underused optimisation. → Escalate By Exception.

Multi-hop

A question whose answer requires chaining several facts across documents. The case where flat retrieval quietly fails and structure pays off. → A Graph Is A Model Not A Database.

Context window

How much the model can hold in view at once. A hard limit, not a target — filling it degrades recall as surely as starving it. → The Attention Budget.

Context engineering

Deciding what reaches the model, in what order and shape, out of a much larger pool. The discipline between prompting and harness-building. → The Attention Budget · Harness Engineering.

Harness

Everything around the model that lets it finish real work: tools, memory, permissions, checks, retries, escalation. The part you own. → Harness Engineering · The Model Is The Easy Part.

Agentic

A system that plans, acts, observes and loops toward a goal rather than answering once. The loop — not the tools — is what makes it agentic. → Keep The Work That Was Done.

Guardrail

A runtime check constraining what a system may accept, do or emit. Seven layers, not one output filter — and distinct from governance, which is the policy layer above. → Guardrails Are Layers.

Hallucination

A confident fabrication with no basis in any source. Its danger is shape: it arrives indistinguishable from a correct answer. → A Confident Error Has No Tell.

Eval / eval suite

A versioned set of test cases scoring model or agent output. Only a gate if it can block a release; otherwise it is a report. → Evals Are The Release Gate.

Golden dataset

Cases with known-good answers, used to measure whether the system is improving. Its overlooked twin is the regression set — previously-fixed failures that prove you have not broken what worked. → Evals Are The Release Gate.

Trajectory accuracy

Scoring the path an agent took, not just its final answer — the check that catches being right for the wrong reasons. → Evals Are The Release Gate.

LLM-as-judge

Using a model to score another model's output against a rubric. Only as good as the rubric's decidability, and never a substitute for an independent check. → Never Marks Its Own Homework.

AI observability

Instrumentation that measures whether answers are correct — grounded, relevant, unbiased, non-drifting — rather than whether the system ran. → Uptime Is Not Correctness.

Drift

Gradual divergence of output quality from its baseline, usually caused upstream by stale sources or changed schemas rather than by the model. → Uptime Is Not Correctness.

Fine-tuning

Further training a base model on task-specific examples, changing its weights. Contrast prompting and RAG, which change only what it sees at inference. → The Model Is The Easy Part.

LoRA

Fine-tuning by training small adapter layers instead of the whole model — most of the benefit at a fraction of the cost.

Token

The sub-word unit a model actually reads and is billed for. Token counts never match word counts, which is why limits and costs surprise people.

Tokenization

Splitting text into tokens. The reason "unbelievable" can cost three units and why exact strings sometimes behave oddly.

Temperature

The sampling-spread dial. Higher does not mean more creative so much as more scattered — on factual work it means more errors.

Attention

The mechanism weighing how much each token matters to each other token. The source of both the model's power and its finite attention budget. → The Attention Budget.

KV cache

Stored attention state that spares the model recomputing earlier tokens. It grows linearly with sequence length and is the main memory bottleneck at inference.

Prefill vs decode

Prefill processes the prompt in parallel and is compute-bound; decode emits tokens one at a time and is memory-bound. Why long prompts and long answers scale differently.

Speculative decoding

A small draft model proposes tokens that a larger model verifies in one pass — same output, less waiting.

Quantization

Reducing weight precision (32-bit to 8- or 4-bit) for smaller, faster, cheaper inference, at some accuracy cost.

Knowledge distillation

Training a smaller student model to reproduce a larger teacher's behaviour.

Mixture of Experts (MoE)

Routing each input to a few specialist sub-networks instead of the whole model — capacity without proportional compute.

RLHF

Reinforcement learning from human feedback: teaching a model which answers people prefer. Why models are agreeable, and why agreeableness needs guarding against. → Flattery Is Failure.

Inference-time compute

Spending more reasoning steps at answer time rather than more training beforehand. The 'think longer' dial.

Primacy en recency

Wat vooraan en achteraan in een reeks staat wordt het best onthouden; het midden het slechtst. Geldt voor mensen (Glanzer & Cunitz, 1966) én voor lange contexten. → The Middle Is Where Things Are Lost.

Episodic memory

Wat er eerder GEBEURDE — samengevatte sessies die later worden teruggehaald voor continuïteit. Iets anders dan feiten. → Retrieval Is Not Memory.

Semantic memory (agent)

De feiten en relaties die een agent over jou of je domein bewaart, meestal als key-value of graaf. Stabiele kennis, geen gespreksgeschiedenis.

Procedural memory

Geleerde toolsequenties: niet wat waar is, maar hoe iets werkte. De laag die van een agent een vakman maakt in plaats van een lezer. → The Demo Agent And The Production Agent.

Working memory (scratchpad)

De state van de taak die nu loopt. Verdwijnt na afloop, tenzij expliciet gepersisteerd.

Context fork

Een skill of subagent in een geïsoleerd contextvenster draaien, zodat de hoofdcontext schoon blijft. Waarom een skill iets anders is dan een lange regel in CLAUDE.md. → The Attention Budget.

Stateless (MCP)

Onafhankelijke requests met metadata in headers, in plaats van sessie-gebonden verbindingen. Maakt serverless en edge mogelijk; vraagt dat connection-state naar externe stores verhuist. → Mcp Is Plumbing Not Magic.

Capability negotiation

Client en server spreken bij het verbinden af wat elk ondersteunt, zodat versieverschillen geen stille fouten worden.

DAG (agent-planning)

Gerichte acyclische graaf van deelstappen met opgeloste afhankelijkheden — het verschil tussen een plan en one-shot redeneren. → Keep The Work That Was Done.

Circuit breaker

Stopt aanroepen naar een falende dienst zodra een drempel is bereikt, om cascadefouten te voorkomen. In agent-lussen het tegengif tegen retry-spiralen. → Keep The Work That Was Done.

Saga (compensatie)

Een gedistribueerde transactie opgedeeld in lokale stappen die elk een compenserende actie hebben. Het antwoord op 'wat als stap 4 van 6 faalt'.

Human-in-the-loop gate

Een poort die alleen opengaat na menselijke goedkeuring, geplaatst op de onomkeerbare stap in plaats van bij elke actie. → Escalate By Exception.

Security & governance

SSRF (Server-Side Request Forgery)

An attack that tricks an agent into fetching internal URLs (e.g. cloud metadata) on the attacker's behalf. Why an agent's reach must be governed. → article Governing What Your AI Can Touch.

Prompt injection

Malicious instructions hidden in content the AI reads, hijacking its next action. The reason least-privilege matters for agents. → article Governing What Your AI Can Touch.

Least privilege

Granting only the minimum access needed, just in time; deny beats ask beats allow. The guard on an agent's power. → article Governing What Your AI Can Touch.

Least privilege (for agents)

Granting an agent only the tools and permissions a task requires — read-only first, write later, human confirmation on the irreversible step. → Content Is An Attack Surface · Guardrails Are Layers.

Governance vs guardrails

Guardrails constrain an action at runtime; governance is the accountability layer above — who decided, who may access, what is auditable. Routinely conflated. → Guardrails Are Layers.

Web & publishing

Static-site generator

A website compiled to plain files at build time (e.g. Astro), served with no runtime server or CMS or database. The mechanism behind a brain that publishes itself. → A Brain That Publishes Itself.

rel=canonical

An HTML link that declares the original home of a piece: publish on your own site first, then point syndicated copies (Medium) back to it, so Google credits you. → One Way Publishing.

Model Outlives the Tool

The principle that your data model / structure survives every tool that renders it (diagramming app, BI tool, the AI itself). Folded into → Rent The Ai Own The Structure.

A volatile, promotable layer: a term that keeps recurring or earns its own article gets promoted to a concept. The coined concepts live in the concept library.