Title: A Living Glossary for Agent-Driven Development

URL Source: https://blog.hem.si/blog/2026-09-13-glossary-for-agent-development/

Markdown Content:
---
title: A Living Glossary for Agent-Driven Development
description: Natural language becomes code changes, so humans and agents need the same words: definitions in source comments, one generated glossary, references that fail the build when they dangle.
image: https://blog.hem.si/blog-placeholder-1.jpg
---

 13.09.2026 

# A Living Glossary for Agent-Driven Development

[Get Site as Markdown ](https://markdown.new/https://blog.hem.si/blog/2026-09-13-glossary-for-agent-development/) 

---

The more I build with agentic coding environments, the more precision matters — even though, or rather exactly because, I work in natural language. Every sentence I write can turn into a source-code change, so vague wording turns into vague diffs.

Describing things exactly is harder than it looks. And a glossary I maintain by hand rots within weeks: I rename something in code and forget the docs, or the docs describe behavior from three versions ago.

How nice could it be if to just type `jl:domain.output.jxlinfo-sidecar` and a feature/view/visuell element could be mentiond without any contradiction.

So I built a small system that is easy to maintain because it mostly maintains itself: keywords are defined once, directly at the implementation in source comments, referenced everywhere else with a fixed syntax, and a generator extracts them into one always up-to-date glossary file. The result feels like a living glossary — I can name, describe, and reference every concept exactly, together with the AI. Dot notation gives me levels, a project-owned prefix keeps keys stable and even addressable across projects. It goes as far as invisible references inside Markdown docs and clickable key links in GitHub issues.

> I’m sure that a lot of people have already built different approaches to living glossaries for agent-driven development. This is just my approach.

I am still experimenting with this; if you built something similar — or better — I want to hear about it.

## Living example

The system runs in production on its own first subject — the glossary below is generated, not hand-written:

* [dhcgn/jxleet docs/GLOSSARY.md](https://github.com/dhcgn/jxleet/blob/main/docs/GLOSSARY.md) (44 keys on 13.09.2026)

What to look at: the `<!-- generated - do not edit -->` header, the facet chapters (`jl:domain`, `jl:view`, `jl:tech`, `jl:build`), and the ref-count column — count 0 marks orphans, high counts mark hot concepts like `jl:view.queue`. Every description there traces back to one source comment at the implementing code.

## Prerequisites

* Go toolchain (the extractor and generator are Go; my repo also uses `go generate`).
* The example repo: [dhcgn/jxleet](https://github.com/dhcgn/jxleet), generated file `docs/GLOSSARY.md` (44 keys on 13.09.2026).
* The reusable skills: [dhcgn/personal-dev-skills](https://github.com/dhcgn/personal-dev-skills/tree/main/.apm/skills) (`glossary-init`, `glossary-create`, `glossary-search`, plus `glossary-create-issue`).
* Optional: `rg` (with `git grep` fallback) for the search scripts.

## How it works

### 1\. Define once, in code, at the implementation

Every user-visible element, process step, and feature gets one `Key` \+ `Description` pair. The key starts with a project prefix (`jl:` for jxleet), then dot-namespaced lowercase segments:

* `jl:domain.route.transcode`
* `jl:domain.output.replace`
* `jl:view.queue`
* `jl:tech.update.notify-only`

The first segment is a mandatory facet (`domain`, `view`, `tech`, `build`) that becomes a chapter in the generated file; facets and their order live in a small hand-edited `docs/GLOSSARY.topics.yaml`. Keys are stable once merged — a rename updates all references in the same change.

Short terms use a single line in the file’s native comment syntax:

```go
// jl:domain.route.transcode=JPEG repacked losslessly with --lossless_jpeg=1; djxl restores the original byte for byte.
```

```yaml
# jl:domain.preset.rule-fallback=Trailing "*" rule; without it unmatched files are skipped and reported.
```

Longer descriptions use a `---`\-fenced block:

```go
// Finalize commits the encoder's temp output ...
//
// ---
// jl:Key: jl:domain.output.replace
// jl:Description: >
//
//	Write a temp file in the target directory, decode-verify it ...
//
// ---
func Finalize(...) ... {
```

Supported comment styles: Go/TypeScript `//`, Svelte `//` \+ `<!-- -->` \+ `/* */`, CSS `/* */`, YAML/PowerShell `#`, Markdown `<!-- -->`. Only comment payloads match — never string literals or plain prose. Test fixtures, vendored dirs, and the generated file itself are excluded.

### 2\. Reference it with `ref:`, everywhere else

Definitions answer “what is it called”; references answer “we mean _that_”:

```go
// Each value names a preset file (ref:jl:domain.preset).
```

```md
Every file takes one `ref:jl:domain.route`, never a bare format name.
```

Code is scanned in comments only; Markdown is scanned as full prose. A reference with no matching definition lands in a **Broken references** section of the generated file — and generation exits non-zero, so dangling pointers can never merge silently. Refs to bare facets (`jl:domain`) never resolve; only full defined keys do.

Docs get the same mechanism as invisible tags — one HTML comment per section, zero visual clutter:

```html
<!-- ref:jl:domain.route.transcode -->
```

### 3\. Generate, and let the gates keep it fresh

A small generator walks the repo, merges definitions, groups by facet chapter, sorts alphabetically, and writes `docs/GLOSSARY.md` with Key + **ref-count** \+ Description tables. The ref-count shows how many `ref:` pointers aim at each key, so orphans (count 0) and hot concepts stand out.

Staleness fails the gate: a test regenerates into memory and diffs against the committed file, so `go test ./...` fails when someone edits a comment without regenerating:

```bash
go generate ./...
go test ./...
```

Nothing here is Go-specific in principle: the generator is a small text walk (collect comment payloads, match key patterns, render tables), so a script in whatever language you prefer — Python, PowerShell, TypeScript — does the same job. For me it was a close call; I picked Go’s built-in `go generate` because the directive lives next to the code it regenerates and needs no extra runner wired into CI.

### 4\. Use it with agents and in issues

Three skills chain the workflow (`glossary-init` → `glossary-create` → `glossary-search`):

* `glossary-init` bootstraps a repo from zero (copy extractor, pick prefix, topics file, seed 3–5 definitions, wire gates, extend `AGENTS.md`).
* `glossary-create` is the daily convention: define once at the anchor file, `ref:` everywhere else.
* `glossary-search` ships read-only finder scripts with identical behavior in both shells — `scripts/glossary-search.sh` and `scripts/glossary-search.ps1` — so humans and agents validate and navigate every key definition and `ref:` the same way:

```bash
scripts/glossary-search.sh defs                                  # all definitions
scripts/glossary-search.sh refs                                  # all ref: mentions
scripts/glossary-search.sh usages jl:domain.output.replace       # everything about one key
scripts/glossary-search.sh keys                                  # unique tokens, diff against the .md table for gaps
```

```powershell
scripts/glossary-search.ps1 defs
scripts/glossary-search.ps1 refs
scripts/glossary-search.ps1 usages jl:domain.output.replace
scripts/glossary-search.ps1 keys
```

Output is plain `file:line:text`; the scripts need only `rg` (preferred) or `git` on `PATH`, and already exclude test fixtures, `node_modules`, and the generated file itself.

The prefix is a parameter (`glossary.New("acme")`), so other repos reuse the same tooling with their own namespace — that is what makes keys referenceable across projects.

In GitHub issues, every key becomes clickable through a repo-scoped code-search link, so the reader sees exactly what I mean instead of my paraphrase:

`https://github.com/search?q=repo%3Adhcgn%2Fjxleet+%22jl%3Aview.queue%22&type=code`

> I know this is something every LSP (Language Server Protocol) is able to handle, this is just my lightweight, language-agnostic approach that works across different shells and CI environments.

## Verification

```bash
go generate ./...
git diff --exit-code -- docs/GLOSSARY.md
go test ./...
```

* No diff: comments and glossary are in sync.
* Tests green: no stale output, no duplicate definitions with differing text, no broken `ref:` pointers.
* Spot check without the generator: pick a key in `docs/GLOSSARY.md` and confirm every pointer resolves — `scripts/glossary-search.sh usages <key>` (bash) or `scripts/glossary-search.ps1 usages <key>` (PowerShell).

## Limitations

* Tested in a Go-led repo (Go, TypeScript/Svelte frontend, YAML, PowerShell, Markdown). Other stacks reuse the parsing ideas but need their own walk.
* Facet chapters are hand-curated in `GLOSSARY.topics.yaml`; keys with an unknown facet render under `Uncategorized` with a warning instead of failing — strictness is a deliberate follow-up.
* Wording of prose is convention, not CI-gated: `AGENTS.md` declares the glossary normative for _naming_ (SHOULD-use keys), agents ask and fix the description on ambiguity.
* Early-stage experiment: the skill set and the extractor will keep evolving; check the repos for the current state.

## Links

* <https://github.com/dhcgn/jxleet/blob/main/docs/GLOSSARY.md>
* <https://github.com/dhcgn/personal-dev-skills/tree/main/.apm/skills>
* <https://agentskills.io>