Corpus Curator
Organizations lose operational knowledge continuously, and most of the attention and budget in the current AI cycle goes to selecting a model rather than to the material the model is meant to work from. That material is spread across systems and people, and maintaining it is slow work that is usually not assigned to anyone. A Corpus Curator is the person who does it.
The discipline
A Corpus Curator does not produce new knowledge. It already exists, spread across reports, dashboards, investigation notes, SOPs, whatever record was kept of past decisions, and the memory of whoever has been there longest. The work is collecting it, determining what is still true, and putting it somewhere the next person who needs it can retrieve it.
It has more in common with engineering than with technical writing. A knowledge base is written on the assumption that a person will read it from the top. A corpus gets queried in fragments and out of order, often by software that has none of the surrounding context the author took for granted. Building for the second case changes both what you write and how you structure it.
The DoubleCs Method
The work generally follows six steps, from locating the source material to establishing a process the team can maintain.
Discovery
Finding where the knowledge actually lives: which reports, which systems, which documents, and which people. In my experience the more useful half of this step is identifying what was never written down at all, because someone has been carrying it and nobody has noticed.
Collection
Pulling the sources into one place. Exports, SOPs, investigation notes, reports, dashboards, whatever record exists of why earlier decisions were made, and the things that currently exist only because a particular person will explain them if asked.
Curation
This is the step that requires judgment. Sources may contradict one another, a written procedure may no longer reflect actual practice, or an important decision rule may never have been documented at all. Deciding what belongs requires understanding how the operation actually works, which is why this step is difficult to reduce to formatting or file organization.
Structure
Organizing it so it can be retrieved. Consistent formats, explicit relationships between related documents, and enough surrounding context that a person or a system reading one piece of it can tell what it is looking at and when it was last confirmed.
Validation
Confirming it is current and correct. A corpus that is confidently wrong causes more damage than no corpus, because people stop verifying what it tells them.
Operational Intelligence
At the end you have a corpus your team maintains without me, and a documented process for keeping it current as people, systems, and procedures change.
Knowledge decays on its own
People leave and take context with them. A system gets replaced and the documents that referenced it stop making sense, though nothing marks them as out of date. The reasoning behind a decision is rarely recorded anywhere, so when the same question returns it gets reconstructed or guessed at. None of this is scheduled and none of it is visible until someone needs an answer that used to be straightforward. Curation is the maintenance work that keeps that decay from accumulating.