How-to guides
Task-shaped recipes. Each one assumes you have a working install and starts
from a concrete goal rather than from a concept — for the concepts, see
Decisions; for the shape of a type, see
Build and extract
- Author a domain schema — the six bundled
schemas are a starting point, not a limit. Covers the YAML, what each field
does to the prompt, and the thing that surprises everyone: a schema shapes
the prompt and validates nothing, so an off-schema entity comes back
unflagged.
- Harden model calls — retry with jitter, rate
limiting, circuit breaking, and caching, all composed over the Cache port
so a single process and a fleet behave the same way.
- Tune ingestion throughput —
chunk_size
and concurrency interact, so tuning either alone gives the wrong answer:
the arithmetic that decides how many calls actually run at once, and the
chunk size below which extraction starts inventing duplicate identities.
Search and read
- Retrieve entities — a query string to ranked
entities, fusing a semantic channel over the vector store with a lexical
one over blocking keys. Read it for what the two score scales mean and for
the recall blocking costs you.
Keep the graph honest
the guide to read second. A populated graph through blocking, scoring,
banding and adjudication, to a merge you can audit and reverse. This is the
step whose absence makes a knowledge graph quietly wrong.
- Query a timeline — temporal extents, the interval
relations inferred between them, and time-sliced reads.
Storage and projections
- Use the write model — the aggregates, the
three events, and how to emit rather than write.
what to do instead of build_graph when you have a real event store.
- Rebuild a projection — wipe and replay, which
is the payoff for extraction writing to no store.
- Index documents without extracting them — build a
corpus of passages with no model call, what an empty entity_ids means, and
the extract-then-index case that is lossy.
- Use the pgvector store — schema, the score
expression, and why there is deliberately no ANN index.
- Implement a store adapter — writing a
third GraphStore or VectorStore, and pointing the compliance suite at it
so you find out whether you got the contract right.
Development
everything the default suite deliberately leaves out, including the two
invocation constraints that produce dozens of failures reading as flakiness.
- Run the ingestion benchmark — wall-clock and
accuracy against a live endpoint, what it refuses and why, and the exit
code table.
- Bootstrap a project with SpecOps — initialize or adopt
SpecOps Project Management as Code with opinionated invariants and Diataxis docs.