Domain schema YAML reference

A domain schema is a YAML file describing the entity types, relationship types

and prompt template that redstring uses when extracting a knowledge graph from

one kind of content. Six are bundled with the package, under

src/redstring/extraction/domains/schemas/; you can also load your own from

any directory.

Every file is parsed with yaml.safe_load and then validated against the

DomainSchema pydantic model in

src/redstring/extraction/domains/models.py. That model is the authority for

everything on this page: field names, defaults, bounds, and the rules that

reject a file. Each model in the hierarchy sets extra="forbid" and

str_strip_whitespace=True, so an unrecognised key is an error and surrounding

whitespace never reaches the graph.

The hierarchy is:


DomainSchema
β”œβ”€β”€ EntityTypeSchema (list, β‰₯1)
β”‚   └── PropertySchema (list)
β”œβ”€β”€ RelationshipTypeSchema (list, β‰₯1)
β”œβ”€β”€ ConfidenceThresholds
└── extraction_prompt_template

A schema prompts the extractor; it does not constrain it. Entity and

relationship types the LLM returns outside the declared set are not discarded β€”

custom is always an acceptable entity type and related_to is always an

acceptable relationship type. See

ADR 0011 for the

reasoning, and note the consequence throughout this reference: most fields here

shape what the model is asked for, and only a few are enforced at load time.

This page describes what each field means and what will be rejected. For a

step-by-step walk through writing a new schema, see

Author a domain schema. For where

domains fit in the extraction call, see the

README.

Scope and audience

This is a reference, not a walkthrough. It is written for someone who is

authoring or reviewing a domain schema YAML file, or debugging a

SchemaLoadError, and wants a per-field answer without reading

models.py. It assumes you already know what you want the schema to say and

covers what the file format allows you to say it with.

In scope:

RelationshipTypeSchema and ConfidenceThresholds β€” type, requiredness,

default, and bounds

errors

exposes it

enforce, and where those are checked instead

(tests/unit/extraction/domains/test_yaml_schemas.py)

and the derived views (DomainSummary, ClassificationResult) built from

one

Out of scope:

Author a domain schema, which also

covers turning the loaded schema into a system prompt via

domain_system_prompt and handing it to extraction.

argued in

ADR 0011. It is

restated here only where it changes what a field does.

contact with them is the prompt text it produces; see the

README for the pipeline as a whole.

you are a caller-side concern. This page describes ClassificationResult

only as a value a schema author will see, never authors.

Two entry points are part of the public surface β€” load_schema_from_file and

load_schema_from_string are exported from redstring, alongside

DomainSchema itself. Everything else named here (the registry,

load_all_schemas, validate_schema_file, SchemaLoadError) is reached by a

dotted path into redstring.extraction.domains, which makes it internal by

the rule in

ADR 0006: documented so you can

use it knowingly, not promised.

Where schema files live and how they are discovered

The bundled directory

The six schemas that ship with the package live in

src/redstring/extraction/domains/schemas/, one file per domain:


academic_research.yaml
business_corporate.yaml
encyclopedia_wiki.yaml
literature_fiction.yaml
news_journalism.yaml
technical_documentation.yaml

That directory is computed once, at import, as `Path(__file__).parent /

"schemas" in loader.py, and exposed as get_schema_directory()`. It is the

default for every loader and registry entry point that takes a directory. There

is no environment variable and no configuration file that changes it: a

different directory is a caller argument (schema_dir=) or nothing. The

registry used to read DOMAIN_SCHEMA_HOT_RELOAD from the environment and no

longer does, for the reason recorded in

tests/unit/test_library_reads_no_environment.py β€” a library's disk-access

behaviour should not depend on the shell that started the process.

The scan

load_all_schemas(schema_dir=None) is the discovery mechanism. Given a

directory it:

  1. returns {} β€” with a warning logged, not an exception β€” if the directory

does not exist; raises SchemaLoadError if the path exists but is not a

directory

  1. globs .yaml and then .yml, non-recursively; a schema in a

subdirectory is not found

  1. sorts the combined list of paths, so load order is deterministic and does

not depend on filesystem order

  1. loads each file through load_schema_from_file, which reads it as UTF-8 and

parses it with yaml.safe_load β€” never yaml.load, so no YAML tag can

construct a Python object

  1. keys the result by the domain_id inside the file

Files with any other extension are ignored entirely. A README.md, a

.yaml.bak, or a schema.json sitting beside your schemas costs nothing.

The filename does not name the domain

The dictionary load_all_schemas returns, and every registry lookup built on

it, is keyed by the domain_id field parsed out of the file β€” not by the

filename. The six bundled files happen to agree with their domain_id, and

following that convention is worth doing for the reader's sake, but nothing

enforces it. A file named mine.yaml declaring domain_id: literature_fiction

loads as literature_fiction.

The consequence to plan around is collisions. Two files in one directory

declaring the same domain_id is an error:


Duplicate domain_id 'literature_fiction' found in /path/to/second.yaml.
Already loaded from another file.

Because the scan is sorted, the file that loads first β€” and therefore the one

reported as the duplicate β€” is the alphabetically later of the two, stably. By

default this raises SchemaLoadError; with ignore_errors=True it is logged

as a warning and the later file is skipped, leaving the first one registered.

Loading one file

load_schema_from_file(file_path, schema_dir=None) takes either form of path:

bundled directory β€” so load_schema_from_file("literature_fiction.yaml")

reads a bundled schema, and the same call with schema_dir=Path("./schemas")

reads yours

Nonexistent paths and directories-passed-as-files each raise SchemaLoadError

before any parsing happens, so you get "Schema file not found: …" rather than a

YAML error. load_schema_from_string(yaml_content, source_name="")

skips the filesystem altogether; source_name appears in the error messages

and in SchemaLoadError.file_path, so give it something recognisable when you

have one.

Your own directory

Nothing about a custom directory is second-class β€” the bundled schemas are

loaded by exactly the calls you would make:


from pathlib import Path

from redstring.extraction.domains import load_all_schemas
from redstring.extraction.domains.registry import DomainSchemaRegistry

# One-shot: a plain dict of domain_id -> DomainSchema
schemas = load_all_schemas(Path("./my-schemas"))

# Or point the registry at it (see "Loading and registry behaviour" below)
registry = DomainSchemaRegistry.get_instance(Path("./my-schemas"))

Two things to know before you do:

registry pointed at ./my-schemas sees only what is in ./my-schemas,

including for the default-schema lookup, which looks for encyclopedia_wiki

by id and falls back to whichever schema loads first. Copy the bundled files

in if you want both.

single malformed file aborts the whole scan with the offending path on

SchemaLoadError.file_path. ignore_errors=True degrades to loading what it

can β€” appropriate for a long-running process reloading schemas, not for

checking a directory is sound. To check one file deliberately, use

validate_schema_file, which returns (is_valid, error_message) rather than

raising.

get_available_domain_ids(schema_dir=None) is the cheap "what is here?" call:

it scans with ignore_errors=True and returns the sorted ids. Its docstring

describes it as only parsing domain_id from each file; it does not β€” it

fully loads and validates every file, so invalid schemas are silently absent

from its result rather than reported.

Minimal complete example

This is the smallest file that loads. Every key shown is required; nothing here

can be removed, and everything not shown has a default.


domain_id: recipes
display_name: Recipes
description: Cooking recipes, their ingredients, and the techniques they use.

entity_types:
  - id: dish
    description: A prepared dish that a recipe produces.
  - id: ingredient
    description: A single ingredient a recipe calls for.

relationship_types:
  - id: uses
    description: A dish uses an ingredient.
    valid_source_types: [dish]
    valid_target_types: [ingredient]

extraction_prompt_template: |
  You are analyzing a cooking recipe.

  Extract the following types of entities:
  {entity_descriptions}

  Extract the following types of relationships:
  {relationship_descriptions}

Loading it:


from pathlib import Path

from redstring import load_schema_from_file

schema = load_schema_from_file(Path("recipes.yaml"))

What the model filled in

The loaded DomainSchema has more on it than the file says, because three

fields are optional and defaulted:

Field Value after load Where the default comes from
version "1.0.0" DomainSchema.version default
confidence_thresholds.entity_extraction 0.6 ConfidenceThresholds default
confidence_thresholds.relationship_extraction 0.5 ConfidenceThresholds default
every entity type's properties [] default_factory=list
every entity type's examples [] default_factory=list
every relationship type's valid_source_types / valid_target_types [] β€” meaning any entity type default_factory=list
every relationship type's bidirectional false RelationshipTypeSchema.bidirectional default
every property's type "string" PropertySchema.type default
every property's required false PropertySchema.required default

So schema.version is "1.0.0" and

schema.confidence_thresholds.entity_extraction is 0.6 for the file above,

and omitting valid_source_types on a relationship is not "unconstrained by

oversight" β€” it is the documented way to say any source is fine.

What is required, and only that

The file above exercises exactly the required set:

and a non-empty description

required keys

min_length=1 on both lists is the whole structural requirement. A schema with

one entity type and one relationship type is valid, and the two-and-one shown

here is already more than the model demands. The "at least 5 of each" rule you

will see in the six bundled files is a repository convention enforced by

tests/unit/extraction/domains/test_yaml_schemas.py, not by DomainSchema β€”

enforced by that test rather than by the model.

Your own schemas are not subject to it.

An example with the optional fields filled in

Everything the format offers, on one entity type and one relationship type:


entity_types:
  - id: ingredient
    description: A single ingredient a recipe calls for.
    properties:
      - name: quantity
        type: string
        description: Amount used, as the recipe writes it.
        required: false
      - name: grams
        type: number
        description: Normalized weight in grams.
    examples:
      - butter
      - saffron

relationship_types:
  - id: substitutes_for
    description: One ingredient can stand in for another.
    valid_source_types: [ingredient]
    valid_target_types: [ingredient]
    bidirectional: true

confidence_thresholds:
  entity_extraction: 0.7
  relationship_extraction: 0.6

version: "1.2.0"

examples is capped at 10 entries per entity type; the eleventh is an error,

not a truncation. type on a property must be one of string, number,

boolean, array, object. version must be three dot-separated integers,

and YAML will read a bare 1.2 as a float, so quote it.

Three things this example is quietly demonstrating

placeholders the prompt builder substitutes.** The template is otherwise free

text. Nothing in DomainSchema checks that they are present β€” a template

without them loads and then produces a prompt that never names your types.

declared in the same file.** This is the one cross-field rule the model

enforces: dish and ingredient are legal above only because both appear in

entity_types. A typo here is a load-time error, not a silently inert

constraint.

extractor is free to return an ingredient as the source of something you

restricted to dish, and to return entity types you never declared β€”

custom and related_to are always accepted. See

ADR 0011. Use

DomainSchema.validate_relationship if you want to check a relationship

against the declared constraints yourself.

A schema is easiest to get right by starting from a bundled one rather than

from this minimum β€”

src/redstring/extraction/domains/schemas/technical_documentation.yaml is the

most fully populated of the six. For the writing process rather than the

format, see Author a domain schema.

Top-level fields (`DomainSchema`)

DomainSchema is the root model: one YAML file is one DomainSchema. It

declares eight fields, six of them required, and model_config sets

extra="forbid" and str_strip_whitespace=True β€” so any ninth key is a load

error, and leading or trailing whitespace on a string field is removed before

validation bounds are applied.

Unlike the three nested models, DomainSchema is not frozen=True.

EntityTypeSchema, PropertySchema and RelationshipTypeSchema are all

frozen and therefore hashable; the root object is mutable in Python. Nothing in

the loader depends on that, and treating a loaded schema as read-only is the

safe habit β€” the registry hands the same instance to every caller.

Field Required Type Default Bounds
domain_id yes string β€” 1-50 chars, pattern ^[a-z][a-z0-9_]*$
display_name yes string β€” 1-100 chars
description yes string β€” 1-500 chars
entity_types yes list of EntityTypeSchema β€” min_length=1
relationship_types yes list of RelationshipTypeSchema β€” min_length=1
extraction_prompt_template yes string β€” min_length=1
confidence_thresholds no ConfidenceThresholds both defaults applied see that model
version no string "1.0.0" pattern ^\d+\.\d+\.\d+$

The subsections below take each in turn.

Which rules live here, and which do not

Three kinds of check apply to a schema file, and only the first is

DomainSchema's:

normalization each nested model's id/name validator performs.

model validator, which requires every entry in a relationship's

valid_source_types and valid_target_types to name an entity type declared

in the same file. There is no other whole-schema rule.

five relationship types, a declared related_to, entity descriptions of at

least 10 characters, a relationship threshold no higher than the entity

threshold, and uniqueness of entity type ids and of relationship type ids.

All of these are asserted in

tests/unit/extraction/domains/test_yaml_schemas.py against the six bundled

files only. A schema of your own declaring the same entity type id twice

loads without complaint; the duplicate simply sits in the list, and

get_entity_type returns the first match.

Two absences are worth stating plainly, because both look like they should be

checked and are not:

only validation. Neither {entity_descriptions} nor

{relationship_descriptions} is required to appear, and no check rejects a

placeholder that the prompt builder does not substitute. See

extraction_prompt_template.

custom for any schema and is_valid_relationship_type accepts

related_to, and neither is consulted at all during extraction β€” nothing

filters the LLM's output down to what you declared. The fields above shape

what the model is asked for. That is the whole of

ADR 0011, and

it is the reason entity_types is bounded below (a prompt needs something to

say) and not above.

Reading the loaded object

The root model exposes six helpers over these fields, all normalizing their

argument with normalize_type_id β€” the same rule that produced the stored

ids β€” before comparing:


schema.get_entity_type_ids()  # ['dish', 'ingredient']
schema.get_relationship_type_ids()  # ['uses']
schema.get_entity_type("Dish")  # EntityTypeSchema | None
schema.get_relationship_type("USES")  # RelationshipTypeSchema | None
schema.is_valid_entity_type("custom")  # True, always
schema.is_valid_relationship_type("related_to")  # True, always

Note that these lookups normalize case and surrounding whitespace only β€” they

do not apply the space- and hyphen-to-underscore rewriting that the id

validators do at load time. get_entity_type("story arc") finds nothing even

though id: story arc in the file would have been stored as story_arc. See

the normalization, exactly and

the one lookup helper.

`domain_id` β€” required, pattern `^[a-z][a-z0-9_]*$`, max 50

The key the schema is known by. It is what load_all_schemas uses as the

dictionary key, what get_domain_schema(...) and registry.has_domain(...)

take, and what DomainSummary.from_schema copies onto the summary. It is not

derived from the filename β€” see [Where schema files live and how they are

discovered](#where-schema-files-live-and-how-they-are-discovered).


domain_id: literature_fiction

Constraints, all declared on the field itself:

Rule Value
Required yes
Type string
Minimum length 1
Maximum length 50
Pattern ^[a-z][a-z0-9_]*$
Whitespace stripped before validation (str_strip_whitespace=True)

The pattern says: **lowercase ASCII letter first, then any run of lowercase

ASCII letters, digits and underscores.** So recipes, literature_fiction,

iso_9001 and x are all accepted. Rejected: Literature_Fiction (uppercase),

3d_printing (leading digit), _internal (leading underscore),

literature-fiction (hyphen), literature fiction (space), cafΓ© (non-ASCII),

and the empty string.

A rejection is a pydantic ValidationError, wrapped by the loader in

SchemaLoadError with the file path attached:


Schema validation failed for /path/to/mine.yaml: 1 validation error for DomainSchema
domain_id
  String should match pattern '^[a-z][a-z0-9_]*$' [type=string_pattern_mismatch, ...]

It is not normalized β€” unlike every other identifier in the file

This is the field's one real surprise, and it runs against the habit the rest

of the format teaches. Entity type ids, relationship type ids, property names

and the entries of valid_source_types / valid_target_types all pass through

a validator that lowercases, rewrites spaces and hyphens to underscores,

collapses repeated underscores and strips the edges β€” so id: Story Arc is

accepted and stored as story_arc. domain_id has no such validator. It has

a pattern instead, and a pattern rejects rather than repairs.

The practical consequence: domain_id: Literature Fiction is a load error,

where id: Literature Fiction on an entity type is not. Write the id in the

form you want it stored. See the normalization, exactly for the rewriting the other fields do.

Whitespace around the value is the exception, and it is handled by the model

config rather than by a validator: domain_id: " recipes " strips to

recipes before the pattern is applied, and loads.

Lookups normalize, so callers get some slack

The registry compares with domain_id.lower().strip() on the argument, not

on the stored key β€” get_schema(" Literature_Fiction ") and

has_domain("LITERATURE_FICTION") both find literature_fiction. That

tolerance is one-directional: it lets a caller pass a scruffy string, and does

nothing to make a scruffy domain_id in a file loadable.

An unknown id from the registry raises SchemaNotFoundError, listing what is

available:


Unknown domain: 'recipies'. Available domains: academic_research, business_corporate, ...

Uniqueness is per-directory, and only checked there

DomainSchema itself has no opinion about uniqueness β€” nothing stops you

constructing two schemas with the same domain_id. The collision is caught by

load_all_schemas, which raises SchemaLoadError (or warns and skips, under

ignore_errors=True) when a second file in the same scan declares an id

already loaded. Two directories, or two separate load_schema_from_file calls,

will happily produce two literature_fiction schemas.

Conventions worth following

None of these are enforced for schemas of your own; the first two are asserted

against the six bundled files by

tests/unit/extraction/domains/test_yaml_schemas.py.

(literature_fiction.yaml declares literature_fiction), and a test asserts

the match for each of them. It makes a stack trace mentioning a path

immediately tell you which domain broke.

breaking change to anyone selecting a domain by name β€” more like a table name

than a label. display_name is the field to edit when the wording is wrong.

over bizcorp. The id shows up in classification output and error messages,

where it is read by people.

`display_name` β€” required, 1-100 chars

The human-readable label for the domain. Unlike domain_id, nothing looks a

schema up by it and nothing parses it: it is free text shown to people.


display_name: Literature & Fiction
Rule Value
Required yes
Type string
Minimum length 1
Maximum length 100
Pattern none
Normalization none
Whitespace stripped before validation (str_strip_whitespace=True)

Those are the only constraints declared on the field. Punctuation, spaces,

mixed case, ampersands and non-ASCII are all fine β€” five of the six bundled

schemas use an ampersand (Literature & Fiction, News & Journalism,

Business & Corporate, Encyclopedia & Wiki), and the sixth is

Technical Documentation. Only two inputs are rejected: the empty string, and

anything over 100 characters after stripping. A whitespace-only value fails

too, because stripping happens first and leaves a zero-length string.

Note the YAML consequence of that punctuation freedom: a value beginning with

& would be read as an anchor, and one containing : as a mapping. Neither

bites the bundled files, but quote the value if it starts with any of

`& * ! % @ { [`` or contains a colon-space.

Where it is actually used

Two places in the package read it, and neither is extraction:

what appears in registry.list_domains() output. See

reading the loaded object.

sorted(self._schemas.values(), key=lambda s: s.display_name). This is a

plain string sort, so it is case-sensitive and orders by code point: a name

starting with a lowercase letter sorts after every capitalised one.

list_domain_ids() sorts by id instead, and the two orderings need not

agree.

It does not reach the prompt. The extraction template is built from entity

and relationship descriptions (see

extraction_prompt_template); no code path

substitutes display_name into it, so rewording it cannot change what the LLM

is asked for. Nor does it participate in lookup, classification or error

messages β€” SchemaNotFoundError lists domain_ids.

Conventions worth following

None of these are enforced. tests/unit/extraction/domains/test_yaml_schemas.py

asserts only that the key is present in each bundled file; it makes no claim

about its content.

set pairs literature_fiction with Literature & Fiction β€” the id is the

machine key, this is the label, and the second exists precisely so the first

can stay ugly and stable.

a target; every bundled name is under 30. It is rendered beside five other

domains in list_domains().

display_name breaks nothing, because no caller selects on it. Changing

domain_id breaks every caller that names the domain.

`description` β€” required, 1-500 chars

One sentence saying what kind of content this domain covers. The field's own

description in models.py states its purpose exactly: *"Domain description

for classification hints."* It is the only free-text field on DomainSchema

that reaches an LLM.


description: Novels, plays, short stories, poetry, and narrative works
Rule Value
Required yes
Type string
Minimum length 1
Maximum length 500
Pattern none
Normalization none
Whitespace stripped before validation (str_strip_whitespace=True)

As with display_name, the only rejections are the empty string, a

whitespace-only value (stripping runs first and leaves length zero), and

anything over 500 characters after stripping. Newlines are permitted, so a

YAML block scalar works β€” but see below for why you probably do not want one.

It is the classifier's entire view of the domain

This is what distinguishes description from display_name, which no LLM ever

sees. ContentClassifier._build_prompt builds the list of candidate domains

by asking the registry for summaries and rendering one line each:


domain_list = "\n".join(f"- {d.domain_id}: {d.description}" for d in domains)

d there is a DomainSummary, and DomainSummary.from_schema copies

description across verbatim. So when a caller extracts with the domain set to

AUTO, the choice between your schema and the five bundled ones is made from

this one string and the domain_id beside it. Nothing else about the schema β€”

not the entity types, not the relationship types, not the prompt template β€”

appears in the classification prompt.

Two consequences follow, and they are the reason to spend a minute on this

field:

bundled schemas all do: `Research papers, academic journals, scientific

studies, and scholarly works; API documentation, code tutorials, software

guides, and technical references; Annual reports, business news, corporate

communications, and financial content`. Each names the artefacts a classifier

would be looking at, because that is the judgement it is being asked to make.

sees all loaded domains in one prompt, so a description that overlaps another

is the failure mode to design against. encyclopedia_wiki is the fallback

and the broadest of the six, so a new domain whose description could plausibly

read as "reference material" will lose to it.

The f"- {d.domain_id}: {d.description}" format is also why a multi-line

description is a poor idea despite being legal: a block scalar's newlines land

inside a bullet list and blur the boundary between one domain's entry and the

next. One line, under about 120 characters, matches every bundled schema.

Where else it surfaces

DomainSummary.from_schema, and therefore in every

registry.list_domains() result. See

reading the loaded object.

prompt_generator from extraction_prompt_template with

{entity_descriptions} and {relationship_descriptions} substituted β€” both

built from the entity type and relationship type descriptions, not this

one. See extraction_prompt_template. Rewording the

domain description cannot change what is extracted once a domain has been

chosen; it can only change whether it is chosen.

Conventions worth following

tests/unit/extraction/domains/test_yaml_schemas.py asserts only that the

description key is present in each bundled file, alongside the other five

required top-level keys. It makes no claim about the content, and there is no

minimum-length convention here β€” the 10-character floor that test enforces

applies to entity type descriptions, not to this field.

unpunctuated noun phrase; consistency matters more than the specific style,

because they are rendered as a list.

above what helps: the longest bundled description is 78 characters.

only as good as its contrast with its neighbours, and adding a seventh domain

can make a description that read well in isolation ambiguous. Nothing checks

this, and a misclassification does not raise β€” it silently extracts with the

wrong entity types.

`entity_types` β€” required, at least 1 (repository schemas require 5)

The list of entity categories the extractor is asked to look for. Each element

is an EntityTypeSchema; the list is what

{entity_descriptions} expands to in the prompt.


entity_types:
  - id: character
    description: A person, being, or personified entity in the narrative
    properties:
      - name: role
        type: string
        description: "Role in story: protagonist, antagonist, supporting, minor"
    examples:
      - Hamlet
      - Lady Macbeth
Rule Value
Required yes
Type list of EntityTypeSchema
Minimum length 1 (min_length=1)
Maximum length none
Uniqueness of id not enforced by the model
Element model frozen=True, extra="forbid", str_strip_whitespace=True

An empty list is the only rejection the field itself makes:


entity_types
  List should have at least 1 item after validation, not 0 [type=too_short, ...]

Per-element rules β€” id normalization and the identifier requirement, the

1-500 character description, the properties list, the 10-entry examples

cap β€” live on EntityTypeSchema and are documented under [Entity type fields

(EntityTypeSchema)](#entity-type-fields-entitytypeschema).

One is enough for the model; five is the bundled convention

min_length=1 is the whole structural requirement, so a schema declaring a

single entity type loads. The "at least 5" in this heading is a repository

convention, asserted only against the six bundled files by

tests/unit/extraction/domains/test_yaml_schemas.py:


assert len(schema.entity_types) >= 5, (
    f"Schema {domain_id} needs at least 5 entity types, has {len(schema.entity_types)}"
)

The bundled files sit well clear of the floor β€” literature_fiction has 7 and

the other five have 9 each β€” and the same test file requires each of their

entity descriptions to be at least 10 characters. Your own schemas are subject

to neither; both are conventions worth borrowing, and neither is a load-time

error. A companion test asserts only that examples is a list on every

bundled entity type, despite its name (test_entity_types_have_examples) β€” it

does not require the list to be non-empty. All six bundled files populate

examples and properties on every entity type anyway.

Duplicate ids load, and the second one is unreachable

Nothing on DomainSchema checks that entity type ids are distinct. Declaring

character twice validates; the list keeps both, and get_entity_type returns

the first match because it scans in order and returns on the first hit.

get_entity_type_ids() will show the duplicate. The prompt will describe the

type twice.

Uniqueness is asserted for the bundled schemas, in the same test file:


entity_ids = [et.id for et in schema.entity_types]
assert len(entity_ids) == len(set(entity_ids)), ...

Watch for a duplicate produced by normalization rather than by copy-paste:

id: Story Arc and id: story-arc both normalize to story_arc, so two

lines that look unrelated in the YAML collide in the loaded schema. See

the normalization, exactly.

It is the target of the one cross-field rule

The set of ids declared here is what validate_relationship_type_references

checks valid_source_types and valid_target_types against β€” it builds

{et.id for et in self.entity_types} and rejects any endpoint naming something

outside it:


Relationship 'uses' references unknown source type: 'dishe'.
Valid types: ['dish', 'ingredient']

Note the direction: entity types are validated by nothing and validate

everything else. Two consequences for editing an existing schema β€” **removing

an entity type breaks every relationship that names it**, at load time and with

a message naming the relationship rather than the removal; and the comparison

is against normalized ids on both sides, so valid_source_types: [Story Arc]

matches id: story_arc and neither spelling has to match the other literally.

What reaches the prompt

domain_system_prompt renders one markdown bullet per entity type, in

declaration order β€” the list is never sorted, so the order you write is the

order the model reads:


- **character**: A person, being, or personified entity in the narrative (examples: Hamlet, Lady Macbeth)
  Properties: role (Role in story: protagonist, antagonist, supporting, minor)

Three things about that rendering are worth knowing while authoring the list:

src/redstring/extraction/prompt_generator.py is 3, while the model permits

  1. Examples 4 through 10 are stored on the schema and never shown to the

extractor; put the most disambiguating ones first.

one, and by name alone when it does not. There is no cap here, so a type with

fifteen properties spends fifteen names of prompt on itself.

the property and describes it; nothing tells the model that grams is a

number or that a property is mandatory. If that matters, say so in the

property's description.

The list does not constrain the result

The extractor is not restricted to what you declare. is_valid_entity_type

accepts custom for any schema, and nothing in the pipeline filters an

extracted entity against get_entity_type_ids() β€” an entity type you never

wrote is not discarded. That is the design in

ADR 0011, and it

is why this field is bounded below and not above: the list exists so the prompt

has something to say.

Conventions worth following

order, and the first bullets are the ones a model weights most.

crowding out the relationship list. A schema with thirty types is usually two

domains.

prompt, and examples do more to disambiguate a type than a longer

description does.

bundled floor, and the string lands verbatim in the prompt after the type

name β€” write it to be read there.

`relationship_types` β€” required, at least 1 (repository schemas require 5)

The list of edge kinds the extractor is asked to look for. Each element is a

RelationshipTypeSchema; the list is what {relationship_descriptions}

expands to in the prompt.


relationship_types:
  - id: loves
    description: Romantic love between characters
    valid_source_types: [character]
    valid_target_types: [character]

  - id: related_to
    description: General relationship
    bidirectional: true
Rule Value
Required yes
Type list of RelationshipTypeSchema
Minimum length 1 (min_length=1)
Maximum length none
Uniqueness of id not enforced by the model
Element model frozen=True, extra="forbid", str_strip_whitespace=True

An empty list is the only rejection the field itself makes:


relationship_types
  List should have at least 1 item after validation, not 0 [type=too_short, ...]

Per-element rules β€” id normalization, the 1-500 character description, the

two endpoint lists and bidirectional β€” live on RelationshipTypeSchema and

are documented under Relationship type fields.

One is enough for the model; five is the bundled convention

min_length=1 is the whole structural requirement. The "at least 5" in this

heading is a repository convention, asserted only against the six bundled

files by tests/unit/extraction/domains/test_yaml_schemas.py:


assert len(schema.relationship_types) >= 5, (
    f"Schema {domain_id} needs at least 5 relationship types, has {len(schema.relationship_types)}"
)

The bundled files clear the floor comfortably: 11 for

technical_documentation, 12 for academic_research and news_journalism,

13 for business_corporate, 14 for encyclopedia_wiki, and 19 for

literature_fiction. Two other conventions from that file apply here and to

nothing else in the format β€” every bundled schema must declare a related_to

relationship type, and every relationship type must have a non-empty

description (no minimum length, unlike the 10-character floor on entity

descriptions). None of the three binds a schema of your own.

`related_to` is the fallback, and declaring it is a convention

is_valid_relationship_type accepts related_to for any schema, declared or

not, and validate_relationship short-circuits on it: an unknown type named

related_to returns (True, None) without any endpoint check, because there

is no RelationshipTypeSchema to check against. So declaring it buys you

nothing at the model level.

What it buys is prompt text. A declared related_to gets a bullet in

{relationship_descriptions} telling the model the escape hatch exists; an

undeclared one is silently accepted after the fact and never offered. All six

bundled schemas declare it the same way β€” a one-line description, no endpoint

constraints (so any pair of entity types is legal), and bidirectional: true

β€” and a test asserts its presence in each. Put it last, where declaration

order puts it last in the prompt.

Duplicate ids load, and the second one is unreachable

Nothing on DomainSchema checks that relationship type ids are distinct.

Declaring cites twice validates; both stay in the list,

get_relationship_type returns the first match, and the prompt describes

the type twice. Uniqueness is asserted for the bundled schemas

(test_no_duplicate_relationship_type_ids), against the loaded model rather

than the raw YAML β€” so it catches collisions produced by normalization as well

as by copy-paste. id: Depends On and id: depends-on both normalize to

depends_on. See the normalization, exactly.

The endpoint lists are the one thing validated across fields

validate_relationship_type_references runs after the whole model is built and

requires every entry of every valid_source_types and valid_target_types to

name an entity type declared in the same file:


Relationship 'uses' references unknown source type: 'dishe'.
Valid types: ['dish', 'ingredient']

Both sides are compared after normalization β€” the endpoint lists get the same

lowercase/underscore rewriting the entity ids do β€” so `valid_source_types:

[Story Arc] matches id: story_arc` without the spellings agreeing literally.

An empty list means any entity type, and that is the default; the check skips

empty strings, so a stray - "" in the list is ignored rather than rejected.

The direction of the dependency matters when editing: relationship types point

at entity types and nothing points back. **Removing an entity type breaks every

relationship that names it**, at load time, with a message naming the

relationship.

What reaches the prompt

domain_system_prompt renders one bullet per relationship type, in

declaration order β€” the list is never sorted:


- **loves**: Romantic love between characters (from: character; to: character)
- **related_to**: General relationship (bidirectional)

The parenthetical is assembled from whichever of the three is present:

from: … when valid_source_types is non-empty, to: … when

valid_target_types is, and the bare word bidirectional when the flag is

set. A relationship with none of them gets no parenthetical at all. Unlike

entity examples, nothing here is capped β€” every relationship type and every

endpoint you list reaches the model, so a list of nineteen spends nineteen

lines of prompt.

The list does not constrain the result

Nothing in the pipeline filters an extracted relationship against

get_relationship_type_ids(), and the endpoint lists are advisory in exactly

the same way: an extractor is free to return loves between two locations.

validate_relationship(relationship_type, source_entity_type, target_entity_type)

exists so a caller can perform that check deliberately, and returns

(is_valid, error_message) rather than raising. See

ADR 0011.

Conventions worth following

depends_on, occurred_in. Every bundled id reads as source verb target,

which is what makes an unconstrained relationship still unambiguous.

empty otherwise.** An empty list is the documented way to say "any", not an

omission, and over-constraining costs you a load error the day you add an

entity type the relationship should have accepted.

sibling_of, competes_with, partners_with, contradicts. It changes the

prompt text and nothing else; no code reverses an edge on the strength of it.

bullet you want read after the specific ones rather than before them.

prompt, so it is the field most able to crowd out the entity descriptions.

`extraction_prompt_template` β€” required, non-empty

The text a model is given before it sees a chunk. It is the schema's only

output that reaches extraction: domain_system_prompt(domain) takes this

string, substitutes two placeholders, and returns the result for

ExtractionPipeline(provider, system_prompt=...).


extraction_prompt_template: |
  You are analyzing a work of literature (novel, play, short story, poem).

  Extract the following types of entities:
  {entity_descriptions}

  Extract the following types of relationships:
  {relationship_descriptions}

  Focus on:
  - Named characters and their roles in the narrative
  - Key themes and motifs
Rule Value
Required yes
Type string
Minimum length 1 (min_length=1)
Maximum length none
Pattern none
Placeholders required none β€” see below
Whitespace stripped before validation (str_strip_whitespace=True)

min_length=1 is the entire validation. The field's docstring in models.py

calls it a "Jinja2-style template", which it is not: no Jinja2 is involved

anywhere in the package, and no template engine of any kind parses it.

Substitution is two literal `str.replace` calls

domain_system_prompt in

src/redstring/extraction/prompt_generator.py does exactly this:


schema.extraction_prompt_template.replace(
    "{entity_descriptions}", _entity_descriptions(schema)
).replace("{relationship_descriptions}", _relationship_descriptions(schema))

Three consequences follow from it being replace rather than str.format or a

template engine, and each is a thing you can rely on:

placeholder, no {domain_id}, no {display_name}. Several tests and the

module docstrings in domains/ use "Extract entities from: {content}" as a

throwaway template; that is a valid schema whose prompt contains the literal

characters {content}, not a supported placeholder. Content reaches the

model as the user message, not through this string.

stray { β€” for instance a JSON example inside your prompt. replace passes

it through untouched, so you may show the model a literal

{"entities": [...]} without escaping anything.

{entity_descriptions} twice renders the entity list twice. Naming it zero

times is legal and renders the template as itself β€” the source comments on

this deliberately: *"a domain whose prompt is entirely prose is a domain

whose author decided the type list was not worth the tokens."* Nothing warns

you, so a typo like {entity_description} produces a prompt that names none

of your entity types and still extracts.

What each placeholder expands to is documented under

extraction_prompt_template; briefly, one markdown

bullet per declared type in declaration order, with the first three examples

and all properties for entities, and endpoint/bidirectional annotations for

relationships.

YAML: use a block scalar, and expect the trailing newline gone

All six bundled schemas write the value as | (literal block scalar), which is

the right choice: it preserves line breaks without requiring quotes or escapes,

and a prompt is multi-line by nature. Two details:

removes it β€” the loaded string ends at the last non-whitespace character.

Leading indentation common to the block is stripped by YAML itself, so the

two-space indent in the file does not reach the model.

Extract entities: names, places` is a YAML error; the block scalar form has

no such problem.

No length rule for your schemas; a 100-character floor for the bundled ones

tests/unit/extraction/domains/test_yaml_schemas.py asserts two things about

each of the six bundled templates and nothing about yours:


assert len(schema.extraction_prompt_template) > 100, ...
assert "{entity_descriptions}" in template, ...
assert "{relationship_descriptions}" in template, ...

That is the only place the two placeholders are required at all β€” the model

does not check for them, so the rule binds the repository's own files and not a

schema you load from your own directory. The bundled templates run 599 to 702

characters, comfortably clear of the floor, and share a shape worth copying:

one sentence naming the content kind, the two placeholder blocks under

imperative headings, then a "Focus on:" list of domain-specific instructions.

Conventions worth following

entity and relationship types you carefully declared reach the model. A

template without them makes the rest of the file inert for extraction, which

is almost never what an author means.

analyzing …". It costs one line and it is what tells the model which reading

to apply to an ambiguous chunk.

bundled order is intro, entities, relationships, then "Focus on:" β€” the

specifics land as elaboration on a type list the model has already read.

pydantic model LlmProvider.extract is given, and a template that describes

a different JSON shape does not change it β€” it only invites output the

mapper cannot read. The deleted generate_json_schema documented in

prompt_generator.py is exactly that failure preserved as a comment.

everything you declared, uncapped for relationships, so the template's own

prose competes with the type lists for the model's attention.

`confidence_thresholds` β€” optional, defaults applied

A pair of floats saying how confident an extraction should be before a caller

takes it seriously. The field is optional; omitting it yields a

ConfidenceThresholds with both defaults, so schema.confidence_thresholds is

never None.


confidence_thresholds:
  entity_extraction: 0.7
  relationship_extraction: 0.6
Rule Value
Required no
Type ConfidenceThresholds (a mapping)
Default ConfidenceThresholds() β€” entity_extraction: 0.6, relationship_extraction: 0.5
Keys entity_extraction, relationship_extraction, both optional
Bounds each 0.0 <= x <= 1.0 (ge=0.0, le=1.0)
Unknown keys rejected (extra="forbid")
Mutability frozen=True β€” unlike DomainSchema itself

Per-key detail is under confidence_thresholds.

Partial mappings are fine

The two keys default independently, so specifying one leaves the other at its

own default:


confidence_thresholds:
  entity_extraction: 0.9    # relationship_extraction is still 0.5

There is no cross-field validator on this model β€” nothing requires

relationship_extraction to be less than or equal to entity_extraction, and

nothing rejects the pair 0.0 / 1.0. Both 0.0 and 1.0 are explicitly valid

bounds, tested as such in tests/unit/extraction/domains/test_models.py.

An out-of-range value is a pydantic ValidationError, wrapped by the loader:


Schema validation failed for /path/to/mine.yaml: 1 validation error for DomainSchema
confidence_thresholds.entity_extraction
  Input should be less than or equal to 1 [type=less_than_equal, input_value=1.1, ...]

Note that these are floats, not percentages: entity_extraction: 70 is a load

error, not seventy percent.

Nothing in the library reads them

This is the field's most important property and the easiest to get wrong.

Grep the package and confidence_thresholds appears in exactly three places:

its definition in models.py, the re-exports in

redstring/__init__.py and extraction/domains/__init__.py, and the tests.

No extraction, merging, mapping or projection code consults it. Setting

entity_extraction: 1.0 does not cause a single entity to be dropped, and

setting 0.0 does not admit one that would otherwise have been filtered.

Two nearby things are separate and easy to confuse with it:

with its own DEFAULT_CONFIDENCE_THRESHOLD, applied to the classifier's

confidence in the domain it picked. It has nothing to do with this field, and

is not read from any schema.

entity gets when the model omits one β€” the midpoint, deliberately not 1.0.

It is a fallback for missing data, not a filter.

So the field is advisory metadata a caller may act on. If you want it

enforced, read it and filter yourself:


threshold = schema.confidence_thresholds.entity_extraction
kept = [e for e in result.entities if e.confidence >= threshold]

This fits the design in

ADR 0011 β€” a

schema describes a domain and prompts for it; it does not police the output β€”

but note the difference in kind from the other fields. entity_types and

extraction_prompt_template at least reach the model as prompt text. These two

numbers reach nothing at all until you use them.

What the bundled schemas set

All six declare the block explicitly rather than relying on defaults:

Domain entity_extraction relationship_extraction
literature_fiction 0.6 0.5
academic_research 0.7 0.6
encyclopedia_wiki 0.7 0.6
technical_documentation 0.7 0.6
business_corporate 0.75 0.65
news_journalism 0.75 0.65

The pattern is a 0.1 gap with the relationship threshold lower, and the level

tracks how much a wrong edge costs: fiction is the most forgiving, financial

and journalistic content the least.

What is checked, and how weakly

tests/unit/extraction/domains/test_yaml_schemas.py has two tests here, and

neither can fail:

0.0..1.0 β€” which the model has already guaranteed with ge/le, so the

assertion is unreachable by construction.

fails when the relationship threshold is the higher of the two. It is a

reporting device, not a gate; a bundled schema inverting the pair would show

up as a skipped test and a green suite.

Treat the "relationship threshold should not exceed the entity threshold"

convention as advice with a reminder attached, not a rule. It binds nothing,

including the bundled files.

Conventions worth following

bundled set. The numbers are a statement about the domain's tolerance for a

wrong extraction, and writing them down is how a reader learns you thought

about it.

cannot be more trustworthy than the two entities it connects, and this is

what the (skipping) test is gesturing at.

reads these, a schema whose thresholds imply strict filtering while the

caller applies none is a lie by omission. Either wire them up at the call

site or leave them at their defaults.

`version` β€” optional, defaults to `1.0.0`, pattern `^\d+\.\d+\.\d+$`

A three-part version string for the schema file itself. It is optional, and

omitting it leaves schema.version == "1.0.0".


version: "1.2.0"
Rule Value
Required no
Type string
Default "1.0.0"
Pattern ^\d+\.\d+\.\d+$
Minimum / maximum length none declared
Normalization none
Whitespace stripped before validation (str_strip_whitespace=True)

The pattern is the whole validation: **three dot-separated runs of digits, and

nothing else.** Accepted: 1.0.0, 0.0.1, 10.20.30, 0.1.0. Rejected, each

of these exercised in tests/unit/extraction/domains/test_models.py: 1.0

(two parts), 1.0.0.0 (four), v1.0.0 (prefix), 1.0.0-beta (semver

pre-release). Despite the field's description calling it "semver format", it is

narrower than semver β€” the -beta and +build suffixes SemVer 2.0.0 permits

are load errors here.

A rejection looks like:


Schema validation failed for /path/to/mine.yaml: 1 validation error for DomainSchema
version
  String should match pattern '^\d+\.\d+\.\d+$' [type=string_pattern_mismatch, input_value='1.0', ...]

Quote it

version: 1.2.0 happens to work β€” YAML reads three-part dotted values as a

string β€” but version: 1.2 is read as the float 1.2, and version: 1.0

as 1.0, neither of which is a string at all:


version
  Input should be a valid string [type=string_type, input_value=1.2, ...]

That error names a type problem rather than a format one, which is confusing

when what you meant was a version. All six bundled schemas write

version: "1.0.0" with quotes; do the same and the failure mode disappears.

Nothing reads it

version appears in exactly three places in the package: its declaration in

models.py, the docstring above it, and the tests. No loader branches on it,

no registry compares it, DomainSummary does not carry it, and no migration

or compatibility check exists to consult it. Bumping it changes nothing about

how the schema loads or what it extracts.

It is therefore **documentation for humans and for whatever versioning

discipline you impose yourself**, in the same category as the confidence

thresholds: recorded on the model, acted on only if a caller chooses to. The

difference is that a caller plausibly will read

confidence_thresholds; there is no code anywhere, in the library or its

tests, that reads version for any purpose but asserting its shape.

The one place it is load-bearing by convention is review. A schema is a prompt,

and a prompt change alters extraction results without altering any code, so the

version is the only marker in the file that says "the graph this produces is

not the graph the previous revision produced." Git history records that too,

but only for people who go looking.

What is checked

tests/unit/extraction/domains/test_yaml_schemas.py::test_schema_version_is_semver

asserts, for each of the six bundled files, that schema.version splits on .

into three parts and that each part is .isdigit(). Both assertions are

already guaranteed by the field's own pattern, so β€” like the confidence-bounds

test beside it β€” this one cannot fail while the model is unchanged. There is no

test that the bundled versions differ from each other, that they ever increase,

or that a change to a schema file is accompanied by a bump.

All six bundled schemas are at 1.0.0 and have never moved.

Conventions worth following

field produces.

for wording.** Anything that changes what the extractor is asked for changes

the graph, and that is the distinction worth recording. Adding an entity type

is a minor; fixing a typo in a description is a patch.

and stored graphs may key on entity and relationship type ids, so removing

one is the schema's breaking change even though nothing in redstring will

say so.

build metadata, so a release process that appends either produces a schema

that will not load.

Entity type fields (`EntityTypeSchema`)

One element of the entity_types list. It names a category of thing the

extractor is asked to find, describes it, and optionally lists properties and

examples. Everything on the model exists to produce prompt text β€” nothing here

filters or validates an extraction result.


- id: character
  description: A person, being, or personified entity in the narrative
  properties:
    - name: role
      type: string
      description: "Role in story: protagonist, antagonist, supporting, minor"
  examples:
    - Hamlet
    - Lady Macbeth
Field Required Type Default Bounds
id yes string β€” 1-100 chars, normalized, must be a valid Python identifier
description yes string β€” 1-500 chars
properties no list of PropertySchema [] none
examples no list of strings [] at most 10 entries

model_config is extra="forbid", frozen=True, str_strip_whitespace=True.

Three consequences:

synonyms; a misspelled descripton fails rather than being ignored.

can put entity types in a set; you cannot repair one after load.

bounds apply, so a value that is entirely spaces fails min_length=1.

The one lookup helper

EntityTypeSchema.get_property(name) returns a PropertySchema or None. It

normalizes its argument the way the name validator does β€” lowercase, strip,

spaces and hyphens to underscores β€” so get_property("Return Type") finds

return_type. It does not collapse repeated underscores or strip the edges,

which the validator does, so a property stored as return_type is not found by

get_property("__return__type__"). This is the more forgiving of the two

lookup styles in the format: DomainSchema.get_entity_type normalizes case and

whitespace only.

Nothing enforces uniqueness or a maximum

Neither EntityTypeSchema nor DomainSchema checks that entity type ids are

distinct, and there is no upper bound on how many you declare or on how many

properties one carries. Uniqueness is a repository convention asserted against

the six bundled files by

tests/unit/extraction/domains/test_yaml_schemas.py; see [entity_types

β€” required, at least 1](#entity_types--required-at-least-1-repository-schemas-require-5)

for what a duplicate does (it loads, and get_entity_type returns the first).

The whole type in one line of prompt

_entity_descriptions in src/redstring/extraction/prompt_generator.py

renders each entity type as one bullet, plus a second indented line when it has

properties:


- **character**: A person, being, or personified entity in the narrative (examples: Hamlet, Lady Macbeth)
  Properties: role (Role in story: protagonist, antagonist, supporting, minor)

Reading that rendering backwards tells you what each field is worth:

MAX_EXAMPLES_PER_TYPE β€” which is 3**, while the model permits 10. Examples

four through ten are stored and never shown.

them, uncapped. A property's type and required flag do not appear

anywhere in the prompt.

An entity type with no examples and no properties renders as the bullet alone,

which is legal and is what the minimal example in this page produces.

Which entity type ids are special

is_valid_entity_type returns True for anything in

get_entity_type_ids() and for the literal custom, whatever the schema

declares. Nothing in the extraction pipeline calls it, so an entity type the

model invents is neither rejected nor recorded as invalid β€” the helper is there

for a caller who wants to ask. See which entity-type ids are special and

ADR 0011.

The other direction is enforced: the ids declared here are the closed set that

valid_source_types and valid_target_types are checked against at load time.

Entity types validate relationships; nothing validates entity types.

Conventions worth following

id is singular, and the type names a kind of thing rather than a collection

of them.

The bundled floor, asserted by test_entity_types_have_descriptions. The

string lands verbatim after the type name in the prompt, so write it to be

read there rather than as a schema comment.

model, so a list of ten is nine characters of YAML per wasted entry. A test

named test_entity_types_have_examples checks only that the field is a

list, so the bundled files' full coverage of examples is habit, not a

gate.

prompt length in proportion to their descriptions and are uncapped, so a type

with fifteen of them crowds out the rest of the schema.

The per-field detail for id, description, properties and examples

follows; for the writing process rather than the format, see

Author a domain schema.

`id` β€” required, normalized, must be a valid Python identifier

The name of the entity type, as the prompt will show it and as

valid_source_types / valid_target_types will refer to it. It is the one

field in the format that is rewritten rather than merely checked.


- id: plot_point
  description: A significant event in the narrative
Rule Value
Required yes
Type string
Minimum length 1 (applied to the input, after whitespace stripping)
Maximum length 100 (applied to the input, before normalization)
Pattern none β€” a normalizing validator instead
Post-condition the normalized value must satisfy str.isidentifier()
Uniqueness not enforced by the model

The normalization, exactly

EntityTypeSchema.validate_entity_type_id in

src/redstring/extraction/domains/models.py runs four steps in this order:

  1. v.lower().strip() β€” lowercase, then trim surrounding whitespace
  2. .replace(" ", "_").replace("-", "_") β€” spaces and hyphens become

underscores

  1. collapse runs: while "__" in normalized: normalized.replace("__", "_")
  2. .strip("_") β€” remove leading and trailing underscores

Then two rejections: an empty result, and a result that is not a valid Python

identifier.

You write Stored as Outcome
character character unchanged
Plot Point plot_point tested in test_entity_type_id_normalization
literary-device literary_device tested in test_entity_type_id_normalization_hyphens
Story Arc (two spaces) story_arc runs collapsed by step 3
__internal__ internal edges stripped by step 4
123invalid 123invalid rejected β€” tested in test_invalid_entity_type_id
--- `` rejected β€” empty after normalization
plot.point plot.point rejected β€” . is not rewritten
cafΓ© cafΓ© accepted β€” isidentifier() is Unicode-aware
class class accepted β€” a keyword is a valid identifier

The last two are the surprises. str.isidentifier() is the whole test, and it

accepts any Unicode identifier character, so cafΓ© and entitΓ© load. It also

accepts Python keywords, because keyword.iskeyword is not consulted β€” an

entity type called class, import or None is legal here. Nothing in

redstring evals or execs an id, so neither is a hazard; they are just not

rejected.

Note also what is not rewritten: only the space and the hyphen become

underscores. A dot, slash, colon or ampersand survives step 2 unchanged and

then fails isidentifier(). So read-only loads and read/only does not.

The two rejection messages

An empty result and an invalid one are distinct errors, both raised as

ValueError inside the validator and surfaced by pydantic as a

ValidationError β€” which the loader wraps in SchemaLoadError with the file

path:


Entity type ID cannot be empty after normalization: '---'
Entity type ID must be a valid identifier: '123invalid' -> '123invalid'

The second message shows both spellings, which is what makes a normalization

failure diagnosable: the arrow tells you what your input became before it was

judged.

The length bounds fail differently, because min_length and max_length are

field constraints applied before the validator runs. id: " " fails as

String should have at least 1 character (whitespace stripping happens first

and leaves nothing), not as a normalization error. And the 100-character

ceiling applies to what you wrote, not to what it normalizes to β€” a

101-character id whose underscores would collapse to 40 characters is still

rejected.

The id you write is not necessarily the id the schema exposes

This is the consequence to keep in mind everywhere else in the file, and it has

three edges:

id: Plot Point means the prompt says plot_point and every lookup key is

plot_point.

It runs normalize_type_id on its argument, so get_entity_type("Character")

and get_entity_type("Plot Point") both find what id: Plot Point stored.

Every lookup on this model shares that one function β€” the same is true of

get_relationship_type, is_valid_entity_type and

is_valid_relationship_type. It was not: they lowercased and stripped only,

so the id you wrote did not find itself (BACKLOG B75, now closed).

RelationshipTypeSchema applies the identical rewriting to every entry of

valid_source_types and valid_target_types, so `valid_source_types:

[Plot Point] matches id: Plot Point` without either spelling being literal.

The cross-field check compares normalized to normalized.

Two ids that look different in YAML can therefore collide: Story Arc,

story-arc and __story_arc__ are one entity type after loading. Nothing

rejects the duplicate β€” the list keeps both entries, get_entity_type returns

the first, and the prompt describes the type twice. See the normalization, exactly for the rule stated once

across all four fields it governs.

It is a prompt token, not a constraint

The id reaches the extractor as the bolded name in one markdown bullet

(- character: …), in declaration order. Nothing downstream filters an

extracted entity against it: is_valid_entity_type accepts custom for any

schema and is not called during extraction at all. The id's only enforced role

is as the target of the relationship endpoint check β€” entity types validate

relationships, and nothing validates entity types. See

ADR 0011.

Conventions worth following

will accept either, but a file whose ids differ from the ids the schema

exposes makes every lookup and every valid_source_types entry a small

translation exercise. All six bundled schemas are already in snake case.

singular lowercase ASCII; the Unicode and keyword allowances exist because

isidentifier() grants them, not because anything wants them.

type ids, so renaming one is the schema's breaking change β€” bump the major

version, and expect nothing in redstring to warn you.

only in case, spacing or hyphenation load as duplicates in silence for your

schemas; the bundled files are protected by

tests/unit/extraction/domains/test_yaml_schemas.py, which asserts

uniqueness against the loaded model and therefore catches this shape.

`description` β€” required, 1-500 chars

What this entity type is. The string lands verbatim in the extraction prompt,

immediately after the bolded type name, so it is instruction text for a model

rather than a comment for a reader.


- id: function
  description: A callable function or method with a specific signature
Rule Value
Required yes
Type string
Minimum length 1 (min_length=1)
Maximum length 500 (max_length=500)
Pattern none
Normalization none β€” unlike id, it is stored exactly as written
Whitespace stripped before validation (str_strip_whitespace=True)

The field's own declaration in models.py describes it as a

"Human-readable description for extraction prompts", which is precisely its

scope. Only three inputs are rejected: a missing key, the empty string, and a

whitespace-only value β€” stripping runs first, so description: " " fails

min_length=1 rather than passing as three spaces. Anything over 500

characters after stripping fails too. Newlines are legal; see below for why

they are a bad idea.

Note that this is a different field from the top-level

description on DomainSchema, which

has the same name and the same bounds but a completely different job: that one

is the classifier's view of the whole domain and never reaches extraction,

while this one reaches extraction and never reaches the classifier.

Exactly where it goes

_entity_line in src/redstring/extraction/prompt_generator.py is the whole

of its use:


line = f"- **{entity_type.id}**: {entity_type.description}"

so the bullet that reaches the model reads:


- **function**: A callable function or method with a specific signature (examples: parse_document, get_entity)

Two things follow. First, a multi-line description breaks the bullet list β€”

the second line is not indented and is not prefixed, so it renders as a

paragraph between two bullets and blurs which type it belongs to. Keep it to

one line. Second, there is no escaping and no markdown processing: the

string is interpolated as-is, so a backtick, an asterisk or a colon in your

description is passed through and read by the model as markdown. That is

usually harmless and occasionally useful (the bundled schemas use a colon in

property descriptions to introduce a value list), but a stray ** will bold

the rest of the line.

A 10-character floor, for the bundled schemas only

tests/unit/extraction/domains/test_yaml_schemas.py::test_entity_types_have_descriptions

asserts, for every entity type in each of the six bundled files:


assert et.description, f"Entity type {et.id} in {domain_id} has no description"
assert len(et.description) >= 10, f"Entity type {et.id} in {domain_id} has too short description"

The first assertion is already guaranteed by min_length=1. The second is

not β€” description: A person is nine characters and loads fine in a schema of

your own. The floor is a repository convention and binds nothing you load from

your own directory. It is the only length convention on any description field:

relationship type descriptions are checked for presence and not for length,

and neither the top-level domain description nor a property description has a

convention at all.

It is prompt text, not metadata

Nothing reads this field except the prompt builder. It does not appear in

DomainSummary, it is not consulted by is_valid_entity_type or

validate_relationship, and no extracted entity is checked against it. Its

entire effect is on what the model is asked to look for, which is the design

recorded in

ADR 0011: the

schema prompts and does not constrain. The practical reading is that editing a

description is a behavioural change to extraction with no test in the

library that can see it β€” which is what the version field exists to record.

Conventions worth following

A callable function or method with a specific signature, not

The function entity type. The id is already printed immediately before it.

band, and the bullet is competing with every other type for the model's

attention.

description is where you tell the model that class is object-oriented and

module is an import path β€” the ids alone do not distinguish them, and

examples only reaches the prompt three entries deep.

continues into the (examples: …) parenthetical, where a period reads

oddly.

`properties` β€” optional list of `PropertySchema`

The structured attributes you want extracted alongside an entity of this type.

Each element is a PropertySchema; the list is optional and defaults to [].


- id: function
  description: A callable function or method with a specific signature
  properties:
    - name: signature
      type: string
      description: Function signature including parameters
    - name: is_async
      type: boolean
      description: Whether the function is asynchronous
Rule Value
Required no
Type list of PropertySchema
Default [] (default_factory=list)
Minimum length none β€” an empty list is valid, and so is omitting the key
Maximum length none
Uniqueness of name not enforced by the model
Element model frozen=True, extra="forbid", str_strip_whitespace=True

The field itself declares no constraints at all: no min_length, no

max_length, no validator. Everything that can reject a properties block is

a per-element rule on PropertySchema β€” a name that is empty or does not

normalize to a valid identifier, a type outside the five allowed literals, a

description over 500 characters, or any key other than those four. Those are

documented under [Property fields

(PropertySchema)](#property-fields-propertyschema).

What a property reaches the model as

_property_hints in src/redstring/extraction/prompt_generator.py is the

whole of the list's use:


", ".join(
    f"{prop.name} ({prop.description})" if prop.description else prop.name for prop in properties
)

and _entity_descriptions emits that as a second, indented line under the

entity bullet β€” only when the list is non-empty:


- **function**: A callable function or method with a specific signature (examples: extract_entities, create_user)
  Properties: signature (Function signature including parameters), is_async (Whether the function is asynchronous)

Three properties of that rendering matter while authoring:

is truncated to the first MAX_EXAMPLES_PER_TYPE (3), nothing here is

dropped. technical_documentation's function type declares five

properties and all five are printed. A type with fifteen spends fifteen

names and fifteen descriptions of prompt on itself, in declaration order.

name and, in parentheses, the description. Nothing tells the model that

parameters is an array or that a property is mandatory. If either matters,

say so in the property's description β€” that is the only channel.

contributes its bare name.

The extractor is not held to it

This is the field's biggest gap between what it looks like and what it does.

The wire model LlmProvider.extract is given carries a free-form

properties: dict[str, Any] per entity, described to the model as *"Any other

attributes the text states about this entity"* β€” it is not built from the

schema, and its keys are not constrained to the names you declare. Nothing in

mapping.py compares an extracted entity's property keys against

EntityTypeSchema.properties: map_extraction copies the dict through with

properties=dict(candidate.properties) and no filtering, renaming, defaulting

or type coercion.

So all four consequences hold at once, and none of them is an oversight β€”

they are ADR 0011

applied at property granularity:

with required: true

declared value type and no coercion runs

required: true is therefore a statement of intent for a reader, not a

guarantee β€” and since required does not even reach the prompt, it is not

currently a hint to the model either. If you need any of this enforced, read

schema.get_entity_type(entity.entity_type).properties and check the

extracted dict yourself.

Looking one up

EntityTypeSchema.get_property(name) is the only accessor, returning a

PropertySchema or None. It normalizes its argument with normalize_type_id

β€” the same function the name validator uses β€” so `get_property("Return

Type"), get_property("return-type") and get_property("__return__type__")`

all find return_type. Duplicate names are not rejected; get_property

returns the first match, scanning in declaration order.

What the bundled schemas do

All six populate properties on every entity type, most heavily in

technical_documentation (function has signature, parameters,

return_type, is_async, docstring). No test asserts any of this: the only

properties-adjacent assertion in

tests/unit/extraction/domains/test_yaml_schemas.py is

test_entity_types_have_examples, which checks examples is a list, and

whose comment ("ensure the property exists and is a list") uses "property" in

the Python sense rather than this one. There is no convention here that a test

enforces β€” not a count, not a naming style, not a required description.

Conventions worth following

costs prompt length in proportion to its description, uncapped, competing

with the entity and relationship lists. Three to five per type is the

bundled range.

is_async alone tells a model less than `is_async (Whether the function is

asynchronous)`. Every property in every bundled schema has one.

number` is invisible to the model and unenforced afterwards, so "Normalized

weight in grams" does more work than the type field does.

not Return Type β€” the validator accepts either, and a file whose names

differ from the stored names makes get_property a translation exercise.

attribute of one entity; a link to another entity belongs in

relationship_types, where the endpoint lists and the graph projection can

see it.

`examples` β€” optional, maximum 10 entries

Sample names of things that are instances of this entity type, used for

few-shot prompting. The list is optional and defaults to [].


- id: character
  description: A person, being, or personified entity in the narrative
  examples:
    - Hamlet
    - Lady Macbeth
    - Jay Gatsby
    - Elizabeth Bennet
Rule Value
Required no
Type list of strings
Default [] (default_factory=list)
Minimum length none β€” an empty list is valid, and so is omitting the key
Maximum length 10 (validate_examples_length)
Per-entry bounds none β€” no minimum, no maximum, no pattern
Normalization none, beyond whitespace stripping on each entry
Uniqueness not enforced

The cap rejects; it does not truncate

EntityTypeSchema.validate_examples_length is three lines and the whole of the

field's validation:


if len(v) > 10:
    raise ValueError(f"Maximum 10 examples allowed, got {len(v)}")

So an eleventh example is a load error, surfaced by pydantic as a

ValidationError and wrapped by the loader in SchemaLoadError with the file

path attached:


Value error, Maximum 10 examples allowed, got 11

Exactly 10 is accepted; both boundaries are pinned as examples in

tests/unit/extraction/domains/test_models.py

(test_entity_type_examples_exactly_10 and test_entity_type_examples_max_10).

Nothing silently drops the tail at load time β€” the truncation happens later, in

the prompt builder, and at a much lower number.

Only the first three reach the model

MAX_EXAMPLES_PER_TYPE in

src/redstring/extraction/prompt_generator.py is 3, and _entity_line

slices with it:


if entity_type.examples:
    examples = ", ".join(entity_type.examples[:MAX_EXAMPLES_PER_TYPE])
    line += f" (examples: {examples})"

giving the parenthetical at the end of the entity bullet:


- **character**: A person, being, or personified entity in the narrative (examples: Hamlet, Lady Macbeth, Jay Gatsby)

The constant carries its own reasoning in the source β€” *"All of them is not

better. The examples are there to disambiguate the type, and a schema listing

twenty of them would spend most of the prompt on one type β€” which reads to the

model as emphasis rather than as illustration."*

The gap between the two numbers is the thing to plan around: **the model

permits 10 and the prompt shows 3.** Entries four through ten are stored on the

loaded schema, are readable through

schema.get_entity_type("character").examples, and are never shown to the

extractor. literature_fiction declares four examples on character and four

on theme; Elizabeth Bennet and mortality are the entries no model sees.

So order the list deliberately β€” put the most disambiguating examples

first, and treat positions four onward as documentation for whoever reads the

YAML rather than as prompt input.

An empty list produces no parenthetical at all: the bullet is `- id:

description` and nothing is lost.

Entries are stripped, and otherwise unchecked

str_strip_whitespace=True on EntityTypeSchema applies to each string in the

list, so - " Hamlet " is stored as Hamlet. Nothing else is done to an

entry. In particular:

renders as (examples: , Hamlet). There is no min_length on the items.

spaces, capitals, punctuation and non-ASCII all survive verbatim, which is

the point. The murder of King Duncan and 1920s New York are bundled

entries.

appears twice in the prompt if it falls in the first three.

builds the parenthetical, so an example containing a comma is

indistinguishable from two examples once the model reads it. Prefer entries

without one.

It is prompt text and nothing else

Nothing reads examples outside the prompt builder. It does not appear in

DomainSummary, it is not consulted by is_valid_entity_type,

get_entity_type or validate_relationship, and no extracted entity is

checked against it. An extractor is free to return entities that resemble none

of them, and returning exactly one of them is not treated as a stronger result.

The list biases what the model looks for; it constrains nothing, per

ADR 0011.

That has a review consequence worth stating: editing this list changes

extraction behaviour with no test in the library able to see the change, which

is one of the things the schema's version

field exists to record.

What is checked

tests/unit/extraction/domains/test_yaml_schemas.py::test_entity_types_have_examples

is the only test naming this field against the bundled schemas, and **it does

not check what its name says**:


for et in schema.entity_types:
    # Not all entity types require examples, but most should have them
    # We'll just ensure the property exists and is a list
    assert isinstance(et.examples, list), (
        f"Entity type {et.id} in {domain_id} examples should be a list"
    )

examples is typed list[str] with a default_factory=list, so the assertion

is true by construction and cannot fail. There is no enforced convention here β€”

not a minimum count, not non-emptiness, not a style. The six bundled files

populate every entity type with two, three or four examples out of habit rather

than under a gate.

Conventions worth following

type sits in the two-to-four range, and only the first three are ever seen.

A fourth is for the reader; a fifth is usually noise.

a protagonist. An example does more to fix the type's boundary than a

longer description does, and it does it by being concrete.

has both class and module, the examples are where you show the

difference β€” the ids do not, and description is competing for the same

line.

parenthetical joined by ", ", so a long example crowds the bullet and one

containing a comma reads as two.

Property fields (`PropertySchema`)

One element of an entity type's properties list. It names a structured

attribute you want extracted alongside an entity of that type, and β€” like

everything else in the format β€” it produces prompt text and nothing more.


properties:
  - name: signature
    type: string
    description: Function signature including parameters
    required: false
Field Required Type Default Bounds
name yes string β€” 1-100 chars, normalized, must be a valid Python identifier
type no one of string, number, boolean, array, object "string" Literal β€” anything else is rejected
description no string or absent None max 500 chars, no minimum
required no boolean false β€”

model_config is extra="forbid", frozen=True, str_strip_whitespace=True,

the same as the other two nested models. So a fifth key is a load error (there

is no default, no enum, no pattern), the object is immutable and

hashable, and whitespace is stripped from name and description before their

bounds apply.

Note that description is the only string field in the whole format that is

optional and nullable: PropertySchema.description is str | None with a

default of None, where the domain, entity type and relationship type

descriptions are all required and non-empty. A property with no description is

legal and renders as a bare name.

Only two of the four fields reach the model

_property_hints in src/redstring/extraction/prompt_generator.py is the

entire consumer of this model:


", ".join(
    f"{prop.name} ({prop.description})" if prop.description else prop.name for prop in properties
)

and _entity_descriptions emits that as a second, indented line under the

entity bullet, only when the list is non-empty:


- **function**: A callable function or method with a specific signature (examples: extract_entities, create_user)
  Properties: signature (Function signature including parameters), parameters (List of parameter names and types), is_async (Whether the function is asynchronous)

type and required appear nowhere in that string. Nothing tells the

model that parameters is an array or that a property is mandatory; the only

channel for either is the description text. That is the single most important

fact about this model, and it is why the conventions below push everything you

want the extractor to know into the description.

Nothing is capped, either β€” unlike examples, which is truncated to the first

three, every property is printed in declaration order.

Nothing enforces the schema afterwards

The wire model an LlmProvider fills in carries a free-form

properties: dict[str, Any] per entity, and map_extraction copies it through

with properties=dict(candidate.properties) β€” no filtering against declared

names, no renaming, no defaulting, no coercion. So:

This is ADR 0011

at property granularity. required is currently the weakest field in the

format: it does not validate, and it does not even reach the prompt, so it is a

note to a human reader. If you need any of this enforced, read

schema.get_entity_type(entity.entity_type).properties and check the extracted

dict at the call site.

Looking one up

EntityTypeSchema.get_property(name) is the only accessor. It normalizes its

argument with normalize_type_id, the same function the name validator runs,

so get_property("Return Type"), get_property("return-type") and

get_property("__return__type__") all find return_type. Duplicate names are

not rejected by anything; get_property returns the first match in declaration

order.

What the bundled schemas do

Across the six bundled files there are 120 properties, and the distribution

is worth knowing before you reach for a field the format offers:

Observation Count
type: string 110
type: array 7
type: boolean 3
type: number / type: object 0 β€” never used
required: written at all 0 β€” never used
properties with no description 0
most properties on one entity type 5 (technical_documentation's function)

No test in tests/unit/extraction/domains/test_yaml_schemas.py asserts

anything about properties: there is no count convention, no naming rule, no

required description. The uniformity above is habit, and the habit is sound β€”

type and required buy nothing at present, so the bundled schemas mostly

leave them at their defaults and spend the effort on descriptions.

The per-field detail for name, type, description and required is the

table at the top of this section.

Relationship type fields (`RelationshipTypeSchema`)

One element of the top-level relationship_types list: an edge kind the

extractor is asked to look for. Like everything else in the format it produces

prompt text β€” a relationship type the model returns that is not declared here

is not discarded (ADR 0011).


relationship_types:
  - id: loves
    description: Romantic love between characters
    valid_source_types: [character]
    valid_target_types: [character]
    bidirectional: false
Field Required Type Default Bounds
id yes string β€” 1-100 chars, normalized, must be a valid Python identifier
description yes string β€” 1-500 chars
valid_source_types no list of string [] β€” meaning any each element normalized; each must name a declared entity type
valid_target_types no list of string [] β€” meaning any same
bidirectional no boolean false β€”

model_config is extra="forbid", frozen=True, str_strip_whitespace=True,

the same as the other two nested models: a sixth key is a load error, the

object is immutable and hashable, and whitespace is stripped before the bounds

apply.

`id` is normalized the same way an entity type id is

Lowercased and stripped, spaces and hyphens replaced with underscores, runs of

underscores collapsed to one, leading and trailing underscores removed. The

result must be a valid Python identifier or the load fails:


Relationship type ID must be a valid identifier: 'is-a?' -> 'is_a?'

So Works At, works-at and works_at are the same relationship type, and

__works__at__ normalizes to works_at as well. Write the normalized form;

relying on the normalizer means the id in your file and the id in the prompt

differ, and it is the normalized one that a caller matches against.

An id that normalizes to nothing at all β€” "___", " " β€” is rejected with a

different message naming the empty result, because "not an identifier" would

be a confusing thing to say about a string the file did contain.

`description` is required here, unlike a property's

1-500 characters, and there is no way to omit it. This is the same rule the

entity type's description carries, and it differs from PropertySchema, whose

description is the one optional-and-nullable string in the format. The

asymmetry is not arbitrary: the description is what

{relationship_descriptions} expands to, so a relationship type without one

would contribute a bare id to the prompt and tell the model nothing about when

to use it.

The two endpoint lists constrain nothing at extraction time

valid_source_types and valid_target_types are lists of entity type ids.

Empty means any, which is the default, and is why omitting them is not the

same as "unconstrained by accident" β€” it is the declared value.

Two things happen to them, and neither is enforcement of the model's output:

transform the ids get. [Main Character] is stored as ["main_character"].

validate_relationship_type_references model validator on DomainSchema β€”

the one cross-field rule in the format. A typo is a load error naming the

offender and listing the valid ids, rather than a constraint that silently

matches nothing:

```

Relationship 'loves' references unknown source type: 'charcter'.

Valid types: ['character', 'location']

```

Note that this is checked against the entity types in the same file, so

it cannot be satisfied by an entity type another schema declares.

What they do not do is filter extraction. DomainSchema.validate_relationship

and the is_valid_source / is_valid_target helpers on this model are

available for a caller who wants to check an edge, and nothing in the

extraction path calls them.

They normalize their argument the same way the loader normalized the list,

through the module's one normalize_type_id, so the string you wrote in the

YAML matches itself:


schema = RelationshipTypeSchema(id="loves", description="…", valid_source_types=["Main Character"])
schema.valid_source_types  # ['main_character']
schema.is_valid_source("Main Character")  # True
schema.is_valid_source("main_character")  # True

That is worth stating because it was not true: the lookup lowercased and

stripped only, so is_valid_source("Main Character") answered False against

a list built from that exact string. A caller passing an EntityTypeSchema.id

never saw it, since those are normalized on load; a caller passing an

Entity.entity_type β€” free-form text straight from the model, where "Main

Character" is an ordinary answer β€” got the wrong result every time.

`bidirectional` is a prompt hint, not a graph property

false by default. Relationships in the graph are directed β€”

Relationship has a source_entity_id and a target_entity_id and nothing

else β€” so setting this does not cause a second edge to be written, and no

reader treats an edge as symmetric because its declared type said so. It

describes the relationship to the model, which is the whole of the format's

job.

Looking one up

DomainSchema.get_relationship_type(id) returns the schema or None, and

is_valid_relationship_type(id) returns a bool. Both normalize the id they are

given the same way the loader does, so a lookup by "Works At" finds

works_at β€” which is the behaviour the endpoint lists above do not have.