Index documents without extracting them

You have a pile of documents and no budget to run a model over all of them.

index_documents splits them into a ChunkStore and asks no model anything,

so the passages are there to search, cite and โ€” later โ€” extract from, at no

per-token cost.


from redstring import InMemoryChunkStore, SourceDocument, index_documents

corpus = InMemoryChunkStore(dimension=768)
report = await index_documents(
    [SourceDocument(id="engine-memo", text=memo_text)],
    store=corpus,
    tenant_id=tenant_id,
)

passages = await corpus.get_by_source("engine-memo", tenant_id)

report is an IndexReport: documents_indexed, chunks_written and

documents_skipped. get_by_source returns the passages in chunk_index

order, ties broken by id.

There is no LlmProvider parameter and no place one could be passed. That is

the point of the function rather than an omission: build the corpus for

everything you hold, then run build_graph over whichever subset is worth

paying for.

Swap InMemoryChunkStore for PostgresChunkStore when the corpus outlives

the process; the port is the same, and the compliance suite in

src/redstring/testing/chunk_store.py is what says so. PostgresChunkStore is not

in redstring.__all__ -- reach it by path,

redstring.chunks.adapters.postgres.PostgresChunkStore, the way

Neo4jGraphStore, PgVectorStore, RedisCache and LangChainProvider are

all reached.

Choosing how documents are split

chunker= takes any Chunker. Omitted, you get a SlidingWindowChunker with

its own defaults โ€” the same splitter ExtractionPipeline uses, so a document

indexed and later extracted is split the same way.

Chunker is exported and is the whole contract: a chunk(text) returning a

ChunkingResult. Your own Chunker needs no dotted import at all, and the

two bundled ones are exported:


from redstring import BoundaryPreferenceChunker, SlidingWindowChunker

BoundaryPreferenceChunker is the one to pass when the passages will be

quoted back to a reader. Both cascade paragraph โ†’ sentence โ†’ word โ†’ hard

cut; it differs in searching the whole window for a boundary rather than its

last 500 characters, and in recognising a sentence that ends the text or is

followed by a closing quote. A chunk that ends mid-sentence produces a

quotation nobody can use, which is the cost SlidingWindowChunker's narrower

search pays at its own default size of 3000.

It is not the default, and the reason is chunk ids: they are

content-addressed over the passage text, so moving a boundary re-keys every

chunk of every document re-ingested with the other chunker. That is not a

migration โ€” replace_source handles it, and the old passages go โ€” but it is

a fact to know before switching a corpus that something else cites into.

Re-indexing with a different chunker replaces that source's passages

wholesale: the new split is a different chunking, so it is recorded and

replace_source deletes the passages the new split does not contain. Chunks

are content-addressed, so passages that survive the re-split keep their ids.

An empty `entity_ids` means no entities, not "not yet"

Every StoredChunk written by this function has an empty entity_ids, and

the type says what that means:

An empty entity_ids means no entities were extracted from this passage. It

does not mean extraction is pending.

There is no third state, and code that reads emptiness as a work queue will be

wrong forever while looking reasonable in review โ€” every passage from this

path is legitimately empty, and so is any passage from extraction that

happened to contain no entities. If you need to know which documents have been

extracted, that question is answered by the event log or by the graph, not by

this field.

What the default guarantees, and what it does not

index_documents takes an optional event_store. Without one it creates an

InMemoryEventStore per call, and the difference is narrower than "indexing

is idempotent" sounds:

repeat within one call repeat across calls
no event_store skipped re-indexed, counted as documents_indexed
event_store given skipped skipped

The aggregate refuses a chunking it has already recorded, and that refusal

lives in its state. With no event store the second call rebuilds the

aggregate from nothing, so there is no recorded chunking for it to refuse

against.

The consequence is a cost rather than a corruption. Chunks are

content-addressed, so the second write produces the identical rows and the

corpus is unchanged. What is lost is the report: a re-run counts every

document as newly indexed, so documents_indexed cannot be read as "work that

needed doing".

Pass an AggregateStore to make the suppression real, and to have a log the

corpus can be rebuilt from:


report = await index_documents(documents, store=corpus, tenant_id=tenant_id, event_store=events)

This is the same trade build_graph makes.

Indexing a document you have already extracted discards its entity links

Both write paths emit the same event and both go through replace_source,

which writes a source's chunking as one operation. So the last write to a

source wins, whole:

land last. The links survive. This is the normal order.

land last. The links are discarded.

That is documented behaviour with a test pinning it, not an accident โ€” but it

is lossy, and nothing warns you. Re-extract the document if you need the links

back. The safe habit is to index a corpus once, up front, and let extraction

be the thing that runs afterwards.

The two paths are deliberately not deduplicated against each other: they

record their chunkings under different keys, because a scheme that called them

the same chunking would make indexing-before-extracting silently emit nothing

and drop every entity link the extraction found, while reporting success.

Related

and driving projections from a log.

ChunkStore of your own against the shared compliance suite.

ids are content-addressed and why replace_source is one call.