Steadlith

Content-defined chunk identities, cache-aware planning, and transactional indexing for RAG corpora that change.

Steadlith is a Python CLI and library for maintaining local retrieval indexes when source documents change. It assigns content-defined identities to chunks, previews each change, reuses cached embeddings, and publishes the resulting SQLite index update as one transaction.

Use Steadlith when a document corpus changes a little at a time and rebuilding every embedding would repeat work. It is not a RAG framework or a remote vector database.

Install Steadlith Follow the quick start

python -m pip install steadlith
steadlith init
steadlith plan

The core workflow is: source files → stable chunk identities → a read-only plan → cached or new embeddings → one transactional SQLite update. Start with steadlith plan; it does not contact an embedding provider or write index state.

What Steadlith provides

Follow the quick start

Install Steadlith, index a local corpus, query it, and verify the result.

Configure a project

Choose source globs, chunking, embeddings, cache paths, and index paths.

Operate an index safely

Plan changes, approve deletions, inspect drift, recover migrations, and compact tombstones.

Use the Python API

Integrate the stable chunker, parameter, chunk, and cache interfaces.

  • Deterministic Rabin content-defined chunking with stable v1 chunk identities.

  • A read-only plan before every index mutation.

  • A content-addressed embedding cache keyed by chunk, model, and provider parameters.

  • Transactional SQLite index publication with stale-plan detection.

  • Tombstones for immediate logical deletion and explicit physical compaction.

  • Offline lexical retrieval, an OpenAI adapter, and a local sentence-transformers adapter.

  • Status, verification, migration, cache, churn, and retrieval commands.

  • Machine-readable JSON for automation.

Steadlith does not claim that every edit changes only a fixed number of chunks. The current TTTD selector is stateful, and an exact counterexample is kept as a regression test. Use the bundled benchmark tools on your own corpus before making cost or quality decisions.

Choose a path

First use

Read Installation, Quick start, and Core model.

Running a local index

Use Indexing, Querying, Verification and recovery, and the CLI reference.

Selecting embeddings

Read Embedding providers and Backends and providers.

Changing chunking or models

Read Migrations, Chunk identity, and Compatibility.

Integrating with Python

Start with the Python API guide and the generated API reference.

Contributing

Read Development setup, Testing, Documentation, and Adapter conformance.

Scope

The supported reference deployment is one logical index per local SQLite database. It is suitable for development, evaluation, command-line workflows, and single-host applications after workload-specific testing. Remote vector databases, replicated serving, multi-tenant namespaces, and high-availability orchestration are outside the current implementation.