Skip to content
SPDF 5.0

LangChain loader (Python)

The same loader for LangChain in Python.

Reviewed Markdown

A LangChain document loader for SPDF files (Semantic Processed Document Format), so that retrieval-augmented answers cite the printed page instead of a chunk number.

An SPDF file is a SQLite database holding one document that has already been read: every passage (fragment) carries its exact anchor (physical page and printed folio, second of a recording, slide, verse…). SpdfLoader turns each passage into a LangChain Document whose metadata holds a ready-made short citation such as (Saorín Ferrer, 2026, p. 1) and a portable anchor URI such as spdf:sha256-…#p=2&f=1&char=15,307, both computed by the official library spdf-format.

Install

pip install spdf-langchain

Requires Python 3.10 or later, langchain-core>=0.3 and spdf-format (standard library only). From a checkout of the repository:

pip install -e python/ -e "integrations/langchain-python[test]"

Usage

from spdf_langchain import SpdfLoader

loader = SpdfLoader("spdf-in-five-pages.spdf", locale="en")   # a file, a folder or a list of both
docs = loader.load()                  # or: for doc in loader.lazy_load(): ...

docs[0].page_content                  # the literal passage, exactly as in the source
docs[0].metadata["citation"]          # '(Saorín Ferrer, 2026, p. 1)'
docs[0].metadata["anchor_uri"]        # 'spdf:sha256-50d9…5f4c#p=2&f=1&char=15,307'
docs[0].id                            # 'sha256-50d9…5f4c:f2-1' (stable across runs)
  • A folder is searched recursively for *.spdf files (hidden files and folders are skipped); a list may mix files and folders. lazy_load() yields the documents one by one, one file open at a time; load(), aload() and alazy_load() come from BaseLoader.
  • Unsafe or invalid files are refused by spdf-format: loading one raises its error (spdf.UnsafeFileError, spdf.NotSpdfError…, all subclasses of spdf.SpdfError) with the validation code (E020…) and the file path in the message. A missing path raises FileNotFoundError.
  • Legacy SPDF 4.0 and 4.1 files (also gzip-wrapped) read like 5.0 ones.
  • Vector stores in recent langchain-core versions take the ids from Document.id; with older ones, pass them yourself: store.add_documents(docs, ids=[d.id for d in docs]).

Options

OptionDefaultMeaning
granularity"fragment""fragment": one document per passage (about 150 to 300 words). "unit": one per page, time span, slide…
locale"en"Locale of citation: "en" or "es" ("es-ES" works; others fall back to English).
with_vectorsNoneId of a vector space stored in the files ("all-MiniLM-L6-v2@384"). Puts the stored vector in metadata["vector"] (a list of floats) and the space id in metadata["vector_space"]. Off by default, because most vector stores expect flat metadata. A file without that space raises spdf.SpdfError.

Metadata

Values are flat scalars (str, int, float, bool), so every vector store accepts them (the only exception is vector, and only if you ask for it). A key whose value would be null is left out (Chroma and others reject None): the cover of a book has no printed_folio key, a page has no t0.

KeyTypeExample (first passage of the English fixture)Notes
sourcestrfixtures/spdf-in-five-pages.spdfPath the file was read from.
spdf_versionstr5.04.0 or 4.1 for legacy files.
spdf_doc_idstrspdf-in-five-pagesThe document id inside the file. Not called doc_id, which LangChain's multi-vector and parent-document retrievers (and LlamaIndex vector stores) use for their own ids.
docrefstrsha256-50d94244…5f4cDocument reference used by anchor URIs (SHA-256 of the original).
titlestrSPDF in five pages
authorsstrSaorín Ferrer
yearint2026
languagestrenBCP 47.
kindstrpdfpdf, epub, audio, video…
fragment_idstrf2-1Fragment granularity only.
unit_idstru2The unit (page…) where the passage starts.
anchor_typestrpagepage, time, section, slide, sheet, web, image, verse, canonical.
physical_pageint2Page anchors: position of the page in the file.
printed_foliostr1The folio as printed ("xiv", "1r").
folio_inferredboolfalseTrue when the folio was deduced, not read; the citation prints it in brackets, p. [3].
sectionstrI. AnchorsHeading path joined with " / ". In unit granularity, the sections that share the unit are joined with " | ".
contextstrSPDF in five pages, I. AnchorsOne line that situates the passage (fragment granularity).
anchorstr{"chars":[15,307],"confidence":1,"physical":2,"printed":"1","source":"read","type":"page"}The start anchor as canonical JSON; json.loads it for the full object.
anchor_endstrEnd anchor (JSON) when the passage crosses into another unit.
t0, t1floatSeconds, for time anchors (recordings).
anchor_uristrspdf:sha256-50d94244…5f4c#p=2&f=1&char=15,307Resolve it with spdf.open(path).locate(uri); parse it with spdf.parse_uri.
citationstr(Saorín Ferrer, 2026, p. 1)Short author-date citation in the chosen locale.
vector, vector_spacelist, strOnly with with_vectors.

page_content is always the literal passage (fragments.text or units.text), never the modernised-spelling search layer, which SPDF forbids quoting.

End-to-end example (no API key)

Retrieval with a toy embedding and an in-memory vector store, then a prompt in which every passage carries its citation; pipe the prompt into any chat model.

import hashlib
import math
import re

from langchain_core.embeddings import Embeddings
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.vectorstores import InMemoryVectorStore

from spdf_langchain import SpdfLoader


class HashingEmbeddings(Embeddings):
    """A toy bag-of-words embedding: no model to download, no API key."""

    def __init__(self, dim: int = 512) -> None:
        self.dim = dim

    def _vec(self, text: str) -> list[float]:
        v = [0.0] * self.dim
        for word in re.findall(r"\w+", text.lower()):
            v[int(hashlib.md5(word.encode()).hexdigest(), 16) % self.dim] += 1.0
        norm = math.sqrt(sum(x * x for x in v)) or 1.0
        return [x / norm for x in v]

    def embed_documents(self, texts: list[str]) -> list[list[float]]:
        return [self._vec(t) for t in texts]

    def embed_query(self, text: str) -> list[float]:
        return self._vec(text)


docs = SpdfLoader("integrations/fixtures/spdf-in-five-pages.spdf", locale="en").load()
store = InMemoryVectorStore.from_documents(docs, embedding=HashingEmbeddings())  # needs numpy

question = "How is a plate without a printed folio cited?"
hits = store.similarity_search(question, k=2)
for d in hits:
    print(d.metadata["citation"], d.metadata["anchor_uri"])

prompt = ChatPromptTemplate.from_messages([
    ("system", "Answer from the passages only. After each claim, copy the citation of its passage."),
    ("human", "{context}\n\nQuestion: {question}"),
])
context = "\n\n".join(f"{d.page_content} {d.metadata['citation']}" for d in hits)
messages = prompt.invoke({"context": context, "question": question})
# answer = chat_model.invoke(messages)

Output:

(Saorín Ferrer, 2026, p. [3]) spdf:sha256-50d94244…5f4c#p=4&f=3&char=0,133
(Saorín Ferrer, 2026, p. 2) spdf:sha256-50d94244…5f4c#p=3&f=2&char=19,258

The plate carries no printed number; its folio is inferred, so the citation prints it in brackets. The context the model receives reads:

Plate I. A page with its folio and a manicule pointing at a passage. This plate carries no printed number; its folio, 3, is inferred. (Saorín Ferrer, 2026, p. [3])

Every unit records who read it: … the citation puts it in brackets. (Saorín Ferrer, 2026, p. 2)

Reusing the vectors stored in the file

SPDF files may ship vectors (f.spaces() in spdf-format lists them, with model, size and any task prefixes). When your embedding model is the one that produced a space, load the vectors instead of embedding every passage again, for example with FAISS:

from langchain_community.vectorstores import FAISS          # pip install langchain-community faiss-cpu
from langchain_huggingface import HuggingFaceEmbeddings      # pip install langchain-huggingface

from spdf_langchain import SpdfLoader

docs = SpdfLoader("library/", with_vectors="all-MiniLM-L6-v2@384").load()
pairs = [(d.page_content, d.metadata.pop("vector")) for d in docs]   # keep the metadata flat
store = FAISS.from_embeddings(
    pairs,
    HuggingFaceEmbeddings(model_name="sentence-transformers/all-MiniLM-L6-v2"),  # embeds queries only
    metadatas=[d.metadata for d in docs],
    ids=[d.id for d in docs],
)
print(store.similarity_search("inferred folio", k=1)[0].metadata["citation"])

Development

cd integrations/langchain-python
uv venv && uv pip install -e ../../python -e ".[test]"
.venv/bin/python -m pytest

The tests use the shared fixtures in integrations/fixtures/.

License

MIT OR Apache-2.0, at your option.