Cargador de LangChain (Python)
El mismo cargador para LangChain en Python.
Revisado Markdown
El README de la integración está en inglés, como su código.
A LangChain document loader for SPDF files (Semantic Processed Document Format), so that retrieval-augmented answers cite the printed page instead of a chunk number.
An SPDF file is a SQLite database holding one document that has already been read: every
passage (fragment) carries its exact anchor (physical page and printed folio, second of a
recording, slide, verse…). SpdfLoader turns each passage into a LangChain Document
whose metadata holds a ready-made short citation such as (Saorín Ferrer, 2026, p. 1) and
a portable anchor URI such as spdf:sha256-…#p=2&f=1&char=15,307, both computed by the
official library spdf-format.
Install
pip install spdf-langchainRequires Python 3.10 or later, langchain-core>=0.3 and spdf-format (standard library
only). From a checkout of the repository:
pip install -e python/ -e "integrations/langchain-python[test]"Usage
from spdf_langchain import SpdfLoader
loader = SpdfLoader("spdf-in-five-pages.spdf", locale="en") # a file, a folder or a list of both
docs = loader.load() # or: for doc in loader.lazy_load(): ...
docs[0].page_content # the literal passage, exactly as in the source
docs[0].metadata["citation"] # '(Saorín Ferrer, 2026, p. 1)'
docs[0].metadata["anchor_uri"] # 'spdf:sha256-50d9…5f4c#p=2&f=1&char=15,307'
docs[0].id # 'sha256-50d9…5f4c:f2-1' (stable across runs)- A folder is searched recursively for
*.spdffiles (hidden files and folders are skipped); a list may mix files and folders.lazy_load()yields the documents one by one, one file open at a time;load(),aload()andalazy_load()come fromBaseLoader. - Unsafe or invalid files are refused by
spdf-format: loading one raises its error (spdf.UnsafeFileError,spdf.NotSpdfError…, all subclasses ofspdf.SpdfError) with the validation code (E020…) and the file path in the message. A missing path raisesFileNotFoundError. - Legacy SPDF 4.0 and 4.1 files (also gzip-wrapped) read like 5.0 ones.
- Vector stores in recent
langchain-coreversions take the ids fromDocument.id; with older ones, pass them yourself:store.add_documents(docs, ids=[d.id for d in docs]).
Options
| Option | Default | Meaning |
|---|---|---|
granularity | "fragment" | "fragment": one document per passage (about 150 to 300 words). "unit": one per page, time span, slide… |
locale | "en" | Locale of citation: "en" or "es" ("es-ES" works; others fall back to English). |
with_vectors | None | Id of a vector space stored in the files ("all-MiniLM-L6-v2@384"). Puts the stored vector in metadata["vector"] (a list of floats) and the space id in metadata["vector_space"]. Off by default, because most vector stores expect flat metadata. A file without that space raises spdf.SpdfError. |
Metadata
Values are flat scalars (str, int, float, bool), so every vector store accepts
them (the only exception is vector, and only if you ask for it). A key whose value
would be null is left out (Chroma and others reject None): the cover of a book has no
printed_folio key, a page has no t0.
| Key | Type | Example (first passage of the English fixture) | Notes |
|---|---|---|---|
source | str | fixtures/spdf-in-five-pages.spdf | Path the file was read from. |
spdf_version | str | 5.0 | 4.0 or 4.1 for legacy files. |
spdf_doc_id | str | spdf-in-five-pages | The document id inside the file. Not called doc_id, which LangChain's multi-vector and parent-document retrievers (and LlamaIndex vector stores) use for their own ids. |
docref | str | sha256-50d94244…5f4c | Document reference used by anchor URIs (SHA-256 of the original). |
title | str | SPDF in five pages | |
authors | str | Saorín Ferrer | |
year | int | 2026 | |
language | str | en | BCP 47. |
kind | str | pdf | pdf, epub, audio, video… |
fragment_id | str | f2-1 | Fragment granularity only. |
unit_id | str | u2 | The unit (page…) where the passage starts. |
anchor_type | str | page | page, time, section, slide, sheet, web, image, verse, canonical. |
physical_page | int | 2 | Page anchors: position of the page in the file. |
printed_folio | str | 1 | The folio as printed ("xiv", "1r"). |
folio_inferred | bool | false | True when the folio was deduced, not read; the citation prints it in brackets, p. [3]. |
section | str | I. Anchors | Heading path joined with " / ". In unit granularity, the sections that share the unit are joined with " | ". |
context | str | SPDF in five pages, I. Anchors | One line that situates the passage (fragment granularity). |
anchor | str | {"chars":[15,307],"confidence":1,"physical":2,"printed":"1","source":"read","type":"page"} | The start anchor as canonical JSON; json.loads it for the full object. |
anchor_end | str | End anchor (JSON) when the passage crosses into another unit. | |
t0, t1 | float | Seconds, for time anchors (recordings). | |
anchor_uri | str | spdf:sha256-50d94244…5f4c#p=2&f=1&char=15,307 | Resolve it with spdf.open(path).locate(uri); parse it with spdf.parse_uri. |
citation | str | (Saorín Ferrer, 2026, p. 1) | Short author-date citation in the chosen locale. |
vector, vector_space | list, str | Only with with_vectors. |
page_content is always the literal passage (fragments.text or units.text), never the
modernised-spelling search layer, which SPDF forbids quoting.
End-to-end example (no API key)
Retrieval with a toy embedding and an in-memory vector store, then a prompt in which every passage carries its citation; pipe the prompt into any chat model.
import hashlib
import math
import re
from langchain_core.embeddings import Embeddings
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.vectorstores import InMemoryVectorStore
from spdf_langchain import SpdfLoader
class HashingEmbeddings(Embeddings):
"""A toy bag-of-words embedding: no model to download, no API key."""
def __init__(self, dim: int = 512) -> None:
self.dim = dim
def _vec(self, text: str) -> list[float]:
v = [0.0] * self.dim
for word in re.findall(r"\w+", text.lower()):
v[int(hashlib.md5(word.encode()).hexdigest(), 16) % self.dim] += 1.0
norm = math.sqrt(sum(x * x for x in v)) or 1.0
return [x / norm for x in v]
def embed_documents(self, texts: list[str]) -> list[list[float]]:
return [self._vec(t) for t in texts]
def embed_query(self, text: str) -> list[float]:
return self._vec(text)
docs = SpdfLoader("integrations/fixtures/spdf-in-five-pages.spdf", locale="en").load()
store = InMemoryVectorStore.from_documents(docs, embedding=HashingEmbeddings()) # needs numpy
question = "How is a plate without a printed folio cited?"
hits = store.similarity_search(question, k=2)
for d in hits:
print(d.metadata["citation"], d.metadata["anchor_uri"])
prompt = ChatPromptTemplate.from_messages([
("system", "Answer from the passages only. After each claim, copy the citation of its passage."),
("human", "{context}\n\nQuestion: {question}"),
])
context = "\n\n".join(f"{d.page_content} {d.metadata['citation']}" for d in hits)
messages = prompt.invoke({"context": context, "question": question})
# answer = chat_model.invoke(messages)Output:
(Saorín Ferrer, 2026, p. [3]) spdf:sha256-50d94244…5f4c#p=4&f=3&char=0,133
(Saorín Ferrer, 2026, p. 2) spdf:sha256-50d94244…5f4c#p=3&f=2&char=19,258The plate carries no printed number; its folio is inferred, so the citation prints it in brackets. The context the model receives reads:
Plate I. A page with its folio and a manicule pointing at a passage. This plate carries no printed number; its folio, 3, is inferred. (Saorín Ferrer, 2026, p. [3])
Every unit records who read it: … the citation puts it in brackets. (Saorín Ferrer, 2026, p. 2)Reusing the vectors stored in the file
SPDF files may ship vectors (f.spaces() in spdf-format lists them, with model, size and
any task prefixes). When your embedding model is the one that produced a space, load the
vectors instead of embedding every passage again, for example with FAISS:
from langchain_community.vectorstores import FAISS # pip install langchain-community faiss-cpu
from langchain_huggingface import HuggingFaceEmbeddings # pip install langchain-huggingface
from spdf_langchain import SpdfLoader
docs = SpdfLoader("library/", with_vectors="all-MiniLM-L6-v2@384").load()
pairs = [(d.page_content, d.metadata.pop("vector")) for d in docs] # keep the metadata flat
store = FAISS.from_embeddings(
pairs,
HuggingFaceEmbeddings(model_name="sentence-transformers/all-MiniLM-L6-v2"), # embeds queries only
metadatas=[d.metadata for d in docs],
ids=[d.id for d in docs],
)
print(store.similarity_search("inferred folio", k=1)[0].metadata["citation"])Development
cd integrations/langchain-python
uv venv && uv pip install -e ../../python -e ".[test]"
.venv/bin/python -m pytestThe tests use the shared fixtures in integrations/fixtures/.
License
MIT OR Apache-2.0, at your option.