Saltar al contenido
SPDF 5.0

Lector de LlamaIndex (Python)

El mismo lector para LlamaIndex en Python.

Revisado Markdown

El README de la integración está en inglés, como su código.

A LlamaIndex reader for SPDF files (Semantic Processed Document Format), so that retrieval-augmented answers cite the printed page instead of a chunk number.

An SPDF file is a SQLite database holding one document that has already been read: every passage (fragment) carries its exact anchor (physical page and printed folio, second of a recording, slide, verse…). SpdfReader turns each passage into a LlamaIndex Document whose metadata holds a ready-made short citation such as (Saorín Ferrer, 2026, p. 1) and a portable anchor URI such as spdf:sha256-…#p=2&f=1&char=15,307, both computed by the official library spdf-format.

Install

pip install spdf-llamaindex

Requires Python 3.10 or later, llama-index-core>=0.12 and spdf-format (standard library only). From a checkout of the repository:

pip install -e python/ -e "integrations/llamaindex-python[test]"

Usage

from spdf_llamaindex import SpdfReader

reader = SpdfReader(locale="en")
docs = reader.load_data("spdf-in-five-pages.spdf")   # a file, a folder or a list of both

docs[0].text                     # the literal passage, exactly as in the source
docs[0].metadata["citation"]     # '(Saorín Ferrer, 2026, p. 1)'
docs[0].metadata["anchor_uri"]   # 'spdf:sha256-50d9…5f4c#p=2&f=1&char=15,307'
docs[0].id_                      # 'sha256-50d9…5f4c:f2-1' (stable across runs)
  • A folder is searched recursively for *.spdf files (hidden files and folders are skipped); a list may mix files and folders. lazy_load_data() yields the documents one by one, one file open at a time.
  • With SimpleDirectoryReader, register the reader for the extension: SimpleDirectoryReader("library/", file_extractor={".spdf": SpdfReader()}). An fsspec filesystem passed as fs= is honoured.
  • extra_info={...} adds metadata to every document (and wins over the reader's keys).
  • Unsafe or invalid files are refused by spdf-format: loading one raises its error (spdf.UnsafeFileError, spdf.NotSpdfError…, all subclasses of spdf.SpdfError) with the validation code (E020…) and the file path in the message. A missing path raises FileNotFoundError.
  • Legacy SPDF 4.0 and 4.1 files (also gzip-wrapped) read like 5.0 ones.

Options

OptionDefaultMeaning
granularity"fragment""fragment": one document per passage (about 150 to 300 words). "unit": one per page, time span, slide…
locale"en"Locale of citation: "en" or "es" ("es-ES" works; others fall back to English).
include_embeddingsNoneId of a vector space stored in the files ("all-MiniLM-L6-v2@384"). Sets Document.embedding from the stored vectors, so that an index whose embedding model matches that space does not embed the passages again. A file without that space raises spdf.SpdfError; a passage without a stored vector keeps embedding=None and is embedded by the index.
excluded_embed_metadata_keysall but title, sectionKeys kept out of the text that is embedded.
excluded_llm_metadata_keysall but title, authors, year, section, citationKeys hidden from the LLM. By default the LLM sees the citation line next to each passage and can copy it into its answer.

Metadata

Values are flat scalars (str, int, float, bool), so every vector store accepts them. A key whose value would be null is left out (Chroma and others reject None): the cover of a book has no printed_folio key, a page has no t0.

KeyTypeExample (first passage of the English fixture)Notes
sourcestrfixtures/spdf-in-five-pages.spdfPath the file was read from.
spdf_versionstr5.04.0 or 4.1 for legacy files.
spdf_doc_idstrspdf-in-five-pagesThe document id inside the file. Not called doc_id: LlamaIndex vector stores overwrite doc_id, document_id and ref_doc_id with the node's reference document id.
docrefstrsha256-50d94244…5f4cDocument reference used by anchor URIs (SHA-256 of the original).
titlestrSPDF in five pages
authorsstrSaorín Ferrer
yearint2026
languagestrenBCP 47.
kindstrpdfpdf, epub, audio, video…
fragment_idstrf2-1Fragment granularity only.
unit_idstru2The unit (page…) where the passage starts.
anchor_typestrpagepage, time, section, slide, sheet, web, image, verse, canonical.
physical_pageint2Page anchors: position of the page in the file.
printed_foliostr1The folio as printed ("xiv", "1r").
folio_inferredboolfalseTrue when the folio was deduced, not read; the citation prints it in brackets, p. [3].
sectionstrI. AnchorsHeading path joined with " / ". In unit granularity, the sections that share the unit are joined with " | ".
contextstrSPDF in five pages, I. AnchorsOne line that situates the passage (fragment granularity).
anchorstr{"chars":[15,307],"confidence":1,"physical":2,"printed":"1","source":"read","type":"page"}The start anchor as canonical JSON; json.loads it for the full object.
anchor_endstrEnd anchor (JSON) when the passage crosses into another unit.
t0, t1floatSeconds, for time anchors (recordings).
anchor_uristrspdf:sha256-50d94244…5f4c#p=2&f=1&char=15,307Resolve it with spdf.open(path).locate(uri); parse it with spdf.parse_uri.
citationstr(Saorín Ferrer, 2026, p. 1)Short author-date citation in the chosen locale.
vector_spacestrall-MiniLM-L6-v2@384Only when include_embeddings attached a vector.

The document text is always the literal passage (fragments.text or units.text), never the modernised-spelling search layer, which SPDF forbids quoting.

End-to-end example (no API key)

A complete retrieval-augmented query with a toy embedding and LlamaIndex's MockLLM; swap them for your models. The sources of the answer carry their citations.

import hashlib
import math
import re

from llama_index.core import VectorStoreIndex
from llama_index.core.embeddings import BaseEmbedding
from llama_index.core.llms import MockLLM

from spdf_llamaindex import SpdfReader


class HashingEmbedding(BaseEmbedding):
    """A toy bag-of-words embedding: no model to download, no API key."""

    dim: int = 512

    def _vec(self, text: str) -> list[float]:
        v = [0.0] * self.dim
        for word in re.findall(r"\w+", text.lower()):
            v[int(hashlib.md5(word.encode()).hexdigest(), 16) % self.dim] += 1.0
        norm = math.sqrt(sum(x * x for x in v)) or 1.0
        return [x / norm for x in v]

    def _get_text_embedding(self, text: str) -> list[float]:
        return self._vec(text)

    def _get_query_embedding(self, query: str) -> list[float]:
        return self._vec(query)

    async def _aget_query_embedding(self, query: str) -> list[float]:
        return self._vec(query)


docs = SpdfReader(locale="en").load_data("integrations/fixtures/spdf-in-five-pages.spdf")
index = VectorStoreIndex(docs, embed_model=HashingEmbedding())

engine = index.as_query_engine(llm=MockLLM(), similarity_top_k=2)
response = engine.query("How is a plate without a printed folio cited?")
for source in response.source_nodes:
    print(source.node.metadata["citation"], source.node.metadata["anchor_uri"])

Output:

(Saorín Ferrer, 2026, p. [3]) spdf:sha256-50d94244…5f4c#p=4&f=3&char=0,133
(Saorín Ferrer, 2026, p. 2) spdf:sha256-50d94244…5f4c#p=3&f=2&char=19,258

The plate carries no printed number; its folio is inferred, so the citation prints it in brackets. What the LLM receives for each passage is:

title: SPDF in five pages
authors: Saorín Ferrer
year: 2026
section: III. Read once, query many
citation: (Saorín Ferrer, 2026, p. [3])

Plate I. A page with its folio and a manicule pointing at a passage. …

Fragments are already passage-sized, so the example passes the documents straight to VectorStoreIndex(docs, …). VectorStoreIndex.from_documents(docs, …) also works (the splitter copies the metadata to every node), but it creates new nodes without the stored embeddings.

Reusing the vectors stored in the file

SPDF files may ship vectors (f.spaces() in spdf-format lists them, with model, size and any task prefixes). When your embedding model is the one that produced a space, load the vectors instead of embedding every passage again:

from llama_index.core import VectorStoreIndex
from llama_index.embeddings.huggingface import HuggingFaceEmbedding  # pip install llama-index-embeddings-huggingface

from spdf_llamaindex import SpdfReader

docs = SpdfReader(include_embeddings="all-MiniLM-L6-v2@384").load_data("library/")
embed = HuggingFaceEmbedding(model_name="sentence-transformers/all-MiniLM-L6-v2")  # local, embeds queries only
index = VectorStoreIndex(docs, embed_model=embed)   # not from_documents: keep the stored vectors
print(index.as_retriever().retrieve("inferred folio")[0].node.metadata["citation"])

Development

cd integrations/llamaindex-python
uv venv && uv pip install -e ../../python -e ".[test]"
.venv/bin/python -m pytest

The tests use the shared fixtures in integrations/fixtures/.

License

MIT OR Apache-2.0, at your option.