Saltar al contenido
SPDF 5.0

SPDF en Python

Cómo instalar y usar la implementación de SPDF en Python (spdf-format): abrir, validar, buscar y citar. Solo sqlite3 de la biblioteca estándar; se importa como spdf.

Revisado Markdown

  • Paquete: spdf-format
  • Instalar: pip install spdf-format
  • Registro: PyPI
  • Nivel: primero
  • CI: CI en marcha
  • Carpeta: python/

El README de la biblioteca está en inglés.

Read, validate, search, cite and write SPDF files from Python.

SPDF (Semantic Processed Document Format) is an open format for documents that have been read once and can be cited forever: every passage carries its exact anchor (printed page or folio, second of a recording, slide, verse, canonical reference), so a citation can only print what the source says. A .spdf file is a plain SQLite 3 database with full-text indexes, optional embedding vectors and CSL-JSON metadata.

This package is a native, independent implementation of SPDF 5.0 and of the legacy 4.0 and 4.1 formats. It is installed as spdf-format and imported as spdf.

  • Python 3.10 or newer, standard library only (sqlite3).
  • Optional extras: numpy (fast vector search), crypto (Ed25519 through cryptography; a pure-Python fallback is included), pandas, arrow.
  • Passes the whole SPDF conformance suite (0.4.1, 341 cases): reader, semantic reader, writer and validator, profiles core, semantic and media.
pip install spdf-format            # or: uv add spdf-format
pip install "spdf-format[numpy]"   # faster vector search

Quick start

import spdf

with spdf.open("quijote.spdf") as f:          # 5.0, or legacy 4.x (gzip-wrapped too)
    doc = f.document                          # spdf.Document, metadata is CSL-JSON
    print(doc.display_title, f.cite())        # El ingenioso hidalgo… (Cervantes Saavedra, 1605)

    for hit in f.search("lanza en astillero"):
        print(f.cite(hit.fragment), hit.anchor_uri)
        # (Cervantes Saavedra, 1605, p. 23)  spdf:sha256-3f2a…#p=29&f=23&char=0,159

    unit = f.unit_by_printed("23")            # the page whose printed folio is 23
    print(unit.text)

Lexical search follows the reference algorithm of the specification: words are OR-ed, quoted phrases ("…", “…”, «…», „…“) are AND-ed, and case and diacritics are folded by the FTS5 index (unicode61 remove_diacritics 2). Queries in Chinese, Japanese or Korean use the trigram index when the file has one, and a substring scan otherwise.

with spdf.open("darwin.spdf") as f:
    space = f.space("embeddinggemma-2@768")   # model, dims, dtype, task prefixes
    qvec = my_model.encode(space.task_prefixes["query"] + "natural selection")
    f.search_vector(qvec, space=space.id, limit=5)
    f.search_hybrid("natural selection", qvec, space=space.id)   # RRF, k = 10
    f.search_vector(qvec, space=space.id, target="unit")          # hits carry unit_id

Every hit has id, target (fragment, unit or figure), score, via, anchor and anchor_uri; fragment_id, unit_id and figure_id give the id for each target.

Validate

report = spdf.validate("quijote.spdf")
report.valid            # True when there are no errors
report.codes            # {"W102"} …
report.to_dict()        # {"valid", "version", "profile", "errors", "warnings"} as in the spec

Validation follows the order of the specification and reports every code it finds: E001 not SQLite, E002 unknown version, E003 gzip-wrapped 5.0 (warning), E010/E011 missing table or column, E012 missing metadata key, E013 not exactly one document, E020 trigger, view or foreign virtual table, E030–E032 vectors and spaces, E040–E042 anchors, E050/E051 metadata, E060 unknown required extension, E070 FTS index out of sync, E080 blob hash, E081/E082 content hash and signature, E090 unit order, and the warnings W100–W110 (W103: a fragment crosses from one kind of matter to another, such as body text into a plate, or from a page with a printed folio to one without; writers should set matter on the units of paged documents and not let fragments cross).

Write

with spdf.Writer("out.spdf") as w:
    w.add_document({
        "id": "quijote", "kind": "pdf", "mime": "application/pdf", "bytes": 1203456,
        "source_sha256": "3f2a…", "source_ref": w.add_blob("original.pdf", "application/pdf", pdf_bytes),
        "metadata": {"type": "book", "title": "El ingenioso hidalgo don Quijote de la Mancha",
                     "author": [{"family": "Cervantes Saavedra", "given": "Miguel de"}],
                     "issued": {"date-parts": [[1605]]}, "language": "es"},
    })
    page = {"type": "page", "physical": 29, "printed": "23", "roman": False, "source": "read"}
    w.add_unit({"id": "u29", "ord": 1, "anchor": page, "text": "En un lugar de la Mancha…",
                "reader": "pdf-text-layer"})
    w.add_fragment({"id": "f1", "unit": "u29", "ord": 1, "text": "En un lugar de la Mancha…",
                    "anchor": {**page, "chars": [0, 24]}})
    w.add_space({"id": "embeddinggemma-2@768", "provider": "google", "model": "embeddinggemma-2",
                 "dims": 768, "modalities": ["text"]})
    w.add_vector("fragment", "f1", "embeddinggemma-2@768", data=vector)   # list, ndarray or bytes

The writer builds the file next to its destination, rebuilds the FTS index (distributed files carry no triggers), fills the required metadata keys (spdf_version, profile, created, generator, document_id), computes content_sha256, can sign it (finalize(sign_key=…)), compacts it with VACUUM, validates it and only then moves it into place. Text is normalized to NFC. i8 and f16 vectors are quantized as the specification says.

spdf.convert_legacy("old.spdf", "new.spdf") converts a 4.x file to 5.0, and spdf.write_source(full_dump, path) rebuilds a file from a full dump (f.full_dump()).

Anchors, URIs and citations

uri = spdf.make_uri("sha256-3f2a…", {"type": "page", "physical": 29, "printed": "21", "chars": [118, 301]})
# 'spdf:sha256-3f2a…#p=29&f=21&char=118,301'
spdf.parse_uri(uri)        # {"docref": "sha256-3f2a…", "locator": {"p": 29, "f": "21", "char": [118, 301]}}

spdf.cite({"type": "page", "physical": 9, "printed": "1r", "foliation": "leaf"}, doc, locale="es")
# '(Cervantes Saavedra, 1605, fol. 1r)'
spdf.cite({"type": "time", "t0": 4160.0, "t1": 4175.5}, doc, locale="en")
# '(Cortázar, 1959, 1:09:20)'

f.locate(reference) resolves an anchor URI, or the URL of a .spdf with an anchor fragment, against the file (SPEC §5.4):

f.locate("https://example.org/quijote.spdf#p=5&pe=6&char=101,278").to_dict()
# {"document": True, "units": ["p5", "p6"], "fragments": ["q4"], "char": [101, 278], "xywh": None}

A quotation is cited by the unit it lies in, never by the start of the fragment that contains it (SPEC §18.2), so a passage on page 211 of a fragment that begins on an unnumbered plate cites page 211:

c = f.cite_passage("m4", "tube N N", locale="es")
c.text   # '(Hooke, 1665, p. 211)'
c.uri    # 'spdf:sha256-…#p=321&f=211&char=0,8'

Citations print only what the anchor says: inferred folios in brackets (p. [21]), unnumbered pages as s. p. / n. pag., leaves and columns (fol. 1r, col. 45), h:mm:ss times, slides, sheets, verses and canonical references. Spanish uses e instead of y before the sound /i/ (Gómez e Iglesias).

Bibliography and other exports

ExportAPICLI
CSL-JSON (Zotero, citeproc, Pandoc)f.to_csl_json(), spdf.bibliography.csl_citation_item()spdf export -f csl
BibTeXf.to_bibtex()spdf export -f bibtex
ALTO 4 XML (page units)f.to_alto()spdf export -f alto
TEI P5 (minimal: header, pb, p, lg/l, u, note)f.to_tei()spdf export -f tei
IIIF Presentation 3 manifestf.to_iiif(base_url)spdf export -f iiif --base-url URL [--images-dir DIR]
JSON Lines of fragmentsspdf.interop.frames.fragment_records(f)spdf export -f jsonl
pandas DataFramef.to_pandas(vectors="space id")
Arrow tablef.to_arrow(vectors="space id")
Canonical dump (JCS)f.dump(), f.dump_json()spdf dump

ALTO has no invented coordinates: SPDF stores text per unit, so blocks and lines carry none (they are optional in ALTO 4). The IIIF manifest paints each canvas with the page image, adds the unit text as a supplementing annotation, figure descriptions as describing annotations on #xywh=percent: regions, and sections as ranges; audio and video become one time-based canvas with a range per unit.

Annotations and collections

User annotations and libraries live outside the documents, as the specification's sidecar files:

from spdf import sidecars

with spdf.open("quijote.spdf") as f:
    hit = f.search("lanza en astillero")[0]
    note = sidecars.annotation(f, hit.fragment, body="Origen del tópico.")   # W3C Web Annotation
sidecars.write_annotations("notas.spdfa.json", [note], label="Notas de lectura")

lib = sidecars.library(["quijote.spdf", "lazarillo.spdf"], "Tesis: fuentes")
sidecars.write_library("fuentes.spdfl.json", lib)

Each annotation targets the document by identity (spdf:sha256-…) with an SpdfAnchorSelector (the anchor URI) and a TextQuoteSelector (exact text with a little context), so it survives a re-reading that shifts offsets.

Command line

spdf validate FILE… [--json]          exit status 1 if a file is invalid
spdf dump FILE [--pretty]             canonical dump (RFC 8785)
spdf info FILE | --env                summary; --env shows the SQLite and FTS5 in use
spdf search FILE QUERY [--vector JSON --space ID] [--mode lexical|vector|hybrid] [--json]
spdf cite FILE [--fragment ID [--quote TEXT] | --unit ID | --uri URI] [--locale es|en] [--bibtex]
spdf export FILE -f csl|bibtex|alto|tei|iiif|jsonl
spdf convert OLD.spdf NEW.spdf        legacy 4.x to 5.0
spdf sign FILE --key KEY / spdf verify FILE [--public-key ed25519:…]
spdf conformance [DIR]                run the conformance suite, print the report

Other packages can add subcommands through the spdf.commands entry point group: the entry point is a callable register(subparsers) that adds an argparse parser and sets func (a handler returning the exit status). The SPDF producer adds spdf build this way:

[project.entry-points."spdf.commands"]
build = "spdf_build.cli:register"

Safety

Files are untrusted input. spdf.open checks the SQLite header before SQLite sees the file, decompresses gzip input to a private temporary file with a size limit (4 GiB by default), opens the database read-only by URI (mode=ro) with query_only, trusted_schema=OFF, cell_size_check and, on Python 3.12 or newer, SQLITE_DBCONFIG_DEFENSIVE; it never loads extensions, refuses triggers, views and virtual tables other than the format's FTS5 indexes (legacy files may keep their three FTS triggers), refuses files that require unknown extensions, and enforces a maximum blob size (512 MiB by default, max_blob_size=).

FTS5 and Python builds

Lexical search, writing and the FTS integrity check need SQLite's FTS5, which depends on how Python was built. It is present in the python.org installers, Homebrew, uv python install (python-build-standalone), conda and the usual Linux distributions; CI checks Linux, macOS and Windows with Python 3.10 to 3.13. spdf info --env tells you what you have. Without FTS5, reading, vector search, citations, exports and validation still work, and the operations that need it raise spdf.Fts5UnavailableError explaining what to do; on Linux, pip install pysqlite3-binary is picked up automatically as a fallback driver.

Conformance

The SPDF conformance suite lives in conformance/ in the repository. Run it with:

spdf conformance path/to/conformance        # prints {"impl", "version", "passed", "failed", "skipped"}

This implementation claims every kind of case: dump, legacy_dump, roundtrip, validate, search_lexical, search_vector, search_hybrid, anchor_uri, cite, quantize, locate, cite_passage, export_csl, export_bibtex and export_structure (checked on its own ALTO, TEI and IIIF output). CI publishes its report as the conformance-python artifact.

Development

cd python
uv sync --group dev
uv run pytest
uv run ruff check src tests && uv run ruff format --check src tests
uv run mypy

License

Code: MIT OR Apache-2.0, at your option. The SPDF specification is CC BY 4.0.