SPDF 5.0 · Semantic Processed Document Format · open standard
Read once.Cite forever.
SPDF is an open file format for documents that have already been read. Every passage carries its exact anchor: the printed page, the folio, the second of a recording, the slide or the verse. A citation can only print what the source says.
spdf:sha256-3f2a9c…#p=29&f=21&char=118,301
(Darwin, 1859, p. 21)
Physical page 29 of the file, printed folio 21, characters 118 to 301 of that page. The citation is computed from the anchor stored at reading time; nothing is guessed.
Why a format
Reading a document well is slow and expensive: OCR, transcription, finding the printed folios, sectioning, embeddings. SPDF stores the result so that nobody has to do it twice, and so that whatever cites from it can be checked.
Anchors
Every passage knows where it is. Fragments are stored with the place they came from: physical page and printed folio (roman, inferred or by leaf), second and word timings in audio and video, slide, sheet range, verse line or a canonical reference such as Stephanus 514a. Anchors serialise as portable URIs.
{"type":"page","physical":29,"printed":"21"}Provenance
Every field says who wrote it. Each unit records which reader produced its text (a PDF text layer, a vision model, a speech recogniser) and with what confidence; the metadata records where each field came from (colophon, title page, catalogue). Inferred folios are cited in brackets.
reader: gemma-4-e4b · confidence: 0.97Read once, query many
The expensive part happens once. OCR, transcription, sectioning and embeddings are paid for when the file is produced. After that it answers lexical, semantic and hybrid queries offline, even on a phone, with SQLite’s own full-text index and vectors from several models side by side.
fts5 unicode61 · f32 | f16 | i8 · RRF k = 10Portability
One file, any language, no server. A .spdf is a plain SQLite 3 database: no custom container, no account. It can be memory-mapped or read over HTTP ranges, and twelve independent implementations open it, all tested against the same conformance suite.
one document = one fileHonest citation
A citation can only print what the source says. Short citations and bibliography (CSL-JSON, BibTeX) are derived from the stored anchor and the CSL record, never generated. Agents get the same guarantee through the MCP server: they search, they quote and they cite with the exact folio, and they cannot invent one.
(Darwin, 1859, p. [21])
Inside a .spdf
A single SQLite file, uncompressed so it can be read by ranges, with no triggers and no views. Readers open it read-only, in defensive mode, and never load extensions. The schema is small enough to learn in an afternoon.
Profiles
- coreText and anchors. Enough to search and cite.
- semanticCore plus vectors from one or more embedding models.
- mediaCore plus audio and video with per-word timings.
- fullAll of the above.
spdf_metaversion, profile, generator, document iddocumentsone CSL-JSON record, with the provenance of each fieldunitsthe citable units: pages, time spans, slides, sheetsfragmentspassages of 150 to 300 words with their anchorsfragments_ftsFTS5 index, accent-insensitivesectionsthe heading treefiguresfigures, plates and frames, with region and descriptionspacesvector spaces, declared as model@dimsvectorslittle-endian f32, f16 or i8blobsthe original and the images, with their SHA-256provenancewhat produced what, with which model, and whenextensionsx_vendor_name tables, required or optional
Twelve implementations, one suite
The Rust implementation is the reference and also exposes a C ABI. The others are native and independent: each one opens, validates, dumps, searches, formats anchors, cites and writes, and each one is checked by the same conformance cases on every commit. Status and CI of every implementation.
Where to start
Validate
Drop a .spdf on the validator: it checks the file against the specification and shows what is inside, in your browser, without uploading anything.
Open the validator →Read
SPDF Reader opens, searches and cites SPDF files on macOS, Windows, Linux, iOS, Android and the web, with local models and no account.
Download the reader →Build
spdf build turns a PDF, a scan, an EPUB or a recording into an SPDF with local models, or with your own API key. Or start from SPDF Commons, a small collection of public-domain works.
Browse SPDF Commons →For agents
Every page of this site has a Markdown twin (the same address ending in .md), /llms.txt indexes them and /llms-full.txt carries the whole specification. The spdf-mcp server lets any agent search a folder of SPDF files and cite with the exact folio.
How agents use SPDF →