Saltar al contenido
SPDF 5.0

RFC 0001: SPDF 5.0

SPDF 5.0 is the first public version of the format. It keeps the idea of the Scholaris versions 4.0 and 4.1 (one document per SQLite file; every passage carries an anchor to the printed page, folio, verse, second or…

Revisado Markdown

Esta RFC solo está en inglés.

  • Status: Accepted (2026-10-07)
  • Authors: José Luis Saorín Ferrer jl@joseluissaorin.com (editor)
  • Created: 2026-10-07
  • Discussion: review by the implementers of the libraries, recorded in the change log of spec/CONTRACT.md (drafts 0 to 1.2)
  • Specification: SPDF 5.0, a new major version. It replaces 4.1 as the current version; 4.0 and 4.1 become legacy versions that every reader still reads
  • Affects: the whole of spec/SPEC.md, spec/schema/spdf-5.0.sql, spec/json-schema/, conformance/
  • Conformance cases: the conformance suite 0.1.0 (220 cases) and 0.2.0 (228 cases, cases_sha256 2d28cb5169e426ce26fec1df6c7481e2e5a3b208f0f14f66411e51043aa442c3)
  • Supersedes / superseded by: none
  • Decision: accepted by the editor on 2026-10-07. This RFC bootstraps the process and was accepted without the public discussion and final comment periods that apply to every later RFC (see governance/RFC-PROCESS.md). It is published so that anyone can see what changed from 4.1 and why; any of its decisions can be revisited by a later RFC.

Summary

SPDF 5.0 is the first public version of the format. It keeps the idea of the Scholaris versions 4.0 and 4.1 (one document per SQLite file; every passage carries an anchor to the printed page, folio, verse, second or slide it comes from; several vector spaces may coexist) and turns it into an open standard: the container is plain uncompressed SQLite, the file identifies itself in its header, identifiers are in English, metadata is CSL-JSON, text offsets are defined, the anchor vocabulary grows, anchors get a portable URI aligned with W3C Media Fragments and RFC 5147, vectors can be stored in half precision or 8 bits, conformance profiles and an extension mechanism are defined, content can be hashed and signed, and distributed files may no longer contain code (triggers, views or foreign virtual tables). Every 5.0 reader still reads 4.0 and 4.1 files.

The normative text is spec/SPEC.md, whose Appendix A summarizes the changes from 4.1; section numbers below refer to it.

Motivation

SPDF 4.x was the internal export format of Scholaris. It worked for one application and one team, but it had properties that stand in the way of a format anyone can implement and archive:

  • Compressed container. A 4.x file is a SQLite database wrapped in gzip. A reader must decompress the whole file before reading a single row, so it cannot read by HTTP ranges or memory-map the file, and must defend itself against decompression bombs. Most of the bulk of a typical file (page images, an embedded original) is already in compressed formats and gains little from a second layer.
  • No identification. 4.x files carry application_id 0. The version is stored in a table (spdf.spdf_version) and only sometimes in user_version, which some platforms did not allow Scholaris to set (the 4.0 sample in the conformance suite has user_version 0). Tools such as file(1) or DROID cannot tell a 4.x file from any other gzip stream.
  • Spanish identifiers. Tables and columns are named in Spanish (documentos, unidades, anio, ancla_fin). That was natural inside Scholaris and is a barrier for everyone else.
  • Bespoke metadata. Document metadata used Scholaris's own JSON structure (MetadatosDocumento), which no reference manager understands.
  • Undefined offsets. Nothing said how to count a position inside a text, and the languages that implement SPDF count differently (UTF-16 code units in JavaScript, code points in Python, bytes in Rust and Go).
  • Missing anchors. Verse numbers, canonical references (Stephanus, Bekker, biblical, CTS) and leaf or column foliation, which are how poetry, classical texts, early printed books and manuscripts are cited, could not be expressed.
  • No way to point at a passage from outside the file. 4.x defined no URI for an anchor.
  • Vectors in f32 only, which makes files with several vector spaces large and is wasteful on phones.
  • SQL code inside distributed files. 4.x files carry three triggers that keep the full-text index in sync. Opening a file that contains SQL code written by someone else is a risk that a document format should not ask readers to take.
  • No integrity, no profiles, no extension mechanism, no rights: nothing to detect tampering, nothing to say what a minimal conforming file is, no way for a vendor to add data without forking the format, and nowhere to say under what terms a file may be shared.

Guide-level explanation

A 5.0 file is a SQLite 3 database that any SQLite tool can open. Its header says what it is: bytes 68 to 71 read SPDF (the application_id) and bytes 60 to 63 hold the version, 500. Inside there is exactly one document:

  • documents: one row with the CSL-JSON record of the work, the SHA-256 of the original file, its media type and size, and its rights;
  • units: the citable units in reading order (pages, time spans, slides, sections, sheets), each with its anchor and its text in NFC;
  • fragments: passages of about 150 to 300 words, each with the anchor of its start and, if it crosses units, of its end, indexed by FTS5 for lexical search;
  • optionally sections, figures, vector spaces and vectors, embedded blobs (page images, the original), provenance and extensions.

An anchor is a small JSON object:

{"type":"page","physical":29,"printed":"21","source":"read","chars":[118,301]}

and the same location as a URI, which names the document by the hash of its original bytes so it survives renaming and copying:

spdf:sha256-3f2a…#p=29&f=21&char=118,301

A citation is computed from the stored anchor and the CSL record, never generated: (Darwin, 1859, p. 21). A producer may also store a SHA-256 of the canonical content and an Ed25519 signature over it, so that a reader can check that the content is the one the signer vouched for.

Normative changes

The changes against 4.1, by area, with the section of spec/SPEC.md that states each rule.

Container and identification (§2, §24)

  1. A 5.0 file is an uncompressed SQLite 3 database holding exactly one document; the header starts at byte 0. Page size 4096, journal mode DELETE and a final VACUUM are recommended. Collections are separate manifests (§17).
  2. PRAGMA application_id = 1397769286 (0x53504446, stored big-endian at offset 68, which reads SPDF in ASCII).
  3. PRAGMA user_version = major × 100 + minor × 10 (5.0 is 500). Readers accept 500 to 599, may warn on a newer minor (W105), and refuse other majors (E002).
  4. Extension .spdf; media type application/vnd.spdf+sqlite3 (registration pending); Uniform Type Identifier com.joseluissaorin.spdf.
  5. Files MUST NOT contain triggers, views, or virtual tables other than the FTS5 tables of the schema. Writers keep the full-text index in sync themselves.

Safe opening (§2.4, §14)

  1. Every reader opens files read-only, with query_only on, trusted_schema off, the defensive flag on and extension loading off; refuses triggers, views and foreign virtual tables (except the three legacy FTS triggers); and bounds the size of any single value (RECOMMENDED 512 MiB) and of decompressed gzip input (RECOMMENDED 4 GiB). Operations that write, such as the FTS5 integrity-check, run on a private copy.

Schema with English identifiers (§3, §20)

  1. Tables: spdf becomes spdf_meta, documentos documents, unidades units, secciones sections, fragmentos fragments, figuras figures, espacios spaces, vectores vectors, procedencia provenance; blobs keeps its name. Every column is renamed as §20.1 lists.
  2. Removed: documentos.estado and documentos.bibliotecas (library membership belongs to collection manifests) and index-only columns.
  3. Added: documents.rights (§16); documents.source_ref, nullable, replacing original; spaces.dtype, spaces.truncated_from, spaces.task_prefixes; blobs.sha256; provenance.model; the extensions table.
  4. spdf_meta has the REQUIRED keys spdf_version, profile, created, generator and document_id; integrity keys are OPTIONAL.
  5. units.ord is numbered from 1 and contiguous (4.x numbered units from 0).

Metadata (§6)

  1. documents.metadata is one CSL-JSON item plus an spdf extension object for what CSL cannot hold: provenance per field, a date range for undated works, the original language, ORCID identifiers. The record can be handed to Zotero, citeproc or Pandoc as it is.

Text and offsets (§7)

  1. All stored text is NFC. Positions inside a text are counted in Unicode code points over the NFC text, end exclusive, which every language can compute the same way.
  2. The literal text is never modernized; the search_text column (introduced in 4.1) holds a modernized-spelling layer used only for search.

Anchors (§4)

  1. New anchor types: verse and canonical (schemes such as stephanus, bekker, bible or cts).
  2. Page anchors gain foliation: page (default), leaf (fol. 1r) or column (col. 45).
  3. Any anchor MAY carry region (fractions 0 to 1 of the unit image) and chars (code point range in the unit's NFC text).
  4. Required members are fixed per type and validated (E040, E041, E042).

Anchor URI (§5)

  1. New URI form spdf:<docref>#<params>, with an ABNF. The preferred docref is sha256- and the hex SHA-256 of the original, which survives renaming and copying.
  2. Parameters have one canonical order and encoding, and format(parse(uri)) reproduces the URI byte for byte.
  3. Where SPDF overlaps with existing standards it uses their syntax: t= and xywh=percent: as in W3C Media Fragments URI 1.0, and char= as in RFC 5147.
  4. Resolution rules say how a reader finds the unit a URI designates.

Vectors (§9)

  1. Vector components are little-endian f32, f16 (IEEE binary16) or i8 (value q/127); the space id records model, dimensions and dtype.
  2. Spaces record Matryoshka truncation and the task prefixes used at encoding time; a compatibility rule says when one query vector serves several spaces; writers quantize with fixed rounding rules.
  3. A file without vectors is valid.

Search, citation and export (§8, §18, §19)

  1. Reference algorithms that conformance tests: lexical search through FTS5 with the unicode61 remove_diacritics 2 tokenizer (terms sent as written, without case folding), a route for Chinese, Japanese and Korean with an optional trigram index and a substring fallback, brute-force vector search, and hybrid search by reciprocal rank fusion with k = 10. Products MAY rank better.
  2. A short citation function for Spanish and English, and exports to CSL-JSON and BibTeX (REQUIRED) and other formats.

Profiles and extensions (§10, §11)

  1. Profiles, declared in spdf_meta.profile: core, semantic, media and full.
  2. Extensions are declared in the extensions table, with tables named x_<vendor>_<name>. A reader that meets a required extension it does not know refuses the file (E060); optional ones are ignored.

Integrity, signature and rights (§12, §13, §16)

  1. A canonical JSON dump of the file (RFC 8785, with fixed rounding and ordering) is the conformance oracle and the basis of integrity.
  2. content_sha256 hashes that dump; signature is an Ed25519 signature (RFC 8032) over it. Because the dump covers blob and vector bytes through their hashes but not the SQLite page layout, the signature survives VACUUM and SQLite version changes.
  3. documents.rights states the licence (SPDX), the access level and the holder.

Validation (§22)

  1. A deterministic check order and a closed list of error and warning codes. Conformance compares the sets of codes, not the messages.

Sidecars (§17)

  1. User annotations live outside the document, in *.spdfa.json files (W3C Web Annotation with an SpdfAnchorSelector and a TextQuoteSelector), so the document stays immutable and sharing it never shares its reader's notes.
  2. Collections are *.spdfl.json manifests that list documents by hash.

Legacy (§20)

  1. Every reader reads 4.0 and 4.1 files: detects gzip, decompresses within the limit, tolerates exactly the triggers fragmentos_ai, fragmentos_ad and fragmentos_au, and presents the 5.0 view, with "legacy": true in the dump and warning W110. A gzip-wrapped 5.0 file is read but flagged (E003, a warning).
  2. Version 3.0 MAY be supported through an importer.

Conformance cases

The conformance suite is the evidence for this RFC. Version 0.1.0 (2026-10-07) published 220 cases; version 0.2.0 (the same day) brought them to 228: anchor_uri 40, cite 79, dump 7, legacy_dump 2, quantize 6, roundtrip 7, search_hybrid 2, search_lexical 39, search_vector 5 and validate 41. They use seven 5.0 files built from public-domain texts, two authentic legacy files (4.0 and 4.1, gzip-wrapped, with FTS triggers), and invalid files that each break one rule. How each expectation is obtained (SQLite as the oracle for lexical search, exact arithmetic for vectors, hand-reviewed URIs and citations) is described in conformance/README.md; the changes are in conformance/CHANGELOG.md.

Backwards compatibility

  • 5.0 is a new major version. A 4.x reader cannot read 5.0 files: the schema is different and the container is no longer gzip. This is intended.
  • 5.0 readers read 4.x. Every conforming 5.0 reader reads 4.0 and 4.1 files through the 5.0 view (§20). Nothing that a 4.x file holds about the document is lost; only the Scholaris shelf state (estado, bibliotecas) is dropped.
  • Writers produce 5.0 only. Scholaris keeps reading its 4.x files and exports 5.0 through the spdf-format library.
  • Sidecars are new and carry their own version (spdf_library: "1.0" in manifests).
  • Nothing is deprecated within 5.0. The obligation to read 4.x lasts for the whole 5.x line; a future major version decides by RFC whether its readers still read 4.x (see governance/VERSIONING.md).

Security and privacy

Sections 14 and 15 of the specification are new in 5.0. In short:

  • Files contain no triggers, views or foreign virtual tables, and readers refuse files that do, so opening a file runs no code written by its author. Readers open read-only, with trusted_schema off, defensive mode on and extensions disabled.
  • The 5.0 container is not compressed. Legacy gzip input is decompressed within a limit to defeat decompression bombs; values and JSON nesting are bounded.
  • Text is light Markdown rendered without raw HTML; remote references are never fetched automatically; blob keys are sanitized before anything is written to disk; text passed to language models is data, not instructions.
  • Vectors can leak the text they were computed from; provenance can leak details of the producer. Rights and confidentiality rules that apply to a text apply to its vectors.
  • content_sha256 and the Ed25519 signature give integrity and attribution of the content, not confidentiality, and say nothing about whether the signer's key should be trusted.

Alternatives

  • Keep the gzip wrapper. Rejected: it prevents HTTP range reads and memory mapping, forces full decompression and adds decompression-bomb handling, for little gain on already-compressed images and originals. Readers still accept it for legacy files.
  • A ZIP package with JSON and a SQLite index inside (as EPUB or OOXML do). Rejected: two levels of parsing, and the full-text index needs SQLite anyway; SQLite alone gives random access, an index and a single file.
  • Pure JSON, Parquet or Arrow. Rejected: no built-in full-text search, and either no random access (JSON) or no good fit for long text with anchors (columnar formats).
  • Keep Spanish identifiers with English aliases. Rejected: two names for every table and column would double the surface of every implementation. The specification keeps a faithful Spanish translation instead.
  • UTF-16 code units, bytes or grapheme clusters for offsets. Rejected: UTF-16 ties the format to JavaScript, bytes to one encoding, and grapheme clusters to a Unicode version. Code points over NFC are stable and cheap everywhere.
  • IIIF-style regions (pct:) or bare fractions in xywh. Rejected in favour of W3C Media Fragments (percent:); in Media Fragments bare numbers mean pixels, so bare fractions would have been misread. Mapping to IIIF Image API regions is trivial.
  • Signing the file bytes, or wrapping the file in JWS or COSE. Rejected: SQLite writes its own version into the header and VACUUM reorders pages, so byte signatures break without any change in content. A signature over the canonical content is stable. An envelope format can still be added later as an extension.
  • Allowing triggers and views. Rejected for safety; writers can keep the index in sync without them.

Unresolved questions

  • The registration of application/vnd.spdf+sqlite3 with IANA, and whether to use the registered +sqlite3 structured syntax suffix instead (application/vnd.spdf+sqlite3). Drafts of this and other registrations are in governance/drafts/.
  • The provisional registration of the spdf URI scheme (RFC 7595), which §5.4 announces.
  • Whether the anchor URI parameters may also be used as the fragment identifier of a URL that points to a .spdf file (https://example.org/x.spdf#p=29), which the media type registration would like to say.
  • How a reader and a validator of an earlier minor version treat an anchor type, a dtype or a validation rule introduced by a later minor (today an unknown anchor type is error E041), so that the compatibility promise holds for validators too.
  • Markup for mathematics and other non-textual content inside unit text, which 5.0 does not specify.
  • A persistent identifier (DOI) for each published version of the specification.

Implementations

This RFC becomes Implemented when two independent implementations pass the whole conformance suite in CI, as their conformance-<folder> artifacts show. The table records what each implementation reported in its own commits on 2026-10-07; several independent implementations already report the whole suite, so the editor will mark the RFC Implemented once their CI artifacts confirm it.

ImplementationFolderReported on 2026-10-07
Rust (reference, also the C ABI)rust/228 of 228 (commit 38c75e5)
TypeScript (spdf-format, npm)js/228 of 228, also in Chromium and Bun (commit e063594)
Python (spdf-format, PyPI)python/228 of 228 (commit 557af8f)
Gogo/228 of 228 (commit 9355f22)
Swiftswift/228 of 228 on macOS and the iOS simulator (commit 54d5525)
PHP and Rubyphp/, ruby/228 of 228 (commit 537202e)
Kotlin/JVM, C#, Rkotlin/, dotnet/, r/in development
Julia, C (over the Rust ABI)julia/, c/planned
Reference producer spdf build (Python)producer/in development
Scholaris (second producer)externalreads 4.x; 5.0 export through spdf-format planned