---
title: "RFC 0001: SPDF 5.0"
description: "SPDF 5.0 is the first public version of the format. It keeps the idea of the Scholaris versions 4.0 and 4.1 (one document per SQLite file; every passage carries an anchor to the printed page, folio, verse, second or…"
url: https://spdf.joseluissaorin.com/es/gobernanza/rfcs/0001
markdown: https://spdf.joseluissaorin.com/es/gobernanza/rfcs/0001.md
lang: es
alternate_en: https://spdf.joseluissaorin.com/governance/rfcs/0001.md
updated: 2026-10-07
author: José Luis Saorín Ferrer (https://joseluissaorin.com)
license: CC-BY-4.0
---

# RFC 0001: SPDF 5.0

> SPDF 5.0 is the first public version of the format. It keeps the idea of the Scholaris versions 4.0 and 4.1 (one document per SQLite file; every passage carries an anchor to the printed page, folio, verse, second or…

- **Status**: Accepted (2026-10-07)
- **Authors**: José Luis Saorín Ferrer <jl@joseluissaorin.com> (editor)
- **Created**: 2026-10-07
- **Discussion**: review by the implementers of the libraries, recorded in the change log of `spec/CONTRACT.md` (drafts 0 to 1.2)
- **Specification**: SPDF 5.0, a new major version. It replaces 4.1 as the current version; 4.0 and 4.1 become legacy versions that every reader still reads
- **Affects**: the whole of `spec/SPEC.md`, `spec/schema/spdf-5.0.sql`, `spec/json-schema/`, `conformance/`
- **Conformance cases**: the conformance suite 0.1.0 (220 cases) and 0.2.0 (228 cases, `cases_sha256` `2d28cb5169e426ce26fec1df6c7481e2e5a3b208f0f14f66411e51043aa442c3`)
- **Supersedes / superseded by**: none
- **Decision**: accepted by the editor on 2026-10-07. This RFC bootstraps the process and was accepted without the public discussion and final comment periods that apply to every later RFC (see [`governance/RFC-PROCESS.md`](../../governance/RFC-PROCESS.md)). It is published so that anyone can see what changed from 4.1 and why; any of its decisions can be revisited by a later RFC.

## Summary

SPDF 5.0 is the first public version of the format. It keeps the idea of the Scholaris
versions 4.0 and 4.1 (one document per SQLite file; every passage carries an anchor to
the printed page, folio, verse, second or slide it comes from; several vector spaces may
coexist) and turns it into an open standard: the container is plain uncompressed
SQLite, the file identifies itself in its header, identifiers are in English, metadata
is CSL-JSON, text offsets are defined, the anchor vocabulary grows, anchors get a
portable URI aligned with W3C Media Fragments and RFC 5147, vectors can be stored in
half precision or 8 bits, conformance profiles and an extension mechanism are defined,
content can be hashed and signed, and distributed files may no longer contain code
(triggers, views or foreign virtual tables). Every 5.0 reader still reads 4.0 and 4.1
files.

The normative text is [`spec/SPEC.md`](../SPEC.md), whose Appendix A summarizes the
changes from 4.1; section numbers below refer to it.

## Motivation

SPDF 4.x was the internal export format of Scholaris. It worked for one application and
one team, but it had properties that stand in the way of a format anyone can implement
and archive:

- **Compressed container.** A 4.x file is a SQLite database wrapped in gzip. A reader
  must decompress the whole file before reading a single row, so it cannot read by HTTP
  ranges or memory-map the file, and must defend itself against decompression bombs.
  Most of the bulk of a typical file (page images, an embedded original) is already in
  compressed formats and gains little from a second layer.
- **No identification.** 4.x files carry `application_id` 0. The version is stored in a
  table (`spdf.spdf_version`) and only sometimes in `user_version`, which some platforms
  did not allow Scholaris to set (the 4.0 sample in the conformance suite has
  `user_version` 0). Tools such as file(1) or DROID cannot tell a 4.x file from any
  other gzip stream.
- **Spanish identifiers.** Tables and columns are named in Spanish (`documentos`,
  `unidades`, `anio`, `ancla_fin`). That was natural inside Scholaris and is a barrier
  for everyone else.
- **Bespoke metadata.** Document metadata used Scholaris's own JSON structure
  (`MetadatosDocumento`), which no reference manager understands.
- **Undefined offsets.** Nothing said how to count a position inside a text, and the
  languages that implement SPDF count differently (UTF-16 code units in JavaScript, code
  points in Python, bytes in Rust and Go).
- **Missing anchors.** Verse numbers, canonical references (Stephanus, Bekker, biblical,
  CTS) and leaf or column foliation, which are how poetry, classical texts, early printed
  books and manuscripts are cited, could not be expressed.
- **No way to point at a passage from outside the file.** 4.x defined no URI for an
  anchor.
- **Vectors in f32 only**, which makes files with several vector spaces large and is
  wasteful on phones.
- **SQL code inside distributed files.** 4.x files carry three triggers that keep the
  full-text index in sync. Opening a file that contains SQL code written by someone else
  is a risk that a document format should not ask readers to take.
- **No integrity, no profiles, no extension mechanism, no rights**: nothing to detect
  tampering, nothing to say what a minimal conforming file is, no way for a vendor to
  add data without forking the format, and nowhere to say under what terms a file may be
  shared.

## Guide-level explanation

A 5.0 file is a SQLite 3 database that any SQLite tool can open. Its header says what it
is: bytes 68 to 71 read `SPDF` (the `application_id`) and bytes 60 to 63 hold the
version, 500. Inside there is exactly one document:

- `documents`: one row with the CSL-JSON record of the work, the SHA-256 of the original
  file, its media type and size, and its rights;
- `units`: the citable units in reading order (pages, time spans, slides, sections,
  sheets), each with its anchor and its text in NFC;
- `fragments`: passages of about 150 to 300 words, each with the anchor of its start and,
  if it crosses units, of its end, indexed by FTS5 for lexical search;
- optionally `sections`, `figures`, vector `spaces` and `vectors`, embedded `blobs`
  (page images, the original), `provenance` and `extensions`.

An anchor is a small JSON object:

```json
{"type":"page","physical":29,"printed":"21","source":"read","chars":[118,301]}
```

and the same location as a URI, which names the document by the hash of its original
bytes so it survives renaming and copying:

```text
spdf:sha256-3f2a…#p=29&f=21&char=118,301
```

A citation is computed from the stored anchor and the CSL record, never generated:
`(Darwin, 1859, p. 21)`. A producer may also store a SHA-256 of the canonical content
and an Ed25519 signature over it, so that a reader can check that the content is the one
the signer vouched for.

## Normative changes

The changes against 4.1, by area, with the section of `spec/SPEC.md` that states each
rule.

### Container and identification (§2, §24)

1. A 5.0 file is an **uncompressed** SQLite 3 database holding exactly one document; the
   header starts at byte 0. Page size 4096, journal mode DELETE and a final `VACUUM` are
   recommended. Collections are separate manifests (§17).
2. `PRAGMA application_id = 1397769286` (0x53504446, stored big-endian at offset 68,
   which reads `SPDF` in ASCII).
3. `PRAGMA user_version` = major × 100 + minor × 10 (5.0 is 500). Readers accept 500 to
   599, may warn on a newer minor (W105), and refuse other majors (E002).
4. Extension `.spdf`; media type `application/vnd.spdf+sqlite3` (registration pending); Uniform
   Type Identifier `com.joseluissaorin.spdf`.
5. Files MUST NOT contain triggers, views, or virtual tables other than the FTS5 tables
   of the schema. Writers keep the full-text index in sync themselves.

### Safe opening (§2.4, §14)

6. Every reader opens files read-only, with `query_only` on, `trusted_schema` off, the
   defensive flag on and extension loading off; refuses triggers, views and foreign
   virtual tables (except the three legacy FTS triggers); and bounds the size of any
   single value (RECOMMENDED 512 MiB) and of decompressed gzip input (RECOMMENDED 4 GiB).
   Operations that write, such as the FTS5 `integrity-check`, run on a private copy.

### Schema with English identifiers (§3, §20)

7. Tables: `spdf` becomes `spdf_meta`, `documentos` `documents`, `unidades` `units`,
   `secciones` `sections`, `fragmentos` `fragments`, `figuras` `figures`, `espacios`
   `spaces`, `vectores` `vectors`, `procedencia` `provenance`; `blobs` keeps its name.
   Every column is renamed as §20.1 lists.
8. Removed: `documentos.estado` and `documentos.bibliotecas` (library membership belongs
   to collection manifests) and index-only columns.
9. Added: `documents.rights` (§16); `documents.source_ref`, nullable, replacing
   `original`; `spaces.dtype`, `spaces.truncated_from`, `spaces.task_prefixes`;
   `blobs.sha256`; `provenance.model`; the `extensions` table.
10. `spdf_meta` has the REQUIRED keys `spdf_version`, `profile`, `created`, `generator`
    and `document_id`; integrity keys are OPTIONAL.
11. `units.ord` is numbered from 1 and contiguous (4.x numbered units from 0).

### Metadata (§6)

12. `documents.metadata` is one CSL-JSON item plus an `spdf` extension object for what
    CSL cannot hold: provenance per field, a date range for undated works, the original
    language, ORCID identifiers. The record can be handed to Zotero, citeproc or Pandoc
    as it is.

### Text and offsets (§7)

13. All stored text is NFC. Positions inside a text are counted in **Unicode code
    points over the NFC text**, end exclusive, which every language can compute the
    same way.
14. The literal text is never modernized; the `search_text` column (introduced in 4.1)
    holds a modernized-spelling layer used only for search.

### Anchors (§4)

15. New anchor types: `verse` and `canonical` (schemes such as `stephanus`, `bekker`,
    `bible` or `cts`).
16. Page anchors gain `foliation`: `page` (default), `leaf` (`fol. 1r`) or `column`
    (`col. 45`).
17. Any anchor MAY carry `region` (fractions 0 to 1 of the unit image) and `chars` (code
    point range in the unit's NFC text).
18. Required members are fixed per type and validated (E040, E041, E042).

### Anchor URI (§5)

19. New URI form `spdf:<docref>#<params>`, with an ABNF. The preferred docref is
    `sha256-` and the hex SHA-256 of the original, which survives renaming and copying.
20. Parameters have one canonical order and encoding, and `format(parse(uri))`
    reproduces the URI byte for byte.
21. Where SPDF overlaps with existing standards it uses their syntax: `t=` and
    `xywh=percent:` as in W3C Media Fragments URI 1.0, and `char=` as in RFC 5147.
22. Resolution rules say how a reader finds the unit a URI designates.

### Vectors (§9)

23. Vector components are little-endian `f32`, `f16` (IEEE binary16) or `i8` (value
    q/127); the space id records model, dimensions and dtype.
24. Spaces record Matryoshka truncation and the task prefixes used at encoding time; a
    compatibility rule says when one query vector serves several spaces; writers
    quantize with fixed rounding rules.
25. A file without vectors is valid.

### Search, citation and export (§8, §18, §19)

26. Reference algorithms that conformance tests: lexical search through FTS5 with the
    `unicode61 remove_diacritics 2` tokenizer (terms sent as written, without case
    folding), a route for Chinese, Japanese and Korean with an optional `trigram` index
    and a substring fallback, brute-force vector search, and hybrid search by reciprocal
    rank fusion with k = 10. Products MAY rank better.
27. A short citation function for Spanish and English, and exports to CSL-JSON and
    BibTeX (REQUIRED) and other formats.

### Profiles and extensions (§10, §11)

28. Profiles, declared in `spdf_meta.profile`: `core`, `semantic`, `media` and `full`.
29. Extensions are declared in the `extensions` table, with tables named
    `x_<vendor>_<name>`. A reader that meets a **required** extension it does not know
    refuses the file (E060); optional ones are ignored.

### Integrity, signature and rights (§12, §13, §16)

30. A canonical JSON dump of the file (RFC 8785, with fixed rounding and ordering) is
    the conformance oracle and the basis of integrity.
31. `content_sha256` hashes that dump; `signature` is an Ed25519 signature (RFC 8032)
    over it. Because the dump covers blob and vector bytes through their hashes but not
    the SQLite page layout, the signature survives `VACUUM` and SQLite version changes.
32. `documents.rights` states the licence (SPDX), the access level and the holder.

### Validation (§22)

33. A deterministic check order and a closed list of error and warning codes.
    Conformance compares the sets of codes, not the messages.

### Sidecars (§17)

34. User annotations live outside the document, in `*.spdfa.json` files (W3C Web
    Annotation with an `SpdfAnchorSelector` and a `TextQuoteSelector`), so the document
    stays immutable and sharing it never shares its reader's notes.
35. Collections are `*.spdfl.json` manifests that list documents by hash.

### Legacy (§20)

36. Every reader reads 4.0 and 4.1 files: detects gzip, decompresses within the limit,
    tolerates exactly the triggers `fragmentos_ai`, `fragmentos_ad` and `fragmentos_au`,
    and presents the 5.0 view, with `"legacy": true` in the dump and warning W110. A
    gzip-wrapped 5.0 file is read but flagged (E003, a warning).
37. Version 3.0 MAY be supported through an importer.

## Conformance cases

The conformance suite is the evidence for this RFC. Version 0.1.0 (2026-10-07) published
220 cases; version 0.2.0 (the same day) brought them to 228: `anchor_uri` 40, `cite` 79,
`dump` 7, `legacy_dump` 2, `quantize` 6, `roundtrip` 7, `search_hybrid` 2,
`search_lexical` 39, `search_vector` 5 and `validate` 41. They use seven 5.0 files built
from public-domain texts, two authentic legacy files (4.0 and 4.1, gzip-wrapped, with
FTS triggers), and invalid files that each break one rule. How each expectation is
obtained (SQLite as the oracle for lexical search, exact arithmetic for vectors,
hand-reviewed URIs and citations) is described in
[`conformance/README.md`](../../conformance/README.md); the changes are in
[`conformance/CHANGELOG.md`](../../conformance/CHANGELOG.md).

## Backwards compatibility

- **5.0 is a new major version.** A 4.x reader cannot read 5.0 files: the schema is
  different and the container is no longer gzip. This is intended.
- **5.0 readers read 4.x.** Every conforming 5.0 reader reads 4.0 and 4.1 files through
  the 5.0 view (§20). Nothing that a 4.x file holds about the document is lost; only the
  Scholaris shelf state (`estado`, `bibliotecas`) is dropped.
- **Writers** produce 5.0 only. Scholaris keeps reading its 4.x files and exports 5.0
  through the `spdf-format` library.
- **Sidecars** are new and carry their own version (`spdf_library: "1.0"` in
  manifests).
- **Nothing is deprecated** within 5.0. The obligation to read 4.x lasts for the whole
  5.x line; a future major version decides by RFC whether its readers still read 4.x
  (see [`governance/VERSIONING.md`](../../governance/VERSIONING.md)).

## Security and privacy

Sections 14 and 15 of the specification are new in 5.0. In short:

- Files contain no triggers, views or foreign virtual tables, and readers refuse files
  that do, so opening a file runs no code written by its author. Readers open read-only,
  with `trusted_schema` off, defensive mode on and extensions disabled.
- The 5.0 container is not compressed. Legacy gzip input is decompressed within a
  limit to defeat decompression bombs; values and JSON nesting are bounded.
- Text is light Markdown rendered without raw HTML; remote references are never fetched
  automatically; blob keys are sanitized before anything is written to disk; text passed
  to language models is data, not instructions.
- Vectors can leak the text they were computed from; provenance can leak details of the
  producer. Rights and confidentiality rules that apply to a text apply to its vectors.
- `content_sha256` and the Ed25519 signature give integrity and attribution of the
  content, not confidentiality, and say nothing about whether the signer's key should be
  trusted.

## Alternatives

- **Keep the gzip wrapper.** Rejected: it prevents HTTP range reads and memory mapping,
  forces full decompression and adds decompression-bomb handling, for little gain on
  already-compressed images and originals. Readers still accept it for legacy files.
- **A ZIP package with JSON and a SQLite index inside** (as EPUB or OOXML do). Rejected:
  two levels of parsing, and the full-text index needs SQLite anyway; SQLite alone gives
  random access, an index and a single file.
- **Pure JSON, Parquet or Arrow.** Rejected: no built-in full-text search, and either no
  random access (JSON) or no good fit for long text with anchors (columnar formats).
- **Keep Spanish identifiers with English aliases.** Rejected: two names for every
  table and column would double the surface of every implementation. The specification
  keeps a faithful Spanish translation instead.
- **UTF-16 code units, bytes or grapheme clusters for offsets.** Rejected: UTF-16 ties
  the format to JavaScript, bytes to one encoding, and grapheme clusters to a Unicode
  version. Code points over NFC are stable and cheap everywhere.
- **IIIF-style regions (`pct:`) or bare fractions in `xywh`.** Rejected in favour of W3C
  Media Fragments (`percent:`); in Media Fragments bare numbers mean pixels, so bare
  fractions would have been misread. Mapping to IIIF Image API regions is trivial.
- **Signing the file bytes, or wrapping the file in JWS or COSE.** Rejected: SQLite
  writes its own version into the header and `VACUUM` reorders pages, so byte signatures
  break without any change in content. A signature over the canonical content is stable.
  An envelope format can still be added later as an extension.
- **Allowing triggers and views.** Rejected for safety; writers can keep the index in
  sync without them.

## Unresolved questions

- The registration of `application/vnd.spdf+sqlite3` with IANA, and whether to use the
  registered `+sqlite3` structured syntax suffix instead (`application/vnd.spdf+sqlite3`).
  Drafts of this and other registrations are in `governance/drafts/`.
- The provisional registration of the `spdf` URI scheme (RFC 7595), which §5.4 announces.
- Whether the anchor URI parameters may also be used as the fragment identifier of a
  URL that points to a `.spdf` file (`https://example.org/x.spdf#p=29`), which the media
  type registration would like to say.
- How a reader and a validator of an earlier minor version treat an anchor type, a dtype
  or a validation rule introduced by a later minor (today an unknown anchor type is
  error E041), so that the compatibility promise holds for validators too.
- Markup for mathematics and other non-textual content inside unit text, which 5.0 does
  not specify.
- A persistent identifier (DOI) for each published version of the specification.

## Implementations

This RFC becomes Implemented when two independent implementations pass the whole
conformance suite in CI, as their `conformance-<folder>` artifacts show. The table
records what each implementation reported in its own commits on 2026-10-07; several
independent implementations already report the whole suite, so the editor will mark the
RFC Implemented once their CI artifacts confirm it.

| Implementation | Folder | Reported on 2026-10-07 |
|---|---|---|
| Rust (reference, also the C ABI) | `rust/` | 228 of 228 (commit `38c75e5`) |
| TypeScript (`spdf-format`, npm) | `js/` | 228 of 228, also in Chromium and Bun (commit `e063594`) |
| Python (`spdf-format`, PyPI) | `python/` | 228 of 228 (commit `557af8f`) |
| Go | `go/` | 228 of 228 (commit `9355f22`) |
| Swift | `swift/` | 228 of 228 on macOS and the iOS simulator (commit `54d5525`) |
| PHP and Ruby | `php/`, `ruby/` | 228 of 228 (commit `537202e`) |
| Kotlin/JVM, C#, R | `kotlin/`, `dotnet/`, `r/` | in development |
| Julia, C (over the Rust ABI) | `julia/`, `c/` | planned |
| Reference producer `spdf build` (Python) | `producer/` | in development |
| Scholaris (second producer) | external | reads 4.x; 5.0 export through `spdf-format` planned |
