Skip to content
SPDF 5.0

SPDF 5.0 specification

The normative specification of the SPDF 5.0 format: container, schema, anchors and their URI, CSL-JSON metadata, search, vectors, integrity, validation, citation and conformance.

Reviewed Markdown

  • Status: Working Draft, 2026-10-07. Stable enough to implement; changes go through the RFC process (spec/rfcs/) and are logged in spec/CONTRACT.md until 5.0 is final.
  • Editor: José Luis Saorín Ferrer.
  • This version: spec/SPEC.md in https://github.com/joseluissaorin/spdf.
  • Spanish translation: SPEC.es.md (faithful; in case of conflict the English text prevails).
  • License: this specification is published under CC BY 4.0. Code in the repository is MIT OR Apache-2.0. Contributors commit not to assert patents against implementations.

Abstract

SPDF is an open, portable file format for documents that have been read once and can be cited forever. A .spdf file holds the text of one document (a printed book, a scan, a recording, a slide deck, a spreadsheet, a web page) as a set of citable units and searchable fragments, and every fragment carries an exact anchor: the printed page or leaf, the second of a recording, the slide, the verse, the canonical reference. A citation produced from an SPDF file can only print what the source says. The container is a plain SQLite 3 database, the metadata is a CSL-JSON item, the full-text index uses tokenizers that ship with every SQLite, and optional embedding vectors from several models can live side by side. Any language with SQLite can read SPDF without special libraries.

Status of this document

This is the first public version of the format (earlier versions, 3.0 to 4.1, were internal to Scholaris and are covered as legacy in §20). The conformance suite in conformance/ is part of the specification: where this text and a conformance case disagree, the disagreement is a bug to be resolved through the RFC process; until it is, implementations follow the conformance case.

Contents

Preface

SPDF was born inside Scholaris, an application written by José Luis Saorín Ferrer to insert verified, page-exact citations into academic writing. Scholaris needed to read a source once (with a PDF text layer, a vision model, or a speech recognizer), keep what it had read, and answer for years afterwards the only question a citation must answer honestly: where exactly does the source say this? The answer had to survive the original file being moved, the reading model being replaced and the search engine being rewritten. The result was a file per document, the Scholaris Processed Document Format, which went through a gzip-compressed JSON-and-SQLite version (3.0) and a Spanish-named SQLite schema (4.0 and 4.1).

Version 5.0 is the first version designed for everyone. It keeps what experience proved right and drops what tied it to one program: identifiers are in English, the container is uncompressed so it can be memory-mapped and read by HTTP ranges, the metadata is plain CSL-JSON so that Zotero, citeproc and Pandoc understand it, and every number in a conformance case comes from an oracle that anyone can rerun. The name became Semantic Processed Document Format; the initials did not change.

Five principles guide every decision in this specification:

  1. Anchors first. Every fragment knows exactly where it comes from: physical page and printed folio, leaf and side, second, slide, verse, canonical reference. Nothing that cannot be anchored is citable.
  2. Provenance. A file says who read each unit and with what confidence, which model produced each vector, and how each metadata field was obtained. Derived data can be recomputed from the original plus the units.
  3. Read once, query many times. Reading a document is expensive (vision models, speech recognition, human correction); querying it must be cheap, offline, and possible from any language with SQLite.
  4. Portability. One file, one document, no server, no proprietary dependency, no compression layer to undo, no code inside the file. Readers in many languages pass the same conformance suite.
  5. Honest citation. A citation prints only what an anchor says. A folio that was inferred is printed in brackets; an unnumbered page is cited as unnumbered; the modernized spelling used for search is never quoted.

1. Conventions and terminology

The key words "MUST", "MUST NOT", "REQUIRED", "SHALL", "SHALL NOT", "SHOULD", "SHOULD NOT", "RECOMMENDED", "NOT RECOMMENDED", "MAY", and "OPTIONAL" in this document are to be interpreted as described in BCP 14 [RFC 2119] [RFC 8174] when, and only when, they appear in all capitals, as shown here.

ABNF follows [RFC 5234]. JSON follows [RFC 8259]; "JSON object", "array", "string" and "number" have their RFC 8259 meanings. SQL follows SQLite's dialect.

  • Document: the work an SPDF file describes (one per file).
  • Original: the bytes the document was read from (PDF, image set, audio, EPUB…).
  • Unit: a citable division of the document: a page or leaf, a time span, a slide, a section, a sheet range. Units are ordered and numbered from 1.
  • Fragment: a searchable, citable passage of roughly 150 to 300 words, with the anchor of its start and, if it crosses units, of its end.
  • Anchor: a JSON object that locates a unit, a fragment or a figure in the document (§4).
  • Anchor URI: the textual form of an anchor, spdf:<docref>#<params> (§5).
  • Space: a vector space, i.e. the model, dimensions and encoding that produced a set of embedding vectors (§9).
  • Reader: software that opens SPDF files and exposes their content. Writer: software that creates SPDF files. Validator: software that checks files against this specification. Producer: a writer that also reads originals (OCR, speech recognition, embeddings).
  • Code point: a Unicode scalar value. Lengths and offsets in this specification count code points, never bytes or UTF-16 code units.
  • NFC: Unicode Normalization Form C [UAX #15].
  • JCS: the JSON Canonicalization Scheme [RFC 8785].

2. Container

2.1 File

An SPDF 5.0 file is a SQLite 3 database file [SQLITE-FORMAT] holding exactly one document. It MUST NOT be wrapped in any compression or archive layer: the database header MUST start at byte 0.

Writers MUST set:

  • PRAGMA application_id = 1397769286 (hexadecimal 0x53504446). SQLite stores it big-endian at byte offset 68 of the header, so bytes 68 to 71 read "SPDF" in ASCII.
  • PRAGMA user_version = 500. The value encodes the specification version as major × 100 + minor × 10 (5.0 → 500, 5.1 → 510).
  • The spdf_meta row spdf_version to "5.0" (§3.2).

Writers SHOULD use a page size of 4096 bytes, the rollback journal in DELETE mode (never leave a -wal or -journal file next to a distributed file), and run VACUUM after the last write so the file has no free pages. Writers SHOULD NOT use auto_vacuum.

A file MUST NOT contain triggers or views, and MUST NOT contain virtual tables other than the FTS5 tables defined in §3. Writers keep the full-text index in sync themselves (for example with INSERT INTO fragments_fts(fragments_fts) VALUES('rebuild') before VACUUM).

2.2 Name and type

The file extension is .spdf. The media type is application/vnd.spdf+sqlite3 (§24). One file holds one document; libraries of documents are described by a separate collection manifest (§17).

2.3 Gzip input

Legacy 4.x files are SQLite databases wrapped in gzip [RFC 1952] (§20). Readers MUST therefore accept a file that starts with the gzip magic bytes 1F 8B, decompress it (to memory or to a temporary file) with a configurable limit on the decompressed size (RECOMMENDED default 4 GiB), and continue with the result. A gzip-wrapped 5.0 file is readable but non-conforming: validators report it as E003 in the warnings list (§22).

2.4 Safe opening

SPDF files come from strangers. Every reader MUST open them as follows, and MUST refuse the file if a step cannot be honoured by its SQLite binding:

  1. Open the database read-only (SQLITE_OPEN_READONLY, or the URI parameter mode=ro). Never open a distributed file read-write in place.
  2. PRAGMA query_only = 1 and PRAGMA trusted_schema = OFF.
  3. Enable SQLITE_DBCONFIG_DEFENSIVE where the binding exposes it, and keep extension loading disabled (sqlite3_enable_load_extension(db, 0); never call load_extension).
  4. Read sqlite_master and refuse the file if it contains a trigger, a view, or a virtual table other than fragments_fts and fragments_fts_trigram declared USING fts5. Legacy 4.x files are allowed exactly the three triggers fragmentos_ai, fragmentos_ad and fragmentos_au (§20), which never fire on a read-only connection.
  5. Enforce a configurable maximum size for any single BLOB or TEXT value read (RECOMMENDED default 512 MiB), for example with sqlite3_limit(db, SQLITE_LIMIT_LENGTH, …).

Readers SHOULD also disable memory-mapped I/O (PRAGMA mmap_size = 0) and enable PRAGMA cell_size_check = ON for files from untrusted sources, and MAY run PRAGMA quick_check before use. Operations that need to write, such as the FTS5 integrity-check command, MUST run on a private copy (for example an in-memory copy made with the backup API), never on the file. §14 explains the threats.

3. Schema

3.1 Overview

The normative schema is the SQL script schema/spdf-5.0.sql, reproduced in full below. Every table in it is REQUIRED, even when empty; only fragments_fts_trigram is OPTIONAL. Column names, types and constraints MUST be as written. Writers MUST NOT add columns to these tables; data that does not fit goes into extension tables (§11). Readers MUST ignore columns they do not know (a later minor version may add OPTIONAL columns, §23).

JSON stored in TEXT columns MUST be valid JSON [RFC 8259] encoded in UTF-8; writers MAY serialize it in any form (the canonical dump re-serializes it, §12). Timestamps are ISO 8601 / RFC 3339 strings in UTC with a Z suffix. Identifiers (id columns) are non-empty strings chosen by the writer; they are opaque, case-sensitive and stable for the life of the file.

PRAGMA application_id = 1397769286;  -- 0x53504446, "SPDF"
PRAGMA user_version = 500;

CREATE TABLE spdf_meta (key TEXT PRIMARY KEY, value TEXT NOT NULL);

CREATE TABLE documents (
  id TEXT PRIMARY KEY, kind TEXT NOT NULL, metadata TEXT NOT NULL,
  source_sha256 TEXT NOT NULL, source_ref TEXT, mime TEXT NOT NULL,
  bytes INTEGER NOT NULL, unit_count INTEGER NOT NULL, duration REAL,
  created TEXT NOT NULL, updated TEXT NOT NULL,
  title TEXT, authors TEXT, year INTEGER, language TEXT, rights TEXT);

CREATE TABLE units (
  id TEXT PRIMARY KEY, document TEXT NOT NULL REFERENCES documents(id),
  ord INTEGER NOT NULL, anchor TEXT NOT NULL, text TEXT NOT NULL DEFAULT '',
  notes TEXT, header TEXT, footer TEXT, image TEXT, thumbnail TEXT,
  reader TEXT NOT NULL, confidence REAL NOT NULL DEFAULT 1, printed TEXT,
  t0 REAL, t1 REAL, words TEXT);
CREATE INDEX units_doc ON units(document, ord);
CREATE INDEX units_printed ON units(document, printed);

CREATE TABLE sections (
  id TEXT PRIMARY KEY, document TEXT NOT NULL, parent TEXT, level INTEGER NOT NULL,
  title TEXT NOT NULL, unit_from TEXT NOT NULL, unit_to TEXT, summary TEXT);

CREATE TABLE fragments (
  n INTEGER PRIMARY KEY, id TEXT NOT NULL UNIQUE, document TEXT NOT NULL,
  unit TEXT NOT NULL, ord INTEGER NOT NULL, text TEXT NOT NULL,
  context TEXT NOT NULL DEFAULT '', section TEXT, anchor TEXT NOT NULL,
  anchor_end TEXT, search_text TEXT);
CREATE INDEX fragments_doc ON fragments(document, ord);
CREATE INDEX fragments_unit ON fragments(unit);

CREATE VIRTUAL TABLE fragments_fts USING fts5(
  text, context, section, search_text,
  content='fragments', content_rowid='n',
  tokenize='unicode61 remove_diacritics 2');
-- OPTIONAL:
-- CREATE VIRTUAL TABLE fragments_fts_trigram USING fts5(
--   text, content='fragments', content_rowid='n', tokenize='trigram');

CREATE TABLE figures (
  id TEXT PRIMARY KEY, document TEXT NOT NULL, unit TEXT NOT NULL,
  image TEXT NOT NULL, caption TEXT, description TEXT, anchor TEXT NOT NULL);

CREATE TABLE spaces (
  id TEXT PRIMARY KEY, provider TEXT NOT NULL, model TEXT NOT NULL, version TEXT,
  dims INTEGER NOT NULL, dtype TEXT NOT NULL DEFAULT 'f32',
  normalized INTEGER NOT NULL DEFAULT 1, truncated_from INTEGER,
  modalities TEXT NOT NULL, task_prefixes TEXT, created TEXT);

CREATE TABLE vectors (
  target TEXT NOT NULL, id TEXT NOT NULL, space TEXT NOT NULL REFERENCES spaces(id),
  document TEXT NOT NULL, data BLOB NOT NULL, PRIMARY KEY (target, id, space));

CREATE TABLE blobs (key TEXT PRIMARY KEY, mime TEXT NOT NULL, sha256 TEXT NOT NULL,
  data BLOB NOT NULL);

CREATE TABLE provenance (
  document TEXT NOT NULL, stage TEXT NOT NULL, provider TEXT, model TEXT,
  detail TEXT, ms INTEGER, at TEXT NOT NULL);

CREATE TABLE extensions (name TEXT PRIMARY KEY, version TEXT NOT NULL,
  required INTEGER NOT NULL DEFAULT 0);

3.2 spdf_meta

Key/value pairs about the file. REQUIRED keys:

keyvalue
spdf_version"5.0"
profilespace-separated profile names, a subset of core semantic media full (§10); always includes core
createdcreation time of the file (UTC)
generatorname/version of the writer, e.g. spdf-producer/0.3.1
document_idequal to documents.id

OPTIONAL keys: content_sha256, signature, signer (§13) and license_note (free text for humans). Other keys MAY be added by later versions or by extensions (prefixed x_<vendor>_); readers MUST ignore keys they do not know.

3.3 documents

Exactly one row.

  • kind: one of pdf (PDF with a usable text layer), scanned_pdf (PDF read by vision), photos (a set of page photographs), image (a single image), audio, video, document (DOCX, ODT, RTF, HTML, Markdown, plain text), epub, slides, sheet, web. Readers MUST accept unknown kinds and treat them as document.
  • metadata: the CSL-JSON item with the spdf extension (§6).
  • source_sha256: lowercase hexadecimal SHA-256 of the original bytes. It identifies the document across copies and is the preferred document reference in anchor URIs.
  • source_ref: where the original is: blob:<key> when shipped inside the file, an absolute URL, or NULL.
  • mime, bytes: media type and size in bytes of the original.
  • unit_count: number of rows in units (a mismatch is warning W102).
  • duration: seconds, for audio and video; NULL otherwise.
  • created, updated: when the document record was created and last changed.
  • title, authors, year, language: denormalized copies for filtering without parsing JSON: the CSL title; the family names (or literal names) of the CSL author list joined with "; "; the first year of issued; the CSL language. They MUST agree with metadata when present.
  • rights: JSON rights object (§16) or NULL.

3.4 units

One row per citable unit, ord = 1, 2, 3… without gaps (E090), in reading order.

  • anchor: the unit's anchor (§4).
  • text: the full text of the unit as read, NFC, light Markdown (§7). Empty string for units without text (a blank page, a photograph).
  • notes: JSON array of strings (footnotes detached from the body) or NULL.
  • header, footer: running heads and feet, kept out of text, or NULL.
  • image, thumbnail: blob:<key> or URL of the unit's image (page, frame, slide) and of its thumbnail, or NULL.
  • reader: what produced text (pdf-text-layer, tesseract-5, gemma-4-e4b, whisper-large-v3-turbo, human…). confidence: 0 to 1.
  • printed: the printed folio of a page unit, copied from its anchor, so readers can "go to page 145" with an index.
  • t0, t1: start and end in seconds, copied from a time anchor; NULL otherwise.
  • words: word timings for audio and video (§7.4) or NULL.

3.5 sections

The heading tree. level starts at 1; parent is the id of the enclosing section or NULL; unit_from and unit_to are the first and last unit ids (unit_to NULL when the section ends with the document); summary is OPTIONAL text in the document language.

3.6 fragments

  • n: a positive integer, unique, stable: it is the rowid the FTS5 index uses (an implicit rowid may change on VACUUM).
  • unit: id of the unit where the fragment starts. ord: reading order within the document (increasing with the position of the fragment in the text).
  • text: the literal passage, NFC, exactly as in the source (never modernized).
  • context: one line that situates the fragment in the work ("Chapter III: the struggle for existence"), used by search; empty string if none.
  • section: JSON array of strings, the heading path, or NULL.
  • anchor: anchor of the start of the fragment. anchor_end: anchor of its end when it crosses into another unit; NULL otherwise.
  • search_text: the modernized-spelling layer (§25.3): text used ONLY for search (aſsi → así, V. M. → vuestra merced). Empty string when it adds nothing; NULL when not computed. It MUST NOT be displayed as the text of the source or quoted.

3.7 fragments_fts and fragments_fts_trigram

fragments_fts is an external-content FTS5 index over fragments with the columns text, context, section and search_text in this order and the tokenizer unicode61 remove_diacritics 2, which every SQLite with FTS5 provides. It MUST be in sync with fragments (E070). fragments_fts_trigram is OPTIONAL, indexes text only with the trigram tokenizer (SQLite 3.34 or later), and SHOULD be present when the document is mostly in Chinese, Japanese or Korean.

3.8 figures

Figures, plates, tables as images, photographs inside a page. image is blob:<key> of a cropped image, or the unit's image together with a region in the anchor. caption is the printed caption, if any; description is a description in the document language (for accessibility and search). anchor normally carries a region.

3.9 spaces and vectors

See §9. vectors.target is fragment, unit or figure and vectors.id is the id of that row; data is the little-endian vector.

3.10 blobs

Binary content shipped inside the file: the original, page images, cropped figures, thumbnails. key is an opaque string (by convention path-like, pages/0001.png), mime its media type, sha256 the lowercase hex SHA-256 of data (E080). Other tables refer to a blob as blob:<key>.

3.11 provenance

One row per production step: stage (reading, transcription, folios, metadata, embedding, figures…), provider, model, detail (JSON object or NULL), ms (duration in milliseconds) and at (UTC timestamp). See §15 for what not to record.

3.12 extensions

See §11.

4. Anchors

4.1 General

An anchor is a JSON object with a string member type. The types defined by this version and their members are:

typeREQUIRED membersOPTIONAL members
pagephysical (integer ≥ 1), printed (string or null)roman (boolean), foliation (page, leaf, column; default page), source (read, inferred, epub, none), confidence (0–1)
timet0, t1 (seconds, 0 ≤ t0 ≤ t1)speaker (string)
sectionpath (array of strings)paragraph (integer ≥ 1), printed (string)
sliden (integer ≥ 1)
sheetsheet (string), row_from, row_to (integers)
weburl (string)path, paragraph, accessed (ISO date)
image
verseline_from (integer)line_to (integer), printed (string)
canonicalscheme (string), ref (string)

Every anchor MAY also carry:

  • region: {"x", "y", "w", "h"}, numbers between 0 and 1, fractions of the width and height of the unit's image, origin at the top left;
  • chars: [start, end], code point offsets into the NFC text of the anchor's unit, 0 ≤ start ≤ end ≤ length, end exclusive (E042);
  • matter: what kind of matter the unit is: body (the text of the work), front (preliminaries: title page, contents, licences, dedication, prologue of an edition), back (index, colophon, appendices of an edition), plate (a plate or fold-out outside the text pages), cover, library (bookplates, stamps, library or digitizer pages, licences of a digital edition) or blank. Absent means body; readers MUST treat values they do not know as body. Writers SHOULD set it on the units of paged documents whenever it is not body.

In this specification an "integer" is a JSON number with an integral value: 10 and 10.0 are the same JSON value and both are integers. An anchor whose JSON is invalid or that lacks or mistypes a REQUIRED member is invalid (E040); an unknown type is E041. Readers MUST preserve members they do not know when they copy anchors.

4.2 Pages, folios and leaves

physical is the 1-based position of the page in the original (the PDF page index, the photo number). printed is the folio exactly as printed on the page ("23", "xiv", "A-3", "1r"), or null when the page carries no number.

  • roman: true marks folios in roman numerals (front matter).
  • foliation describes what the printed numbers count: page (each page numbered), leaf (each leaf numbered, sides recto and verso, printed as "1r", "1v"), or column (columns numbered, as in some dictionaries and early printed books).
  • source says how printed was obtained: read (seen on the page), inferred (deduced from neighbouring pages, e.g. an unnumbered verso), epub (from an EPUB page list), none (no folio; printed is null).
  • An inferred folio is cited in brackets, p. [21]; a page without folio is cited as unnumbered (§18). A producer MUST NOT invent folios: if no evidence supports a number, printed is null and source is none.

4.3 Time, sections, verses and canonical references

Time anchors locate recordings in seconds from the start of the original; speaker names who speaks. Section anchors locate unpaginated text (EPUB, DOCX, HTML) by heading path and paragraph number, and MAY add the equivalent printed page when the edition provides a page list. Verse anchors count lines of verse (line_from, line_to), as printed editions number them. Canonical anchors use a citation system that is independent of any edition: stephanus (Plato), bekker (Aristotle), bible (book chapter:verse), cts (a CTS URN [CTS]), or any other documented scheme; schemes are lowercase ASCII.

4.4 Start and end

A fragment's anchor locates its start; anchor_end, when present, locates its end and has the same type. A citation of the whole fragment then prints a range (pp. 145-146).

  • The end unit of a fragment is the first unit after its start unit (in ord order) whose anchor equals anchor_end once chars and region are removed from both.
  • chars in anchor gives the part of the fragment that lies in the start unit, and chars in anchor_end the part that lies in the end unit (usually [0, b]). Writers SHOULD set both on crossing fragments, so that readers know which unit each part of the passage comes from.
  • Writers SHOULD NOT let a fragment cross from a unit of one matter to a unit of another (body text into a plate, a cover, a library page or a licence), nor from a page with a printed folio to a page without one: the citation of such a fragment would mix locators of different natures. Validators report such fragments as W103.

5. Anchor URI

5.1 Syntax

An anchor URI names a place in a document independently of any file:

spdf:sha256-3f2a…c9#p=29&f=21&char=118,301

The document reference is sha256- followed by the 64 lowercase hex digits of documents.source_sha256 (RECOMMENDED: it is the same for every copy of the document), or the percent-encoded document id. The fragment is a list of parameters. The parameters reuse W3C Media Fragments syntax [MEDIA-FRAGMENTS] for time (t=) and space (xywh=) and RFC 5147 [RFC 5147] syntax for character ranges (char=), so tools that know those standards can interpret them.

The canonical form is defined by this ABNF [RFC 5234]:

spdf-uri    = "spdf:" docref [ "#" params ]
docref      = hash-ref / id-ref
hash-ref    = "sha256-" 64lhex
lhex        = DIGIT / %x61-66                     ; 0-9 a-f
id-ref      = 1*vchar                             ; percent-encoded document id
params      = param *( "&" param )
param       = p / pe / f / fe / t / s / para / sl / sh / rows / v / ref / char / xywh
p           = "p=" posint                         ; physical page
pe          = "pe=" posint                        ; physical end page
f           = "f=" value                          ; printed folio
fe          = "fe=" value                         ; printed end folio
t           = "t=" number [ "," number ]          ; seconds, Media Fragments npt
s           = "s=" value *( "/" value )           ; section path
para        = "para=" uint                        ; paragraph
sl          = "sl=" posint                        ; slide
sh          = "sh=" value                         ; sheet name
rows        = "rows=" uint "-" uint               ; sheet rows
v           = "v=" uint [ "-" uint ]              ; verse lines
ref         = "ref=" value ":" value              ; canonical scheme ":" reference
char        = "char=" uint "," uint               ; code points, RFC 5147 style
xywh        = "xywh=percent:" number "," number "," number "," number
value       = *vchar
vchar       = unreserved / pct-encoded
unreserved  = ALPHA / DIGIT / "-" / "." / "_" / "~"
pct-encoded = "%" HEXDIG HEXDIG                   ; uppercase in the canonical form
posint      = %x31-39 *DIGIT
uint        = "0" / posint
number      = uint [ "." 1*DIGIT ]

In the canonical form, parameters appear at most once and in the order of the param rule above (p, pe, f, fe, t, s, para, sl, sh, rows, v, ref, char, xywh); values are UTF-8 strings in which every byte other than an unreserved character is percent-encoded with uppercase hex digits; in s the separators between path elements are literal / and a / inside an element is %2F; in ref the first literal : separates the scheme from the reference, and colons inside them are %3A. Numbers use the shortest decimal form of ECMAScript (4160, 4175.5, 0.125), never an exponent.

5.2 From an anchor to a URI

Formatting maps an anchor (and optionally an end anchor) to parameters:

anchorparameters
pagep = physical; f = printed if not null; with an end page: pe = its physical if different, fe = its printed if not null and different from printed
timet = t0, then t1 (or the end anchor's t1)
section, webs = path if not empty; para = paragraph; f = printed; fe as for pages
slidesl = n
sheetsh = sheet; rows = row_from-row_to
versev = line_from, or line_from-line_to when line_to is present and different; f = printed
canonicalref = scheme:ref
imagenone
anychar = chars; xywh = region × 100, as percent:

t values are rounded to 6 decimal places. xywh values are fractions × 100 rounded to 4 decimal places (0.125 → 12.5, 0.333333 → 33.3333). A URI without parameters (spdf:<docref>) designates the whole document.

5.3 Parsing

Parsing returns the document reference and a locator object with one member per parameter present: p, pe, para, sl (integers); f, fe, sh (strings); t (array of one or two numbers); s (array of strings); rows (two integers); v (one or two integers); ref (object with scheme and ref); char (two integers); xywh (four fractions, the percent values divided by 100 and rounded to 6 decimals).

Parsers MUST accept percent-encoding with lowercase hex digits, parameters in any order, unencoded non-ASCII characters (IRI form [RFC 3987]), the npt: prefix and the clock forms h:mm:ss[.f] and mm:ss[.f] in t. Parsers MUST ignore parameters whose names they do not know. Parsers MUST reject: a scheme other than spdf:; an empty document reference; a repeated parameter; malformed numbers; p, pe or sl equal to 0; a char or t range whose end precedes its start; xywh without the percent: unit (pixel coordinates cannot be resolved without the image); percent-encoding that does not decode to valid UTF-8.

Formatting a parsed locator MUST give back the canonical URI byte for byte. The conformance suite checks format, parse and round trip for every anchor type.

5.4 Resolution

locate(file, reference) resolves an anchor URI, or the URL of an SPDF resource with a fragment identifier (§24), against a file, and returns:

{"document": true, "units": ["p5", "p6"], "fragments": ["q4"], "char": [101, 278], "xywh": null}
  1. Reference. An spdf: URI is parsed as in §5.3; document is true when its document reference is sha256- followed by the file's source_sha256, or the file's document id. Any other reference (an https: URL, a file path) designates the file itself: document is true and the text after its first #, if any, is parsed as the parameter list of §5.3. When document is false, units and fragments are empty (implementations MAY report this as an error instead; conformance runners map such an error to document: false).
  2. Rule. The first parameter present in the order p, f, t, sl, v, ref, s, sh selects the predicate below. Without any of them (no fragment, or only char and xywh) the reference designates the whole document and units and fragments are empty.
  3. Predicate on an anchor (members absent from the anchor never match):
    • p: a page anchor with p ≤ physical ≤ pe (pe defaults to p);
    • f: printed equal to f (for units, the units.printed column);
    • t: a time anchor with t0 ≤ t < t1, where t is the first value of the parameter; the last unit with a time anchor (in ord order) also matches when t equals its t1;
    • sl: a slide anchor with n = sl;
    • v: a verse anchor with line_from ≤ v ≤ line_to (line_to defaults to line_from), where v is the first value of the parameter;
    • ref: a canonical anchor with the same scheme and ref;
    • s: a section or web anchor whose path starts with the elements of s; when para is present, the path must equal s and paragraph must equal para;
    • sh: a sheet anchor with sheet = sh and, when rows is present, row_from ≤ a ≤ row_to for its first value a.
  4. Matches. units are the ids of the units whose anchor matches, in ord order. fragments are the ids of the fragments whose start anchor or whose anchor_end matches, in n order (a fragment that ends on a page is found from that page). When no unit matches but some fragments do, units are the distinct start units of those fragments, in ord order.
  5. Characters. char refers to the text of the first unit of units. When char = [c, d] is present, a fragment is kept in fragments only if its start unit is that unit and its anchor has chars = [a, b] that overlap the range, or its end unit (§4.4) is that unit and its anchor_end has chars that overlap it; [a, b] overlaps [c, d] when a < d and c < b (for c < d), or when a ≤ c < b (for c = d).
  6. char and xywh are copied from the locator, or null.

Several units may match (two pages printed "1", a verse number repeated in two poems): locate returns them all and the reader lets the user choose; p always disambiguates pages, which is why formatted URIs carry it.

The spdf URI scheme is intended for provisional registration [RFC 7595]; the request is drafted in governance/drafts/uri-scheme-spdf.md.

6. Metadata

6.1 CSL-JSON item

documents.metadata is one CSL-JSON item [CSL-JSON] describing the document as it should be cited: at least type (a CSL type such as book, article-journal, chapter, thesis, speech, interview, broadcast, motion_picture, webpage, dataset, graphic) and title (E051 if either is missing). Common members: author, editor, translator, interviewer (arrays of names {family, given} or {literal}, with the CSL particles non-dropping-particle and dropping-particle when needed), issued ({"date-parts": [[year, month, day]]}), original-date, title-short, original-title, container-title, collection-title, publisher, publisher-place, volume, issue, page, edition, DOI, ISBN, ISSN, URL, accessed, language (BCP 47), abstract, note. The id member is OPTIONAL inside the file; exports set it (§19).

Writers MUST NOT invent metadata. A field that cannot be supported by the original or by a cited external source is omitted.

6.2 The spdf extension object

The member spdf of the item holds what CSL cannot express. All its members are OPTIONAL:

"spdf": {
  "provenance": {"title": {"source": "title-page", "confidence": 0.99},
                 "issued": {"source": "colophon", "confidence": 0.95}},
  "undated": {"from": 1600, "to": 1610, "basis": "printer active years"},
  "original_language": "fr",
  "subtitle": "con anotaciones de Fernando de Herrera",
  "orcid": {"Foucault, Michel": "0000-0000-0000-0000"}
}
  • provenance: per CSL field, where the value came from (reading, title-page, colophon, crossref, openalex, wikidata, user, epub, pdf, …) and a confidence between 0 and 1.
  • undated: for works without a printed date, a plausible range (from, to, years, negative for BCE) and the evidence (basis). It MUST NOT be copied into issued: a citation prints "s. f." / "n.d." (§18).
  • original_language: BCP 47 tag of the original language of a translation.
  • subtitle: the subtitle when the CSL title is "Title: Subtitle".
  • orcid: ORCID identifiers by name ("Family, Given").

Other members MAY be added by extensions with the prefix x_<vendor>_.

7. Text, normalization and offsets

7.1 Encoding and normalization

All text is UTF-8 in NFC. Writers MUST normalize to NFC before storing and before computing offsets. Writers MUST NOT store U+0000, unpaired surrogates or noncharacters, and SHOULD NOT store other control characters except U+0009 (tab) and U+000A (line feed). Lines end with U+000A only.

7.2 Offsets

chars offsets (§4.1) and every length in this specification count code points of the NFC text. Implementations whose strings are UTF-16 (JavaScript, Java, C#, Swift's NSString) MUST convert: a character outside the Basic Multilingual Plane counts as one code point but two UTF-16 units.

7.3 Light Markdown

units.text MAY use this subset of CommonMark [COMMONMARK]: paragraphs separated by a blank line; # to ###### headings; *emphasis* and **strong**; - and 1. lists; > quotations; tables in GitHub style; footnote markers [^1] whose text goes to notes. Readers MUST NOT interpret raw HTML in text; they display it as text. Offsets count the stored characters, markup included. Fragments SHOULD keep the markup of their source unit so that fragments.text is a substring of the unit's text whenever the fragment does not cross units.

Speaker turns in transcripts start with the label **Name:** followed by one space (**Neil Armstrong:** Houston, Tranquility Base here.).

7.4 Word timings

units.words is the JSON object {"v": 1, "t0": <seconds>, "cs": [start, duration, start, duration, …]}: one pair of integers per word, in centiseconds from t0 (which equals the unit's t0). The words are the maximal runs of non-whitespace characters of the unit's text after removing the speaker labels (**Name:**); cs therefore holds exactly twice as many integers as there are words. Readers use it to highlight the word being spoken and to turn a char range into a time range.

The reference search defines what conformance tests: results that every implementation returns identically from the same file. Products MAY rank better (stopwords, query expansion, reranking, filters); they MUST still offer the reference behaviour to pass the suite, and SHOULD label the difference in their documentation.

A result item is {"fragment_id", "score", "via", "anchor", "anchor_uri"} where via lists the contributing methods ("lexical", "vector") in that order, anchor is the fragment's anchor and anchor_uri the URI formatted from anchor and anchor_end with the sha256- document reference.

Given a query string and a limit:

  1. Normalize: q = NFC(query).
  2. Phrases: scan q from left to right. An opening mark " (U+0022), “ (U+201C), « (U+00AB) or „ (U+201E) opens a phrase that the next ", ” (U+201D), » (U+00BB) or, for „, “ or ” respectively closes. The text between the marks is the phrase. An opening mark without a closing mark is treated as a separator.
  3. Words: maximal runs of characters whose Unicode general category is a letter (L), a mark (M) or a number (N). A phrase term is the phrase's words joined with one space; phrases without words are dropped.
  4. Terms: if there is at least one phrase term, the terms are the phrase terms (loose words outside quotes are discarded) and the operator is AND. Otherwise the terms are the words of q and the operator is OR. Duplicate terms are removed, keeping the first, comparing them by the key lower(remove_Mn(NFD(term))) (Unicode default lowercase, after removing nonspacing marks); the key is used only to detect duplicates. No stopwords are removed. Without terms the result is empty.
  5. MATCH string: every term, as written (no case folding, no decomposition), is an FTS5 string: " + the term with each " doubled + "; the strings are joined with AND or OR. The tokenizer folds case and diacritics itself; folding the query beforehand would break matches (Straße, fin).
  6. Query:
    SELECT f.n, f.id, bm25(fragments_fts, 1.0, 0.5, 0.5, 1.0) AS r
      FROM fragments_fts JOIN fragments f ON f.n = fragments_fts.rowid
     WHERE fragments_fts MATCH ?1 ORDER BY r, f.n LIMIT ?2
    The score is −r. Because search_text is the fourth indexed column, a query in modern spelling finds old spelling without any special step. Since n is the rowid of the index, implementations MAY rank inside the index alone (SELECT rowid, bm25(…) FROM fragments_fts WHERE fragments_fts MATCH ?1 ORDER BY 2, 1 LIMIT ?2) and look up the fragment ids of the returned rows only; the result is identical. Readers that fetch files by HTTP ranges SHOULD do so, as it avoids reading every matching fragment.
  7. CJK route: if q contains a code point in one of the ranges U+2E80–U+2FDF, U+3040–U+30FF, U+3100–U+312F, U+3130–U+318F, U+31A0–U+31FF, U+3400–U+4DBF, U+4E00–U+9FFF, U+A960–U+A97F, U+AC00–U+D7AF, U+F900–U+FAFF, U+FF66–U+FF9F or U+20000–U+3FFFF, step 6 is replaced:
    • if fragments_fts_trigram exists and every term has at least 3 code points, the same MATCH string runs against fragments_fts_trigram, ordered by bm25(fragments_fts_trigram) then n; the score is −bm25;
    • otherwise (no trigram index, or a term shorter than 3 code points, which a trigram index cannot match) the substring fallback runs on fragments.text: for each fragment, hits = the number of terms t with instr(text, t) > 0; fragments with hits ≥ 1 (with the OR operator) or hits = number of terms (with AND) are returned ordered by hits descending, then n; the score is hits.

Given a space, a target (fragment by default, or unit, figure), a query vector of dims numbers and a limit: compare the query with every vector of that space and target by brute force. Every component is converted to an IEEE 754 binary64 number (f32 and f16 exactly; i8 as q/127). The query is used as given, not normalized. When the space has normalized = 1 the score is the dot product; otherwise it is the cosine similarity. Results are ordered by score descending, then by fragment n, unit ord or figure id. A result item for the target unit or figure carries unit_id or figure_id instead of fragment_id, and the anchor URI of the unit's or figure's own anchor. Products MAY use approximate indexes; the reference is exhaustive.

Run the lexical search and the vector search (target fragment), each with depth max(limit, 50), and fuse them by reciprocal rank fusion [RRF] with k = 10: score = Σ 1/(10 + rank) over the lists that contain the fragment, rank starting at 1. Order by score descending, then n; keep limit results. The constant 10 was measured in Scholaris: the classic 60 flattens short, good lists.

8.4 Comparison in conformance

Result order MUST match exactly; scores MUST match within an absolute tolerance of 1e-6. The lexical scores are those of SQLite's own bm25(), which is the oracle.

9. Vector spaces

9.1 Spaces

A row of spaces describes how a set of vectors was produced:

  • id: <model>@<dims> for f32 vectors and <model>@<dims>:<dtype> otherwise (embeddinggemma-2@768, embeddinggemma-2@256:i8). A suffix +<variant> MAY follow to separate vectors of the same model computed from different inputs (legacy Scholaris uses +contexto).
  • provider (who ran the model: local, google, inferbox…), model, version.
  • dims: number of components.
  • dtype: f32 (IEEE 754 binary32), f16 (binary16) or i8 (signed byte; the value is q/127). Other values are invalid (E032).
  • normalized: 1 if every stored vector has unit Euclidean norm (before quantization).
  • truncated_from: for Matryoshka truncation [MRL], the original dimension (768 for a vector cut to 256); NULL otherwise. Truncated vectors SHOULD be renormalized before storing, with normalized = 1.
  • modalities: JSON array of the input modalities the model accepts (text, image, audio, video, pdf).
  • task_prefixes: JSON object with the prefixes or instructions used at encoding time, {"document": "…", "query": "…"}, so that a reader can encode queries the same way; NULL if none.

A file MAY hold several spaces; a file without spaces is valid (profile core).

9.2 Vectors

vectors.data is the vector as dims little-endian values of the space's dtype, so its length is dims × 4, 2 or 1 bytes (E030). Every vector refers to a space in spaces (E031). Writers quantize as follows: f32 → f16 with IEEE round-to-nearest-even; f32 → i8 with q = clamp(round_half_away_from_zero(v × 127), −127, 127). The value −128 is not used. A value that does not fit the dtype (a finite f32 above 65504 that would round to infinity in f16, a non-finite value) is an error for the writer, never silently stored.

9.3 Compatibility between spaces and quantizations

Two spaces are compatible, and one query vector serves both, when provider, model, version, dims, normalized, truncated_from and task_prefixes are equal; dtype may differ. A reader holding a model MAY therefore search an f32 space and its i8 copy with the same query. Spaces that differ in any other field are not comparable: readers MUST NOT mix scores across incompatible spaces, and MUST NOT compare vectors of different dimensions. A Matryoshka space (truncated_from = 768, dims = 256) is compatible with a query only if the query was truncated to the same dimensions and renormalized.

10. Profiles

spdf_meta.profile declares which promises a file makes. Profiles are cumulative labels; a file lists every profile it satisfies.

profilerequirements
coreREQUIRED in every file. All tables of §3; at least one unit; every unit, fragment and figure anchored; text in NFC; FTS index in sync.
semanticAt least one space, and vectors for every fragment in at least one space. A semantic file without vectors raises W100.
mediakind is audio or video; units carry time anchors and t0/t1; duration is set; words SHOULD be present. A media file without time anchors raises W101.
fullsemantic and, for audio and video, media; for paged kinds, page images (units.image) and figures where the original has them.

Readers MUST NOT refuse a file because of its profile; profiles tell readers what to expect and validators what to check.

11. Extensions

Data that this specification does not define goes into extension tables named x_<vendor>_<name> (lowercase ASCII letters, digits and _; <vendor> is a name the author controls, e.g. x_scholaris_claims). Each extension in use is declared in the extensions table with its name (<vendor>_<name> or the table prefix), a version, and required:

  • required = 0: readers that do not know the extension ignore it.
  • required = 1: the file cannot be understood without it; a reader that does not know it MUST refuse the file with E060.

Extensions MUST NOT change the meaning of the core tables, MUST NOT add columns to them, and SHOULD NOT be required. Extension tables are not part of the canonical dump. An extension that proves useful to several implementations becomes part of the core through the RFC process (§23).

12. Canonical dump

The canonical dump is a JSON view of a file that every implementation produces identically. It is the oracle of the conformance suite and the input of integrity hashing.

{"spdf_version": "5.0",            // legacy files: "4.0"/"4.1" and "legacy": true
 "meta": {"<key>": "<value>", …},  // every spdf_meta row
 "fts": {"tokenizer": "unicode61 remove_diacritics 2", "trigram": false},
 "document": {"id", "kind", "metadata", "source_sha256", "source_ref", "mime", "bytes",
              "unit_count", "duration", "created", "updated", "title", "authors",
              "year", "language", "rights"},
 "units": [{"id", "ord", "anchor", "text", "notes", "header", "footer", "image",
            "thumbnail", "reader", "confidence", "printed", "t0", "t1", "words"}],
 "sections": [{"id", "parent", "level", "title", "unit_from", "unit_to", "summary"}],
 "fragments": [{"n", "id", "unit", "ord", "text", "context", "section", "anchor",
                "anchor_end", "search_text"}],
 "figures": [{"id", "unit", "image", "caption", "description", "anchor"}],
 "spaces": [{"id", "provider", "model", "version", "dims", "dtype", "normalized",
             "truncated_from", "modalities", "task_prefixes", "created"}],
 "vectors": {"<space id>": {"count": 6, "sha256": "<hex>"}},
 "blobs": [{"key", "mime", "bytes", "sha256"}],
 "provenance": [{"stage", "provider", "model", "detail", "ms", "at"}],
 "extensions": [{"name", "version", "required"}]}

Rules:

  1. Every member listed is present. SQL NULL becomes null; INTEGER a JSON integer; REAL a JSON number; TEXT a string. Columns that hold JSON (metadata, rights, anchor, anchor_end, notes, words, section, modalities, task_prefixes, detail) are parsed and embedded as JSON values. The document column of child tables is omitted. Booleans stored as integers (normalized, required) stay integers.
  2. Every number that is not an integer, including those inside parsed JSON, is rounded to 6 decimal places (round half to even on its exact binary value); a result of −0 becomes
  3. Order: units by ord; fragments by n; sections, figures and spaces by id; blobs by key; extensions by name (code point order, which is SQLite's BINARY collation over UTF-8); provenance by the UTF-8 bytes of the JCS serialization of each entry. fts.trigram is true if and only if fragments_fts_trigram exists; fts.tokenizer is the value of the tokenize option of fragments_fts as declared, without its quotes and with runs of whitespace collapsed to one space (unicode61 remove_diacritics 2; unicode61, the FTS5 default, if absent).
  4. vectors has one member per distinct vectors.space: count is the number of rows and sha256 the hex SHA-256 of their data blobs concatenated in order of target, then id.
  5. blobs[].bytes and blobs[].sha256 are computed from data, not copied from the sha256 column.
  6. Serialization, whenever bytes matter (hashing), is JCS [RFC 8785]: no whitespace, object members sorted by the UTF-16 code units of their names, numbers in the ECMAScript form (1, not 1.0; 0.000001, not 1e-6), strings in UTF-8 with only ", \ and U+0000–U+001F escaped.

For legacy files the dump is the 5.0 view defined in §20, with the legacy spdf_version and "legacy": true.

13. Integrity and signatures

spdf_meta.content_sha256 (OPTIONAL) is the lowercase hex SHA-256 of the JCS serialization of the canonical dump from which the members meta.content_sha256, meta.signature and meta.signer have been removed. It covers all content except extension tables, and it is independent of SQLite's page layout, so two writers that store the same content produce the same hash.

spdf_meta.signature (OPTIONAL, requires content_sha256 and signer) is the standard base64 encoding, with padding, of an Ed25519 signature [RFC 8032] over the ASCII bytes of the string spdf-content-sha256: followed by the hex content_sha256. spdf_meta.signer is ed25519: followed by the standard base64 encoding of the 32-byte public key.

Validators that meet content_sha256 MUST recompute it (E081 on mismatch) and, if a signature is present, MUST verify it (E082 on failure). A valid signature proves that the holder of the key produced this content; it says nothing about whether the key is trustworthy. Readers SHOULD show who signed (the key, or a name the user associated with it) and MUST NOT present an unknown key as trusted.

14. Security considerations

An SPDF file is a database written by someone else. Opening it is parsing untrusted input with a complex engine. The threats, and the rules of this specification that answer them:

  • Code in the schema. Triggers, views and virtual tables can run SQL or call modules when the database is used. Files MUST NOT contain them (§2.1); readers MUST refuse them, open read-only with query_only, trusted_schema = OFF and the defensive flag, and never load extensions (§2.4). The legacy FTS triggers are tolerated only because they never fire on a read-only connection.
  • Malformed databases. SQLite is robust against corrupt files but recommends extra care for untrusted ones [SQLITE-SECURITY]: disable memory-mapped I/O, enable cell_size_check, set length limits, consider quick_check.
  • Decompression bombs. Gzip input (legacy) MUST be decompressed with a size limit (§2.3).
  • Oversized values. Blobs, texts and JSON values MUST be bounded; JSON parsers SHOULD limit nesting depth (RECOMMENDED 64).
  • Query injection. User text never reaches FTS5 as syntax: every term is a quoted FTS5 string (§8.1). SQL is always parameterized. Implementations MAY cap the number of terms (RECOMMENDED 64) to bound query cost.
  • Paths. Blob keys are opaque strings, not file names. A reader that extracts blobs to disk MUST sanitize them (no absolute paths, no .., no device names).
  • Remote references. source_ref, image, thumbnail, URL and web anchors may point to the network. Readers MUST NOT fetch them automatically: fetching discloses that the file was opened and can reach internal services. Fetch only on a user action, and show the address first.
  • Active content. text is light Markdown; readers MUST NOT render raw HTML from it, and MUST escape text before inserting it into HTML. Images from blobs are untrusted input to image decoders; SVG MUST NOT be rendered with scripts enabled.
  • Forged provenance. Provenance, confidence and metadata are claims made by the writer. Only a signature from a trusted key (§13) attributes them.
  • Model inputs. Text read from a file may contain instructions aimed at language models ("ignore the previous instructions…"). Applications that pass SPDF text to a model MUST treat it as data, not as instructions.

15. Privacy considerations

  • Vectors can leak text. Embeddings can be inverted: published attacks reconstruct most of a short input from its vector [VEC2TEXT]. Distributing the vectors of a text is close to distributing the text. Rights and confidentiality rules that apply to the text apply to its vectors (§16); a writer asked to strip the text of a restricted document MUST strip its vectors too.
  • Provenance can leak the producer. Writers SHOULD NOT record local file paths, user names, machine names, account identifiers, API keys or prompts containing personal data in provenance.detail or generator.
  • People in documents. Interviews and recordings name speakers and may contain personal data. Producers SHOULD let users remove or pseudonymize speaker names, and readers SHOULD NOT index speaker names into shared services without consent.
  • Annotations are personal. User annotations live outside the file, in .spdfa.json sidecars (§17), so that sharing a document never shares its reader's notes.
  • Opening is observable only if a reader fetches remote references; see §14.

16. Rights

documents.rights is NULL or a JSON object:

{"license": "CC-BY-4.0", "access": "open", "holder": "Universidad de La Laguna",
 "note": "Text and images under CC BY 4.0; page scans courtesy of the library."}
  • license: an SPDX license identifier or expression [SPDX] (CC-BY-4.0, CC0-1.0), or the URL of a license or rights statement (for the public domain, https://creativecommons.org/publicdomain/mark/1.0/; for rights statements, http://rightsstatements.org/vocab/InC/1.0/).
  • access: open (anyone may receive the file), restricted (only the audience the holder allows: a class, a library), or private (personal copy).
  • holder: the rights holder, or null. note: free text.

SPDF does not grant rights. A file made from a copyrighted work is a copy of that work, including its vectors (§15). Producers SHOULD fill rights when they know them, SHOULD default access to private when they do not, and readers SHOULD show rights before sharing a file. Public-domain status depends on jurisdiction; note is the place to say which.

17. Annotations and collections

17.1 Annotations: .spdfa.json

User annotations (highlights, notes, tags) are stored outside the document, in a file with the extension .spdfa.json, as a W3C Web Annotation [WEB-ANNOTATION] AnnotationCollection in JSON-LD:

{"@context": "http://www.w3.org/ns/anno.jsonld",
 "type": "AnnotationCollection", "spdf_annotations": "1.0",
 "label": "Notas de lectura",
 "first": {"type": "AnnotationPage", "items": [
   {"id": "urn:uuid:7b0c…", "type": "Annotation", "motivation": "commenting",
    "created": "2026-10-07T09:00:00Z",
    "body": {"type": "TextualBody", "value": "Origen del tópico.", "format": "text/plain", "language": "es"},
    "target": {"source": "spdf:sha256-3f2a…c9",
               "selector": [
                 {"type": "SpdfAnchorSelector", "value": "spdf:sha256-3f2a…c9#p=29&f=21&char=118,301"},
                 {"type": "TextQuoteSelector", "exact": "En un lugar de la Mancha",
                  "prefix": "", "suffix": ", de cuyo nombre"}]}}]}}

target.source is the anchor URI without fragment. The SpdfAnchorSelector carries the full anchor URI; the TextQuoteSelector [WEB-ANNOTATION] lets the annotation survive a re-reading that shifts offsets. Readers that cannot resolve the anchor SHOULD fall back to the quote. The member spdf_annotations gives the version of this profile. Selectors of other types (a FragmentSelector conforming to Media Fragments for time and region) MAY be added for tools that do not know SPDF.

17.2 Collections: .spdfl.json

A library is a manifest, not a container:

{"spdf_library": "1.0", "name": "Tesis: fuentes", "created": "2026-10-07T00:00:00Z",
 "items": [{"sha256": "3f2a…c9", "title": "El ingenioso hidalgo…", "authors": "Cervantes Saavedra",
            "year": 1605, "url": "https://example.org/quijote.spdf", "file_sha256": "…"}]}

sha256 is the document's source_sha256 (the identity used by anchor URIs); url and file_sha256 (SHA-256 of the .spdf file bytes) are OPTIONAL and let a reader fetch and check a copy. Items are ordered as the user ordered them.

JSON Schemas for both sidecars are in json-schema/.

18. Short citation

cite(anchor, anchor_end, metadata, locale) produces an author-date citation in parentheses, the form most styles share, so that every implementation prints the same locator. Full bibliographies and other styles are produced from the CSL-JSON item with a CSL processor (§19).

18.1 Citing an anchor

( names ", " year [ ", " locator ] ")"

The locales es and en are defined; a locale is matched by its primary language subtag (es-ES → es), and any other locale falls back to en.

Names, from CSL author. The name of a person is literal if present; otherwise the non-dropping-particle, a space and family; otherwise given. With one author, that name; with two, A y B (es) or A and B (en), where Spanish writes e instead of y when the second name begins with the sound /i/ (i, í, hi or hí not followed by a vowel: Gómez e Iglesias, Gómez e Hidalgo, but Gómez y Hierro); with three or more, A et al. in both locales. Without authors, the title-short, or else the title up to its first colon, trimmed.

Year: the first year of issued; negative years are written as 375 a. C. (es) or 375 BC (en). Without a year, s. f. (es) or n.d. (en). The spdf.undated range is not printed in a short citation.

Locator:

anchoresen
page, folio readp. 145p. 145
page, roman foliop. xivp. xiv
page, inferred foliop. [21]p. [21]
page without folios. p.n. pag.
page range (end with another folio)pp. 145-146, pp. 20-[21]same
leaf / leaf rangefol. 1r, fol. [2v], fols. 1r-[1v]same
column / column rangecol. 45, cols. 45-46same
time (t0, floor seconds)1:09:20, 0:42same
time range (end anchor's t1)0:12-0:24same
section or web with printedas a pageas a page
section or web§ 3.2 El panóptico, párr. 4§ 3.2 El panóptico, para. 4
slidediap. 3slide 3
sheetDatos, filas 4-9, Datos, fila 4Datos, rows 4-9, Datos, row 4
versev. 1234, vv. 1234-1240same
canonical514asame
image(no locator)(no locator)

Times are written h:mm:ss from one hour on and m:ss below (hours are not wrapped: ground elapsed time 109:24:48). A range is printed only when both ends have a printed folio and the folios differ; brackets mark each inferred end separately. An end without a printed folio never takes part in a range: the citation prints the folio of the other end alone (p. 211, never pp. s. p.-211), and s. p. / n. pag. only when neither end has one. Labels always come from the foliation of the end that is printed: an unnumbered page followed by leaf Ir gives fol. Ir; a range whose two printed ends have different foliations labels each end (p. xiv-fol. 1r); section and web anchors count as pages. The locator is omitted when it would be empty, giving (Hooke, 1665).

18.2 Citing a passage

A citation MUST locate the passage it quotes, not the fragment that happens to contain it. cite_passage(fragment, quote, locale) cites a quotation taken from a fragment:

  1. Split the fragment into its parts: the text of the start unit between the two values of anchor.chars, and, for a crossing fragment, the text of the end unit (§4.4) between the two values of anchor_end.chars (the whole text of a unit when chars is absent).
  2. If the quotation lies in the start part, cite the start unit's anchor, with chars giving the position of the quotation in that unit. Otherwise, if it lies in the end part, cite the end unit's anchor alone, with its chars. Otherwise, if it spans both parts, cite the range from the start unit's anchor to the end unit's anchor, without chars, under the range rule above (an end without a folio does not count).
  3. The result is the short citation of §18 and the anchor URI of the cited anchor or range (§5).

Readers and citation tools MUST NOT cite a passage with the start anchor of its fragment when the passage is not in the start unit: a quotation from the second page of a fragment that begins on an unnumbered plate cites the folio of the second page.

19. Exports

Implementations MUST export CSL-JSON and BibTeX as defined in §19.1 to §19.3, and MAY export the other formats of §19.4. Exports never invent data: fields absent from the file are absent from the export. An export takes one or several documents, in order.

19.1 Keys

Every exported document gets a key, used as the CSL id and as the BibTeX key:

  1. Take the first name of the CSL author list: its family, else its literal, else its given. Fold it: decompose with NFKD, keep only the ASCII letters A–Z and a–z, lowercase. (Cervantes Saavedra → cervantessaavedra.)
  2. If that is empty (no author, or no ASCII letter in the name), fold in the same way the first whitespace-separated word of title-short, or of title when there is no title-short. (Lazarillo de Tormes → lazarillo.)
  3. If that is still empty, use anon.
  4. Append the first year of issued in decimal (negative years keep their sign), or nd when there is none: cervantessaavedra1605, anonnd.
  5. When the same key occurs more than once in one export, every occurrence gets a suffix in export order: a, b, … z, aa, ab…

19.2 CSL-JSON

The CSL-JSON export is a JSON array with one item per document: the metadata item without its spdf member, with id set to the key. It is compared as JSON.

A citation of a passage adds to the item the CSL label and locator of an anchor and optional end anchor, so that a CSL processor can print it in any style:

anchorlabellocator
page, foliation page / leaf / columnpage / folio / columnthe folio as in §18: 145, [21], xiv, 1r, ranges 145-146, 1r-[1v]; no label or locator when printed is null
section or web with printedpageas for pages
section or web with paragraphparagraphthe paragraph number
other section or web with a pathsectionthe last element of the path
timetimestamp1:09:20, ranges 0:12-0:24 (as in §18)
verseverse1234 or 1234-1240
canonicalsectionthe ref
sheetline4 or 4-9
slide, imagenonenone (CSL has no slide locator; the short citation of §18 prints it)

19.3 BibTeX

The BibTeX export is text with one entry per document, in export order, separated by one empty line:

@book{cervantessaavedra1605,
  author = {Cervantes Saavedra, Miguel de},
  title = {{El} ingenioso hidalgo don {Quijote} de la {Mancha}},
  year = {1605},
  publisher = {Juan de la Cuesta},
  address = {Madrid},
  language = {es}
}
  • Entry type from the CSL type: book → book; article-journal, article-magazine, article-newspaper → article; chapter → incollection; paper-conference → inproceedings; thesis → phdthesis; report → techreport; anything else → misc.
  • Fields, in this order, each only when its source is present and not empty: author (CSL author), editor (editor), title, year (first year of issued), journal for article entries or else booktitle (container-title), publisher, address (publisher-place), series (collection-title), volume, number (issue), pages (page), edition, doi (DOI), isbn (ISBN), url (URL), language, note.
  • Values are written {…} in UTF-8. In every value, \ becomes \textbackslash{}, { becomes \{ and } becomes \}; nothing else is escaped.
  • Names: a literal name is written in braces, {National Aeronautics and Space Administration}; otherwise the family name (preceded by the non-dropping-particle and a space, if any) and the given name are written Family, Given, or in braces when only one of them exists. Names are joined with and.
  • Capitals: in title and in journal/booktitle, every whitespace-separated word that contains an uppercase letter (Unicode general category Lu) is wrapped in braces, after escaping, so that styles cannot lowercase it: {El} ingenioso hidalgo don {Quijote}.
  • Comparison: two exports are equal when, after removing leading and trailing whitespace from every line and dropping empty lines, their lines are identical.

19.4 Other formats

  • ALTO [ALTO] (MAY): ALTO 4, one Page per page unit, with PHYSICAL_IMG_NR = physical and PRINTED_IMG_NR = printed only when printed is not null and its source is not inferred (ALTO records printed numbers, and an inferred folio is not printed); one TextBlock per paragraph and one TextLine per line; coordinates only when the producer has them (from an extension), never invented.
  • TEI [TEI] (MAY, minimal): teiHeader from the metadata (titleStmt, publicationStmt with the rights, sourceDesc with the CSL fields), and a body with a <pb/> before each page unit, whose n is the folio as cited in §18 without its label (n="ii", n="[iv]", n="1r"; no n for unnumbered pages) and whose facs is the unit image, if any; <p> for paragraphs, <lg>/<l n> for verse, <u who> for speaker turns, and <note place="foot"> for notes.
  • IIIF Presentation 3 [IIIF] (MAY): a Manifest with one Canvas per unit, in ord order; the label of a page canvas is {"none": [n]} with n as the TEI n, and page canvases of unnumbered pages have no label; the unit image is the painting annotation and the text a supplementing annotation; audio and video are one time-based canvas with duration and a Range per unit or section; sections become structures; anchors with a region become #xywh=percent: targets.
  • Web Annotation (MAY): citations and search results as annotations with the selectors of §17.1.

The conformance suite checks, for paged documents, the page sequence of these exports: the PHYSICAL_IMG_NR/PRINTED_IMG_NR pairs of ALTO, the n of each TEI pb and the label of each IIIF page canvas, in order.

20. Legacy formats

20.1 SPDF 4.0 and 4.1

Scholaris 4.x files MUST be readable by every reader. They are SQLite databases, usually wrapped in gzip, with Spanish identifiers. Detection, after decompressing: a table spdf (clave, valor) whose row spdf_version starts with 4., or user_version 400 or 410 together with a table documentos. application_id is 0. The schema is reproduced verbatim in schema/spdf-4.1.sql and schema/spdf-4.0.sql (4.0 lacks unidades.palabras and fragmentos.texto_busqueda). Legacy files contain the triggers fragmentos_ai, fragmentos_ad and fragmentos_au, tolerated by §2.4.

Readers present legacy files through the 5.0 view:

  • Tables: spdf → spdf_meta (clave → key, valor → value), documentos → documents, unidades → units, secciones → sections, fragmentos → fragments, figuras → figures, espacios → spaces, vectores → vectors, blobs → blobs, procedencia → provenance; no extensions.
  • Columns: tipo → kind, metadatos → metadata, huella → source_sha256, original → source_ref, unidades → unit_count, duracion → duration, creado → created, actualizado → updated, titulo → title, autores → authors, anio → year, idioma → language; orden → ord, ancla → anchor, texto → text, notas → notes, cabecera → header, pie → footer, imagen → image, miniatura → thumbnail, lector → reader, confianza → confidence, impresa → printed, palabras → words; padre → parent, nivel → level, unidad_desde → unit_from, unidad_hasta → unit_to, resumen → summary; unidad → unit, contexto → context, seccion → section, ancla_fin → anchor_end, texto_busqueda → search_text; in figures pie → caption, descripcion → description; proveedor → provider, modelo → model, normalizado → normalized, modalidades → modalities (texto → text, imagen → image); objetivo → target (fragmento → fragment, unidad → unit, figura → figure), espacio → space, valores → data; clave → key, datos → data; fase → stage, detalle → detail, cuando → at.
  • Kinds: pdf, pdf_escaneado → scanned_pdf, fotos → photos, imagen → image, audio, video, documento → document, epub, presentacion → slides, hoja → sheet, web.
  • Anchors: tipo → type (pagina → page, tiempo → time, seccion → section, diapositiva → slide, hoja → sheet, web, imagen → image), fisica → physical, impresa → printed, romana → roman, origen → source (leido → read, deducido → inferred, epub, ninguno → none), confianza → confidence, hablante → speaker, ruta → path, parrafo → paragraph, n, hoja → sheet, filaDesde → row_from, filaHasta → row_to, consultada → accessed, region. Unknown members are kept as they are.
  • Metadata (MetadatosDocumento → CSL-JSON): titulo → title, or "titulo: subtitulo" with title-short = titulo and spdf.subtitle = subtitulo; tituloOriginal → original-title; autores, editores, traductores, entrevistadores ({nombre, apellidos, orcid}) → author, editor, translator, interviewer ({family: apellidos, given: nombre}, empty parts omitted; ORCID to spdf.orcid under "apellidos, nombre"); fecha → issued with its full date parts when its year equals anio or there is no anio, otherwise anio → issued; anioOriginal → original-date; editorial → publisher; lugar → publisher-place; revista, else contenedor → container-title; coleccion → collection-title; volumen → volume; numero → issue; paginas → page; edicion → edition; doi → DOI; isbn → ISBN; url → URL; idioma → language; resumen → abstract; idiomaOriginal → spdf.original_language; sinFecha {desde, hasta, fundamento} → spdf.undated {from, to, basis}; procedencia → spdf.provenance, with field names mapped as above and fuente → source (lectura → reading, usuario → user, colofon → colophon, impresores → printers, others unchanged), confianza → confidence. tipoCSL → type; without it, the type is article-journal when revista is present, otherwise by kind: audio and presentacion → speech, video → motion_picture, web → webpage, hoja → dataset, imagen and fotos → graphic, anything else → book. Empty strings, nulls and empty arrays are omitted.
  • Other rules: spdf_meta keys creado → created and generador → generator, others unchanged; documentos.estado and documentos.bibliotecas are dropped; rights is null; spaces get dtype f32, truncated_from and task_prefixes null and created from creado; provenance.model is null; blob hashes are computed; units are renumbered ord = 1, 2, 3… in order of (orden, id), because 4.x numbers units from 0; fragments.ord keeps orden. References in original, imagen and miniatura: an empty string becomes null (in figures it stays an empty string), a value equal to a key of blobs becomes blob:<key>, any other value is kept as an opaque reference.

The 4.x anchor has no foliation; leaf folios did not exist in 4.x.

20.2 SPDF 3.0 and earlier

Scholaris v1 to v3 wrote gzip-wrapped SQLite databases with the tables metadata (key, value, including schema_version) and chunks, among others. Readers MAY import them; importing is a conversion with losses (cross-modal links, scenes and some embeddings have no place in 5.0) and the importer SHOULD report what it dropped. Version 3.0 is documented historically in the Scholaris repository; this specification does not define it.

21. Conformance

21.1 Product classes

  • A conforming reader opens files safely (§2.4), reads 5.0 and legacy 4.x files, produces the canonical dump (§12), parses and formats anchor URIs (§5), produces short citations (§18), runs the reference lexical search (§8.1) and exports CSL-JSON and BibTeX (§19). A semantic reader also runs the reference vector and hybrid search.
  • A conforming writer produces files that validate without errors or warnings for the profiles they declare and whose canonical dump equals the dump the writer was given (round trip).
  • A conforming validator reports exactly the codes of §22 for the validation cases of the suite.

21.2 Levels

An implementation states its class and the profiles it covers, for example "reader and writer, profiles core and semantic". Its claim is backed by the conformance suite: it passes every case of the kinds its class requires (dump, legacy_dump, anchor_uri, cite, cite_passage, search_lexical, validate, locate, export_csl, export_bibtex for readers; plus search_vector and search_hybrid for semantic readers; plus roundtrip and quantize for writers; export_structure for implementations that export ALTO, TEI or IIIF), with the suite version it was tested against. Partial implementations MAY exist but MUST NOT call themselves conforming.

21.3 The suite

The suite (conformance/ in the repository) is normative for behaviour. Its protocol (case format, runner report, CI convention) is in conformance/README.md. Each suite release has a version and a manifest with the number of cases and their hash.

22. Validation

22.1 Procedure

A validator checks a file in this order; a step marked stop ends validation:

  1. If the file starts with 1F 8B, decompress it (§2.3).
  2. If the result is not a SQLite database: E001, stop.
  3. Determine the version: application_id 1397769286 with user_version 500–599 is 5.x; the legacy detection of §20.1 is 4.x; anything else: E002, stop. For 4.x, report W110 and check only: the tables spdf, documentos, unidades, fragmentos, fragmentos_fts exist (E010 each) and there is no trigger or view other than the three tolerated triggers (E020); stop.
  4. A gzip-wrapped 5.x file: E003 in the warnings. A minor version above 0: W105; for such a file, unknown anchor types (E041) and unknown dtypes (E032) are reported in the warnings instead of the errors, because a later minor version may define them.
  5. Triggers, views and foreign virtual tables: E020 for each.
  6. Required tables (E010 each) and required columns (E011 each).
  7. spdf_meta keys (E012 each).
  8. documents holds exactly one row (E013); metadata and rights are valid JSON (E050); metadata has string type and title (E051).
  9. Required extensions unknown to the validator (E060).
  10. units.ord is 1…N (E090); unit_count equals N (W102).
  11. Anchors of units, fragments (start and end) and figures (E040, E041, E042); fragments that cross matter or a folio boundary (W103, §4.4).
  12. Spaces: dtype (E032). Vectors: known space (E031), length (E030).
  13. FTS index in sync: run INSERT INTO fragments_fts(fragments_fts, rank) VALUES('integrity-check', 1) (and the same on fragments_fts_trigram) on a private copy; an error is E070.
  14. Blobs: stored sha256 equals the computed one (E080).
  15. If there were no errors so far and content_sha256 is present: recompute it (E081); if it matches and signature is present, verify it (E082).
  16. Profile warnings: W100, W101.

The result is a JSON object (schema in json-schema/validation-result.schema.json):

{"valid": false, "version": "5.0", "profile": ["core"],
 "errors": [{"code": "E090", "message": "units.ord is not 1..N", "where": "units"}],
 "warnings": []}

valid is true if and only if errors is empty. version is null when unknown. Messages are free text; conformance compares the sets of codes.

22.2 Codes

codemeaning
E001not a SQLite database (or bad gzip)
E002unknown application_id or version
E003gzip-wrapped 5.0 file (reported as a warning)
E010missing required table
E011missing required column
E012missing required spdf_meta key
E013documents does not hold exactly one row
E020trigger, view or foreign virtual table present
E030vector length ≠ dims × dtype size
E031vector refers to an unknown space
E032unknown dtype
E040invalid anchor (bad JSON, missing or mistyped required member)
E041unknown anchor type
E042chars out of range
E050invalid metadata or rights JSON
E051metadata without string type and title
E060unknown required extension
E070FTS index out of sync
E080blob sha256 mismatch
E081content_sha256 mismatch
E082signature does not verify
E090units.ord not contiguous from 1
W100profile semantic without vectors
W101profile media without time anchors
W102unit_count ≠ number of units
W103fragment crosses between units of different matter, or between a page with a printed folio and one without
W105newer minor version than the validator's
W110legacy 4.x file

Codes are never reused with another meaning. New codes are added by minor versions.

23. Versioning and compatibility

The specification is versioned MAJOR.MINOR; editorial corrections do not change the version. user_version encodes it (§2.1).

  • A minor version (5.1, 5.2…) only adds OPTIONAL things: tables, columns, spdf_meta keys, anchor members or types, metadata members, validation warnings or errors for things that were already forbidden. A 5.0 reader reads every 5.x file, ignoring what it does not know; it MAY warn (W105). Validators report the anchor types and dtypes of a newer minor version as warnings, not errors (§22.1). A 5.x writer that uses nothing new SHOULD write user_version 500.
  • A major version (6.0) may change or remove things. Readers MUST refuse majors they do not know (E002) and SHOULD keep reading older majors (as 5.0 reads 4.x).
  • Deprecation: a feature is deprecated in a minor version, with the reason and the replacement, and removed no earlier than the next major and at least 24 months later.
  • Promise: a file that conforms to 5.0 will be readable by every conforming reader of any later 5.x version, and its anchor URIs will keep resolving.
  • The conformance suite and each library have their own version numbers; the suite's manifest says which specification version it tests.

Changes are proposed and decided through the RFC process in spec/rfcs/ and governance/.

24. Media type and file identification

  • Media type: application/vnd.spdf+sqlite3 (registration with IANA in preparation; template in governance/drafts/iana-media-type.md). The structured syntax suffix +sqlite3 tells generic tools that the file is a SQLite 3 database. OPTIONAL parameter version ("5.0"). Encoding: binary. Legacy 4.x files are gzip data and have no registered type of their own.
  • Fragment identifiers: for a resource of type application/vnd.spdf+sqlite3, the fragment identifier is the params rule of §5.1, with the meaning it has in an anchor URI for the document in that resource: https://example.org/quijote.spdf#p=5&f=1r.
  • Extension: .spdf. Sidecars: .spdfa.json and .spdfl.json, served as application/json (or application/ld+json for annotations).
  • Magic numbers: bytes 0–15 are 53 51 4C 69 74 65 20 66 6F 72 6D 61 74 20 33 00 ("SQLite format 3" and a NUL); bytes 68–71 are 53 50 44 46 ("SPDF"); bytes 60–63 hold user_version big-endian (00 00 01 F4 for 5.0). Legacy 4.x files start with 1F 8B and cannot be told from other gzip files without decompressing.
  • Uniform Type Identifier (Apple platforms): com.joseluissaorin.spdf, conforming to public.data and public.database, until a vendor-neutral identifier is agreed.

25. Internationalization

25.1 Languages and scripts

Language tags are BCP 47 [BCP 47]: es, en-GB, la, grc (Ancient Greek), lzh (Literary Chinese), ar. documents.language is the main language; fragments in other languages need no tagging in this version. Text is stored in logical order, whatever its direction.

25.2 Right-to-left text

Arabic, Hebrew, Syriac and other right-to-left scripts are stored in logical order, without bidirectional control characters except those present in the source. Readers display them with the Unicode Bidirectional Algorithm [UAX #9] and SHOULD isolate user-supplied strings (dir="auto"). Anchor URIs percent-encode such text, so they are direction-neutral; when an IRI form is displayed, readers SHOULD isolate it. chars offsets count code points in logical order.

25.3 Old texts

The text of a unit or fragment is the text of the source, never modernized: long s (ſ), u/v and i/j alternations, abbreviations and tildes stay as printed. The search_text layer carries a modernized form used only by search. The unicode61 tokenizer with remove_diacritics 2 already folds case, Latin diacritics and ſ; it does not fold Greek accents and breathings, ligatures such as æ and œ, or ß, so producers SHOULD put the folded forms they need in search_text (for polytonic Greek, the text without diacritics). Citations quote text, never search_text.

25.4 Chinese, Japanese and Korean

The unicode61 tokenizer treats a run of Han characters as one token. Files whose text is mostly CJK SHOULD include fragments_fts_trigram; the reference search then matches substrings of three or more characters by trigram and shorter ones by substring (§8.1). Producers MAY add a segmented form (words separated by spaces) to search_text.

25.5 Numbers and folios

Printed folios are stored as printed, in any script ("xiv", "٣٤", "三"). Readers MUST NOT convert them for citation; they MAY offer conversions for navigation.

References

Normative

  • [BCP 47] Phillips, A., Davis, M., "Tags for Identifying Languages", BCP 47, RFC 5646.
  • [COMMONMARK] CommonMark Spec, version 0.31.2, https://spec.commonmark.org/0.31.2/.
  • [CSL-JSON] Citation Style Language, CSL-JSON schema, https://github.com/citation-style-language/schema.
  • [MEDIA-FRAGMENTS] W3C, "Media Fragments URI 1.0 (basic)", Recommendation, 2012.
  • [RFC 1952] Deutsch, P., "GZIP file format specification version 4.3".
  • [RFC 2119] Bradner, S., "Key words for use in RFCs to Indicate Requirement Levels".
  • [RFC 3986] Berners-Lee, T., et al., "Uniform Resource Identifier (URI): Generic Syntax".
  • [RFC 3987] Duerst, M., Suignard, M., "Internationalized Resource Identifiers (IRIs)".
  • [RFC 5147] Wilde, E., Duerst, M., "URI Fragment Identifiers for the text/plain Media Type".
  • [RFC 5234] Crocker, D., Overell, P., "Augmented BNF for Syntax Specifications: ABNF".
  • [RFC 8032] Josefsson, S., Liusvaara, I., "Edwards-Curve Digital Signature Algorithm (EdDSA)".
  • [RFC 8174] Leiba, B., "Ambiguity of Uppercase vs Lowercase in RFC 2119 Key Words".
  • [RFC 8259] Bray, T., "The JavaScript Object Notation (JSON) Data Interchange Format".
  • [RFC 8785] Rundgren, A., et al., "JSON Canonicalization Scheme (JCS)".
  • [SQLITE-FORMAT] SQLite, "Database File Format", https://www.sqlite.org/fileformat.html.
  • [SQLITE-FTS5] SQLite, "SQLite FTS5 Extension", https://www.sqlite.org/fts5.html.
  • [UAX #15] Unicode Standard Annex #15, "Unicode Normalization Forms".
  • [WEB-ANNOTATION] W3C, "Web Annotation Data Model", Recommendation, 2017.

Informative

  • [ALTO] Library of Congress, "ALTO: Technical Metadata for Layout and Text Objects", version 4.
  • [CTS] "Canonical Text Services" protocol and URN scheme, http://cite-architecture.github.io/.
  • [IIIF] IIIF Consortium, "IIIF Presentation API 3.0".
  • [MRL] Kusupati, A., et al., "Matryoshka Representation Learning", NeurIPS 2022.
  • [RFC 6838] Freed, N., Klensin, J., Hansen, T., "Media Type Specifications and Registration Procedures".
  • [RFC 7595] Thaler, D., et al., "Guidelines and Registration Procedures for URI Schemes".
  • [RRF] Cormack, G. V., Clarke, C. L. A., Büttcher, S., "Reciprocal Rank Fusion outperforms Condorcet and individual rank learning methods", SIGIR 2009.
  • [SPDX] SPDX License List, https://spdx.org/licenses/.
  • [SQLITE-SECURITY] SQLite, "Defense Against The Dark Arts", https://www.sqlite.org/security.html.
  • [TEI] TEI Consortium, "TEI P5: Guidelines for Electronic Text Encoding and Interchange".
  • [UAX #9] Unicode Standard Annex #9, "Unicode Bidirectional Algorithm".
  • [VEC2TEXT] Morris, J. X., et al., "Text Embeddings Reveal (Almost) As Much As Text", EMNLP 2023.

Appendix A. Changes from SPDF 4.1

  • Uncompressed container with application_id and user_version; gzip only for legacy.
  • English identifiers; legacy files read through the 5.0 view.
  • Metadata as a CSL-JSON item with the spdf extension object.
  • New anchor types verse and canonical; foliation (leaves and columns); chars and region on any anchor.
  • Anchor URI with ABNF, aligned with W3C Media Fragments and RFC 5147.
  • spaces.dtype (f32, f16, i8), truncated_from, task_prefixes; space compatibility.
  • Profiles, extensions, rights, blob hashes, provenance model.
  • No triggers or views in distributed files; safe opening procedure.
  • Canonical dump, integrity hash and Ed25519 signatures.
  • Units numbered from 1.
  • Removed: documentos.estado and documentos.bibliotecas (library membership belongs to collection manifests).