SPDF 5.0 specification
The normative specification of the SPDF 5.0 format: container, schema, anchors and their URI, CSL-JSON metadata, search, vectors, integrity, validation, citation and conformance.
Reviewed Markdown
- Status: Working Draft, 2026-10-07. Stable enough to implement; changes go through
the RFC process (
spec/rfcs/) and are logged inspec/CONTRACT.mduntil 5.0 is final. - Editor: José Luis Saorín Ferrer.
- This version:
spec/SPEC.mdin https://github.com/joseluissaorin/spdf. - Spanish translation:
SPEC.es.md(faithful; in case of conflict the English text prevails). - License: this specification is published under CC BY 4.0. Code in the repository is MIT OR Apache-2.0. Contributors commit not to assert patents against implementations.
Abstract
SPDF is an open, portable file format for documents that have been read once and can
be cited forever. A .spdf file holds the text of one document (a printed book, a scan,
a recording, a slide deck, a spreadsheet, a web page) as a set of citable units and
searchable fragments, and every fragment carries an exact anchor: the printed page or
leaf, the second of a recording, the slide, the verse, the canonical reference. A citation
produced from an SPDF file can only print what the source says. The container is a plain
SQLite 3 database, the metadata is a CSL-JSON item, the full-text index uses tokenizers
that ship with every SQLite, and optional embedding vectors from several models can live
side by side. Any language with SQLite can read SPDF without special libraries.
Status of this document
This is the first public version of the format (earlier versions, 3.0 to 4.1, were
internal to Scholaris and are covered as legacy in §20). The conformance suite
in conformance/ is part of the specification: where this text and a conformance case
disagree, the disagreement is a bug to be resolved through the RFC process; until it is,
implementations follow the conformance case.
Contents
- Preface
- 1. Conventions and terminology
- 2. Container
- 3. Schema
- 4. Anchors
- 5. Anchor URI
- 6. Metadata
- 7. Text, normalization and offsets
- 8. Reference search
- 9. Vector spaces
- 10. Profiles
- 11. Extensions
- 12. Canonical dump
- 13. Integrity and signatures
- 14. Security considerations
- 15. Privacy considerations
- 16. Rights
- 17. Annotations and collections
- 18. Short citation
- 19. Exports
- 20. Legacy formats
- 21. Conformance
- 22. Validation
- 23. Versioning and compatibility
- 24. Media type and file identification
- 25. Internationalization
- References
- Appendix A. Changes from SPDF 4.1
Preface
SPDF was born inside Scholaris, an application written by José Luis Saorín Ferrer to insert verified, page-exact citations into academic writing. Scholaris needed to read a source once (with a PDF text layer, a vision model, or a speech recognizer), keep what it had read, and answer for years afterwards the only question a citation must answer honestly: where exactly does the source say this? The answer had to survive the original file being moved, the reading model being replaced and the search engine being rewritten. The result was a file per document, the Scholaris Processed Document Format, which went through a gzip-compressed JSON-and-SQLite version (3.0) and a Spanish-named SQLite schema (4.0 and 4.1).
Version 5.0 is the first version designed for everyone. It keeps what experience proved right and drops what tied it to one program: identifiers are in English, the container is uncompressed so it can be memory-mapped and read by HTTP ranges, the metadata is plain CSL-JSON so that Zotero, citeproc and Pandoc understand it, and every number in a conformance case comes from an oracle that anyone can rerun. The name became Semantic Processed Document Format; the initials did not change.
Five principles guide every decision in this specification:
- Anchors first. Every fragment knows exactly where it comes from: physical page and printed folio, leaf and side, second, slide, verse, canonical reference. Nothing that cannot be anchored is citable.
- Provenance. A file says who read each unit and with what confidence, which model produced each vector, and how each metadata field was obtained. Derived data can be recomputed from the original plus the units.
- Read once, query many times. Reading a document is expensive (vision models, speech recognition, human correction); querying it must be cheap, offline, and possible from any language with SQLite.
- Portability. One file, one document, no server, no proprietary dependency, no compression layer to undo, no code inside the file. Readers in many languages pass the same conformance suite.
- Honest citation. A citation prints only what an anchor says. A folio that was inferred is printed in brackets; an unnumbered page is cited as unnumbered; the modernized spelling used for search is never quoted.
1. Conventions and terminology
The key words "MUST", "MUST NOT", "REQUIRED", "SHALL", "SHALL NOT", "SHOULD", "SHOULD NOT", "RECOMMENDED", "NOT RECOMMENDED", "MAY", and "OPTIONAL" in this document are to be interpreted as described in BCP 14 [RFC 2119] [RFC 8174] when, and only when, they appear in all capitals, as shown here.
ABNF follows [RFC 5234]. JSON follows [RFC 8259]; "JSON object", "array", "string" and "number" have their RFC 8259 meanings. SQL follows SQLite's dialect.
- Document: the work an SPDF file describes (one per file).
- Original: the bytes the document was read from (PDF, image set, audio, EPUB…).
- Unit: a citable division of the document: a page or leaf, a time span, a slide, a section, a sheet range. Units are ordered and numbered from 1.
- Fragment: a searchable, citable passage of roughly 150 to 300 words, with the anchor of its start and, if it crosses units, of its end.
- Anchor: a JSON object that locates a unit, a fragment or a figure in the document (§4).
- Anchor URI: the textual form of an anchor,
spdf:<docref>#<params>(§5). - Space: a vector space, i.e. the model, dimensions and encoding that produced a set of embedding vectors (§9).
- Reader: software that opens SPDF files and exposes their content. Writer: software that creates SPDF files. Validator: software that checks files against this specification. Producer: a writer that also reads originals (OCR, speech recognition, embeddings).
- Code point: a Unicode scalar value. Lengths and offsets in this specification count code points, never bytes or UTF-16 code units.
- NFC: Unicode Normalization Form C [UAX #15].
- JCS: the JSON Canonicalization Scheme [RFC 8785].
2. Container
2.1 File
An SPDF 5.0 file is a SQLite 3 database file [SQLITE-FORMAT] holding exactly one document. It MUST NOT be wrapped in any compression or archive layer: the database header MUST start at byte 0.
Writers MUST set:
PRAGMA application_id = 1397769286(hexadecimal 0x53504446). SQLite stores it big-endian at byte offset 68 of the header, so bytes 68 to 71 read "SPDF" in ASCII.PRAGMA user_version = 500. The value encodes the specification version as major × 100 + minor × 10 (5.0 → 500, 5.1 → 510).- The
spdf_metarowspdf_versionto"5.0"(§3.2).
Writers SHOULD use a page size of 4096 bytes, the rollback journal in DELETE mode (never
leave a -wal or -journal file next to a distributed file), and run VACUUM after
the last write so the file has no free pages. Writers SHOULD NOT use auto_vacuum.
A file MUST NOT contain triggers or views, and MUST NOT contain virtual tables other
than the FTS5 tables defined in §3. Writers keep the full-text index in sync
themselves (for example with INSERT INTO fragments_fts(fragments_fts) VALUES('rebuild')
before VACUUM).
2.2 Name and type
The file extension is .spdf. The media type is application/vnd.spdf+sqlite3
(§24). One file holds one document; libraries of documents are described
by a separate collection manifest (§17).
2.3 Gzip input
Legacy 4.x files are SQLite databases wrapped in gzip [RFC 1952] (§20).
Readers MUST therefore accept a file that starts with the gzip magic bytes 1F 8B,
decompress it (to memory or to a temporary file) with a configurable limit on the
decompressed size (RECOMMENDED default 4 GiB), and continue with the result. A gzip-wrapped
5.0 file is readable but non-conforming: validators report it as E003 in the warnings
list (§22).
2.4 Safe opening
SPDF files come from strangers. Every reader MUST open them as follows, and MUST refuse the file if a step cannot be honoured by its SQLite binding:
- Open the database read-only (
SQLITE_OPEN_READONLY, or the URI parametermode=ro). Never open a distributed file read-write in place. PRAGMA query_only = 1andPRAGMA trusted_schema = OFF.- Enable
SQLITE_DBCONFIG_DEFENSIVEwhere the binding exposes it, and keep extension loading disabled (sqlite3_enable_load_extension(db, 0); never callload_extension). - Read
sqlite_masterand refuse the file if it contains a trigger, a view, or a virtual table other thanfragments_ftsandfragments_fts_trigramdeclaredUSING fts5. Legacy 4.x files are allowed exactly the three triggersfragmentos_ai,fragmentos_adandfragmentos_au(§20), which never fire on a read-only connection. - Enforce a configurable maximum size for any single BLOB or TEXT value read (RECOMMENDED
default 512 MiB), for example with
sqlite3_limit(db, SQLITE_LIMIT_LENGTH, …).
Readers SHOULD also disable memory-mapped I/O (PRAGMA mmap_size = 0) and enable
PRAGMA cell_size_check = ON for files from untrusted sources, and MAY run
PRAGMA quick_check before use. Operations that need to write, such as the FTS5
integrity-check command, MUST run on a private copy (for example an in-memory copy made
with the backup API), never on the file. §14 explains the threats.
3. Schema
3.1 Overview
The normative schema is the SQL script schema/spdf-5.0.sql,
reproduced in full below. Every table in it is REQUIRED, even when empty; only
fragments_fts_trigram is OPTIONAL. Column names, types and constraints MUST be as
written. Writers MUST NOT add columns to these tables; data that does not fit goes into
extension tables (§11). Readers MUST ignore columns they do not know (a
later minor version may add OPTIONAL columns, §23).
JSON stored in TEXT columns MUST be valid JSON [RFC 8259] encoded in UTF-8; writers MAY
serialize it in any form (the canonical dump re-serializes it, §12). Timestamps
are ISO 8601 / RFC 3339 strings in UTC with a Z suffix. Identifiers (id columns) are
non-empty strings chosen by the writer; they are opaque, case-sensitive and stable for the
life of the file.
PRAGMA application_id = 1397769286; -- 0x53504446, "SPDF"
PRAGMA user_version = 500;
CREATE TABLE spdf_meta (key TEXT PRIMARY KEY, value TEXT NOT NULL);
CREATE TABLE documents (
id TEXT PRIMARY KEY, kind TEXT NOT NULL, metadata TEXT NOT NULL,
source_sha256 TEXT NOT NULL, source_ref TEXT, mime TEXT NOT NULL,
bytes INTEGER NOT NULL, unit_count INTEGER NOT NULL, duration REAL,
created TEXT NOT NULL, updated TEXT NOT NULL,
title TEXT, authors TEXT, year INTEGER, language TEXT, rights TEXT);
CREATE TABLE units (
id TEXT PRIMARY KEY, document TEXT NOT NULL REFERENCES documents(id),
ord INTEGER NOT NULL, anchor TEXT NOT NULL, text TEXT NOT NULL DEFAULT '',
notes TEXT, header TEXT, footer TEXT, image TEXT, thumbnail TEXT,
reader TEXT NOT NULL, confidence REAL NOT NULL DEFAULT 1, printed TEXT,
t0 REAL, t1 REAL, words TEXT);
CREATE INDEX units_doc ON units(document, ord);
CREATE INDEX units_printed ON units(document, printed);
CREATE TABLE sections (
id TEXT PRIMARY KEY, document TEXT NOT NULL, parent TEXT, level INTEGER NOT NULL,
title TEXT NOT NULL, unit_from TEXT NOT NULL, unit_to TEXT, summary TEXT);
CREATE TABLE fragments (
n INTEGER PRIMARY KEY, id TEXT NOT NULL UNIQUE, document TEXT NOT NULL,
unit TEXT NOT NULL, ord INTEGER NOT NULL, text TEXT NOT NULL,
context TEXT NOT NULL DEFAULT '', section TEXT, anchor TEXT NOT NULL,
anchor_end TEXT, search_text TEXT);
CREATE INDEX fragments_doc ON fragments(document, ord);
CREATE INDEX fragments_unit ON fragments(unit);
CREATE VIRTUAL TABLE fragments_fts USING fts5(
text, context, section, search_text,
content='fragments', content_rowid='n',
tokenize='unicode61 remove_diacritics 2');
-- OPTIONAL:
-- CREATE VIRTUAL TABLE fragments_fts_trigram USING fts5(
-- text, content='fragments', content_rowid='n', tokenize='trigram');
CREATE TABLE figures (
id TEXT PRIMARY KEY, document TEXT NOT NULL, unit TEXT NOT NULL,
image TEXT NOT NULL, caption TEXT, description TEXT, anchor TEXT NOT NULL);
CREATE TABLE spaces (
id TEXT PRIMARY KEY, provider TEXT NOT NULL, model TEXT NOT NULL, version TEXT,
dims INTEGER NOT NULL, dtype TEXT NOT NULL DEFAULT 'f32',
normalized INTEGER NOT NULL DEFAULT 1, truncated_from INTEGER,
modalities TEXT NOT NULL, task_prefixes TEXT, created TEXT);
CREATE TABLE vectors (
target TEXT NOT NULL, id TEXT NOT NULL, space TEXT NOT NULL REFERENCES spaces(id),
document TEXT NOT NULL, data BLOB NOT NULL, PRIMARY KEY (target, id, space));
CREATE TABLE blobs (key TEXT PRIMARY KEY, mime TEXT NOT NULL, sha256 TEXT NOT NULL,
data BLOB NOT NULL);
CREATE TABLE provenance (
document TEXT NOT NULL, stage TEXT NOT NULL, provider TEXT, model TEXT,
detail TEXT, ms INTEGER, at TEXT NOT NULL);
CREATE TABLE extensions (name TEXT PRIMARY KEY, version TEXT NOT NULL,
required INTEGER NOT NULL DEFAULT 0);3.2 spdf_meta
Key/value pairs about the file. REQUIRED keys:
| key | value |
|---|---|
spdf_version | "5.0" |
profile | space-separated profile names, a subset of core semantic media full (§10); always includes core |
created | creation time of the file (UTC) |
generator | name/version of the writer, e.g. spdf-producer/0.3.1 |
document_id | equal to documents.id |
OPTIONAL keys: content_sha256, signature, signer (§13) and
license_note (free text for humans). Other keys MAY be added by later versions or by
extensions (prefixed x_<vendor>_); readers MUST ignore keys they do not know.
3.3 documents
Exactly one row.
kind: one ofpdf(PDF with a usable text layer),scanned_pdf(PDF read by vision),photos(a set of page photographs),image(a single image),audio,video,document(DOCX, ODT, RTF, HTML, Markdown, plain text),epub,slides,sheet,web. Readers MUST accept unknown kinds and treat them asdocument.metadata: the CSL-JSON item with thespdfextension (§6).source_sha256: lowercase hexadecimal SHA-256 of the original bytes. It identifies the document across copies and is the preferred document reference in anchor URIs.source_ref: where the original is:blob:<key>when shipped inside the file, an absolute URL, or NULL.mime,bytes: media type and size in bytes of the original.unit_count: number of rows inunits(a mismatch is warning W102).duration: seconds, for audio and video; NULL otherwise.created,updated: when the document record was created and last changed.title,authors,year,language: denormalized copies for filtering without parsing JSON: the CSLtitle; the family names (or literal names) of the CSLauthorlist joined with"; "; the first year ofissued; the CSLlanguage. They MUST agree withmetadatawhen present.rights: JSON rights object (§16) or NULL.
3.4 units
One row per citable unit, ord = 1, 2, 3… without gaps (E090), in reading order.
anchor: the unit's anchor (§4).text: the full text of the unit as read, NFC, light Markdown (§7). Empty string for units without text (a blank page, a photograph).notes: JSON array of strings (footnotes detached from the body) or NULL.header,footer: running heads and feet, kept out oftext, or NULL.image,thumbnail:blob:<key>or URL of the unit's image (page, frame, slide) and of its thumbnail, or NULL.reader: what producedtext(pdf-text-layer,tesseract-5,gemma-4-e4b,whisper-large-v3-turbo,human…).confidence: 0 to 1.printed: the printed folio of a page unit, copied from its anchor, so readers can "go to page 145" with an index.t0,t1: start and end in seconds, copied from a time anchor; NULL otherwise.words: word timings for audio and video (§7.4) or NULL.
3.5 sections
The heading tree. level starts at 1; parent is the id of the enclosing section or
NULL; unit_from and unit_to are the first and last unit ids (unit_to NULL when the
section ends with the document); summary is OPTIONAL text in the document language.
3.6 fragments
n: a positive integer, unique, stable: it is the rowid the FTS5 index uses (an implicit rowid may change onVACUUM).unit: id of the unit where the fragment starts.ord: reading order within the document (increasing with the position of the fragment in the text).text: the literal passage, NFC, exactly as in the source (never modernized).context: one line that situates the fragment in the work ("Chapter III: the struggle for existence"), used by search; empty string if none.section: JSON array of strings, the heading path, or NULL.anchor: anchor of the start of the fragment.anchor_end: anchor of its end when it crosses into another unit; NULL otherwise.search_text: the modernized-spelling layer (§25.3): text used ONLY for search (aſsi→así,V. M.→vuestra merced). Empty string when it adds nothing; NULL when not computed. It MUST NOT be displayed as the text of the source or quoted.
3.7 fragments_fts and fragments_fts_trigram
fragments_fts is an external-content FTS5 index over fragments with the columns
text, context, section and search_text in this order and the tokenizer
unicode61 remove_diacritics 2, which every SQLite with FTS5 provides. It MUST be in sync
with fragments (E070). fragments_fts_trigram is OPTIONAL, indexes text only with the
trigram tokenizer (SQLite 3.34 or later), and SHOULD be present when the document is
mostly in Chinese, Japanese or Korean.
3.8 figures
Figures, plates, tables as images, photographs inside a page. image is blob:<key> of
a cropped image, or the unit's image together with a region in the anchor. caption
is the printed caption, if any; description is a description in the document language
(for accessibility and search). anchor normally carries a region.
3.9 spaces and vectors
See §9. vectors.target is fragment, unit or figure and vectors.id
is the id of that row; data is the little-endian vector.
3.10 blobs
Binary content shipped inside the file: the original, page images, cropped figures,
thumbnails. key is an opaque string (by convention path-like, pages/0001.png), mime
its media type, sha256 the lowercase hex SHA-256 of data (E080). Other tables refer
to a blob as blob:<key>.
3.11 provenance
One row per production step: stage (reading, transcription, folios,
metadata, embedding, figures…), provider, model, detail (JSON object or
NULL), ms (duration in milliseconds) and at (UTC timestamp). See §15 for
what not to record.
3.12 extensions
See §11.
4. Anchors
4.1 General
An anchor is a JSON object with a string member type. The types defined by this version
and their members are:
| type | REQUIRED members | OPTIONAL members |
|---|---|---|
page | physical (integer ≥ 1), printed (string or null) | roman (boolean), foliation (page, leaf, column; default page), source (read, inferred, epub, none), confidence (0–1) |
time | t0, t1 (seconds, 0 ≤ t0 ≤ t1) | speaker (string) |
section | path (array of strings) | paragraph (integer ≥ 1), printed (string) |
slide | n (integer ≥ 1) | |
sheet | sheet (string), row_from, row_to (integers) | |
web | url (string) | path, paragraph, accessed (ISO date) |
image | ||
verse | line_from (integer) | line_to (integer), printed (string) |
canonical | scheme (string), ref (string) |
Every anchor MAY also carry:
region:{"x", "y", "w", "h"}, numbers between 0 and 1, fractions of the width and height of the unit's image, origin at the top left;chars:[start, end], code point offsets into the NFCtextof the anchor's unit,0 ≤ start ≤ end ≤ length, end exclusive (E042);matter: what kind of matter the unit is:body(the text of the work),front(preliminaries: title page, contents, licences, dedication, prologue of an edition),back(index, colophon, appendices of an edition),plate(a plate or fold-out outside the text pages),cover,library(bookplates, stamps, library or digitizer pages, licences of a digital edition) orblank. Absent meansbody; readers MUST treat values they do not know asbody. Writers SHOULD set it on the units of paged documents whenever it is notbody.
In this specification an "integer" is a JSON number with an integral value: 10 and
10.0 are the same JSON value and both are integers. An anchor whose JSON is invalid or
that lacks or mistypes a REQUIRED member is invalid (E040); an unknown type is E041. Readers MUST preserve members they do not know when
they copy anchors.
4.2 Pages, folios and leaves
physical is the 1-based position of the page in the original (the PDF page index, the
photo number). printed is the folio exactly as printed on the page ("23", "xiv", "A-3",
"1r"), or null when the page carries no number.
roman: truemarks folios in roman numerals (front matter).foliationdescribes what the printed numbers count:page(each page numbered),leaf(each leaf numbered, sidesrecto andverso, printed as"1r","1v"), orcolumn(columns numbered, as in some dictionaries and early printed books).sourcesays howprintedwas obtained:read(seen on the page),inferred(deduced from neighbouring pages, e.g. an unnumbered verso),epub(from an EPUB page list),none(no folio;printedis null).- An inferred folio is cited in brackets,
p. [21]; a page without folio is cited as unnumbered (§18). A producer MUST NOT invent folios: if no evidence supports a number,printedis null andsourceisnone.
4.3 Time, sections, verses and canonical references
Time anchors locate recordings in seconds from the start of the original; speaker names
who speaks. Section anchors locate unpaginated text (EPUB, DOCX, HTML) by heading path and
paragraph number, and MAY add the equivalent printed page when the edition provides a page
list. Verse anchors count lines of verse (line_from, line_to), as printed editions
number them. Canonical anchors use a citation system that is independent of any edition:
stephanus (Plato), bekker (Aristotle), bible (book chapter:verse), cts (a CTS URN
[CTS]), or any other documented scheme; schemes are lowercase ASCII.
4.4 Start and end
A fragment's anchor locates its start; anchor_end, when present, locates its end and
has the same type. A citation of the whole fragment then prints a range
(pp. 145-146).
- The end unit of a fragment is the first unit after its start unit (in
ordorder) whose anchor equalsanchor_endoncecharsandregionare removed from both. charsinanchorgives the part of the fragment that lies in the start unit, andcharsinanchor_endthe part that lies in the end unit (usually[0, b]). Writers SHOULD set both on crossing fragments, so that readers know which unit each part of the passage comes from.- Writers SHOULD NOT let a fragment cross from a unit of one
matterto a unit of another (body text into a plate, a cover, a library page or a licence), nor from a page with a printed folio to a page without one: the citation of such a fragment would mix locators of different natures. Validators report such fragments as W103.
5. Anchor URI
5.1 Syntax
An anchor URI names a place in a document independently of any file:
spdf:sha256-3f2a…c9#p=29&f=21&char=118,301The document reference is sha256- followed by the 64 lowercase hex digits of
documents.source_sha256 (RECOMMENDED: it is the same for every copy of the document), or
the percent-encoded document id. The fragment is a list of parameters. The parameters
reuse W3C Media Fragments syntax [MEDIA-FRAGMENTS] for time (t=) and space (xywh=) and
RFC 5147 [RFC 5147] syntax for character ranges (char=), so tools that know those
standards can interpret them.
The canonical form is defined by this ABNF [RFC 5234]:
spdf-uri = "spdf:" docref [ "#" params ]
docref = hash-ref / id-ref
hash-ref = "sha256-" 64lhex
lhex = DIGIT / %x61-66 ; 0-9 a-f
id-ref = 1*vchar ; percent-encoded document id
params = param *( "&" param )
param = p / pe / f / fe / t / s / para / sl / sh / rows / v / ref / char / xywh
p = "p=" posint ; physical page
pe = "pe=" posint ; physical end page
f = "f=" value ; printed folio
fe = "fe=" value ; printed end folio
t = "t=" number [ "," number ] ; seconds, Media Fragments npt
s = "s=" value *( "/" value ) ; section path
para = "para=" uint ; paragraph
sl = "sl=" posint ; slide
sh = "sh=" value ; sheet name
rows = "rows=" uint "-" uint ; sheet rows
v = "v=" uint [ "-" uint ] ; verse lines
ref = "ref=" value ":" value ; canonical scheme ":" reference
char = "char=" uint "," uint ; code points, RFC 5147 style
xywh = "xywh=percent:" number "," number "," number "," number
value = *vchar
vchar = unreserved / pct-encoded
unreserved = ALPHA / DIGIT / "-" / "." / "_" / "~"
pct-encoded = "%" HEXDIG HEXDIG ; uppercase in the canonical form
posint = %x31-39 *DIGIT
uint = "0" / posint
number = uint [ "." 1*DIGIT ]In the canonical form, parameters appear at most once and in the order of the param
rule above (p, pe, f, fe, t, s, para, sl, sh, rows, v, ref,
char, xywh); values are UTF-8 strings in which every byte other than an unreserved
character is percent-encoded with uppercase hex digits; in s the separators between path
elements are literal / and a / inside an element is %2F; in ref the first literal
: separates the scheme from the reference, and colons inside them are %3A. Numbers
use the shortest decimal form of ECMAScript (4160, 4175.5, 0.125), never an exponent.
5.2 From an anchor to a URI
Formatting maps an anchor (and optionally an end anchor) to parameters:
| anchor | parameters |
|---|---|
page | p = physical; f = printed if not null; with an end page: pe = its physical if different, fe = its printed if not null and different from printed |
time | t = t0, then t1 (or the end anchor's t1) |
section, web | s = path if not empty; para = paragraph; f = printed; fe as for pages |
slide | sl = n |
sheet | sh = sheet; rows = row_from-row_to |
verse | v = line_from, or line_from-line_to when line_to is present and different; f = printed |
canonical | ref = scheme:ref |
image | none |
| any | char = chars; xywh = region × 100, as percent: |
t values are rounded to 6 decimal places. xywh values are fractions × 100 rounded to 4
decimal places (0.125 → 12.5, 0.333333 → 33.3333). A URI without parameters
(spdf:<docref>) designates the whole document.
5.3 Parsing
Parsing returns the document reference and a locator object with one member per
parameter present: p, pe, para, sl (integers); f, fe, sh (strings); t
(array of one or two numbers); s (array of strings); rows (two integers); v (one or
two integers); ref (object with scheme and ref); char (two integers); xywh (four
fractions, the percent values divided by 100 and rounded to 6 decimals).
Parsers MUST accept percent-encoding with lowercase hex digits, parameters in any order,
unencoded non-ASCII characters (IRI form [RFC 3987]), the npt: prefix and the clock
forms h:mm:ss[.f] and mm:ss[.f] in t. Parsers MUST ignore parameters whose names
they do not know. Parsers MUST reject: a scheme other than spdf:; an empty document
reference; a repeated parameter; malformed numbers; p, pe or sl equal to 0; a
char or t range whose end precedes its start; xywh without the percent: unit
(pixel coordinates cannot be resolved without the image); percent-encoding that does not
decode to valid UTF-8.
Formatting a parsed locator MUST give back the canonical URI byte for byte. The conformance suite checks format, parse and round trip for every anchor type.
5.4 Resolution
locate(file, reference) resolves an anchor URI, or the URL of an SPDF resource with a
fragment identifier (§24), against a file, and returns:
{"document": true, "units": ["p5", "p6"], "fragments": ["q4"], "char": [101, 278], "xywh": null}- Reference. An
spdf:URI is parsed as in §5.3;documentis true when its document reference issha256-followed by the file'ssource_sha256, or the file's document id. Any other reference (anhttps:URL, a file path) designates the file itself:documentis true and the text after its first#, if any, is parsed as the parameter list of §5.3. Whendocumentis false,unitsandfragmentsare empty (implementations MAY report this as an error instead; conformance runners map such an error todocument: false). - Rule. The first parameter present in the order
p,f,t,sl,v,ref,s,shselects the predicate below. Without any of them (no fragment, or onlycharandxywh) the reference designates the whole document andunitsandfragmentsare empty. - Predicate on an anchor (members absent from the anchor never match):
p: apageanchor withp ≤ physical ≤ pe(pedefaults top);f:printedequal tof(for units, theunits.printedcolumn);t: atimeanchor witht0 ≤ t < t1, wheretis the first value of the parameter; the last unit with a time anchor (inordorder) also matches whentequals itst1;sl: aslideanchor withn = sl;v: averseanchor withline_from ≤ v ≤ line_to(line_todefaults toline_from), wherevis the first value of the parameter;ref: acanonicalanchor with the sameschemeandref;s: asectionorwebanchor whosepathstarts with the elements ofs; whenparais present, thepathmust equalsandparagraphmust equalpara;sh: asheetanchor withsheet = shand, whenrowsis present,row_from ≤ a ≤ row_tofor its first valuea.
- Matches.
unitsare the ids of the units whose anchor matches, inordorder.fragmentsare the ids of the fragments whose startanchoror whoseanchor_endmatches, innorder (a fragment that ends on a page is found from that page). When no unit matches but some fragments do,unitsare the distinct start units of those fragments, inordorder. - Characters.
charrefers to the text of the first unit ofunits. Whenchar=[c, d]is present, a fragment is kept infragmentsonly if its start unit is that unit and itsanchorhaschars=[a, b]that overlap the range, or its end unit (§4.4) is that unit and itsanchor_endhascharsthat overlap it;[a, b]overlaps[c, d]whena < dandc < b(forc < d), or whena ≤ c < b(forc = d). charandxywhare copied from the locator, or null.
Several units may match (two pages printed "1", a verse number repeated in two poems):
locate returns them all and the reader lets the user choose; p always disambiguates
pages, which is why formatted URIs carry it.
The spdf URI scheme is intended for provisional registration [RFC 7595]; the request is
drafted in governance/drafts/uri-scheme-spdf.md.
6. Metadata
6.1 CSL-JSON item
documents.metadata is one CSL-JSON item [CSL-JSON] describing the document as it should
be cited: at least type (a CSL type such as book, article-journal, chapter,
thesis, speech, interview, broadcast, motion_picture, webpage, dataset,
graphic) and title (E051 if either is missing). Common members: author, editor,
translator, interviewer (arrays of names {family, given} or {literal}, with the
CSL particles non-dropping-particle and dropping-particle when needed), issued
({"date-parts": [[year, month, day]]}), original-date, title-short,
original-title, container-title, collection-title, publisher, publisher-place,
volume, issue, page, edition, DOI, ISBN, ISSN, URL, accessed,
language (BCP 47), abstract, note. The id member is OPTIONAL inside the file;
exports set it (§19).
Writers MUST NOT invent metadata. A field that cannot be supported by the original or by a cited external source is omitted.
6.2 The spdf extension object
The member spdf of the item holds what CSL cannot express. All its members are OPTIONAL:
"spdf": {
"provenance": {"title": {"source": "title-page", "confidence": 0.99},
"issued": {"source": "colophon", "confidence": 0.95}},
"undated": {"from": 1600, "to": 1610, "basis": "printer active years"},
"original_language": "fr",
"subtitle": "con anotaciones de Fernando de Herrera",
"orcid": {"Foucault, Michel": "0000-0000-0000-0000"}
}provenance: per CSL field, where the value came from (reading,title-page,colophon,crossref,openalex,wikidata,user,epub,pdf, …) and a confidence between 0 and 1.undated: for works without a printed date, a plausible range (from,to, years, negative for BCE) and the evidence (basis). It MUST NOT be copied intoissued: a citation prints "s. f." / "n.d." (§18).original_language: BCP 47 tag of the original language of a translation.subtitle: the subtitle when the CSLtitleis "Title: Subtitle".orcid: ORCID identifiers by name ("Family, Given").
Other members MAY be added by extensions with the prefix x_<vendor>_.
7. Text, normalization and offsets
7.1 Encoding and normalization
All text is UTF-8 in NFC. Writers MUST normalize to NFC before storing and before computing offsets. Writers MUST NOT store U+0000, unpaired surrogates or noncharacters, and SHOULD NOT store other control characters except U+0009 (tab) and U+000A (line feed). Lines end with U+000A only.
7.2 Offsets
chars offsets (§4.1) and every length in this specification count code
points of the NFC text. Implementations whose strings are UTF-16 (JavaScript, Java, C#,
Swift's NSString) MUST convert: a character outside the Basic Multilingual Plane counts
as one code point but two UTF-16 units.
7.3 Light Markdown
units.text MAY use this subset of CommonMark [COMMONMARK]: paragraphs separated by a
blank line; # to ###### headings; *emphasis* and **strong**; - and 1. lists;
> quotations; tables in GitHub style; footnote markers [^1] whose text goes to
notes. Readers MUST NOT interpret raw HTML in text; they display it as text.
Offsets count the stored characters, markup included. Fragments SHOULD keep the markup of
their source unit so that fragments.text is a substring of the unit's text whenever
the fragment does not cross units.
Speaker turns in transcripts start with the label **Name:** followed by one space
(**Neil Armstrong:** Houston, Tranquility Base here.).
7.4 Word timings
units.words is the JSON object {"v": 1, "t0": <seconds>, "cs": [start, duration, start, duration, …]}: one pair of integers per word, in centiseconds from t0 (which
equals the unit's t0). The words are the maximal runs of non-whitespace characters of the
unit's text after removing the speaker labels (**Name:**); cs therefore holds
exactly twice as many integers as there are words. Readers use it to highlight the word
being spoken and to turn a char range into a time range.
8. Reference search
The reference search defines what conformance tests: results that every implementation returns identically from the same file. Products MAY rank better (stopwords, query expansion, reranking, filters); they MUST still offer the reference behaviour to pass the suite, and SHOULD label the difference in their documentation.
A result item is {"fragment_id", "score", "via", "anchor", "anchor_uri"} where via
lists the contributing methods ("lexical", "vector") in that order, anchor is the
fragment's anchor and anchor_uri the URI formatted from anchor and anchor_end with
the sha256- document reference.
8.1 Lexical search
Given a query string and a limit:
- Normalize:
q= NFC(query). - Phrases: scan
qfrom left to right. An opening mark"(U+0022),“(U+201C),«(U+00AB) or„(U+201E) opens a phrase that the next",”(U+201D),»(U+00BB) or, for„,“or”respectively closes. The text between the marks is the phrase. An opening mark without a closing mark is treated as a separator. - Words: maximal runs of characters whose Unicode general category is a letter (L), a mark (M) or a number (N). A phrase term is the phrase's words joined with one space; phrases without words are dropped.
- Terms: if there is at least one phrase term, the terms are the phrase terms (loose
words outside quotes are discarded) and the operator is
AND. Otherwise the terms are the words ofqand the operator isOR. Duplicate terms are removed, keeping the first, comparing them by the keylower(remove_Mn(NFD(term)))(Unicode default lowercase, after removing nonspacing marks); the key is used only to detect duplicates. No stopwords are removed. Without terms the result is empty. - MATCH string: every term, as written (no case folding, no decomposition), is
an FTS5 string:
"+ the term with each"doubled +"; the strings are joined withANDorOR. The tokenizer folds case and diacritics itself; folding the query beforehand would break matches (Straße,fin). - Query:The score is −r. Because
SELECT f.n, f.id, bm25(fragments_fts, 1.0, 0.5, 0.5, 1.0) AS r FROM fragments_fts JOIN fragments f ON f.n = fragments_fts.rowid WHERE fragments_fts MATCH ?1 ORDER BY r, f.n LIMIT ?2search_textis the fourth indexed column, a query in modern spelling finds old spelling without any special step. Sincenis the rowid of the index, implementations MAY rank inside the index alone (SELECT rowid, bm25(…) FROM fragments_fts WHERE fragments_fts MATCH ?1 ORDER BY 2, 1 LIMIT ?2) and look up the fragment ids of the returned rows only; the result is identical. Readers that fetch files by HTTP ranges SHOULD do so, as it avoids reading every matching fragment. - CJK route: if
qcontains a code point in one of the ranges U+2E80–U+2FDF, U+3040–U+30FF, U+3100–U+312F, U+3130–U+318F, U+31A0–U+31FF, U+3400–U+4DBF, U+4E00–U+9FFF, U+A960–U+A97F, U+AC00–U+D7AF, U+F900–U+FAFF, U+FF66–U+FF9F or U+20000–U+3FFFF, step 6 is replaced:- if
fragments_fts_trigramexists and every term has at least 3 code points, the same MATCH string runs againstfragments_fts_trigram, ordered bybm25(fragments_fts_trigram)thenn; the score is −bm25; - otherwise (no trigram index, or a term shorter than 3 code points, which a trigram
index cannot match) the substring fallback runs on
fragments.text: for each fragment,hits= the number of termstwithinstr(text, t) > 0; fragments withhits≥ 1 (with theORoperator) orhits= number of terms (withAND) are returned ordered byhitsdescending, thenn; the score ishits.
- if
8.2 Vector search
Given a space, a target (fragment by default, or unit, figure), a query vector of
dims numbers and a limit: compare the query with every vector of that space and target
by brute force. Every component is converted to an IEEE 754 binary64 number (f32 and f16
exactly; i8 as q/127). The query is used as given, not normalized. When the space has
normalized = 1 the score is the dot product; otherwise it is the cosine similarity.
Results are ordered by score descending, then by fragment n, unit ord or figure id.
A result item for the target unit or figure carries unit_id or figure_id instead
of fragment_id, and the anchor URI of the unit's or figure's own anchor. Products MAY
use approximate indexes; the reference is exhaustive.
8.3 Hybrid search
Run the lexical search and the vector search (target fragment), each with depth
max(limit, 50), and fuse them by reciprocal rank fusion [RRF] with k = 10:
score = Σ 1/(10 + rank) over the lists that contain the fragment, rank starting at 1.
Order by score descending, then n; keep limit results. The constant 10 was measured
in Scholaris: the classic 60 flattens short, good lists.
8.4 Comparison in conformance
Result order MUST match exactly; scores MUST match within an absolute tolerance of 1e-6.
The lexical scores are those of SQLite's own bm25(), which is the oracle.
9. Vector spaces
9.1 Spaces
A row of spaces describes how a set of vectors was produced:
id:<model>@<dims>forf32vectors and<model>@<dims>:<dtype>otherwise (embeddinggemma-2@768,embeddinggemma-2@256:i8). A suffix+<variant>MAY follow to separate vectors of the same model computed from different inputs (legacy Scholaris uses+contexto).provider(who ran the model:local,google,inferbox…),model,version.dims: number of components.dtype:f32(IEEE 754 binary32),f16(binary16) ori8(signed byte; the value is q/127). Other values are invalid (E032).normalized: 1 if every stored vector has unit Euclidean norm (before quantization).truncated_from: for Matryoshka truncation [MRL], the original dimension (768for a vector cut to 256); NULL otherwise. Truncated vectors SHOULD be renormalized before storing, withnormalized = 1.modalities: JSON array of the input modalities the model accepts (text,image,audio,video,pdf).task_prefixes: JSON object with the prefixes or instructions used at encoding time,{"document": "…", "query": "…"}, so that a reader can encode queries the same way; NULL if none.
A file MAY hold several spaces; a file without spaces is valid (profile core).
9.2 Vectors
vectors.data is the vector as dims little-endian values of the space's dtype, so its
length is dims × 4, 2 or 1 bytes (E030). Every vector refers to a space in spaces
(E031). Writers quantize as follows: f32 → f16 with IEEE round-to-nearest-even;
f32 → i8 with q = clamp(round_half_away_from_zero(v × 127), −127, 127). The value −128
is not used. A value that does not fit the dtype (a finite f32 above 65504 that would round
to infinity in f16, a non-finite value) is an error for the writer, never silently
stored.
9.3 Compatibility between spaces and quantizations
Two spaces are compatible, and one query vector serves both, when provider,
model, version, dims, normalized, truncated_from and task_prefixes are equal;
dtype may differ. A reader holding a model MAY therefore search an f32 space and its
i8 copy with the same query. Spaces that differ in any other field are not comparable:
readers MUST NOT mix scores across incompatible spaces, and MUST NOT compare vectors of
different dimensions. A Matryoshka space (truncated_from = 768, dims = 256) is
compatible with a query only if the query was truncated to the same dimensions and
renormalized.
10. Profiles
spdf_meta.profile declares which promises a file makes. Profiles are cumulative labels;
a file lists every profile it satisfies.
| profile | requirements |
|---|---|
core | REQUIRED in every file. All tables of §3; at least one unit; every unit, fragment and figure anchored; text in NFC; FTS index in sync. |
semantic | At least one space, and vectors for every fragment in at least one space. A semantic file without vectors raises W100. |
media | kind is audio or video; units carry time anchors and t0/t1; duration is set; words SHOULD be present. A media file without time anchors raises W101. |
full | semantic and, for audio and video, media; for paged kinds, page images (units.image) and figures where the original has them. |
Readers MUST NOT refuse a file because of its profile; profiles tell readers what to expect and validators what to check.
11. Extensions
Data that this specification does not define goes into extension tables named
x_<vendor>_<name> (lowercase ASCII letters, digits and _; <vendor> is a name the
author controls, e.g. x_scholaris_claims). Each extension in use is declared in the
extensions table with its name (<vendor>_<name> or the table prefix), a version,
and required:
required = 0: readers that do not know the extension ignore it.required = 1: the file cannot be understood without it; a reader that does not know it MUST refuse the file with E060.
Extensions MUST NOT change the meaning of the core tables, MUST NOT add columns to them, and SHOULD NOT be required. Extension tables are not part of the canonical dump. An extension that proves useful to several implementations becomes part of the core through the RFC process (§23).
12. Canonical dump
The canonical dump is a JSON view of a file that every implementation produces identically. It is the oracle of the conformance suite and the input of integrity hashing.
{"spdf_version": "5.0", // legacy files: "4.0"/"4.1" and "legacy": true
"meta": {"<key>": "<value>", …}, // every spdf_meta row
"fts": {"tokenizer": "unicode61 remove_diacritics 2", "trigram": false},
"document": {"id", "kind", "metadata", "source_sha256", "source_ref", "mime", "bytes",
"unit_count", "duration", "created", "updated", "title", "authors",
"year", "language", "rights"},
"units": [{"id", "ord", "anchor", "text", "notes", "header", "footer", "image",
"thumbnail", "reader", "confidence", "printed", "t0", "t1", "words"}],
"sections": [{"id", "parent", "level", "title", "unit_from", "unit_to", "summary"}],
"fragments": [{"n", "id", "unit", "ord", "text", "context", "section", "anchor",
"anchor_end", "search_text"}],
"figures": [{"id", "unit", "image", "caption", "description", "anchor"}],
"spaces": [{"id", "provider", "model", "version", "dims", "dtype", "normalized",
"truncated_from", "modalities", "task_prefixes", "created"}],
"vectors": {"<space id>": {"count": 6, "sha256": "<hex>"}},
"blobs": [{"key", "mime", "bytes", "sha256"}],
"provenance": [{"stage", "provider", "model", "detail", "ms", "at"}],
"extensions": [{"name", "version", "required"}]}Rules:
- Every member listed is present. SQL NULL becomes
null; INTEGER a JSON integer; REAL a JSON number; TEXT a string. Columns that hold JSON (metadata,rights,anchor,anchor_end,notes,words,section,modalities,task_prefixes,detail) are parsed and embedded as JSON values. Thedocumentcolumn of child tables is omitted. Booleans stored as integers (normalized,required) stay integers. - Every number that is not an integer, including those inside parsed JSON, is rounded to
6 decimal places (round half to even on its exact binary value); a result of −0 becomes
- Order:
unitsbyord;fragmentsbyn;sections,figuresandspacesbyid;blobsbykey;extensionsbyname(code point order, which is SQLite'sBINARYcollation over UTF-8);provenanceby the UTF-8 bytes of the JCS serialization of each entry.fts.trigramis true if and only iffragments_fts_trigramexists;fts.tokenizeris the value of thetokenizeoption offragments_ftsas declared, without its quotes and with runs of whitespace collapsed to one space (unicode61 remove_diacritics 2;unicode61, the FTS5 default, if absent). vectorshas one member per distinctvectors.space:countis the number of rows andsha256the hex SHA-256 of theirdatablobs concatenated in order oftarget, thenid.blobs[].bytesandblobs[].sha256are computed fromdata, not copied from thesha256column.- Serialization, whenever bytes matter (hashing), is JCS [RFC 8785]: no whitespace, object
members sorted by the UTF-16 code units of their names, numbers in the ECMAScript form
(
1, not1.0;0.000001, not1e-6), strings in UTF-8 with only",\and U+0000–U+001F escaped.
For legacy files the dump is the 5.0 view defined in §20, with the legacy
spdf_version and "legacy": true.
13. Integrity and signatures
spdf_meta.content_sha256 (OPTIONAL) is the lowercase hex SHA-256 of the JCS
serialization of the canonical dump from which the members meta.content_sha256,
meta.signature and meta.signer have been removed. It covers all content except
extension tables, and it is independent of SQLite's page layout, so two writers that store
the same content produce the same hash.
spdf_meta.signature (OPTIONAL, requires content_sha256 and signer) is the standard
base64 encoding, with padding, of an Ed25519 signature [RFC 8032] over the ASCII bytes of
the string spdf-content-sha256: followed by the hex content_sha256. spdf_meta.signer
is ed25519: followed by the standard base64 encoding of the 32-byte public key.
Validators that meet content_sha256 MUST recompute it (E081 on mismatch) and, if a
signature is present, MUST verify it (E082 on failure). A valid signature proves that the
holder of the key produced this content; it says nothing about whether the key is
trustworthy. Readers SHOULD show who signed (the key, or a name the user associated with
it) and MUST NOT present an unknown key as trusted.
14. Security considerations
An SPDF file is a database written by someone else. Opening it is parsing untrusted input with a complex engine. The threats, and the rules of this specification that answer them:
- Code in the schema. Triggers, views and virtual tables can run SQL or call modules
when the database is used. Files MUST NOT contain them (§2.1); readers MUST
refuse them, open read-only with
query_only,trusted_schema = OFFand the defensive flag, and never load extensions (§2.4). The legacy FTS triggers are tolerated only because they never fire on a read-only connection. - Malformed databases. SQLite is robust against corrupt files but recommends extra
care for untrusted ones [SQLITE-SECURITY]: disable memory-mapped I/O, enable
cell_size_check, set length limits, considerquick_check. - Decompression bombs. Gzip input (legacy) MUST be decompressed with a size limit (§2.3).
- Oversized values. Blobs, texts and JSON values MUST be bounded; JSON parsers SHOULD limit nesting depth (RECOMMENDED 64).
- Query injection. User text never reaches FTS5 as syntax: every term is a quoted FTS5 string (§8.1). SQL is always parameterized. Implementations MAY cap the number of terms (RECOMMENDED 64) to bound query cost.
- Paths. Blob keys are opaque strings, not file names. A reader that extracts blobs to
disk MUST sanitize them (no absolute paths, no
.., no device names). - Remote references.
source_ref,image,thumbnail,URLandwebanchors may point to the network. Readers MUST NOT fetch them automatically: fetching discloses that the file was opened and can reach internal services. Fetch only on a user action, and show the address first. - Active content.
textis light Markdown; readers MUST NOT render raw HTML from it, and MUST escape text before inserting it into HTML. Images from blobs are untrusted input to image decoders; SVG MUST NOT be rendered with scripts enabled. - Forged provenance. Provenance, confidence and metadata are claims made by the writer. Only a signature from a trusted key (§13) attributes them.
- Model inputs. Text read from a file may contain instructions aimed at language models ("ignore the previous instructions…"). Applications that pass SPDF text to a model MUST treat it as data, not as instructions.
15. Privacy considerations
- Vectors can leak text. Embeddings can be inverted: published attacks reconstruct most of a short input from its vector [VEC2TEXT]. Distributing the vectors of a text is close to distributing the text. Rights and confidentiality rules that apply to the text apply to its vectors (§16); a writer asked to strip the text of a restricted document MUST strip its vectors too.
- Provenance can leak the producer. Writers SHOULD NOT record local file paths, user
names, machine names, account identifiers, API keys or prompts containing personal data
in
provenance.detailorgenerator. - People in documents. Interviews and recordings name speakers and may contain
personal data. Producers SHOULD let users remove or pseudonymize
speakernames, and readers SHOULD NOT index speaker names into shared services without consent. - Annotations are personal. User annotations live outside the file, in
.spdfa.jsonsidecars (§17), so that sharing a document never shares its reader's notes. - Opening is observable only if a reader fetches remote references; see §14.
16. Rights
documents.rights is NULL or a JSON object:
{"license": "CC-BY-4.0", "access": "open", "holder": "Universidad de La Laguna",
"note": "Text and images under CC BY 4.0; page scans courtesy of the library."}license: an SPDX license identifier or expression [SPDX] (CC-BY-4.0,CC0-1.0), or the URL of a license or rights statement (for the public domain,https://creativecommons.org/publicdomain/mark/1.0/; for rights statements,http://rightsstatements.org/vocab/InC/1.0/).access:open(anyone may receive the file),restricted(only the audience the holder allows: a class, a library), orprivate(personal copy).holder: the rights holder, or null.note: free text.
SPDF does not grant rights. A file made from a copyrighted work is a copy of that work,
including its vectors (§15). Producers SHOULD fill rights when they know
them, SHOULD default access to private when they do not, and readers SHOULD show
rights before sharing a file. Public-domain status depends on jurisdiction; note is
the place to say which.
17. Annotations and collections
17.1 Annotations: .spdfa.json
User annotations (highlights, notes, tags) are stored outside the document, in a file
with the extension .spdfa.json, as a W3C Web Annotation [WEB-ANNOTATION]
AnnotationCollection in JSON-LD:
{"@context": "http://www.w3.org/ns/anno.jsonld",
"type": "AnnotationCollection", "spdf_annotations": "1.0",
"label": "Notas de lectura",
"first": {"type": "AnnotationPage", "items": [
{"id": "urn:uuid:7b0c…", "type": "Annotation", "motivation": "commenting",
"created": "2026-10-07T09:00:00Z",
"body": {"type": "TextualBody", "value": "Origen del tópico.", "format": "text/plain", "language": "es"},
"target": {"source": "spdf:sha256-3f2a…c9",
"selector": [
{"type": "SpdfAnchorSelector", "value": "spdf:sha256-3f2a…c9#p=29&f=21&char=118,301"},
{"type": "TextQuoteSelector", "exact": "En un lugar de la Mancha",
"prefix": "", "suffix": ", de cuyo nombre"}]}}]}}target.source is the anchor URI without fragment. The SpdfAnchorSelector carries the
full anchor URI; the TextQuoteSelector [WEB-ANNOTATION] lets the annotation survive a
re-reading that shifts offsets. Readers that cannot resolve the anchor SHOULD fall back to
the quote. The member spdf_annotations gives the version of this profile. Selectors of
other types (a FragmentSelector conforming to Media Fragments for time and region) MAY
be added for tools that do not know SPDF.
17.2 Collections: .spdfl.json
A library is a manifest, not a container:
{"spdf_library": "1.0", "name": "Tesis: fuentes", "created": "2026-10-07T00:00:00Z",
"items": [{"sha256": "3f2a…c9", "title": "El ingenioso hidalgo…", "authors": "Cervantes Saavedra",
"year": 1605, "url": "https://example.org/quijote.spdf", "file_sha256": "…"}]}sha256 is the document's source_sha256 (the identity used by anchor URIs); url and
file_sha256 (SHA-256 of the .spdf file bytes) are OPTIONAL and let a reader fetch and
check a copy. Items are ordered as the user ordered them.
JSON Schemas for both sidecars are in json-schema/.
18. Short citation
cite(anchor, anchor_end, metadata, locale) produces an author-date citation in
parentheses, the form most styles share, so that every implementation prints the same
locator. Full bibliographies and other styles are produced from the CSL-JSON item with a
CSL processor (§19).
18.1 Citing an anchor
( names ", " year [ ", " locator ] ")"The locales es and en are defined; a locale is matched by its primary language subtag
(es-ES → es), and any other locale falls back to en.
Names, from CSL author. The name of a person is literal if present; otherwise the
non-dropping-particle, a space and family; otherwise given. With one author, that
name; with two, A y B (es) or A and B (en), where Spanish writes e instead of y
when the second name begins with the sound /i/ (i, í, hi or hí not followed by a
vowel: Gómez e Iglesias, Gómez e Hidalgo, but Gómez y Hierro); with three or more,
A et al. in both locales. Without authors, the title-short, or else the title up to
its first colon, trimmed.
Year: the first year of issued; negative years are written as 375 a. C. (es) or
375 BC (en). Without a year, s. f. (es) or n.d. (en). The spdf.undated range is
not printed in a short citation.
Locator:
| anchor | es | en |
|---|---|---|
| page, folio read | p. 145 | p. 145 |
| page, roman folio | p. xiv | p. xiv |
| page, inferred folio | p. [21] | p. [21] |
| page without folio | s. p. | n. pag. |
| page range (end with another folio) | pp. 145-146, pp. 20-[21] | same |
| leaf / leaf range | fol. 1r, fol. [2v], fols. 1r-[1v] | same |
| column / column range | col. 45, cols. 45-46 | same |
| time (t0, floor seconds) | 1:09:20, 0:42 | same |
| time range (end anchor's t1) | 0:12-0:24 | same |
section or web with printed | as a page | as a page |
| section or web | § 3.2 El panóptico, párr. 4 | § 3.2 El panóptico, para. 4 |
| slide | diap. 3 | slide 3 |
| sheet | Datos, filas 4-9, Datos, fila 4 | Datos, rows 4-9, Datos, row 4 |
| verse | v. 1234, vv. 1234-1240 | same |
| canonical | 514a | same |
| image | (no locator) | (no locator) |
Times are written h:mm:ss from one hour on and m:ss below (hours are not wrapped:
ground elapsed time 109:24:48). A range is printed only when both ends have a printed
folio and the folios differ; brackets mark each inferred end separately. An end without
a printed folio never takes part in a range: the citation prints the folio of the other
end alone (p. 211, never pp. s. p.-211), and s. p. / n. pag. only when neither end
has one. Labels always come from the foliation of the end that is printed: an unnumbered
page followed by leaf Ir gives fol. Ir; a range whose two printed ends have different
foliations labels each end (p. xiv-fol. 1r); section and web anchors count as pages. The locator is omitted when it would be empty, giving (Hooke, 1665).
18.2 Citing a passage
A citation MUST locate the passage it quotes, not the fragment that happens to contain
it. cite_passage(fragment, quote, locale) cites a quotation taken from a fragment:
- Split the fragment into its parts: the text of the start unit between the two values
of
anchor.chars, and, for a crossing fragment, the text of the end unit (§4.4) between the two values ofanchor_end.chars(the whole text of a unit whencharsis absent). - If the quotation lies in the start part, cite the start unit's anchor, with
charsgiving the position of the quotation in that unit. Otherwise, if it lies in the end part, cite the end unit's anchor alone, with itschars. Otherwise, if it spans both parts, cite the range from the start unit's anchor to the end unit's anchor, withoutchars, under the range rule above (an end without a folio does not count). - The result is the short citation of §18 and the anchor URI of the cited anchor or range (§5).
Readers and citation tools MUST NOT cite a passage with the start anchor of its
fragment when the passage is not in the start unit: a quotation from the second page of
a fragment that begins on an unnumbered plate cites the folio of the second page.
19. Exports
Implementations MUST export CSL-JSON and BibTeX as defined in §19.1 to §19.3, and MAY export the other formats of §19.4. Exports never invent data: fields absent from the file are absent from the export. An export takes one or several documents, in order.
19.1 Keys
Every exported document gets a key, used as the CSL id and as the BibTeX key:
- Take the first name of the CSL
authorlist: itsfamily, else itsliteral, else itsgiven. Fold it: decompose with NFKD, keep only the ASCII lettersA–Zanda–z, lowercase. (Cervantes Saavedra→cervantessaavedra.) - If that is empty (no author, or no ASCII letter in the name), fold in the same way the
first whitespace-separated word of
title-short, or oftitlewhen there is notitle-short. (Lazarillo de Tormes→lazarillo.) - If that is still empty, use
anon. - Append the first year of
issuedin decimal (negative years keep their sign), orndwhen there is none:cervantessaavedra1605,anonnd. - When the same key occurs more than once in one export, every occurrence gets a
suffix in export order:
a,b, …z,aa,ab…
19.2 CSL-JSON
The CSL-JSON export is a JSON array with one item per document: the metadata item
without its spdf member, with id set to the key. It is compared as JSON.
A citation of a passage adds to the item the CSL label and locator of an anchor and
optional end anchor, so that a CSL processor can print it in any style:
| anchor | label | locator |
|---|---|---|
page, foliation page / leaf / column | page / folio / column | the folio as in §18: 145, [21], xiv, 1r, ranges 145-146, 1r-[1v]; no label or locator when printed is null |
section or web with printed | page | as for pages |
section or web with paragraph | paragraph | the paragraph number |
other section or web with a path | section | the last element of the path |
time | timestamp | 1:09:20, ranges 0:12-0:24 (as in §18) |
verse | verse | 1234 or 1234-1240 |
canonical | section | the ref |
sheet | line | 4 or 4-9 |
slide, image | none | none (CSL has no slide locator; the short citation of §18 prints it) |
19.3 BibTeX
The BibTeX export is text with one entry per document, in export order, separated by one empty line:
@book{cervantessaavedra1605,
author = {Cervantes Saavedra, Miguel de},
title = {{El} ingenioso hidalgo don {Quijote} de la {Mancha}},
year = {1605},
publisher = {Juan de la Cuesta},
address = {Madrid},
language = {es}
}- Entry type from the CSL
type:book→book;article-journal,article-magazine,article-newspaper→article;chapter→incollection;paper-conference→inproceedings;thesis→phdthesis;report→techreport; anything else →misc. - Fields, in this order, each only when its source is present and not empty:
author(CSLauthor),editor(editor),title,year(first year ofissued),journalforarticleentries or elsebooktitle(container-title),publisher,address(publisher-place),series(collection-title),volume,number(issue),pages(page),edition,doi(DOI),isbn(ISBN),url(URL),language,note. - Values are written
{…}in UTF-8. In every value,\becomes\textbackslash{},{becomes\{and}becomes\}; nothing else is escaped. - Names: a
literalname is written in braces,{National Aeronautics and Space Administration}; otherwise the family name (preceded by thenon-dropping-particleand a space, if any) and thegivenname are writtenFamily, Given, or in braces when only one of them exists. Names are joined withand. - Capitals: in
titleand injournal/booktitle, every whitespace-separated word that contains an uppercase letter (Unicode general category Lu) is wrapped in braces, after escaping, so that styles cannot lowercase it:{El} ingenioso hidalgo don {Quijote}. - Comparison: two exports are equal when, after removing leading and trailing whitespace from every line and dropping empty lines, their lines are identical.
19.4 Other formats
- ALTO [ALTO] (MAY): ALTO 4, one
Pageper page unit, withPHYSICAL_IMG_NR=physicalandPRINTED_IMG_NR=printedonly whenprintedis not null and itssourceis notinferred(ALTO records printed numbers, and an inferred folio is not printed); oneTextBlockper paragraph and oneTextLineper line; coordinates only when the producer has them (from an extension), never invented. - TEI [TEI] (MAY, minimal):
teiHeaderfrom the metadata (titleStmt,publicationStmtwith the rights,sourceDescwith the CSL fields), and abodywith a<pb/>before each page unit, whosenis the folio as cited in §18 without its label (n="ii",n="[iv]",n="1r"; nonfor unnumbered pages) and whosefacsis the unit image, if any;<p>for paragraphs,<lg>/<l n>for verse,<u who>for speaker turns, and<note place="foot">for notes. - IIIF Presentation 3 [IIIF] (MAY): a
Manifestwith oneCanvasper unit, inordorder; thelabelof a page canvas is{"none": [n]}withnas the TEIn, and page canvases of unnumbered pages have nolabel; the unit image is the painting annotation and the text asupplementingannotation; audio and video are one time-based canvas withdurationand aRangeper unit or section; sections becomestructures; anchors with a region become#xywh=percent:targets. - Web Annotation (MAY): citations and search results as annotations with the selectors of §17.1.
The conformance suite checks, for paged documents, the page sequence of these exports:
the PHYSICAL_IMG_NR/PRINTED_IMG_NR pairs of ALTO, the n of each TEI pb and the
label of each IIIF page canvas, in order.
20. Legacy formats
20.1 SPDF 4.0 and 4.1
Scholaris 4.x files MUST be readable by every reader. They are SQLite databases, usually
wrapped in gzip, with Spanish identifiers. Detection, after decompressing: a table spdf
(clave, valor) whose row spdf_version starts with 4., or user_version 400 or
410 together with a table documentos. application_id is 0. The schema is reproduced
verbatim in schema/spdf-4.1.sql and
schema/spdf-4.0.sql (4.0 lacks unidades.palabras and
fragmentos.texto_busqueda). Legacy files contain the triggers fragmentos_ai,
fragmentos_ad and fragmentos_au, tolerated by §2.4.
Readers present legacy files through the 5.0 view:
- Tables:
spdf→spdf_meta(clave→key,valor→value),documentos→documents,unidades→units,secciones→sections,fragmentos→fragments,figuras→figures,espacios→spaces,vectores→vectors,blobs→blobs,procedencia→provenance; no extensions. - Columns:
tipo→kind,metadatos→metadata,huella→source_sha256,original→source_ref,unidades→unit_count,duracion→duration,creado→created,actualizado→updated,titulo→title,autores→authors,anio→year,idioma→language;orden→ord,ancla→anchor,texto→text,notas→notes,cabecera→header,pie→footer,imagen→image,miniatura→thumbnail,lector→reader,confianza→confidence,impresa→printed,palabras→words;padre→parent,nivel→level,unidad_desde→unit_from,unidad_hasta→unit_to,resumen→summary;unidad→unit,contexto→context,seccion→section,ancla_fin→anchor_end,texto_busqueda→search_text; in figurespie→caption,descripcion→description;proveedor→provider,modelo→model,normalizado→normalized,modalidades→modalities(texto→text,imagen→image);objetivo→target(fragmento→fragment,unidad→unit,figura→figure),espacio→space,valores→data;clave→key,datos→data;fase→stage,detalle→detail,cuando→at. - Kinds:
pdf,pdf_escaneado→scanned_pdf,fotos→photos,imagen→image,audio,video,documento→document,epub,presentacion→slides,hoja→sheet,web. - Anchors:
tipo→type(pagina→page,tiempo→time,seccion→section,diapositiva→slide,hoja→sheet,web,imagen→image),fisica→physical,impresa→printed,romana→roman,origen→source(leido→read,deducido→inferred,epub,ninguno→none),confianza→confidence,hablante→speaker,ruta→path,parrafo→paragraph,n,hoja→sheet,filaDesde→row_from,filaHasta→row_to,consultada→accessed,region. Unknown members are kept as they are. - Metadata (
MetadatosDocumento→ CSL-JSON):titulo→title, or"titulo: subtitulo"withtitle-short=tituloandspdf.subtitle=subtitulo;tituloOriginal→original-title;autores,editores,traductores,entrevistadores({nombre, apellidos, orcid}) →author,editor,translator,interviewer({family: apellidos, given: nombre}, empty parts omitted; ORCID tospdf.orcidunder"apellidos, nombre");fecha→issuedwith its full date parts when its year equalsanioor there is noanio, otherwiseanio→issued;anioOriginal→original-date;editorial→publisher;lugar→publisher-place;revista, elsecontenedor→container-title;coleccion→collection-title;volumen→volume;numero→issue;paginas→page;edicion→edition;doi→DOI;isbn→ISBN;url→URL;idioma→language;resumen→abstract;idiomaOriginal→spdf.original_language;sinFecha{desde, hasta, fundamento}→spdf.undated{from, to, basis};procedencia→spdf.provenance, with field names mapped as above andfuente→source(lectura→reading,usuario→user,colofon→colophon,impresores→printers, others unchanged),confianza→confidence.tipoCSL→type; without it, the type isarticle-journalwhenrevistais present, otherwise by kind:audioandpresentacion→speech,video→motion_picture,web→webpage,hoja→dataset,imagenandfotos→graphic, anything else →book. Empty strings, nulls and empty arrays are omitted. - Other rules:
spdf_metakeyscreado→createdandgenerador→generator, others unchanged;documentos.estadoanddocumentos.bibliotecasare dropped;rightsis null; spaces getdtypef32,truncated_fromandtask_prefixesnull andcreatedfromcreado;provenance.modelis null; blob hashes are computed; units are renumberedord= 1, 2, 3… in order of (orden,id), because 4.x numbers units from 0;fragments.ordkeepsorden. References inoriginal,imagenandminiatura: an empty string becomes null (in figures it stays an empty string), a value equal to a key ofblobsbecomesblob:<key>, any other value is kept as an opaque reference.
The 4.x anchor has no foliation; leaf folios did not exist in 4.x.
20.2 SPDF 3.0 and earlier
Scholaris v1 to v3 wrote gzip-wrapped SQLite databases with the tables metadata (key,
value, including schema_version) and chunks, among others. Readers MAY import them;
importing is a conversion with losses (cross-modal links, scenes and some embeddings have
no place in 5.0) and the importer SHOULD report what it dropped. Version 3.0 is documented
historically in the Scholaris repository; this specification does not define it.
21. Conformance
21.1 Product classes
- A conforming reader opens files safely (§2.4), reads 5.0 and legacy 4.x files, produces the canonical dump (§12), parses and formats anchor URIs (§5), produces short citations (§18), runs the reference lexical search (§8.1) and exports CSL-JSON and BibTeX (§19). A semantic reader also runs the reference vector and hybrid search.
- A conforming writer produces files that validate without errors or warnings for the profiles they declare and whose canonical dump equals the dump the writer was given (round trip).
- A conforming validator reports exactly the codes of §22 for the validation cases of the suite.
21.2 Levels
An implementation states its class and the profiles it covers, for example "reader and
writer, profiles core and semantic". Its claim is backed by the conformance suite: it
passes every case of the kinds its class requires (dump, legacy_dump, anchor_uri,
cite, cite_passage, search_lexical, validate, locate, export_csl,
export_bibtex for readers; plus search_vector and search_hybrid for semantic readers; plus roundtrip
and quantize for writers; export_structure for implementations that export ALTO, TEI
or IIIF), with the suite
version it was tested against. Partial implementations MAY exist but MUST NOT call
themselves conforming.
21.3 The suite
The suite (conformance/ in the repository) is normative for behaviour. Its protocol
(case format, runner report, CI convention) is in conformance/README.md. Each suite
release has a version and a manifest with the number of cases and their hash.
22. Validation
22.1 Procedure
A validator checks a file in this order; a step marked stop ends validation:
- If the file starts with
1F 8B, decompress it (§2.3). - If the result is not a SQLite database: E001, stop.
- Determine the version:
application_id1397769286 withuser_version500–599 is 5.x; the legacy detection of §20.1 is 4.x; anything else: E002, stop. For 4.x, report W110 and check only: the tablesspdf,documentos,unidades,fragmentos,fragmentos_ftsexist (E010 each) and there is no trigger or view other than the three tolerated triggers (E020); stop. - A gzip-wrapped 5.x file: E003 in the warnings. A minor version above 0: W105; for such a file, unknown anchor types (E041) and unknown dtypes (E032) are reported in the warnings instead of the errors, because a later minor version may define them.
- Triggers, views and foreign virtual tables: E020 for each.
- Required tables (E010 each) and required columns (E011 each).
spdf_metakeys (E012 each).documentsholds exactly one row (E013);metadataandrightsare valid JSON (E050);metadatahas stringtypeandtitle(E051).- Required extensions unknown to the validator (E060).
units.ordis 1…N (E090);unit_countequals N (W102).- Anchors of units, fragments (start and end) and figures (E040, E041, E042);
fragments that cross
matteror a folio boundary (W103, §4.4). - Spaces:
dtype(E032). Vectors: known space (E031), length (E030). - FTS index in sync: run
INSERT INTO fragments_fts(fragments_fts, rank) VALUES('integrity-check', 1)(and the same onfragments_fts_trigram) on a private copy; an error is E070. - Blobs: stored
sha256equals the computed one (E080). - If there were no errors so far and
content_sha256is present: recompute it (E081); if it matches andsignatureis present, verify it (E082). - Profile warnings: W100, W101.
The result is a JSON object (schema in json-schema/validation-result.schema.json):
{"valid": false, "version": "5.0", "profile": ["core"],
"errors": [{"code": "E090", "message": "units.ord is not 1..N", "where": "units"}],
"warnings": []}valid is true if and only if errors is empty. version is null when unknown.
Messages are free text; conformance compares the sets of codes.
22.2 Codes
| code | meaning |
|---|---|
| E001 | not a SQLite database (or bad gzip) |
| E002 | unknown application_id or version |
| E003 | gzip-wrapped 5.0 file (reported as a warning) |
| E010 | missing required table |
| E011 | missing required column |
| E012 | missing required spdf_meta key |
| E013 | documents does not hold exactly one row |
| E020 | trigger, view or foreign virtual table present |
| E030 | vector length ≠ dims × dtype size |
| E031 | vector refers to an unknown space |
| E032 | unknown dtype |
| E040 | invalid anchor (bad JSON, missing or mistyped required member) |
| E041 | unknown anchor type |
| E042 | chars out of range |
| E050 | invalid metadata or rights JSON |
| E051 | metadata without string type and title |
| E060 | unknown required extension |
| E070 | FTS index out of sync |
| E080 | blob sha256 mismatch |
| E081 | content_sha256 mismatch |
| E082 | signature does not verify |
| E090 | units.ord not contiguous from 1 |
| W100 | profile semantic without vectors |
| W101 | profile media without time anchors |
| W102 | unit_count ≠ number of units |
| W103 | fragment crosses between units of different matter, or between a page with a printed folio and one without |
| W105 | newer minor version than the validator's |
| W110 | legacy 4.x file |
Codes are never reused with another meaning. New codes are added by minor versions.
23. Versioning and compatibility
The specification is versioned MAJOR.MINOR; editorial corrections do not change the
version. user_version encodes it (§2.1).
- A minor version (5.1, 5.2…) only adds OPTIONAL things: tables, columns,
spdf_metakeys, anchor members or types, metadata members, validation warnings or errors for things that were already forbidden. A 5.0 reader reads every 5.x file, ignoring what it does not know; it MAY warn (W105). Validators report the anchor types and dtypes of a newer minor version as warnings, not errors (§22.1). A 5.x writer that uses nothing new SHOULD writeuser_version500. - A major version (6.0) may change or remove things. Readers MUST refuse majors they do not know (E002) and SHOULD keep reading older majors (as 5.0 reads 4.x).
- Deprecation: a feature is deprecated in a minor version, with the reason and the replacement, and removed no earlier than the next major and at least 24 months later.
- Promise: a file that conforms to 5.0 will be readable by every conforming reader of any later 5.x version, and its anchor URIs will keep resolving.
- The conformance suite and each library have their own version numbers; the suite's manifest says which specification version it tests.
Changes are proposed and decided through the RFC process in spec/rfcs/ and
governance/.
24. Media type and file identification
- Media type:
application/vnd.spdf+sqlite3(registration with IANA in preparation; template ingovernance/drafts/iana-media-type.md). The structured syntax suffix+sqlite3tells generic tools that the file is a SQLite 3 database. OPTIONAL parameterversion("5.0"). Encoding: binary. Legacy 4.x files are gzip data and have no registered type of their own. - Fragment identifiers: for a resource of type
application/vnd.spdf+sqlite3, the fragment identifier is theparamsrule of §5.1, with the meaning it has in an anchor URI for the document in that resource:https://example.org/quijote.spdf#p=5&f=1r. - Extension:
.spdf. Sidecars:.spdfa.jsonand.spdfl.json, served asapplication/json(orapplication/ld+jsonfor annotations). - Magic numbers: bytes 0–15 are
53 51 4C 69 74 65 20 66 6F 72 6D 61 74 20 33 00("SQLite format 3" and a NUL); bytes 68–71 are53 50 44 46("SPDF"); bytes 60–63 holduser_versionbig-endian (00 00 01 F4for 5.0). Legacy 4.x files start with1F 8Band cannot be told from other gzip files without decompressing. - Uniform Type Identifier (Apple platforms):
com.joseluissaorin.spdf, conforming topublic.dataandpublic.database, until a vendor-neutral identifier is agreed.
25. Internationalization
25.1 Languages and scripts
Language tags are BCP 47 [BCP 47]: es, en-GB, la, grc (Ancient Greek), lzh
(Literary Chinese), ar. documents.language is the main language; fragments in other
languages need no tagging in this version. Text is stored in logical order, whatever its
direction.
25.2 Right-to-left text
Arabic, Hebrew, Syriac and other right-to-left scripts are stored in logical order,
without bidirectional control characters except those present in the source. Readers
display them with the Unicode Bidirectional Algorithm [UAX #9] and SHOULD isolate
user-supplied strings (dir="auto"). Anchor URIs percent-encode such text, so they are
direction-neutral; when an IRI form is displayed, readers SHOULD isolate it. chars
offsets count code points in logical order.
25.3 Old texts
The text of a unit or fragment is the text of the source, never modernized: long s
(ſ), u/v and i/j alternations, abbreviations and tildes stay as printed. The
search_text layer carries a modernized form used only by search. The unicode61
tokenizer with remove_diacritics 2 already folds case, Latin diacritics and ſ; it does
not fold Greek accents and breathings, ligatures such as æ and œ, or ß, so producers
SHOULD put the folded forms they need in search_text (for polytonic Greek, the text
without diacritics). Citations quote text, never search_text.
25.4 Chinese, Japanese and Korean
The unicode61 tokenizer treats a run of Han characters as one token. Files whose text is
mostly CJK SHOULD include fragments_fts_trigram; the reference search then matches
substrings of three or more characters by trigram and shorter ones by substring
(§8.1). Producers MAY add a segmented form (words separated by spaces) to
search_text.
25.5 Numbers and folios
Printed folios are stored as printed, in any script ("xiv", "٣٤", "三"). Readers
MUST NOT convert them for citation; they MAY offer conversions for navigation.
References
Normative
- [BCP 47] Phillips, A., Davis, M., "Tags for Identifying Languages", BCP 47, RFC 5646.
- [COMMONMARK] CommonMark Spec, version 0.31.2, https://spec.commonmark.org/0.31.2/.
- [CSL-JSON] Citation Style Language, CSL-JSON schema, https://github.com/citation-style-language/schema.
- [MEDIA-FRAGMENTS] W3C, "Media Fragments URI 1.0 (basic)", Recommendation, 2012.
- [RFC 1952] Deutsch, P., "GZIP file format specification version 4.3".
- [RFC 2119] Bradner, S., "Key words for use in RFCs to Indicate Requirement Levels".
- [RFC 3986] Berners-Lee, T., et al., "Uniform Resource Identifier (URI): Generic Syntax".
- [RFC 3987] Duerst, M., Suignard, M., "Internationalized Resource Identifiers (IRIs)".
- [RFC 5147] Wilde, E., Duerst, M., "URI Fragment Identifiers for the text/plain Media Type".
- [RFC 5234] Crocker, D., Overell, P., "Augmented BNF for Syntax Specifications: ABNF".
- [RFC 8032] Josefsson, S., Liusvaara, I., "Edwards-Curve Digital Signature Algorithm (EdDSA)".
- [RFC 8174] Leiba, B., "Ambiguity of Uppercase vs Lowercase in RFC 2119 Key Words".
- [RFC 8259] Bray, T., "The JavaScript Object Notation (JSON) Data Interchange Format".
- [RFC 8785] Rundgren, A., et al., "JSON Canonicalization Scheme (JCS)".
- [SQLITE-FORMAT] SQLite, "Database File Format", https://www.sqlite.org/fileformat.html.
- [SQLITE-FTS5] SQLite, "SQLite FTS5 Extension", https://www.sqlite.org/fts5.html.
- [UAX #15] Unicode Standard Annex #15, "Unicode Normalization Forms".
- [WEB-ANNOTATION] W3C, "Web Annotation Data Model", Recommendation, 2017.
Informative
- [ALTO] Library of Congress, "ALTO: Technical Metadata for Layout and Text Objects", version 4.
- [CTS] "Canonical Text Services" protocol and URN scheme, http://cite-architecture.github.io/.
- [IIIF] IIIF Consortium, "IIIF Presentation API 3.0".
- [MRL] Kusupati, A., et al., "Matryoshka Representation Learning", NeurIPS 2022.
- [RFC 6838] Freed, N., Klensin, J., Hansen, T., "Media Type Specifications and Registration Procedures".
- [RFC 7595] Thaler, D., et al., "Guidelines and Registration Procedures for URI Schemes".
- [RRF] Cormack, G. V., Clarke, C. L. A., Büttcher, S., "Reciprocal Rank Fusion outperforms Condorcet and individual rank learning methods", SIGIR 2009.
- [SPDX] SPDX License List, https://spdx.org/licenses/.
- [SQLITE-SECURITY] SQLite, "Defense Against The Dark Arts", https://www.sqlite.org/security.html.
- [TEI] TEI Consortium, "TEI P5: Guidelines for Electronic Text Encoding and Interchange".
- [UAX #9] Unicode Standard Annex #9, "Unicode Bidirectional Algorithm".
- [VEC2TEXT] Morris, J. X., et al., "Text Embeddings Reveal (Almost) As Much As Text", EMNLP 2023.
Appendix A. Changes from SPDF 4.1
- Uncompressed container with
application_idanduser_version; gzip only for legacy. - English identifiers; legacy files read through the 5.0 view.
- Metadata as a CSL-JSON item with the
spdfextension object. - New anchor types
verseandcanonical;foliation(leaves and columns);charsandregionon any anchor. - Anchor URI with ABNF, aligned with W3C Media Fragments and RFC 5147.
spaces.dtype(f32,f16,i8),truncated_from,task_prefixes; space compatibility.- Profiles, extensions,
rights, blob hashes, provenancemodel. - No triggers or views in distributed files; safe opening procedure.
- Canonical dump, integrity hash and Ed25519 signatures.
- Units numbered from 1.
- Removed:
documentos.estadoanddocumentos.bibliotecas(library membership belongs to collection manifests).