# SPDF: the specification and every page of https://spdf.joseluissaorin.com Source: https://spdf.joseluissaorin.com/llms.txt · Generated: 2026-10-07 · Licence: CC BY 4.0 # ===== ENGLISH ===== --- # SPDF 5.0 specification URL: https://spdf.joseluissaorin.com/spec > The normative specification of the SPDF 5.0 format: container, schema, anchors and their URI, CSL-JSON metadata, search, vectors, integrity, validation, citation and conformance. - **Status:** Working Draft, 2026-10-07. Stable enough to implement; changes go through the RFC process (`spec/rfcs/`) and are logged in `spec/CONTRACT.md` until 5.0 is final. - **Editor:** José Luis Saorín Ferrer. - **This version:** `spec/SPEC.md` in . - **Spanish translation:** [`SPEC.es.md`](SPEC.es.md) (faithful; in case of conflict the English text prevails). - **License:** this specification is published under [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/). Code in the repository is MIT OR Apache-2.0. Contributors commit not to assert patents against implementations. ## Abstract SPDF is an open, portable file format for documents that have been **read once and can be cited forever**. A `.spdf` file holds the text of one document (a printed book, a scan, a recording, a slide deck, a spreadsheet, a web page) as a set of citable units and searchable fragments, and every fragment carries an exact **anchor**: the printed page or leaf, the second of a recording, the slide, the verse, the canonical reference. A citation produced from an SPDF file can only print what the source says. The container is a plain SQLite 3 database, the metadata is a CSL-JSON item, the full-text index uses tokenizers that ship with every SQLite, and optional embedding vectors from several models can live side by side. Any language with SQLite can read SPDF without special libraries. ## Status of this document This is the first public version of the format (earlier versions, 3.0 to 4.1, were internal to Scholaris and are covered as legacy in [§20](#legacy)). The conformance suite in `conformance/` is part of the specification: where this text and a conformance case disagree, the disagreement is a bug to be resolved through the RFC process; until it is, implementations follow the conformance case. ## Contents - [Preface](#preface) - [1. Conventions and terminology](#terminology) - [2. Container](#container) - [3. Schema](#schema) - [4. Anchors](#anchors) - [5. Anchor URI](#anchor-uri) - [6. Metadata](#metadata) - [7. Text, normalization and offsets](#text-normalization) - [8. Reference search](#search) - [9. Vector spaces](#vectors) - [10. Profiles](#profiles) - [11. Extensions](#extensions) - [12. Canonical dump](#dump) - [13. Integrity and signatures](#integrity) - [14. Security considerations](#security) - [15. Privacy considerations](#privacy) - [16. Rights](#rights) - [17. Annotations and collections](#annotations) - [18. Short citation](#citation) - [19. Exports](#exports) - [20. Legacy formats](#legacy) - [21. Conformance](#conformance) - [22. Validation](#validation) - [23. Versioning and compatibility](#versioning) - [24. Media type and file identification](#media-type) - [25. Internationalization](#i18n) - [References](#references) - [Appendix A. Changes from SPDF 4.1](#changes) ## Preface SPDF was born inside Scholaris, an application written by José Luis Saorín Ferrer to insert verified, page-exact citations into academic writing. Scholaris needed to read a source once (with a PDF text layer, a vision model, or a speech recognizer), keep what it had read, and answer for years afterwards the only question a citation must answer honestly: *where exactly does the source say this?* The answer had to survive the original file being moved, the reading model being replaced and the search engine being rewritten. The result was a file per document, the *Scholaris Processed Document Format*, which went through a gzip-compressed JSON-and-SQLite version (3.0) and a Spanish-named SQLite schema (4.0 and 4.1). Version 5.0 is the first version designed for everyone. It keeps what experience proved right and drops what tied it to one program: identifiers are in English, the container is uncompressed so it can be memory-mapped and read by HTTP ranges, the metadata is plain CSL-JSON so that Zotero, citeproc and Pandoc understand it, and every number in a conformance case comes from an oracle that anyone can rerun. The name became *Semantic Processed Document Format*; the initials did not change. Five principles guide every decision in this specification: 1. **Anchors first.** Every fragment knows exactly where it comes from: physical page and printed folio, leaf and side, second, slide, verse, canonical reference. Nothing that cannot be anchored is citable. 2. **Provenance.** A file says who read each unit and with what confidence, which model produced each vector, and how each metadata field was obtained. Derived data can be recomputed from the original plus the units. 3. **Read once, query many times.** Reading a document is expensive (vision models, speech recognition, human correction); querying it must be cheap, offline, and possible from any language with SQLite. 4. **Portability.** One file, one document, no server, no proprietary dependency, no compression layer to undo, no code inside the file. Readers in many languages pass the same conformance suite. 5. **Honest citation.** A citation prints only what an anchor says. A folio that was inferred is printed in brackets; an unnumbered page is cited as unnumbered; the modernized spelling used for search is never quoted. ## 1. Conventions and terminology The key words "MUST", "MUST NOT", "REQUIRED", "SHALL", "SHALL NOT", "SHOULD", "SHOULD NOT", "RECOMMENDED", "NOT RECOMMENDED", "MAY", and "OPTIONAL" in this document are to be interpreted as described in BCP 14 [RFC 2119] [RFC 8174] when, and only when, they appear in all capitals, as shown here. ABNF follows [RFC 5234]. JSON follows [RFC 8259]; "JSON object", "array", "string" and "number" have their RFC 8259 meanings. SQL follows SQLite's dialect. - **Document**: the work an SPDF file describes (one per file). - **Original**: the bytes the document was read from (PDF, image set, audio, EPUB…). - **Unit**: a citable division of the document: a page or leaf, a time span, a slide, a section, a sheet range. Units are ordered and numbered from 1. - **Fragment**: a searchable, citable passage of roughly 150 to 300 words, with the anchor of its start and, if it crosses units, of its end. - **Anchor**: a JSON object that locates a unit, a fragment or a figure in the document ([§4](#anchors)). - **Anchor URI**: the textual form of an anchor, `spdf:#` ([§5](#anchor-uri)). - **Space**: a vector space, i.e. the model, dimensions and encoding that produced a set of embedding vectors ([§9](#vectors)). - **Reader**: software that opens SPDF files and exposes their content. **Writer**: software that creates SPDF files. **Validator**: software that checks files against this specification. **Producer**: a writer that also reads originals (OCR, speech recognition, embeddings). - **Code point**: a Unicode scalar value. Lengths and offsets in this specification count code points, never bytes or UTF-16 code units. - **NFC**: Unicode Normalization Form C [UAX #15]. - **JCS**: the JSON Canonicalization Scheme [RFC 8785]. ## 2. Container ### 2.1 File An SPDF 5.0 file is a SQLite 3 database file [SQLITE-FORMAT] holding exactly one document. It MUST NOT be wrapped in any compression or archive layer: the database header MUST start at byte 0. Writers MUST set: - `PRAGMA application_id = 1397769286` (hexadecimal 0x53504446). SQLite stores it big-endian at byte offset 68 of the header, so bytes 68 to 71 read "SPDF" in ASCII. - `PRAGMA user_version = 500`. The value encodes the specification version as major × 100 + minor × 10 (5.0 → 500, 5.1 → 510). - The `spdf_meta` row `spdf_version` to `"5.0"` ([§3.2](#schema)). Writers SHOULD use a page size of 4096 bytes, the rollback journal in `DELETE` mode (never leave a `-wal` or `-journal` file next to a distributed file), and run `VACUUM` after the last write so the file has no free pages. Writers SHOULD NOT use `auto_vacuum`. A file MUST NOT contain triggers or views, and MUST NOT contain virtual tables other than the FTS5 tables defined in [§3](#schema). Writers keep the full-text index in sync themselves (for example with `INSERT INTO fragments_fts(fragments_fts) VALUES('rebuild')` before `VACUUM`). ### 2.2 Name and type The file extension is `.spdf`. The media type is `application/vnd.spdf+sqlite3` ([§24](#media-type)). One file holds one document; libraries of documents are described by a separate collection manifest ([§17](#annotations)). ### 2.3 Gzip input Legacy 4.x files are SQLite databases wrapped in gzip [RFC 1952] ([§20](#legacy)). Readers MUST therefore accept a file that starts with the gzip magic bytes `1F 8B`, decompress it (to memory or to a temporary file) with a configurable limit on the decompressed size (RECOMMENDED default 4 GiB), and continue with the result. A gzip-wrapped 5.0 file is readable but non-conforming: validators report it as `E003` in the warnings list ([§22](#validation)). ### 2.4 Safe opening SPDF files come from strangers. Every reader MUST open them as follows, and MUST refuse the file if a step cannot be honoured by its SQLite binding: 1. Open the database read-only (`SQLITE_OPEN_READONLY`, or the URI parameter `mode=ro`). Never open a distributed file read-write in place. 2. `PRAGMA query_only = 1` and `PRAGMA trusted_schema = OFF`. 3. Enable `SQLITE_DBCONFIG_DEFENSIVE` where the binding exposes it, and keep extension loading disabled (`sqlite3_enable_load_extension(db, 0)`; never call `load_extension`). 4. Read `sqlite_master` and refuse the file if it contains a trigger, a view, or a virtual table other than `fragments_fts` and `fragments_fts_trigram` declared `USING fts5`. Legacy 4.x files are allowed exactly the three triggers `fragmentos_ai`, `fragmentos_ad` and `fragmentos_au` ([§20](#legacy)), which never fire on a read-only connection. 5. Enforce a configurable maximum size for any single BLOB or TEXT value read (RECOMMENDED default 512 MiB), for example with `sqlite3_limit(db, SQLITE_LIMIT_LENGTH, …)`. Readers SHOULD also disable memory-mapped I/O (`PRAGMA mmap_size = 0`) and enable `PRAGMA cell_size_check = ON` for files from untrusted sources, and MAY run `PRAGMA quick_check` before use. Operations that need to write, such as the FTS5 `integrity-check` command, MUST run on a private copy (for example an in-memory copy made with the backup API), never on the file. [§14](#security) explains the threats. ## 3. Schema ### 3.1 Overview The normative schema is the SQL script [`schema/spdf-5.0.sql`](schema/spdf-5.0.sql), reproduced in full below. Every table in it is REQUIRED, even when empty; only `fragments_fts_trigram` is OPTIONAL. Column names, types and constraints MUST be as written. Writers MUST NOT add columns to these tables; data that does not fit goes into extension tables ([§11](#extensions)). Readers MUST ignore columns they do not know (a later minor version may add OPTIONAL columns, [§23](#versioning)). JSON stored in TEXT columns MUST be valid JSON [RFC 8259] encoded in UTF-8; writers MAY serialize it in any form (the canonical dump re-serializes it, [§12](#dump)). Timestamps are ISO 8601 / RFC 3339 strings in UTC with a `Z` suffix. Identifiers (`id` columns) are non-empty strings chosen by the writer; they are opaque, case-sensitive and stable for the life of the file. ```sql PRAGMA application_id = 1397769286; -- 0x53504446, "SPDF" PRAGMA user_version = 500; CREATE TABLE spdf_meta (key TEXT PRIMARY KEY, value TEXT NOT NULL); CREATE TABLE documents ( id TEXT PRIMARY KEY, kind TEXT NOT NULL, metadata TEXT NOT NULL, source_sha256 TEXT NOT NULL, source_ref TEXT, mime TEXT NOT NULL, bytes INTEGER NOT NULL, unit_count INTEGER NOT NULL, duration REAL, created TEXT NOT NULL, updated TEXT NOT NULL, title TEXT, authors TEXT, year INTEGER, language TEXT, rights TEXT); CREATE TABLE units ( id TEXT PRIMARY KEY, document TEXT NOT NULL REFERENCES documents(id), ord INTEGER NOT NULL, anchor TEXT NOT NULL, text TEXT NOT NULL DEFAULT '', notes TEXT, header TEXT, footer TEXT, image TEXT, thumbnail TEXT, reader TEXT NOT NULL, confidence REAL NOT NULL DEFAULT 1, printed TEXT, t0 REAL, t1 REAL, words TEXT); CREATE INDEX units_doc ON units(document, ord); CREATE INDEX units_printed ON units(document, printed); CREATE TABLE sections ( id TEXT PRIMARY KEY, document TEXT NOT NULL, parent TEXT, level INTEGER NOT NULL, title TEXT NOT NULL, unit_from TEXT NOT NULL, unit_to TEXT, summary TEXT); CREATE TABLE fragments ( n INTEGER PRIMARY KEY, id TEXT NOT NULL UNIQUE, document TEXT NOT NULL, unit TEXT NOT NULL, ord INTEGER NOT NULL, text TEXT NOT NULL, context TEXT NOT NULL DEFAULT '', section TEXT, anchor TEXT NOT NULL, anchor_end TEXT, search_text TEXT); CREATE INDEX fragments_doc ON fragments(document, ord); CREATE INDEX fragments_unit ON fragments(unit); CREATE VIRTUAL TABLE fragments_fts USING fts5( text, context, section, search_text, content='fragments', content_rowid='n', tokenize='unicode61 remove_diacritics 2'); -- OPTIONAL: -- CREATE VIRTUAL TABLE fragments_fts_trigram USING fts5( -- text, content='fragments', content_rowid='n', tokenize='trigram'); CREATE TABLE figures ( id TEXT PRIMARY KEY, document TEXT NOT NULL, unit TEXT NOT NULL, image TEXT NOT NULL, caption TEXT, description TEXT, anchor TEXT NOT NULL); CREATE TABLE spaces ( id TEXT PRIMARY KEY, provider TEXT NOT NULL, model TEXT NOT NULL, version TEXT, dims INTEGER NOT NULL, dtype TEXT NOT NULL DEFAULT 'f32', normalized INTEGER NOT NULL DEFAULT 1, truncated_from INTEGER, modalities TEXT NOT NULL, task_prefixes TEXT, created TEXT); CREATE TABLE vectors ( target TEXT NOT NULL, id TEXT NOT NULL, space TEXT NOT NULL REFERENCES spaces(id), document TEXT NOT NULL, data BLOB NOT NULL, PRIMARY KEY (target, id, space)); CREATE TABLE blobs (key TEXT PRIMARY KEY, mime TEXT NOT NULL, sha256 TEXT NOT NULL, data BLOB NOT NULL); CREATE TABLE provenance ( document TEXT NOT NULL, stage TEXT NOT NULL, provider TEXT, model TEXT, detail TEXT, ms INTEGER, at TEXT NOT NULL); CREATE TABLE extensions (name TEXT PRIMARY KEY, version TEXT NOT NULL, required INTEGER NOT NULL DEFAULT 0); ``` ### 3.2 `spdf_meta` Key/value pairs about the file. REQUIRED keys: | key | value | |---|---| | `spdf_version` | `"5.0"` | | `profile` | space-separated profile names, a subset of `core semantic media full` ([§10](#profiles)); always includes `core` | | `created` | creation time of the file (UTC) | | `generator` | `name/version` of the writer, e.g. `spdf-producer/0.3.1` | | `document_id` | equal to `documents.id` | OPTIONAL keys: `content_sha256`, `signature`, `signer` ([§13](#integrity)) and `license_note` (free text for humans). Other keys MAY be added by later versions or by extensions (prefixed `x__`); readers MUST ignore keys they do not know. ### 3.3 `documents` Exactly one row. - `kind`: one of `pdf` (PDF with a usable text layer), `scanned_pdf` (PDF read by vision), `photos` (a set of page photographs), `image` (a single image), `audio`, `video`, `document` (DOCX, ODT, RTF, HTML, Markdown, plain text), `epub`, `slides`, `sheet`, `web`. Readers MUST accept unknown kinds and treat them as `document`. - `metadata`: the CSL-JSON item with the `spdf` extension ([§6](#metadata)). - `source_sha256`: lowercase hexadecimal SHA-256 of the original bytes. It identifies the document across copies and is the preferred document reference in anchor URIs. - `source_ref`: where the original is: `blob:` when shipped inside the file, an absolute URL, or NULL. - `mime`, `bytes`: media type and size in bytes of the original. - `unit_count`: number of rows in `units` (a mismatch is warning W102). - `duration`: seconds, for audio and video; NULL otherwise. - `created`, `updated`: when the document record was created and last changed. - `title`, `authors`, `year`, `language`: denormalized copies for filtering without parsing JSON: the CSL `title`; the family names (or literal names) of the CSL `author` list joined with `"; "`; the first year of `issued`; the CSL `language`. They MUST agree with `metadata` when present. - `rights`: JSON rights object ([§16](#rights)) or NULL. ### 3.4 `units` One row per citable unit, `ord` = 1, 2, 3… without gaps (E090), in reading order. - `anchor`: the unit's anchor ([§4](#anchors)). - `text`: the full text of the unit as read, NFC, light Markdown ([§7](#text-normalization)). Empty string for units without text (a blank page, a photograph). - `notes`: JSON array of strings (footnotes detached from the body) or NULL. - `header`, `footer`: running heads and feet, kept out of `text`, or NULL. - `image`, `thumbnail`: `blob:` or URL of the unit's image (page, frame, slide) and of its thumbnail, or NULL. - `reader`: what produced `text` (`pdf-text-layer`, `tesseract-5`, `gemma-4-e4b`, `whisper-large-v3-turbo`, `human`…). `confidence`: 0 to 1. - `printed`: the printed folio of a page unit, copied from its anchor, so readers can "go to page 145" with an index. - `t0`, `t1`: start and end in seconds, copied from a time anchor; NULL otherwise. - `words`: word timings for audio and video ([§7.4](#text-normalization)) or NULL. ### 3.5 `sections` The heading tree. `level` starts at 1; `parent` is the id of the enclosing section or NULL; `unit_from` and `unit_to` are the first and last unit ids (`unit_to` NULL when the section ends with the document); `summary` is OPTIONAL text in the document language. ### 3.6 `fragments` - `n`: a positive integer, unique, stable: it is the rowid the FTS5 index uses (an implicit rowid may change on `VACUUM`). - `unit`: id of the unit where the fragment starts. `ord`: reading order within the document (increasing with the position of the fragment in the text). - `text`: the literal passage, NFC, exactly as in the source (never modernized). - `context`: one line that situates the fragment in the work ("Chapter III: the struggle for existence"), used by search; empty string if none. - `section`: JSON array of strings, the heading path, or NULL. - `anchor`: anchor of the start of the fragment. `anchor_end`: anchor of its end when it crosses into another unit; NULL otherwise. - `search_text`: the modernized-spelling layer ([§25.3](#i18n)): text used ONLY for search (`aſsi` → `así`, `V. M.` → `vuestra merced`). Empty string when it adds nothing; NULL when not computed. It MUST NOT be displayed as the text of the source or quoted. ### 3.7 `fragments_fts` and `fragments_fts_trigram` `fragments_fts` is an external-content FTS5 index over `fragments` with the columns `text`, `context`, `section` and `search_text` in this order and the tokenizer `unicode61 remove_diacritics 2`, which every SQLite with FTS5 provides. It MUST be in sync with `fragments` (E070). `fragments_fts_trigram` is OPTIONAL, indexes `text` only with the `trigram` tokenizer (SQLite 3.34 or later), and SHOULD be present when the document is mostly in Chinese, Japanese or Korean. ### 3.8 `figures` Figures, plates, tables as images, photographs inside a page. `image` is `blob:` of a cropped image, or the unit's image together with a `region` in the anchor. `caption` is the printed caption, if any; `description` is a description in the document language (for accessibility and search). `anchor` normally carries a `region`. ### 3.9 `spaces` and `vectors` See [§9](#vectors). `vectors.target` is `fragment`, `unit` or `figure` and `vectors.id` is the id of that row; `data` is the little-endian vector. ### 3.10 `blobs` Binary content shipped inside the file: the original, page images, cropped figures, thumbnails. `key` is an opaque string (by convention path-like, `pages/0001.png`), `mime` its media type, `sha256` the lowercase hex SHA-256 of `data` (E080). Other tables refer to a blob as `blob:`. ### 3.11 `provenance` One row per production step: `stage` (`reading`, `transcription`, `folios`, `metadata`, `embedding`, `figures`…), `provider`, `model`, `detail` (JSON object or NULL), `ms` (duration in milliseconds) and `at` (UTC timestamp). See [§15](#privacy) for what not to record. ### 3.12 `extensions` See [§11](#extensions). ## 4. Anchors ### 4.1 General An anchor is a JSON object with a string member `type`. The types defined by this version and their members are: | type | REQUIRED members | OPTIONAL members | |---|---|---| | `page` | `physical` (integer ≥ 1), `printed` (string or null) | `roman` (boolean), `foliation` (`page`, `leaf`, `column`; default `page`), `source` (`read`, `inferred`, `epub`, `none`), `confidence` (0–1) | | `time` | `t0`, `t1` (seconds, 0 ≤ t0 ≤ t1) | `speaker` (string) | | `section` | `path` (array of strings) | `paragraph` (integer ≥ 1), `printed` (string) | | `slide` | `n` (integer ≥ 1) | | | `sheet` | `sheet` (string), `row_from`, `row_to` (integers) | | | `web` | `url` (string) | `path`, `paragraph`, `accessed` (ISO date) | | `image` | | | | `verse` | `line_from` (integer) | `line_to` (integer), `printed` (string) | | `canonical` | `scheme` (string), `ref` (string) | | Every anchor MAY also carry: - `region`: `{"x", "y", "w", "h"}`, numbers between 0 and 1, fractions of the width and height of the unit's image, origin at the top left; - `chars`: `[start, end]`, code point offsets into the NFC `text` of the anchor's unit, `0 ≤ start ≤ end ≤ length`, end exclusive (E042); - `matter`: what kind of matter the unit is: `body` (the text of the work), `front` (preliminaries: title page, contents, licences, dedication, prologue of an edition), `back` (index, colophon, appendices of an edition), `plate` (a plate or fold-out outside the text pages), `cover`, `library` (bookplates, stamps, library or digitizer pages, licences of a digital edition) or `blank`. Absent means `body`; readers MUST treat values they do not know as `body`. Writers SHOULD set it on the units of paged documents whenever it is not `body`. In this specification an "integer" is a JSON number with an integral value: `10` and `10.0` are the same JSON value and both are integers. An anchor whose JSON is invalid or that lacks or mistypes a REQUIRED member is invalid (E040); an unknown `type` is E041. Readers MUST preserve members they do not know when they copy anchors. ### 4.2 Pages, folios and leaves `physical` is the 1-based position of the page in the original (the PDF page index, the photo number). `printed` is the folio exactly as printed on the page ("23", "xiv", "A-3", "1r"), or null when the page carries no number. - `roman: true` marks folios in roman numerals (front matter). - `foliation` describes what the printed numbers count: `page` (each page numbered), `leaf` (each leaf numbered, sides `r`ecto and `v`erso, printed as `"1r"`, `"1v"`), or `column` (columns numbered, as in some dictionaries and early printed books). - `source` says how `printed` was obtained: `read` (seen on the page), `inferred` (deduced from neighbouring pages, e.g. an unnumbered verso), `epub` (from an EPUB page list), `none` (no folio; `printed` is null). - An inferred folio is cited in brackets, `p. [21]`; a page without folio is cited as unnumbered ([§18](#citation)). A producer MUST NOT invent folios: if no evidence supports a number, `printed` is null and `source` is `none`. ### 4.3 Time, sections, verses and canonical references Time anchors locate recordings in seconds from the start of the original; `speaker` names who speaks. Section anchors locate unpaginated text (EPUB, DOCX, HTML) by heading path and paragraph number, and MAY add the equivalent printed page when the edition provides a page list. Verse anchors count lines of verse (`line_from`, `line_to`), as printed editions number them. Canonical anchors use a citation system that is independent of any edition: `stephanus` (Plato), `bekker` (Aristotle), `bible` (book chapter:verse), `cts` (a CTS URN [CTS]), or any other documented scheme; schemes are lowercase ASCII. ### 4.4 Start and end A fragment's `anchor` locates its start; `anchor_end`, when present, locates its end and has the same `type`. A citation of the whole fragment then prints a range (`pp. 145-146`). - The **end unit** of a fragment is the first unit after its start unit (in `ord` order) whose anchor equals `anchor_end` once `chars` and `region` are removed from both. - `chars` in `anchor` gives the part of the fragment that lies in the start unit, and `chars` in `anchor_end` the part that lies in the end unit (usually `[0, b]`). Writers SHOULD set both on crossing fragments, so that readers know which unit each part of the passage comes from. - Writers SHOULD NOT let a fragment cross from a unit of one `matter` to a unit of another (body text into a plate, a cover, a library page or a licence), nor from a page with a printed folio to a page without one: the citation of such a fragment would mix locators of different natures. Validators report such fragments as W103. ## 5. Anchor URI ### 5.1 Syntax An anchor URI names a place in a document independently of any file: ``` spdf:sha256-3f2a…c9#p=29&f=21&char=118,301 ``` The document reference is `sha256-` followed by the 64 lowercase hex digits of `documents.source_sha256` (RECOMMENDED: it is the same for every copy of the document), or the percent-encoded document id. The fragment is a list of parameters. The parameters reuse W3C Media Fragments syntax [MEDIA-FRAGMENTS] for time (`t=`) and space (`xywh=`) and RFC 5147 [RFC 5147] syntax for character ranges (`char=`), so tools that know those standards can interpret them. The canonical form is defined by this ABNF [RFC 5234]: ```abnf spdf-uri = "spdf:" docref [ "#" params ] docref = hash-ref / id-ref hash-ref = "sha256-" 64lhex lhex = DIGIT / %x61-66 ; 0-9 a-f id-ref = 1*vchar ; percent-encoded document id params = param *( "&" param ) param = p / pe / f / fe / t / s / para / sl / sh / rows / v / ref / char / xywh p = "p=" posint ; physical page pe = "pe=" posint ; physical end page f = "f=" value ; printed folio fe = "fe=" value ; printed end folio t = "t=" number [ "," number ] ; seconds, Media Fragments npt s = "s=" value *( "/" value ) ; section path para = "para=" uint ; paragraph sl = "sl=" posint ; slide sh = "sh=" value ; sheet name rows = "rows=" uint "-" uint ; sheet rows v = "v=" uint [ "-" uint ] ; verse lines ref = "ref=" value ":" value ; canonical scheme ":" reference char = "char=" uint "," uint ; code points, RFC 5147 style xywh = "xywh=percent:" number "," number "," number "," number value = *vchar vchar = unreserved / pct-encoded unreserved = ALPHA / DIGIT / "-" / "." / "_" / "~" pct-encoded = "%" HEXDIG HEXDIG ; uppercase in the canonical form posint = %x31-39 *DIGIT uint = "0" / posint number = uint [ "." 1*DIGIT ] ``` In the canonical form, parameters appear at most once and in the order of the `param` rule above (`p`, `pe`, `f`, `fe`, `t`, `s`, `para`, `sl`, `sh`, `rows`, `v`, `ref`, `char`, `xywh`); values are UTF-8 strings in which every byte other than an unreserved character is percent-encoded with uppercase hex digits; in `s` the separators between path elements are literal `/` and a `/` inside an element is `%2F`; in `ref` the first literal `:` separates the scheme from the reference, and colons inside them are `%3A`. Numbers use the shortest decimal form of ECMAScript (`4160`, `4175.5`, `0.125`), never an exponent. ### 5.2 From an anchor to a URI Formatting maps an anchor (and optionally an end anchor) to parameters: | anchor | parameters | |---|---| | `page` | `p` = `physical`; `f` = `printed` if not null; with an end page: `pe` = its `physical` if different, `fe` = its `printed` if not null and different from `printed` | | `time` | `t` = `t0`, then `t1` (or the end anchor's `t1`) | | `section`, `web` | `s` = `path` if not empty; `para` = `paragraph`; `f` = `printed`; `fe` as for pages | | `slide` | `sl` = `n` | | `sheet` | `sh` = `sheet`; `rows` = `row_from`-`row_to` | | `verse` | `v` = `line_from`, or `line_from`-`line_to` when `line_to` is present and different; `f` = `printed` | | `canonical` | `ref` = `scheme`:`ref` | | `image` | none | | any | `char` = `chars`; `xywh` = `region` × 100, as `percent:` | `t` values are rounded to 6 decimal places. `xywh` values are fractions × 100 rounded to 4 decimal places (`0.125` → `12.5`, `0.333333` → `33.3333`). A URI without parameters (`spdf:`) designates the whole document. ### 5.3 Parsing Parsing returns the document reference and a **locator** object with one member per parameter present: `p`, `pe`, `para`, `sl` (integers); `f`, `fe`, `sh` (strings); `t` (array of one or two numbers); `s` (array of strings); `rows` (two integers); `v` (one or two integers); `ref` (object with `scheme` and `ref`); `char` (two integers); `xywh` (four fractions, the percent values divided by 100 and rounded to 6 decimals). Parsers MUST accept percent-encoding with lowercase hex digits, parameters in any order, unencoded non-ASCII characters (IRI form [RFC 3987]), the `npt:` prefix and the clock forms `h:mm:ss[.f]` and `mm:ss[.f]` in `t`. Parsers MUST ignore parameters whose names they do not know. Parsers MUST reject: a scheme other than `spdf:`; an empty document reference; a repeated parameter; malformed numbers; `p`, `pe` or `sl` equal to 0; a `char` or `t` range whose end precedes its start; `xywh` without the `percent:` unit (pixel coordinates cannot be resolved without the image); percent-encoding that does not decode to valid UTF-8. Formatting a parsed locator MUST give back the canonical URI byte for byte. The conformance suite checks format, parse and round trip for every anchor type. ### 5.4 Resolution `locate(file, reference)` resolves an anchor URI, or the URL of an SPDF resource with a fragment identifier ([§24](#media-type)), against a file, and returns: ```json {"document": true, "units": ["p5", "p6"], "fragments": ["q4"], "char": [101, 278], "xywh": null} ``` 1. **Reference.** An `spdf:` URI is parsed as in [§5.3](#anchor-uri); `document` is true when its document reference is `sha256-` followed by the file's `source_sha256`, or the file's document id. Any other reference (an `https:` URL, a file path) designates the file itself: `document` is true and the text after its first `#`, if any, is parsed as the parameter list of §5.3. When `document` is false, `units` and `fragments` are empty (implementations MAY report this as an error instead; conformance runners map such an error to `document: false`). 2. **Rule.** The first parameter present in the order `p`, `f`, `t`, `sl`, `v`, `ref`, `s`, `sh` selects the predicate below. Without any of them (no fragment, or only `char` and `xywh`) the reference designates the whole document and `units` and `fragments` are empty. 3. **Predicate** on an anchor (members absent from the anchor never match): - `p`: a `page` anchor with `p ≤ physical ≤ pe` (`pe` defaults to `p`); - `f`: `printed` equal to `f` (for units, the `units.printed` column); - `t`: a `time` anchor with `t0 ≤ t < t1`, where `t` is the first value of the parameter; the last unit with a time anchor (in `ord` order) also matches when `t` equals its `t1`; - `sl`: a `slide` anchor with `n = sl`; - `v`: a `verse` anchor with `line_from ≤ v ≤ line_to` (`line_to` defaults to `line_from`), where `v` is the first value of the parameter; - `ref`: a `canonical` anchor with the same `scheme` and `ref`; - `s`: a `section` or `web` anchor whose `path` starts with the elements of `s`; when `para` is present, the `path` must equal `s` and `paragraph` must equal `para`; - `sh`: a `sheet` anchor with `sheet = sh` and, when `rows` is present, `row_from ≤ a ≤ row_to` for its first value `a`. 4. **Matches.** `units` are the ids of the units whose anchor matches, in `ord` order. `fragments` are the ids of the fragments whose start `anchor` or whose `anchor_end` matches, in `n` order (a fragment that ends on a page is found from that page). When no unit matches but some fragments do, `units` are the distinct start units of those fragments, in `ord` order. 5. **Characters.** `char` refers to the text of the first unit of `units`. When `char` = `[c, d]` is present, a fragment is kept in `fragments` only if its start unit is that unit and its `anchor` has `chars` = `[a, b]` that overlap the range, or its end unit ([§4.4](#anchors)) is that unit and its `anchor_end` has `chars` that overlap it; `[a, b]` overlaps `[c, d]` when `a < d` and `c < b` (for `c < d`), or when `a ≤ c < b` (for `c = d`). 6. `char` and `xywh` are copied from the locator, or null. Several units may match (two pages printed "1", a verse number repeated in two poems): `locate` returns them all and the reader lets the user choose; `p` always disambiguates pages, which is why formatted URIs carry it. The `spdf` URI scheme is intended for provisional registration [RFC 7595]; the request is drafted in `governance/drafts/uri-scheme-spdf.md`. ## 6. Metadata ### 6.1 CSL-JSON item `documents.metadata` is one CSL-JSON item [CSL-JSON] describing the document as it should be cited: at least `type` (a CSL type such as `book`, `article-journal`, `chapter`, `thesis`, `speech`, `interview`, `broadcast`, `motion_picture`, `webpage`, `dataset`, `graphic`) and `title` (E051 if either is missing). Common members: `author`, `editor`, `translator`, `interviewer` (arrays of names `{family, given}` or `{literal}`, with the CSL particles `non-dropping-particle` and `dropping-particle` when needed), `issued` (`{"date-parts": [[year, month, day]]}`), `original-date`, `title-short`, `original-title`, `container-title`, `collection-title`, `publisher`, `publisher-place`, `volume`, `issue`, `page`, `edition`, `DOI`, `ISBN`, `ISSN`, `URL`, `accessed`, `language` (BCP 47), `abstract`, `note`. The `id` member is OPTIONAL inside the file; exports set it ([§19](#exports)). Writers MUST NOT invent metadata. A field that cannot be supported by the original or by a cited external source is omitted. ### 6.2 The `spdf` extension object The member `spdf` of the item holds what CSL cannot express. All its members are OPTIONAL: ```json "spdf": { "provenance": {"title": {"source": "title-page", "confidence": 0.99}, "issued": {"source": "colophon", "confidence": 0.95}}, "undated": {"from": 1600, "to": 1610, "basis": "printer active years"}, "original_language": "fr", "subtitle": "con anotaciones de Fernando de Herrera", "orcid": {"Foucault, Michel": "0000-0000-0000-0000"} } ``` - `provenance`: per CSL field, where the value came from (`reading`, `title-page`, `colophon`, `crossref`, `openalex`, `wikidata`, `user`, `epub`, `pdf`, …) and a confidence between 0 and 1. - `undated`: for works without a printed date, a plausible range (`from`, `to`, years, negative for BCE) and the evidence (`basis`). It MUST NOT be copied into `issued`: a citation prints "s. f." / "n.d." ([§18](#citation)). - `original_language`: BCP 47 tag of the original language of a translation. - `subtitle`: the subtitle when the CSL `title` is "Title: Subtitle". - `orcid`: ORCID identifiers by name ("Family, Given"). Other members MAY be added by extensions with the prefix `x__`. ## 7. Text, normalization and offsets ### 7.1 Encoding and normalization All text is UTF-8 in NFC. Writers MUST normalize to NFC before storing and before computing offsets. Writers MUST NOT store U+0000, unpaired surrogates or noncharacters, and SHOULD NOT store other control characters except U+0009 (tab) and U+000A (line feed). Lines end with U+000A only. ### 7.2 Offsets `chars` offsets ([§4.1](#anchors)) and every length in this specification count code points of the NFC text. Implementations whose strings are UTF-16 (JavaScript, Java, C#, Swift's `NSString`) MUST convert: a character outside the Basic Multilingual Plane counts as one code point but two UTF-16 units. ### 7.3 Light Markdown `units.text` MAY use this subset of CommonMark [COMMONMARK]: paragraphs separated by a blank line; `#` to `######` headings; `*emphasis*` and `**strong**`; `-` and `1.` lists; `>` quotations; tables in GitHub style; footnote markers `[^1]` whose text goes to `notes`. Readers MUST NOT interpret raw HTML in `text`; they display it as text. Offsets count the stored characters, markup included. Fragments SHOULD keep the markup of their source unit so that `fragments.text` is a substring of the unit's `text` whenever the fragment does not cross units. Speaker turns in transcripts start with the label `**Name:**` followed by one space (`**Neil Armstrong:** Houston, Tranquility Base here.`). ### 7.4 Word timings `units.words` is the JSON object `{"v": 1, "t0": , "cs": [start, duration, start, duration, …]}`: one pair of integers per word, in centiseconds from `t0` (which equals the unit's `t0`). The words are the maximal runs of non-whitespace characters of the unit's `text` after removing the speaker labels (`**Name:**`); `cs` therefore holds exactly twice as many integers as there are words. Readers use it to highlight the word being spoken and to turn a `char` range into a time range. ## 8. Reference search The reference search defines what conformance tests: results that every implementation returns identically from the same file. Products MAY rank better (stopwords, query expansion, reranking, filters); they MUST still offer the reference behaviour to pass the suite, and SHOULD label the difference in their documentation. A result item is `{"fragment_id", "score", "via", "anchor", "anchor_uri"}` where `via` lists the contributing methods (`"lexical"`, `"vector"`) in that order, `anchor` is the fragment's anchor and `anchor_uri` the URI formatted from `anchor` and `anchor_end` with the `sha256-` document reference. ### 8.1 Lexical search Given a query string and a limit: 1. **Normalize**: `q` = NFC(query). 2. **Phrases**: scan `q` from left to right. An opening mark `"` (U+0022), `“` (U+201C), `«` (U+00AB) or `„` (U+201E) opens a phrase that the next `"`, `”` (U+201D), `»` (U+00BB) or, for `„`, `“` or `”` respectively closes. The text between the marks is the phrase. An opening mark without a closing mark is treated as a separator. 3. **Words**: maximal runs of characters whose Unicode general category is a letter (L), a mark (M) or a number (N). A phrase term is the phrase's words joined with one space; phrases without words are dropped. 4. **Terms**: if there is at least one phrase term, the terms are the phrase terms (loose words outside quotes are discarded) and the operator is `AND`. Otherwise the terms are the words of `q` and the operator is `OR`. Duplicate terms are removed, keeping the first, comparing them by the key `lower(remove_Mn(NFD(term)))` (Unicode default lowercase, after removing nonspacing marks); the key is used only to detect duplicates. No stopwords are removed. Without terms the result is empty. 5. **MATCH string**: every term, **as written** (no case folding, no decomposition), is an FTS5 string: `"` + the term with each `"` doubled + `"`; the strings are joined with ` AND ` or ` OR `. The tokenizer folds case and diacritics itself; folding the query beforehand would break matches (`Straße`, `fin`). 6. **Query**: ```sql SELECT f.n, f.id, bm25(fragments_fts, 1.0, 0.5, 0.5, 1.0) AS r FROM fragments_fts JOIN fragments f ON f.n = fragments_fts.rowid WHERE fragments_fts MATCH ?1 ORDER BY r, f.n LIMIT ?2 ``` The score is −r. Because `search_text` is the fourth indexed column, a query in modern spelling finds old spelling without any special step. Since `n` is the rowid of the index, implementations MAY rank inside the index alone (`SELECT rowid, bm25(…) FROM fragments_fts WHERE fragments_fts MATCH ?1 ORDER BY 2, 1 LIMIT ?2`) and look up the fragment ids of the returned rows only; the result is identical. Readers that fetch files by HTTP ranges SHOULD do so, as it avoids reading every matching fragment. 7. **CJK route**: if `q` contains a code point in one of the ranges U+2E80–U+2FDF, U+3040–U+30FF, U+3100–U+312F, U+3130–U+318F, U+31A0–U+31FF, U+3400–U+4DBF, U+4E00–U+9FFF, U+A960–U+A97F, U+AC00–U+D7AF, U+F900–U+FAFF, U+FF66–U+FF9F or U+20000–U+3FFFF, step 6 is replaced: - if `fragments_fts_trigram` exists and every term has at least 3 code points, the same MATCH string runs against `fragments_fts_trigram`, ordered by `bm25(fragments_fts_trigram)` then `n`; the score is −bm25; - otherwise (no trigram index, or a term shorter than 3 code points, which a trigram index cannot match) the **substring fallback** runs on `fragments.text`: for each fragment, `hits` = the number of terms `t` with `instr(text, t) > 0`; fragments with `hits` ≥ 1 (with the `OR` operator) or `hits` = number of terms (with `AND`) are returned ordered by `hits` descending, then `n`; the score is `hits`. ### 8.2 Vector search Given a space, a target (`fragment` by default, or `unit`, `figure`), a query vector of `dims` numbers and a limit: compare the query with every vector of that space and target by brute force. Every component is converted to an IEEE 754 binary64 number (f32 and f16 exactly; i8 as q/127). The query is used as given, not normalized. When the space has `normalized = 1` the score is the dot product; otherwise it is the cosine similarity. Results are ordered by score descending, then by fragment `n`, unit `ord` or figure `id`. A result item for the target `unit` or `figure` carries `unit_id` or `figure_id` instead of `fragment_id`, and the anchor URI of the unit's or figure's own anchor. Products MAY use approximate indexes; the reference is exhaustive. ### 8.3 Hybrid search Run the lexical search and the vector search (target `fragment`), each with depth `max(limit, 50)`, and fuse them by reciprocal rank fusion [RRF] with k = 10: score = Σ 1/(10 + rank) over the lists that contain the fragment, rank starting at 1. Order by score descending, then `n`; keep `limit` results. The constant 10 was measured in Scholaris: the classic 60 flattens short, good lists. ### 8.4 Comparison in conformance Result order MUST match exactly; scores MUST match within an absolute tolerance of 1e-6. The lexical scores are those of SQLite's own `bm25()`, which is the oracle. ## 9. Vector spaces ### 9.1 Spaces A row of `spaces` describes how a set of vectors was produced: - `id`: `@` for `f32` vectors and `@:` otherwise (`embeddinggemma-2@768`, `embeddinggemma-2@256:i8`). A suffix `+` MAY follow to separate vectors of the same model computed from different inputs (legacy Scholaris uses `+contexto`). - `provider` (who ran the model: `local`, `google`, `inferbox`…), `model`, `version`. - `dims`: number of components. - `dtype`: `f32` (IEEE 754 binary32), `f16` (binary16) or `i8` (signed byte; the value is q/127). Other values are invalid (E032). - `normalized`: 1 if every stored vector has unit Euclidean norm (before quantization). - `truncated_from`: for Matryoshka truncation [MRL], the original dimension (`768` for a vector cut to 256); NULL otherwise. Truncated vectors SHOULD be renormalized before storing, with `normalized = 1`. - `modalities`: JSON array of the input modalities the model accepts (`text`, `image`, `audio`, `video`, `pdf`). - `task_prefixes`: JSON object with the prefixes or instructions used at encoding time, `{"document": "…", "query": "…"}`, so that a reader can encode queries the same way; NULL if none. A file MAY hold several spaces; a file without spaces is valid (profile `core`). ### 9.2 Vectors `vectors.data` is the vector as `dims` little-endian values of the space's `dtype`, so its length is `dims` × 4, 2 or 1 bytes (E030). Every vector refers to a space in `spaces` (E031). Writers quantize as follows: f32 → f16 with IEEE round-to-nearest-even; f32 → i8 with `q = clamp(round_half_away_from_zero(v × 127), −127, 127)`. The value −128 is not used. A value that does not fit the dtype (a finite f32 above 65504 that would round to infinity in f16, a non-finite value) is an error for the writer, never silently stored. ### 9.3 Compatibility between spaces and quantizations Two spaces are **compatible**, and one query vector serves both, when `provider`, `model`, `version`, `dims`, `normalized`, `truncated_from` and `task_prefixes` are equal; `dtype` may differ. A reader holding a model MAY therefore search an `f32` space and its `i8` copy with the same query. Spaces that differ in any other field are not comparable: readers MUST NOT mix scores across incompatible spaces, and MUST NOT compare vectors of different dimensions. A Matryoshka space (`truncated_from` = 768, `dims` = 256) is compatible with a query only if the query was truncated to the same dimensions and renormalized. ## 10. Profiles `spdf_meta.profile` declares which promises a file makes. Profiles are cumulative labels; a file lists every profile it satisfies. | profile | requirements | |---|---| | `core` | REQUIRED in every file. All tables of [§3](#schema); at least one unit; every unit, fragment and figure anchored; text in NFC; FTS index in sync. | | `semantic` | At least one space, and vectors for every fragment in at least one space. A `semantic` file without vectors raises W100. | | `media` | `kind` is `audio` or `video`; units carry `time` anchors and `t0`/`t1`; `duration` is set; `words` SHOULD be present. A `media` file without time anchors raises W101. | | `full` | `semantic` and, for audio and video, `media`; for paged kinds, page images (`units.image`) and figures where the original has them. | Readers MUST NOT refuse a file because of its profile; profiles tell readers what to expect and validators what to check. ## 11. Extensions Data that this specification does not define goes into **extension tables** named `x__` (lowercase ASCII letters, digits and `_`; `` is a name the author controls, e.g. `x_scholaris_claims`). Each extension in use is declared in the `extensions` table with its `name` (`_` or the table prefix), a `version`, and `required`: - `required = 0`: readers that do not know the extension ignore it. - `required = 1`: the file cannot be understood without it; a reader that does not know it MUST refuse the file with E060. Extensions MUST NOT change the meaning of the core tables, MUST NOT add columns to them, and SHOULD NOT be required. Extension tables are not part of the canonical dump. An extension that proves useful to several implementations becomes part of the core through the RFC process ([§23](#versioning)). ## 12. Canonical dump The canonical dump is a JSON view of a file that every implementation produces identically. It is the oracle of the conformance suite and the input of integrity hashing. ```jsonc {"spdf_version": "5.0", // legacy files: "4.0"/"4.1" and "legacy": true "meta": {"": "", …}, // every spdf_meta row "fts": {"tokenizer": "unicode61 remove_diacritics 2", "trigram": false}, "document": {"id", "kind", "metadata", "source_sha256", "source_ref", "mime", "bytes", "unit_count", "duration", "created", "updated", "title", "authors", "year", "language", "rights"}, "units": [{"id", "ord", "anchor", "text", "notes", "header", "footer", "image", "thumbnail", "reader", "confidence", "printed", "t0", "t1", "words"}], "sections": [{"id", "parent", "level", "title", "unit_from", "unit_to", "summary"}], "fragments": [{"n", "id", "unit", "ord", "text", "context", "section", "anchor", "anchor_end", "search_text"}], "figures": [{"id", "unit", "image", "caption", "description", "anchor"}], "spaces": [{"id", "provider", "model", "version", "dims", "dtype", "normalized", "truncated_from", "modalities", "task_prefixes", "created"}], "vectors": {"": {"count": 6, "sha256": ""}}, "blobs": [{"key", "mime", "bytes", "sha256"}], "provenance": [{"stage", "provider", "model", "detail", "ms", "at"}], "extensions": [{"name", "version", "required"}]} ``` Rules: 1. Every member listed is present. SQL NULL becomes `null`; INTEGER a JSON integer; REAL a JSON number; TEXT a string. Columns that hold JSON (`metadata`, `rights`, `anchor`, `anchor_end`, `notes`, `words`, `section`, `modalities`, `task_prefixes`, `detail`) are parsed and embedded as JSON values. The `document` column of child tables is omitted. Booleans stored as integers (`normalized`, `required`) stay integers. 2. Every number that is not an integer, including those inside parsed JSON, is rounded to 6 decimal places (round half to even on its exact binary value); a result of −0 becomes 0. 3. Order: `units` by `ord`; `fragments` by `n`; `sections`, `figures` and `spaces` by `id`; `blobs` by `key`; `extensions` by `name` (code point order, which is SQLite's `BINARY` collation over UTF-8); `provenance` by the UTF-8 bytes of the JCS serialization of each entry. `fts.trigram` is true if and only if `fragments_fts_trigram` exists; `fts.tokenizer` is the value of the `tokenize` option of `fragments_fts` as declared, without its quotes and with runs of whitespace collapsed to one space (`unicode61 remove_diacritics 2`; `unicode61`, the FTS5 default, if absent). 4. `vectors` has one member per distinct `vectors.space`: `count` is the number of rows and `sha256` the hex SHA-256 of their `data` blobs concatenated in order of `target`, then `id`. 5. `blobs[].bytes` and `blobs[].sha256` are computed from `data`, not copied from the `sha256` column. 6. Serialization, whenever bytes matter (hashing), is JCS [RFC 8785]: no whitespace, object members sorted by the UTF-16 code units of their names, numbers in the ECMAScript form (`1`, not `1.0`; `0.000001`, not `1e-6`), strings in UTF-8 with only `"`, `\` and U+0000–U+001F escaped. For legacy files the dump is the 5.0 view defined in [§20](#legacy), with the legacy `spdf_version` and `"legacy": true`. ## 13. Integrity and signatures `spdf_meta.content_sha256` (OPTIONAL) is the lowercase hex SHA-256 of the JCS serialization of the canonical dump from which the members `meta.content_sha256`, `meta.signature` and `meta.signer` have been removed. It covers all content except extension tables, and it is independent of SQLite's page layout, so two writers that store the same content produce the same hash. `spdf_meta.signature` (OPTIONAL, requires `content_sha256` and `signer`) is the standard base64 encoding, with padding, of an Ed25519 signature [RFC 8032] over the ASCII bytes of the string `spdf-content-sha256:` followed by the hex `content_sha256`. `spdf_meta.signer` is `ed25519:` followed by the standard base64 encoding of the 32-byte public key. Validators that meet `content_sha256` MUST recompute it (E081 on mismatch) and, if a signature is present, MUST verify it (E082 on failure). A valid signature proves that the holder of the key produced this content; it says nothing about whether the key is trustworthy. Readers SHOULD show who signed (the key, or a name the user associated with it) and MUST NOT present an unknown key as trusted. ## 14. Security considerations An SPDF file is a database written by someone else. Opening it is parsing untrusted input with a complex engine. The threats, and the rules of this specification that answer them: - **Code in the schema.** Triggers, views and virtual tables can run SQL or call modules when the database is used. Files MUST NOT contain them ([§2.1](#container)); readers MUST refuse them, open read-only with `query_only`, `trusted_schema = OFF` and the defensive flag, and never load extensions ([§2.4](#container)). The legacy FTS triggers are tolerated only because they never fire on a read-only connection. - **Malformed databases.** SQLite is robust against corrupt files but recommends extra care for untrusted ones [SQLITE-SECURITY]: disable memory-mapped I/O, enable `cell_size_check`, set length limits, consider `quick_check`. - **Decompression bombs.** Gzip input (legacy) MUST be decompressed with a size limit ([§2.3](#container)). - **Oversized values.** Blobs, texts and JSON values MUST be bounded; JSON parsers SHOULD limit nesting depth (RECOMMENDED 64). - **Query injection.** User text never reaches FTS5 as syntax: every term is a quoted FTS5 string ([§8.1](#search)). SQL is always parameterized. Implementations MAY cap the number of terms (RECOMMENDED 64) to bound query cost. - **Paths.** Blob keys are opaque strings, not file names. A reader that extracts blobs to disk MUST sanitize them (no absolute paths, no `..`, no device names). - **Remote references.** `source_ref`, `image`, `thumbnail`, `URL` and `web` anchors may point to the network. Readers MUST NOT fetch them automatically: fetching discloses that the file was opened and can reach internal services. Fetch only on a user action, and show the address first. - **Active content.** `text` is light Markdown; readers MUST NOT render raw HTML from it, and MUST escape text before inserting it into HTML. Images from blobs are untrusted input to image decoders; SVG MUST NOT be rendered with scripts enabled. - **Forged provenance.** Provenance, confidence and metadata are claims made by the writer. Only a signature from a trusted key ([§13](#integrity)) attributes them. - **Model inputs.** Text read from a file may contain instructions aimed at language models ("ignore the previous instructions…"). Applications that pass SPDF text to a model MUST treat it as data, not as instructions. ## 15. Privacy considerations - **Vectors can leak text.** Embeddings can be inverted: published attacks reconstruct most of a short input from its vector [VEC2TEXT]. Distributing the vectors of a text is close to distributing the text. Rights and confidentiality rules that apply to the text apply to its vectors ([§16](#rights)); a writer asked to strip the text of a restricted document MUST strip its vectors too. - **Provenance can leak the producer.** Writers SHOULD NOT record local file paths, user names, machine names, account identifiers, API keys or prompts containing personal data in `provenance.detail` or `generator`. - **People in documents.** Interviews and recordings name speakers and may contain personal data. Producers SHOULD let users remove or pseudonymize `speaker` names, and readers SHOULD NOT index speaker names into shared services without consent. - **Annotations are personal.** User annotations live outside the file, in `.spdfa.json` sidecars ([§17](#annotations)), so that sharing a document never shares its reader's notes. - **Opening is observable** only if a reader fetches remote references; see [§14](#security). ## 16. Rights `documents.rights` is NULL or a JSON object: ```json {"license": "CC-BY-4.0", "access": "open", "holder": "Universidad de La Laguna", "note": "Text and images under CC BY 4.0; page scans courtesy of the library."} ``` - `license`: an SPDX license identifier or expression [SPDX] (`CC-BY-4.0`, `CC0-1.0`), or the URL of a license or rights statement (for the public domain, `https://creativecommons.org/publicdomain/mark/1.0/`; for rights statements, `http://rightsstatements.org/vocab/InC/1.0/`). - `access`: `open` (anyone may receive the file), `restricted` (only the audience the holder allows: a class, a library), or `private` (personal copy). - `holder`: the rights holder, or null. `note`: free text. SPDF does not grant rights. A file made from a copyrighted work is a copy of that work, including its vectors ([§15](#privacy)). Producers SHOULD fill `rights` when they know them, SHOULD default `access` to `private` when they do not, and readers SHOULD show `rights` before sharing a file. Public-domain status depends on jurisdiction; `note` is the place to say which. ## 17. Annotations and collections ### 17.1 Annotations: `.spdfa.json` User annotations (highlights, notes, tags) are stored outside the document, in a file with the extension `.spdfa.json`, as a W3C Web Annotation [WEB-ANNOTATION] `AnnotationCollection` in JSON-LD: ```json {"@context": "http://www.w3.org/ns/anno.jsonld", "type": "AnnotationCollection", "spdf_annotations": "1.0", "label": "Notas de lectura", "first": {"type": "AnnotationPage", "items": [ {"id": "urn:uuid:7b0c…", "type": "Annotation", "motivation": "commenting", "created": "2026-10-07T09:00:00Z", "body": {"type": "TextualBody", "value": "Origen del tópico.", "format": "text/plain", "language": "es"}, "target": {"source": "spdf:sha256-3f2a…c9", "selector": [ {"type": "SpdfAnchorSelector", "value": "spdf:sha256-3f2a…c9#p=29&f=21&char=118,301"}, {"type": "TextQuoteSelector", "exact": "En un lugar de la Mancha", "prefix": "", "suffix": ", de cuyo nombre"}]}}]}} ``` `target.source` is the anchor URI without fragment. The `SpdfAnchorSelector` carries the full anchor URI; the `TextQuoteSelector` [WEB-ANNOTATION] lets the annotation survive a re-reading that shifts offsets. Readers that cannot resolve the anchor SHOULD fall back to the quote. The member `spdf_annotations` gives the version of this profile. Selectors of other types (a `FragmentSelector` conforming to Media Fragments for time and region) MAY be added for tools that do not know SPDF. ### 17.2 Collections: `.spdfl.json` A library is a manifest, not a container: ```json {"spdf_library": "1.0", "name": "Tesis: fuentes", "created": "2026-10-07T00:00:00Z", "items": [{"sha256": "3f2a…c9", "title": "El ingenioso hidalgo…", "authors": "Cervantes Saavedra", "year": 1605, "url": "https://example.org/quijote.spdf", "file_sha256": "…"}]} ``` `sha256` is the document's `source_sha256` (the identity used by anchor URIs); `url` and `file_sha256` (SHA-256 of the `.spdf` file bytes) are OPTIONAL and let a reader fetch and check a copy. Items are ordered as the user ordered them. JSON Schemas for both sidecars are in [`json-schema/`](json-schema/). ## 18. Short citation `cite(anchor, anchor_end, metadata, locale)` produces an author-date citation in parentheses, the form most styles share, so that every implementation prints the same locator. Full bibliographies and other styles are produced from the CSL-JSON item with a CSL processor ([§19](#exports)). ### 18.1 Citing an anchor ``` ( names ", " year [ ", " locator ] ")" ``` The locales `es` and `en` are defined; a locale is matched by its primary language subtag (`es-ES` → `es`), and any other locale falls back to `en`. **Names**, from CSL `author`. The name of a person is `literal` if present; otherwise the `non-dropping-particle`, a space and `family`; otherwise `given`. With one author, that name; with two, `A y B` (es) or `A and B` (en), where Spanish writes `e` instead of `y` when the second name begins with the sound /i/ (`i`, `í`, `hi` or `hí` not followed by a vowel: `Gómez e Iglesias`, `Gómez e Hidalgo`, but `Gómez y Hierro`); with three or more, `A et al.` in both locales. Without authors, the `title-short`, or else the `title` up to its first colon, trimmed. **Year**: the first year of `issued`; negative years are written as `375 a. C.` (es) or `375 BC` (en). Without a year, `s. f.` (es) or `n.d.` (en). The `spdf.undated` range is not printed in a short citation. **Locator**: | anchor | es | en | |---|---|---| | page, folio read | `p. 145` | `p. 145` | | page, roman folio | `p. xiv` | `p. xiv` | | page, inferred folio | `p. [21]` | `p. [21]` | | page without folio | `s. p.` | `n. pag.` | | page range (end with another folio) | `pp. 145-146`, `pp. 20-[21]` | same | | leaf / leaf range | `fol. 1r`, `fol. [2v]`, `fols. 1r-[1v]` | same | | column / column range | `col. 45`, `cols. 45-46` | same | | time (t0, floor seconds) | `1:09:20`, `0:42` | same | | time range (end anchor's t1) | `0:12-0:24` | same | | section or web with `printed` | as a page | as a page | | section or web | `§ 3.2 El panóptico, párr. 4` | `§ 3.2 El panóptico, para. 4` | | slide | `diap. 3` | `slide 3` | | sheet | `Datos, filas 4-9`, `Datos, fila 4` | `Datos, rows 4-9`, `Datos, row 4` | | verse | `v. 1234`, `vv. 1234-1240` | same | | canonical | `514a` | same | | image | (no locator) | (no locator) | Times are written `h:mm:ss` from one hour on and `m:ss` below (hours are not wrapped: ground elapsed time `109:24:48`). A range is printed only when both ends have a printed folio and the folios differ; brackets mark each inferred end separately. An end without a printed folio never takes part in a range: the citation prints the folio of the other end alone (`p. 211`, never `pp. s. p.-211`), and `s. p.` / `n. pag.` only when neither end has one. Labels always come from the foliation of the end that is printed: an unnumbered page followed by leaf Ir gives `fol. Ir`; a range whose two printed ends have different foliations labels each end (`p. xiv-fol. 1r`); `section` and `web` anchors count as pages. The locator is omitted when it would be empty, giving `(Hooke, 1665)`. ### 18.2 Citing a passage A citation MUST locate the passage it quotes, not the fragment that happens to contain it. `cite_passage(fragment, quote, locale)` cites a quotation taken from a fragment: 1. Split the fragment into its parts: the text of the start unit between the two values of `anchor.chars`, and, for a crossing fragment, the text of the end unit ([§4.4](#anchors)) between the two values of `anchor_end.chars` (the whole text of a unit when `chars` is absent). 2. If the quotation lies in the start part, cite the start unit's anchor, with `chars` giving the position of the quotation in that unit. Otherwise, if it lies in the end part, cite the end unit's anchor alone, with its `chars`. Otherwise, if it spans both parts, cite the range from the start unit's anchor to the end unit's anchor, without `chars`, under the range rule above (an end without a folio does not count). 3. The result is the short citation of §18 and the anchor URI of the cited anchor or range ([§5](#anchor-uri)). Readers and citation tools MUST NOT cite a passage with the start `anchor` of its fragment when the passage is not in the start unit: a quotation from the second page of a fragment that begins on an unnumbered plate cites the folio of the second page. ## 19. Exports Implementations MUST export CSL-JSON and BibTeX as defined in §19.1 to §19.3, and MAY export the other formats of §19.4. Exports never invent data: fields absent from the file are absent from the export. An export takes one or several documents, in order. ### 19.1 Keys Every exported document gets a key, used as the CSL `id` and as the BibTeX key: 1. Take the first name of the CSL `author` list: its `family`, else its `literal`, else its `given`. Fold it: decompose with NFKD, keep only the ASCII letters `A`–`Z` and `a`–`z`, lowercase. (`Cervantes Saavedra` → `cervantessaavedra`.) 2. If that is empty (no author, or no ASCII letter in the name), fold in the same way the first whitespace-separated word of `title-short`, or of `title` when there is no `title-short`. (`Lazarillo de Tormes` → `lazarillo`.) 3. If that is still empty, use `anon`. 4. Append the first year of `issued` in decimal (negative years keep their sign), or `nd` when there is none: `cervantessaavedra1605`, `anonnd`. 5. When the same key occurs more than once in one export, every occurrence gets a suffix in export order: `a`, `b`, … `z`, `aa`, `ab`… ### 19.2 CSL-JSON The CSL-JSON export is a JSON array with one item per document: the `metadata` item without its `spdf` member, with `id` set to the key. It is compared as JSON. A citation of a passage adds to the item the CSL `label` and `locator` of an anchor and optional end anchor, so that a CSL processor can print it in any style: | anchor | `label` | `locator` | |---|---|---| | `page`, foliation `page` / `leaf` / `column` | `page` / `folio` / `column` | the folio as in [§18](#citation): `145`, `[21]`, `xiv`, `1r`, ranges `145-146`, `1r-[1v]`; no label or locator when `printed` is null | | `section` or `web` with `printed` | `page` | as for pages | | `section` or `web` with `paragraph` | `paragraph` | the paragraph number | | other `section` or `web` with a path | `section` | the last element of the path | | `time` | `timestamp` | `1:09:20`, ranges `0:12-0:24` (as in §18) | | `verse` | `verse` | `1234` or `1234-1240` | | `canonical` | `section` | the `ref` | | `sheet` | `line` | `4` or `4-9` | | `slide`, `image` | none | none (CSL has no slide locator; the short citation of §18 prints it) | ### 19.3 BibTeX The BibTeX export is text with one entry per document, in export order, separated by one empty line: ```bibtex @book{cervantessaavedra1605, author = {Cervantes Saavedra, Miguel de}, title = {{El} ingenioso hidalgo don {Quijote} de la {Mancha}}, year = {1605}, publisher = {Juan de la Cuesta}, address = {Madrid}, language = {es} } ``` - **Entry type** from the CSL `type`: `book` → `book`; `article-journal`, `article-magazine`, `article-newspaper` → `article`; `chapter` → `incollection`; `paper-conference` → `inproceedings`; `thesis` → `phdthesis`; `report` → `techreport`; anything else → `misc`. - **Fields**, in this order, each only when its source is present and not empty: `author` (CSL `author`), `editor` (`editor`), `title`, `year` (first year of `issued`), `journal` for `article` entries or else `booktitle` (`container-title`), `publisher`, `address` (`publisher-place`), `series` (`collection-title`), `volume`, `number` (`issue`), `pages` (`page`), `edition`, `doi` (`DOI`), `isbn` (`ISBN`), `url` (`URL`), `language`, `note`. - **Values** are written `{…}` in UTF-8. In every value, `\` becomes `\textbackslash{}`, `{` becomes `\{` and `}` becomes `\}`; nothing else is escaped. - **Names**: a `literal` name is written in braces, `{National Aeronautics and Space Administration}`; otherwise the family name (preceded by the `non-dropping-particle` and a space, if any) and the `given` name are written `Family, Given`, or in braces when only one of them exists. Names are joined with ` and `. - **Capitals**: in `title` and in `journal`/`booktitle`, every whitespace-separated word that contains an uppercase letter (Unicode general category Lu) is wrapped in braces, after escaping, so that styles cannot lowercase it: `{El} ingenioso hidalgo don {Quijote}`. - **Comparison**: two exports are equal when, after removing leading and trailing whitespace from every line and dropping empty lines, their lines are identical. ### 19.4 Other formats - **ALTO** [ALTO] (MAY): ALTO 4, one `Page` per page unit, with `PHYSICAL_IMG_NR` = `physical` and `PRINTED_IMG_NR` = `printed` only when `printed` is not null and its `source` is not `inferred` (ALTO records printed numbers, and an inferred folio is not printed); one `TextBlock` per paragraph and one `TextLine` per line; coordinates only when the producer has them (from an extension), never invented. - **TEI** [TEI] (MAY, minimal): `teiHeader` from the metadata (`titleStmt`, `publicationStmt` with the rights, `sourceDesc` with the CSL fields), and a `body` with a `` before each page unit, whose `n` is the folio as cited in §18 without its label (`n="ii"`, `n="[iv]"`, `n="1r"`; no `n` for unnumbered pages) and whose `facs` is the unit image, if any; `

` for paragraphs, ``/`` for verse, `` for speaker turns, and `` for notes. - **IIIF Presentation 3** [IIIF] (MAY): a `Manifest` with one `Canvas` per unit, in `ord` order; the `label` of a page canvas is `{"none": [n]}` with `n` as the TEI `n`, and page canvases of unnumbered pages have no `label`; the unit image is the painting annotation and the text a `supplementing` annotation; audio and video are one time-based canvas with `duration` and a `Range` per unit or section; sections become `structures`; anchors with a region become `#xywh=percent:` targets. - **Web Annotation** (MAY): citations and search results as annotations with the selectors of [§17.1](#annotations). The conformance suite checks, for paged documents, the page sequence of these exports: the `PHYSICAL_IMG_NR`/`PRINTED_IMG_NR` pairs of ALTO, the `n` of each TEI `pb` and the `label` of each IIIF page canvas, in order. ## 20. Legacy formats ### 20.1 SPDF 4.0 and 4.1 Scholaris 4.x files MUST be readable by every reader. They are SQLite databases, usually wrapped in gzip, with Spanish identifiers. Detection, after decompressing: a table `spdf` (`clave`, `valor`) whose row `spdf_version` starts with `4.`, or `user_version` 400 or 410 together with a table `documentos`. `application_id` is 0. The schema is reproduced verbatim in [`schema/spdf-4.1.sql`](schema/spdf-4.1.sql) and [`schema/spdf-4.0.sql`](schema/spdf-4.0.sql) (4.0 lacks `unidades.palabras` and `fragmentos.texto_busqueda`). Legacy files contain the triggers `fragmentos_ai`, `fragmentos_ad` and `fragmentos_au`, tolerated by [§2.4](#container). Readers present legacy files through the **5.0 view**: - **Tables**: `spdf` → `spdf_meta` (`clave` → `key`, `valor` → `value`), `documentos` → `documents`, `unidades` → `units`, `secciones` → `sections`, `fragmentos` → `fragments`, `figuras` → `figures`, `espacios` → `spaces`, `vectores` → `vectors`, `blobs` → `blobs`, `procedencia` → `provenance`; no extensions. - **Columns**: `tipo` → `kind`, `metadatos` → `metadata`, `huella` → `source_sha256`, `original` → `source_ref`, `unidades` → `unit_count`, `duracion` → `duration`, `creado` → `created`, `actualizado` → `updated`, `titulo` → `title`, `autores` → `authors`, `anio` → `year`, `idioma` → `language`; `orden` → `ord`, `ancla` → `anchor`, `texto` → `text`, `notas` → `notes`, `cabecera` → `header`, `pie` → `footer`, `imagen` → `image`, `miniatura` → `thumbnail`, `lector` → `reader`, `confianza` → `confidence`, `impresa` → `printed`, `palabras` → `words`; `padre` → `parent`, `nivel` → `level`, `unidad_desde` → `unit_from`, `unidad_hasta` → `unit_to`, `resumen` → `summary`; `unidad` → `unit`, `contexto` → `context`, `seccion` → `section`, `ancla_fin` → `anchor_end`, `texto_busqueda` → `search_text`; in figures `pie` → `caption`, `descripcion` → `description`; `proveedor` → `provider`, `modelo` → `model`, `normalizado` → `normalized`, `modalidades` → `modalities` (`texto` → `text`, `imagen` → `image`); `objetivo` → `target` (`fragmento` → `fragment`, `unidad` → `unit`, `figura` → `figure`), `espacio` → `space`, `valores` → `data`; `clave` → `key`, `datos` → `data`; `fase` → `stage`, `detalle` → `detail`, `cuando` → `at`. - **Kinds**: `pdf`, `pdf_escaneado` → `scanned_pdf`, `fotos` → `photos`, `imagen` → `image`, `audio`, `video`, `documento` → `document`, `epub`, `presentacion` → `slides`, `hoja` → `sheet`, `web`. - **Anchors**: `tipo` → `type` (`pagina` → `page`, `tiempo` → `time`, `seccion` → `section`, `diapositiva` → `slide`, `hoja` → `sheet`, `web`, `imagen` → `image`), `fisica` → `physical`, `impresa` → `printed`, `romana` → `roman`, `origen` → `source` (`leido` → `read`, `deducido` → `inferred`, `epub`, `ninguno` → `none`), `confianza` → `confidence`, `hablante` → `speaker`, `ruta` → `path`, `parrafo` → `paragraph`, `n`, `hoja` → `sheet`, `filaDesde` → `row_from`, `filaHasta` → `row_to`, `consultada` → `accessed`, `region`. Unknown members are kept as they are. - **Metadata** (`MetadatosDocumento` → CSL-JSON): `titulo` → `title`, or `"titulo: subtitulo"` with `title-short` = `titulo` and `spdf.subtitle` = `subtitulo`; `tituloOriginal` → `original-title`; `autores`, `editores`, `traductores`, `entrevistadores` (`{nombre, apellidos, orcid}`) → `author`, `editor`, `translator`, `interviewer` (`{family: apellidos, given: nombre}`, empty parts omitted; ORCID to `spdf.orcid` under `"apellidos, nombre"`); `fecha` → `issued` with its full date parts when its year equals `anio` or there is no `anio`, otherwise `anio` → `issued`; `anioOriginal` → `original-date`; `editorial` → `publisher`; `lugar` → `publisher-place`; `revista`, else `contenedor` → `container-title`; `coleccion` → `collection-title`; `volumen` → `volume`; `numero` → `issue`; `paginas` → `page`; `edicion` → `edition`; `doi` → `DOI`; `isbn` → `ISBN`; `url` → `URL`; `idioma` → `language`; `resumen` → `abstract`; `idiomaOriginal` → `spdf.original_language`; `sinFecha` `{desde, hasta, fundamento}` → `spdf.undated` `{from, to, basis}`; `procedencia` → `spdf.provenance`, with field names mapped as above and `fuente` → `source` (`lectura` → `reading`, `usuario` → `user`, `colofon` → `colophon`, `impresores` → `printers`, others unchanged), `confianza` → `confidence`. `tipoCSL` → `type`; without it, the type is `article-journal` when `revista` is present, otherwise by kind: `audio` and `presentacion` → `speech`, `video` → `motion_picture`, `web` → `webpage`, `hoja` → `dataset`, `imagen` and `fotos` → `graphic`, anything else → `book`. Empty strings, nulls and empty arrays are omitted. - **Other rules**: `spdf_meta` keys `creado` → `created` and `generador` → `generator`, others unchanged; `documentos.estado` and `documentos.bibliotecas` are dropped; `rights` is null; spaces get `dtype` `f32`, `truncated_from` and `task_prefixes` null and `created` from `creado`; `provenance.model` is null; blob hashes are computed; units are renumbered `ord` = 1, 2, 3… in order of (`orden`, `id`), because 4.x numbers units from 0; `fragments.ord` keeps `orden`. References in `original`, `imagen` and `miniatura`: an empty string becomes null (in figures it stays an empty string), a value equal to a key of `blobs` becomes `blob:`, any other value is kept as an opaque reference. The 4.x anchor has no `foliation`; leaf folios did not exist in 4.x. ### 20.2 SPDF 3.0 and earlier Scholaris v1 to v3 wrote gzip-wrapped SQLite databases with the tables `metadata` (key, value, including `schema_version`) and `chunks`, among others. Readers MAY import them; importing is a conversion with losses (cross-modal links, scenes and some embeddings have no place in 5.0) and the importer SHOULD report what it dropped. Version 3.0 is documented historically in the Scholaris repository; this specification does not define it. ## 21. Conformance ### 21.1 Product classes - A **conforming reader** opens files safely ([§2.4](#container)), reads 5.0 and legacy 4.x files, produces the canonical dump ([§12](#dump)), parses and formats anchor URIs ([§5](#anchor-uri)), produces short citations ([§18](#citation)), runs the reference lexical search ([§8.1](#search)) and exports CSL-JSON and BibTeX ([§19](#exports)). A **semantic reader** also runs the reference vector and hybrid search. - A **conforming writer** produces files that validate without errors or warnings for the profiles they declare and whose canonical dump equals the dump the writer was given (round trip). - A **conforming validator** reports exactly the codes of [§22](#validation) for the validation cases of the suite. ### 21.2 Levels An implementation states its class and the profiles it covers, for example "reader and writer, profiles core and semantic". Its claim is backed by the conformance suite: it passes every case of the kinds its class requires (`dump`, `legacy_dump`, `anchor_uri`, `cite`, `cite_passage`, `search_lexical`, `validate`, `locate`, `export_csl`, `export_bibtex` for readers; plus `search_vector` and `search_hybrid` for semantic readers; plus `roundtrip` and `quantize` for writers; `export_structure` for implementations that export ALTO, TEI or IIIF), with the suite version it was tested against. Partial implementations MAY exist but MUST NOT call themselves conforming. ### 21.3 The suite The suite (`conformance/` in the repository) is normative for behaviour. Its protocol (case format, runner report, CI convention) is in `conformance/README.md`. Each suite release has a version and a manifest with the number of cases and their hash. ## 22. Validation ### 22.1 Procedure A validator checks a file in this order; a step marked *stop* ends validation: 1. If the file starts with `1F 8B`, decompress it ([§2.3](#container)). 2. If the result is not a SQLite database: E001, *stop*. 3. Determine the version: `application_id` 1397769286 with `user_version` 500–599 is 5.x; the legacy detection of [§20.1](#legacy) is 4.x; anything else: E002, *stop*. For 4.x, report W110 and check only: the tables `spdf`, `documentos`, `unidades`, `fragmentos`, `fragmentos_fts` exist (E010 each) and there is no trigger or view other than the three tolerated triggers (E020); *stop*. 4. A gzip-wrapped 5.x file: E003 in the warnings. A minor version above 0: W105; for such a file, unknown anchor types (E041) and unknown dtypes (E032) are reported in the warnings instead of the errors, because a later minor version may define them. 5. Triggers, views and foreign virtual tables: E020 for each. 6. Required tables (E010 each) and required columns (E011 each). 7. `spdf_meta` keys (E012 each). 8. `documents` holds exactly one row (E013); `metadata` and `rights` are valid JSON (E050); `metadata` has string `type` and `title` (E051). 9. Required extensions unknown to the validator (E060). 10. `units.ord` is 1…N (E090); `unit_count` equals N (W102). 11. Anchors of units, fragments (start and end) and figures (E040, E041, E042); fragments that cross `matter` or a folio boundary (W103, [§4.4](#anchors)). 12. Spaces: `dtype` (E032). Vectors: known space (E031), length (E030). 13. FTS index in sync: run `INSERT INTO fragments_fts(fragments_fts, rank) VALUES('integrity-check', 1)` (and the same on `fragments_fts_trigram`) on a private copy; an error is E070. 14. Blobs: stored `sha256` equals the computed one (E080). 15. If there were no errors so far and `content_sha256` is present: recompute it (E081); if it matches and `signature` is present, verify it (E082). 16. Profile warnings: W100, W101. The result is a JSON object (schema in [`json-schema/validation-result.schema.json`](json-schema/validation-result.schema.json)): ```json {"valid": false, "version": "5.0", "profile": ["core"], "errors": [{"code": "E090", "message": "units.ord is not 1..N", "where": "units"}], "warnings": []} ``` `valid` is true if and only if `errors` is empty. `version` is null when unknown. Messages are free text; conformance compares the sets of codes. ### 22.2 Codes | code | meaning | |---|---| | E001 | not a SQLite database (or bad gzip) | | E002 | unknown `application_id` or version | | E003 | gzip-wrapped 5.0 file (reported as a warning) | | E010 | missing required table | | E011 | missing required column | | E012 | missing required `spdf_meta` key | | E013 | `documents` does not hold exactly one row | | E020 | trigger, view or foreign virtual table present | | E030 | vector length ≠ dims × dtype size | | E031 | vector refers to an unknown space | | E032 | unknown dtype | | E040 | invalid anchor (bad JSON, missing or mistyped required member) | | E041 | unknown anchor type | | E042 | `chars` out of range | | E050 | invalid metadata or rights JSON | | E051 | metadata without string `type` and `title` | | E060 | unknown required extension | | E070 | FTS index out of sync | | E080 | blob sha256 mismatch | | E081 | `content_sha256` mismatch | | E082 | signature does not verify | | E090 | `units.ord` not contiguous from 1 | | W100 | profile `semantic` without vectors | | W101 | profile `media` without time anchors | | W102 | `unit_count` ≠ number of units | | W103 | fragment crosses between units of different `matter`, or between a page with a printed folio and one without | | W105 | newer minor version than the validator's | | W110 | legacy 4.x file | Codes are never reused with another meaning. New codes are added by minor versions. ## 23. Versioning and compatibility The specification is versioned MAJOR.MINOR; editorial corrections do not change the version. `user_version` encodes it ([§2.1](#container)). - A **minor** version (5.1, 5.2…) only adds OPTIONAL things: tables, columns, `spdf_meta` keys, anchor members or types, metadata members, validation warnings or errors for things that were already forbidden. A 5.0 reader reads every 5.x file, ignoring what it does not know; it MAY warn (W105). Validators report the anchor types and dtypes of a newer minor version as warnings, not errors ([§22.1](#validation)). A 5.x writer that uses nothing new SHOULD write `user_version` 500. - A **major** version (6.0) may change or remove things. Readers MUST refuse majors they do not know (E002) and SHOULD keep reading older majors (as 5.0 reads 4.x). - **Deprecation**: a feature is deprecated in a minor version, with the reason and the replacement, and removed no earlier than the next major and at least 24 months later. - **Promise**: a file that conforms to 5.0 will be readable by every conforming reader of any later 5.x version, and its anchor URIs will keep resolving. - The conformance suite and each library have their own version numbers; the suite's manifest says which specification version it tests. Changes are proposed and decided through the RFC process in `spec/rfcs/` and `governance/`. ## 24. Media type and file identification - Media type: `application/vnd.spdf+sqlite3` (registration with IANA in preparation; template in `governance/drafts/iana-media-type.md`). The structured syntax suffix `+sqlite3` tells generic tools that the file is a SQLite 3 database. OPTIONAL parameter `version` (`"5.0"`). Encoding: binary. Legacy 4.x files are gzip data and have no registered type of their own. - Fragment identifiers: for a resource of type `application/vnd.spdf+sqlite3`, the fragment identifier is the `params` rule of [§5.1](#anchor-uri), with the meaning it has in an anchor URI for the document in that resource: `https://example.org/quijote.spdf#p=5&f=1r`. - Extension: `.spdf`. Sidecars: `.spdfa.json` and `.spdfl.json`, served as `application/json` (or `application/ld+json` for annotations). - Magic numbers: bytes 0–15 are `53 51 4C 69 74 65 20 66 6F 72 6D 61 74 20 33 00` ("SQLite format 3" and a NUL); bytes 68–71 are `53 50 44 46` ("SPDF"); bytes 60–63 hold `user_version` big-endian (`00 00 01 F4` for 5.0). Legacy 4.x files start with `1F 8B` and cannot be told from other gzip files without decompressing. - Uniform Type Identifier (Apple platforms): `com.joseluissaorin.spdf`, conforming to `public.data` and `public.database`, until a vendor-neutral identifier is agreed. ## 25. Internationalization ### 25.1 Languages and scripts Language tags are BCP 47 [BCP 47]: `es`, `en-GB`, `la`, `grc` (Ancient Greek), `lzh` (Literary Chinese), `ar`. `documents.language` is the main language; fragments in other languages need no tagging in this version. Text is stored in logical order, whatever its direction. ### 25.2 Right-to-left text Arabic, Hebrew, Syriac and other right-to-left scripts are stored in logical order, without bidirectional control characters except those present in the source. Readers display them with the Unicode Bidirectional Algorithm [UAX #9] and SHOULD isolate user-supplied strings (`dir="auto"`). Anchor URIs percent-encode such text, so they are direction-neutral; when an IRI form is displayed, readers SHOULD isolate it. `chars` offsets count code points in logical order. ### 25.3 Old texts The `text` of a unit or fragment is the text of the source, never modernized: long s (`ſ`), `u`/`v` and `i`/`j` alternations, abbreviations and tildes stay as printed. The `search_text` layer carries a modernized form used only by search. The `unicode61` tokenizer with `remove_diacritics 2` already folds case, Latin diacritics and `ſ`; it does not fold Greek accents and breathings, ligatures such as `æ` and `œ`, or `ß`, so producers SHOULD put the folded forms they need in `search_text` (for polytonic Greek, the text without diacritics). Citations quote `text`, never `search_text`. ### 25.4 Chinese, Japanese and Korean The `unicode61` tokenizer treats a run of Han characters as one token. Files whose text is mostly CJK SHOULD include `fragments_fts_trigram`; the reference search then matches substrings of three or more characters by trigram and shorter ones by substring ([§8.1](#search)). Producers MAY add a segmented form (words separated by spaces) to `search_text`. ### 25.5 Numbers and folios Printed folios are stored as printed, in any script (`"xiv"`, `"٣٤"`, `"三"`). Readers MUST NOT convert them for citation; they MAY offer conversions for navigation. ## References ### Normative - [BCP 47] Phillips, A., Davis, M., "Tags for Identifying Languages", BCP 47, RFC 5646. - [COMMONMARK] CommonMark Spec, version 0.31.2, . - [CSL-JSON] Citation Style Language, CSL-JSON schema, . - [MEDIA-FRAGMENTS] W3C, "Media Fragments URI 1.0 (basic)", Recommendation, 2012. - [RFC 1952] Deutsch, P., "GZIP file format specification version 4.3". - [RFC 2119] Bradner, S., "Key words for use in RFCs to Indicate Requirement Levels". - [RFC 3986] Berners-Lee, T., et al., "Uniform Resource Identifier (URI): Generic Syntax". - [RFC 3987] Duerst, M., Suignard, M., "Internationalized Resource Identifiers (IRIs)". - [RFC 5147] Wilde, E., Duerst, M., "URI Fragment Identifiers for the text/plain Media Type". - [RFC 5234] Crocker, D., Overell, P., "Augmented BNF for Syntax Specifications: ABNF". - [RFC 8032] Josefsson, S., Liusvaara, I., "Edwards-Curve Digital Signature Algorithm (EdDSA)". - [RFC 8174] Leiba, B., "Ambiguity of Uppercase vs Lowercase in RFC 2119 Key Words". - [RFC 8259] Bray, T., "The JavaScript Object Notation (JSON) Data Interchange Format". - [RFC 8785] Rundgren, A., et al., "JSON Canonicalization Scheme (JCS)". - [SQLITE-FORMAT] SQLite, "Database File Format", . - [SQLITE-FTS5] SQLite, "SQLite FTS5 Extension", . - [UAX #15] Unicode Standard Annex #15, "Unicode Normalization Forms". - [WEB-ANNOTATION] W3C, "Web Annotation Data Model", Recommendation, 2017. ### Informative - [ALTO] Library of Congress, "ALTO: Technical Metadata for Layout and Text Objects", version 4. - [CTS] "Canonical Text Services" protocol and URN scheme, . - [IIIF] IIIF Consortium, "IIIF Presentation API 3.0". - [MRL] Kusupati, A., et al., "Matryoshka Representation Learning", NeurIPS 2022. - [RFC 6838] Freed, N., Klensin, J., Hansen, T., "Media Type Specifications and Registration Procedures". - [RFC 7595] Thaler, D., et al., "Guidelines and Registration Procedures for URI Schemes". - [RRF] Cormack, G. V., Clarke, C. L. A., Büttcher, S., "Reciprocal Rank Fusion outperforms Condorcet and individual rank learning methods", SIGIR 2009. - [SPDX] SPDX License List, . - [SQLITE-SECURITY] SQLite, "Defense Against The Dark Arts", . - [TEI] TEI Consortium, "TEI P5: Guidelines for Electronic Text Encoding and Interchange". - [UAX #9] Unicode Standard Annex #9, "Unicode Bidirectional Algorithm". - [VEC2TEXT] Morris, J. X., et al., "Text Embeddings Reveal (Almost) As Much As Text", EMNLP 2023. ## Appendix A. Changes from SPDF 4.1 - Uncompressed container with `application_id` and `user_version`; gzip only for legacy. - English identifiers; legacy files read through the 5.0 view. - Metadata as a CSL-JSON item with the `spdf` extension object. - New anchor types `verse` and `canonical`; `foliation` (leaves and columns); `chars` and `region` on any anchor. - Anchor URI with ABNF, aligned with W3C Media Fragments and RFC 5147. - `spaces.dtype` (`f32`, `f16`, `i8`), `truncated_from`, `task_prefixes`; space compatibility. - Profiles, extensions, `rights`, blob hashes, provenance `model`. - No triggers or views in distributed files; safe opening procedure. - Canonical dump, integrity hash and Ed25519 signatures. - Units numbered from 1. - Removed: `documentos.estado` and `documentos.bibliotecas` (library membership belongs to collection manifests). --- # SPDF: documents read once, cited forever URL: https://spdf.joseluissaorin.com/ > SPDF is an open file format for documents that have already been read. Every passage carries its exact anchor (printed page, folio, second, slide, verse), so a citation can only print what the source says. # Read once. Cite forever. [Read the specification](https://spdf.joseluissaorin.com/spec.md) · [Validate a file](https://spdf.joseluissaorin.com/validator.md) · [Open the reader](https://spdf.joseluissaorin.com/reader/) Anchor URI: `spdf:sha256-a156ce5ac9858e29e74bdc90c424fa1b77621bde7204d43faf18c2e726243ab0#p=41&f=17r&char=0,1424` → (Galilei, 1610, fol. 17r) Physical page 41 of the file, printed folio 17r, characters 0 to 1424 of that page, in *Sidereus nuncius*. The citation is computed from the anchor stored at reading time; nothing is guessed. [Open it in the validator](https://spdf.joseluissaorin.com/validator#url=/commons/files/galilei-sidereus-nuncius-1610.spdf). ## Why a format Reading a document well is slow and expensive: OCR, transcription, finding the printed folios, sectioning, embeddings. SPDF stores the result so that nobody has to do it twice, and so that whatever cites from it can be checked. ### I. Anchors Every passage knows where it is. Fragments are stored with the place they came from: physical page and printed folio (roman, inferred or by leaf), second and word timings in audio and video, slide, sheet range, verse line or a canonical reference such as Stephanus 514a. Anchors serialise as portable URIs. `{"type":"page","physical":29,"printed":"21"}` ### II. Provenance Every field says who wrote it. Each unit records which reader produced its text (a PDF text layer, a vision model, a speech recogniser) and with what confidence; the metadata records where each field came from (colophon, title page, catalogue). Inferred folios are cited in brackets. `reader: gemma-4-e4b · confidence: 0.97` ### III. Read once, query many The expensive part happens once. OCR, transcription, sectioning and embeddings are paid for when the file is produced. After that it answers lexical, semantic and hybrid queries offline, even on a phone, with SQLite’s own full-text index and vectors from several models side by side. `fts5 unicode61 · f32 | f16 | i8 · RRF k = 10` ### IV. Portability One file, any language, no server. A .spdf is a plain SQLite 3 database: no custom container, no account. It can be memory-mapped or read over HTTP ranges, and twelve independent implementations open it, all tested against the same conformance suite. `one document = one file` ### V. Honest citation A citation can only print what the source says. Short citations and bibliography (CSL-JSON, BibTeX) are derived from the stored anchor and the CSL record, never generated. Agents get the same guarantee through the MCP server: they search, they quote and they cite with the exact folio, and they cannot invent one. `(Saorín Ferrer, 2026, p. [3])` ## Inside a .spdf A single SQLite file, uncompressed so it can be read by ranges, with no triggers and no views. Readers open it read-only, in defensive mode, and never load extensions. The schema is small enough to learn in an afternoon. | Table | What it holds | | --- | --- | | `spdf_meta` | version, profile, generator, document id | | `documents` | one CSL-JSON record, with the provenance of each field | | `units` | the citable units: pages, time spans, slides, sheets | | `fragments` | passages of 150 to 300 words with their anchors | | `fragments_fts` | FTS5 index, accent-insensitive | | `sections` | the heading tree | | `figures` | figures, plates and frames, with region and description | | `spaces` | vector spaces, declared as model@dims | | `vectors` | little-endian f32, f16 or i8 | | `blobs` | the original and the images, with their SHA-256 | | `provenance` | what produced what, with which model, and when | | `extensions` | x_vendor_name tables, required or optional | ### Profiles - `core`: Text and anchors. Enough to search and cite. - `semantic`: Core plus vectors from one or more embedding models. - `media`: Core plus audio and video with per-word timings. - `full`: All of the above. ## Twelve implementations, one suite The Rust implementation is the reference and also exposes a C ABI. The others are native and independent: each one opens, validates, dumps, searches, formats anchors, cites and writes, and each one is checked by the same conformance cases on every commit. - **Rust** (`spdf`): CI running. https://spdf.joseluissaorin.com/docs/rust.md - **TypeScript** (`spdf-format`): CI running. https://spdf.joseluissaorin.com/docs/js.md - **Python** (`spdf-format`): CI running. https://spdf.joseluissaorin.com/docs/python.md - **Swift** (`SPDF`): CI running. https://spdf.joseluissaorin.com/docs/swift.md - **Kotlin / JVM** (`io.github.joseluissaorin:spdf`): CI running. https://spdf.joseluissaorin.com/docs/kotlin.md - **Go** (`github.com/joseluissaorin/spdf/go`): CI running. https://spdf.joseluissaorin.com/docs/go.md - **C# / .NET** (`Spdf.Format`): CI failing. https://spdf.joseluissaorin.com/docs/dotnet.md - **PHP** (`joseluissaorin/spdf`): CI running. https://spdf.joseluissaorin.com/docs/php.md - **Ruby** (`spdf-format`): CI running. https://spdf.joseluissaorin.com/docs/ruby.md - **R** (`spdf`): CI running. https://spdf.joseluissaorin.com/docs/r.md - **Julia** (`SPDF.jl`): CI running. https://spdf.joseluissaorin.com/docs/julia.md - **C** (`libspdf`): CI running. https://spdf.joseluissaorin.com/docs/c.md ## Where to start ### Validate Drop a .spdf on the validator: it checks the file against the specification and shows what is inside, in your browser, without uploading anything. https://spdf.joseluissaorin.com/validator.md ### Read SPDF Reader opens, searches and cites SPDF files on macOS, Windows, Linux, iOS, Android and the web, with local models and no account. https://spdf.joseluissaorin.com/download.md ### Build spdf build turns a PDF, a scan, an EPUB or a recording into an SPDF with local models, or with your own API key. Or start from SPDF Commons, a small collection of public-domain works. https://spdf.joseluissaorin.com/commons.md ### For agents Every page of this site has a Markdown twin (the same address ending in .md), /llms.txt indexes them and /llms-full.txt carries the whole specification. The spdf-mcp server lets any agent search a folder of SPDF files and cite with the exact folio. https://spdf.joseluissaorin.com/agents.md --- # Implementations URL: https://spdf.joseluissaorin.com/implementations > Twelve native, independent implementations of SPDF 5.0, all checked by the same conformance suite on every commit, and what each one's CI says today. SPDF is not one library with bindings. It is a specification with several **native, independent implementations**, each written in the idiom of its language and each checked against the same conformance cases. The Rust implementation is the reference and also exposes a C ABI for anyone who would rather not deal with SQLite directly. The table is rebuilt from the continuous integration of the repository every time this site is published. Nothing here is typed by hand: if a run is red, it says red. | Implementation | Package | Tier | CI on main | Last run | | --- | --- | --- | --- | --- | | Rust | `spdf` | first tier | CI running | 2026-10-07 | | TypeScript | `spdf-format` | first tier | CI running | 2026-10-07 | | Python | `spdf-format` | first tier | CI running | 2026-10-07 | | Swift | `SPDF` | first tier | CI running | 2026-10-07 | | Kotlin / JVM | `io.github.joseluissaorin:spdf` | first tier | CI running | 2026-10-07 | | Go | `github.com/joseluissaorin/spdf/go` | first tier | CI running | 2026-10-07 | | C# / .NET | `Spdf.Format` | first tier | CI failing | 2026-10-07 | | PHP | `joseluissaorin/spdf` | second tier | CI running | 2026-10-07 | | Ruby | `spdf-format` | second tier | CI running | 2026-10-07 | | R | `spdf` | second tier | CI running | 2026-10-07 | | Julia | `SPDF.jl` | second tier | CI running | 2026-10-07 | | C | `libspdf` | second tier | CI running | 2026-10-07 | | Reference producer (spdf build) | `producer/` | — | CI passing · 16/16 cases | 2026-10-07 | | SPDF Reader | `reader/` | — | CI running | 2026-10-07 | | Conformance suite | `conformance/` | — | CI running | 2026-10-07 | | Integrations | `integrations/` | — | CI running | 2026-10-07 | | This site | `site/` | — | CI running | 2026-10-07 | Checked on 2026-10-07. JSON: https://spdf.joseluissaorin.com/status.json ## Product classes What each implementation claims, by the product classes of the [specification](https://spdf.joseluissaorin.com/spec.md#conformance) (reader, semantic reader, writer, validator), as recorded in the repository README by the specification's editor. The CI column above is what the machines check today; this table is what is declared. | Implementation | Folder | Reader | Semantic reader | Writer | Validator | ALTO / TEI / IIIF | Suite 0.3.0 | Suite 0.4.0 | |---|---|---|---|---|---|---|---|---| | Rust (reference) | `rust` | yes | yes | yes | yes | untested | 229/229 | pending | | TypeScript | `js` | yes | yes | yes | yes | untested | 229/229 | 309/309 | | Python | `python` | yes | yes | yes | yes | untested | 229/229 | pending | | Swift | `swift` | yes | yes | yes | yes | untested | 229/229 | pending | | Kotlin / JVM | `kotlin` | yes | yes | yes | yes | untested | 229/229 | pending | | Go | `go` | yes | yes | yes | yes | untested | 229/229 | pending | | C# | `dotnet` | yes | yes | yes | yes | untested | 229/229 | pending | | PHP | `php` | yes | yes | yes | yes | untested | 229/229 | 309/309 | | Ruby | `ruby` | yes | yes | yes | yes | untested | 229/229 | 309/309 | | R | `r` | yes | yes | yes | yes | untested | 229/229 | 309/309 | | Julia | `julia` | yes | yes | yes | yes | untested | 229/229 | 309/309 | | C (Rust C ABI) | `c` | yes | yes | yes | yes | untested | 229/229 | 309/309 | | Producer `spdf build` | `producer` | n/a | n/a | outputs validate | n/a | n/a | own checks 16/16 | pending | ## What every implementation does Every implementation, in every language, does the same eight things, and the conformance suite checks all of them: 1. **Opens safely**: read-only, `query_only`, `trusted_schema=OFF`, defensive mode where the binding allows it, never loading extensions, rejecting files with triggers or views, and with bounded blob and decompression sizes. 2. **Validates** a file and reports the error and warning codes of the specification ([§ Validation](https://spdf.joseluissaorin.com/spec.md#validation)). 3. **Reads** SPDF 5.0 and the legacy 4.0 and 4.1 files from Scholaris, which are usually gzip-wrapped and use Spanish identifiers. 4. **Dumps** a file to canonical JSON (RFC 8785), the oracle every other implementation is compared against. 5. **Searches**: lexical (FTS5), vector (brute force, f32, f16 or i8) and hybrid (reciprocal rank fusion with k = 10). 6. **Formats and parses anchor URIs**, byte for byte, in both directions. 7. **Cites**: short author-date citations in English and Spanish, and bibliography as CSL-JSON and BibTeX. 8. **Writes**: builds a valid SPDF from scratch, and round-trips a dump. ## How conformance is checked The suite lives in `conformance/`: source dumps, generated files, legacy files, deliberately broken files and one JSON case per check. A case has an `id`, a `kind` (`dump`, `validate`, `search_lexical`, `search_vector`, `search_hybrid`, `anchor_uri`, `cite`, `legacy_dump`, `roundtrip`), an input and the expected result. Each implementation ships a runner that executes every case and prints one line of JSON: ```json {"impl": "rust", "version": "5.0.0", "passed": ["dump-001", "…"], "failed": [], "skipped": []} ``` The CI of each implementation fails on any failed case and uploads that JSON as the artifact `conformance-`, which is what the table above counts. ## Tiers - **First tier**: Rust, TypeScript, Python, Swift, Kotlin/JVM, Go and C#. Released together with every version of the specification. - **Second tier**: PHP, Ruby, R, Julia and C. The same suite, released when they are ready. ## Producers Reading a document is the producer's job, not the library's. There are two independent producers, and their interoperability is a condition for declaring the format stable: - `spdf build`, the reference producer, in Python, with local models (EmbeddingGemma 2, Gemma 4, Whisper) or your own API key. - [Scholaris](https://scholaris.joseluissaorin.com), in TypeScript, where the format was born. --- # Documentation URL: https://spdf.joseluissaorin.com/docs > How to open, validate, search and cite SPDF files in Rust, TypeScript, Python, Swift, Kotlin, Go, C#, PHP, Ruby, R, Julia and C. Every implementation offers the same operations with the names and habits of its language. Pick yours; each page has the install line, a first example and the full README of the library. - [Rust](https://spdf.joseluissaorin.com/docs/rust.md): `cargo add spdf`. Reference implementation; also provides the C ABI. - [TypeScript](https://spdf.joseluissaorin.com/docs/js.md): `npm install spdf-format`. Node (node:sqlite), Bun and the browser (sqlite-wasm). Powers the validator on this site. - [Python](https://spdf.joseluissaorin.com/docs/python.md): `pip install spdf-format`. Standard library sqlite3 only; imported as spdf. - [Swift](https://spdf.joseluissaorin.com/docs/swift.md): `.package(url: "https://github.com/joseluissaorin/spdf", from: "5.0.0")`. Swift Package Manager, from the repository. Apple platforms and Linux. - [Kotlin / JVM](https://spdf.joseluissaorin.com/docs/kotlin.md): `implementation("io.github.joseluissaorin:spdf:5.0.0")`. Kotlin and Java on the JVM and Android. - [Go](https://spdf.joseluissaorin.com/docs/go.md): `go get github.com/joseluissaorin/spdf/go`. Go module in the monorepo. - [C# / .NET](https://spdf.joseluissaorin.com/docs/dotnet.md): `dotnet add package Spdf.Format`. .NET with Microsoft.Data.Sqlite. - [PHP](https://spdf.joseluissaorin.com/docs/php.md): `composer require joseluissaorin/spdf`. PDO SQLite. - [Ruby](https://spdf.joseluissaorin.com/docs/ruby.md): `gem install spdf-format`. On the sqlite3 gem. - [R](https://spdf.joseluissaorin.com/docs/r.md): `remotes::install_github("joseluissaorin/spdf", subdir = "r")`. On RSQLite; fragments as data frames. - [Julia](https://spdf.joseluissaorin.com/docs/julia.md): `pkg> add SPDF`. On SQLite.jl. - [C](https://spdf.joseluissaorin.com/docs/c.md): `#include "spdf.h" /* link with -lspdf */`. C ABI over the Rust core, for C, C++ and any FFI. ## The same operations everywhere | Operation | What it returns | | --- | --- | | open | A read-only handle on a 5.0 or legacy 4.x file, opened safely | | validate | `{valid, version, profile, errors, warnings}` with the codes of the specification | | dump | The canonical JSON of the file (RFC 8785) | | search (lexical, vector, hybrid) | `{fragment_id, score, via, anchor, anchor_uri}` items | | anchor URI | Format and parse `spdf:` URIs, byte for byte | | cite | `(Family, Year, locator)` in English or Spanish | | export | CSL-JSON and BibTeX | | write | A new, valid SPDF built from your own data | ## Without any library An SPDF file is a SQLite database. Any tool that speaks SQLite reads it; only the safety rules and the anchors need care. ```sh sqlite3 -readonly darwin-origin.spdf "SELECT key, value FROM spdf_meta" sqlite3 -readonly darwin-origin.spdf \ "SELECT f.id, u.printed, substr(f.text, 1, 80) FROM fragments_fts JOIN fragments f ON f.n = fragments_fts.rowid JOIN units u ON u.id = f.unit WHERE fragments_fts MATCH 'selection' LIMIT 5" ``` Legacy 4.x files from Scholaris are gzip-wrapped: `gzip -dc old.spdf > old.sqlite` first. --- # How to cite URL: https://spdf.joseluissaorin.com/cite > How to cite the SPDF specification in your own work, and how SPDF cites the documents it holds, with the exact folio, second or verse. ## Citing the specification If SPDF is useful in your research, please cite the specification itself, with the version you used. Suggested citation: > Saorín Ferrer, José Luis. 2026. *SPDF: Semantic Processed Document Format. Specification, version 5.0.* https://spdf.joseluissaorin.com/spec ```bibtex @techreport{spdf-5.0, author = {Saor{\'\i}n Ferrer, Jos{\'e} Luis}, title = {{SPDF}: Semantic Processed Document Format. Specification, version 5.0}, year = {2026}, institution = {spdf.joseluissaorin.com}, url = {https://spdf.joseluissaorin.com/spec}, note = {CC BY 4.0} } ``` ```json [ { "id": "spdf-5.0", "type": "report", "title": "SPDF: Semantic Processed Document Format. Specification, version 5.0", "author": [ { "family": "Saorín Ferrer", "given": "José Luis" } ], "issued": { "date-parts": [ [ 2026 ] ] }, "URL": "https://spdf.joseluissaorin.com/spec", "version": "5.0", "language": "en" } ] ``` ```yaml cff-version: 1.2.0 message: "If you use SPDF, please cite its specification as below." title: "SPDF: Semantic Processed Document Format. Specification" version: "5.0" date-released: 2026-10-07 authors: - family-names: "Saorín Ferrer" given-names: "José Luis" url: "https://spdf.joseluissaorin.com/spec" license: CC-BY-4.0 ``` A DOI will be minted for each released version of the specification; until then, please cite the URL and the version. ## How SPDF cites what it holds The point of the format is that a citation is **computed from the stored anchor, never generated**. Every implementation has the same `cite` function, tested by the conformance suite, which takes an anchor, the document's CSL record and a locale: | Anchor | English | Spanish | | --- | --- | --- | | page, printed folio read from the page | `(Darwin, 1859, p. 21)` | `(Darwin, 1859, p. 21)` | | page, folio inferred from its neighbours | `(Darwin, 1859, p. [21])` | `(Darwin, 1859, p. [21])` | | page, roman folio | `(Woolf, 1929, p. xiv)` | `(Woolf, 1929, p. xiv)` | | leaf foliation | `(Cervantes, 1605, fol. 1r)` | `(Cervantes, 1605, fol. 1r)` | | a page without a printed number | `(Darwin, 1859, n. pag.)` | `(Darwin, 1859, s. p.)` | | range | `(Darwin, 1859, pp. 21-22)` | `(Darwin, 1859, pp. 21-22)` | | time in a recording | `(Cortázar, 1977, 1:09:20)` | `(Cortázar, 1977, 1:09:20)` | | slide | `(Gould, 2024, slide 3)` | `(Gould, 2024, diap. 3)` | | verse | `(Milton, 1667, vv. 234-240)` | `(Milton, 1667, vv. 234-240)` | | canonical reference (CSL year −375) | `(Plato, 375 BC, 514a)` | `(Plato, 375 a. C., 514a)` | Two authors are joined with *and* in English and *y* in Spanish (*e* before the sound /i/, as the Spanish norm asks); three or more become *et al.* A page without a printed folio is never cited by its position in the file disguised as a page number. Full bibliographic references are exported as **CSL-JSON** (always) and **BibTeX**, so any CSL style (Chicago, APA, MLA, ISO 690…) can be applied with citeproc, Zotero or Pandoc. ## Anchor URIs Every passage can be pointed at with a portable URI that survives renaming and copying the file, because it names the document by the SHA-256 of its original bytes: ```text spdf:sha256-3f2a9c…#p=29&f=21&char=118,301 ``` The parameters follow W3C Media Fragments and RFC 5147 where they overlap: `p` physical page, `f` printed folio, `t` seconds, `s` section path, `sl` slide, `v` verse, `ref` canonical reference, `char` character range, `xywh` region in percent. The full grammar is in the [specification](https://spdf.joseluissaorin.com/spec.md#anchor-uri). ## Citing from your writing tools - **Pandoc**: write `[@spdf:sha256-3f2a9c…#p=29]` in Markdown and the [Pandoc filter](https://spdf.joseluissaorin.com/integrations.md#pandoc) turns it into a real citation with the printed folio, in any CSL style. - **Zotero**: the [Zotero plugin](https://spdf.joseluissaorin.com/integrations.md#zotero) imports the CSL record of an SPDF as an item, attaches the file and copies a citation with the folio. - **Agents**: the [MCP server](https://spdf.joseluissaorin.com/integrations.md#mcp) gives any agent a `cite` tool that returns the citation, the anchor URI and the quoted text together, so it cannot cite a page that does not say what it claims. --- # SPDF for agents URL: https://spdf.joseluissaorin.com/agents > How language models and agents read this site and use SPDF files: Markdown twins, llms.txt, the MCP server, and the rules for citing without inventing. This site is written to be read by people and by machines alike. Everything a person can read here, an agent can fetch as plain text. ## Reading this site - Every page has a **Markdown twin**: the same address ending in `.md` (the home page is `/index.md`). Pages also answer in Markdown when asked with `Accept: text/markdown`, and to `curl` and `wget` by default. - [`/llms.txt`](https://spdf.joseluissaorin.com/llms.txt) lists every page with a one-line description, in English and Spanish. - [`/llms-full.txt`](https://spdf.joseluissaorin.com/llms-full.txt) carries the **whole specification** and every page of the site in one file. - [`/status.json`](https://spdf.joseluissaorin.com/status.json) has the CI status and conformance counts of every implementation, as JSON. - [`/sitemap.xml`](https://spdf.joseluissaorin.com/sitemap.xml) lists every page with its language alternates. `robots.txt` welcomes search engines and AI crawlers, including for training. - Pages carry schema.org JSON-LD: the specification as `TechArticle`, the implementations as `SoftwareSourceCode`, SPDF Commons as `Dataset`. ## Using SPDF files from an agent The [MCP server](https://spdf.joseluissaorin.com/integrations.md#mcp) `spdf-mcp` points at a folder of `.spdf` files and gives any MCP client (Claude, ChatGPT, Cursor, Zed, your own agent) these tools: | Tool | What it does | | --- | --- | | `list_documents` | The documents in the folder, with title, authors, year, kind and number of units | | `search` | Lexical search (or hybrid, when the files carry vectors and a query vector is given) over every document, with anchors | | `read_passage` | The literal text of a fragment, a unit (page, time span, slide) or a range, by id, printed folio or anchor URI | | `cite` | The short citation with the exact folio or second, the anchor URI and the quoted text, in English or Spanish | | `list_figures` | Figures, plates and frames with caption, description and anchor; optionally the image itself | | `get_metadata` | The CSL-JSON record and BibTeX of a document | ```sh npx spdf-mcp ~/Library/SPDF # stdio npx spdf-mcp ~/Library/SPDF --http 8765 # Streamable HTTP, optional ``` ## Rules for citing without inventing 1. **Quote from the file, cite from the anchor.** Take the text of a passage from `read_passage` or the search result, and its citation from `cite`. Never type a page number yourself. 2. **Printed folio, not position.** A page has a physical position in the file and, usually, a printed folio. Cite the printed folio; `cite` already does. If a page has no printed folio, the citation says `n. pag.` (`s. p.` in Spanish): do not replace it with the position. 3. **Brackets mean inferred.** `p. [21]` means the folio was deduced from the neighbouring pages, not read on the page. Keep the brackets. 4. **Keep the anchor URI.** Put it next to the claim (in a footnote, a link or a comment) so a human can open the exact passage with any SPDF reader. 5. **Literal text is literal.** Fragments keep the spelling of the source. The modernised layer (`search_text`) exists only to find them; never quote from it. 6. **If the file does not say it, do not cite it.** A search result is a candidate, not evidence: read the passage before attributing a claim to it. --- # Integrations URL: https://spdf.joseluissaorin.com/integrations > SPDF in the tools people already use: an MCP server for agents, loaders for LlamaIndex and LangChain in Python and JavaScript, a Zotero plugin and a Pandoc filter that turns SPDF anchors into citations with the printed folio. The format is only as useful as the places it reaches. These integrations live in the `integrations/` folder of the repository, each with its tests and its README, and all of them sit on the official libraries: none reimplements the format. - [MCP server: spdf-mcp](https://spdf.joseluissaorin.com/integrations/mcp.md): Any agent searches a folder of SPDF files and cites with the exact folio. - [LangChain.js loader](https://spdf.joseluissaorin.com/integrations/langchain-js.md): Passages as LangChain documents with citation and anchor URI. - [LlamaIndex.TS reader](https://spdf.joseluissaorin.com/integrations/llamaindex-js.md): Passages as LlamaIndex documents, with the stored vectors if you want them. - [LangChain loader (Python)](https://spdf.joseluissaorin.com/integrations/langchain-python.md): The same loader for LangChain in Python. - [LlamaIndex reader (Python)](https://spdf.joseluissaorin.com/integrations/llamaindex-python.md): The same reader for LlamaIndex in Python. - [Zotero 7 and 8 plugin](https://spdf.joseluissaorin.com/integrations/zotero.md): Import an SPDF as an item, attach it, copy a citation with the folio. - [Pandoc filter](https://spdf.joseluissaorin.com/integrations/pandoc.md): SPDF anchors in Markdown become citations with the printed folio, in any CSL style. ## MCP server `spdf-mcp` is a [Model Context Protocol](https://modelcontextprotocol.io) server in TypeScript over `spdf-format`. Point it at a folder of `.spdf` files and any agent can list the documents, search them, read a passage, look at the figures and cite with the exact folio, without being able to invent one. ```sh npx spdf-mcp ~/Library/SPDF # stdio npx spdf-mcp ~/Library/SPDF --http 8765 # Streamable HTTP ``` In Claude Code: `claude mcp add spdf -- npx spdf-mcp ~/Library/SPDF`. In any client that reads a JSON configuration: ```json { "mcpServers": { "spdf": { "command": "npx", "args": ["spdf-mcp", "/path/to/library"] } } } ``` Tools: `list_documents`, `search`, `read_passage`, `cite`, `list_figures`, `get_metadata`. See [SPDF for agents](https://spdf.joseluissaorin.com/agents.md) for the rules they follow. ## LlamaIndex and LangChain Loaders that turn every fragment of an SPDF into a document of the framework, **with its anchor and its citation in the metadata**, so retrieval-augmented answers can cite the printed page instead of a chunk number. ```py from spdf_llamaindex import SpdfReader # pip install spdf-llamaindex docs = SpdfReader(locale="en").load_data("darwin-origin.spdf") docs[0].metadata["citation"] # '(Darwin, 1859, p. 21)' docs[0].metadata["anchor_uri"] # 'spdf:sha256-…#p=29&f=21' ``` ```py from spdf_langchain import SpdfLoader # pip install spdf-langchain for doc in SpdfLoader("library/", locale="es").lazy_load(): print(doc.metadata["citation"], doc.page_content[:60]) ``` ```js import { SpdfLoader } from 'spdf-langchain'; // npm install spdf-langchain const docs = await new SpdfLoader('darwin-origin.spdf').load(); ``` ```js import { SpdfReader } from 'spdf-llamaindex'; // npm install spdf-llamaindex const docs = await new SpdfReader().loadData('darwin-origin.spdf'); ``` ## Zotero A plugin for Zotero 7 and 8 that brings SPDF into a reference library: - **Import an SPDF as an item**: the CSL-JSON record inside the file becomes a Zotero item, with the file attached. - **Attach an SPDF** to an existing item. - **Copy a citation with the folio**: pick a page or paste an anchor URI and get `(Darwin, 1859, p. 21)` on the clipboard, with the anchor URI alongside. Install it from the [`.xpi` file](https://spdf.joseluissaorin.com/zotero/spdf-zotero.xpi.md) (Tools → Plugins → Install Plugin From File); Zotero then updates it from this site. It has been tested outside Zotero (92 tests over Zotero fakes and the real fixtures); the checks still to do in a real Zotero 7 and 8 are listed in [its page](https://spdf.joseluissaorin.com/integrations/zotero.md). ## Pandoc A Lua filter for Pandoc that turns SPDF anchors in your Markdown into real citations, in any CSL style: ```markdown Darwin calls it a struggle for life [@spdf:sha256-3f2a9c…#p=29]. ``` ```sh pandoc essay.md --lua-filter spdf.lua -M spdf-library=library/ --citeproc -o essay.docx ``` The filter reads the SPDF files in the library folder, adds their CSL records to the bibliography, **replaces the physical page with the printed folio** (`p=29` becomes page 21) and warns if the page you cite does not exist. Citeproc then formats it in Chicago, APA, MLA or whatever style you choose. Why Pandoc rather than Calibre: academic writing in Markdown already goes through Pandoc and citeproc, and the step that matters most for the honesty of a citation (turning a position in a file into the folio a reader will find on paper) belongs exactly there. Calibre is a library for reading; the [reader](https://spdf.joseluissaorin.com/download.md) already covers that. --- # Download SPDF Reader URL: https://spdf.joseluissaorin.com/download > SPDF Reader opens, searches and cites SPDF files on macOS, Windows, Linux, iOS, Android and the web, with local models, offline and without an account. Free. **SPDF Reader** is the free reader of the format: open a file, read it page by page or second by second, search it by words or by meaning, and copy a citation with the exact folio. It runs the same interface everywhere, built with Tauri 2 on a Rust core, and the models it uses for semantic search run on your device. - **macOS** (Apple silicon, .dmg): Coming soon - **Windows** (Windows 10 and 11, .msi): Coming soon - **Linux** (AppImage and .deb): Coming soon - **Android** (.apk, Android 10 or later): Coming soon - **iOS and iPadOS** (App Store): Coming soon - **Web**: https://spdf.joseluissaorin.com/reader/ ## In your browser The [web reader](https://spdf.joseluissaorin.com/reader/) is the same application compiled for the web. It runs **entirely in your browser**: files are read with SQLite in WebAssembly and stored in your browser's private storage, and semantic search runs on your graphics card with WebGPU. Nothing is sent anywhere, except the one-time download of the embedding model from Hugging Face and, only if you ask for it and give your own key, Gemini. ## What it does - Opens SPDF 5.0 files and the legacy 4.x files from Scholaris. - Shows each page next to its text, with the printed folio, or plays the recording with its transcript word by word. - Lexical, semantic and hybrid search across your whole library. - Copies citations in English or Spanish, CSL-JSON and BibTeX, with the anchor URI. - Saves your notes outside the file, as W3C Web Annotations (`.spdfa.json`), so the file itself never changes. --- # SPDF Commons URL: https://spdf.joseluissaorin.com/commons > A small, careful collection of public-domain works already read into SPDF, in several languages and of several kinds (scanned books, EPUB, LibriVox recordings), free to download, test and cite. **SPDF Commons** is a small collection of public-domain works already read into SPDF: scanned books with their printed folios, EPUBs with their page lists, and LibriVox recordings with their transcripts timed word by word. They are here to be downloaded, opened, searched and cited, to test implementations against real documents, and to show what the format holds. | Work | Language | Kind | Units | Size | Download | | --- | --- | --- | ---: | ---: | --- | | El garrote mas bien dado, y alcalde de Zalamea (Calderón de la Barca, Pedro, 1746) | Spanish | scanned book | 32 | 5.7 MB | https://spdf.joseluissaorin.com/commons/files/calderon-alcalde-de-zalamea-1746.spdf | | Sidereus nuncius: Magna, longeqve admirabilia spectacula pandens, suspiciendaque proponens vnicuique, præsertim verò philosophis, atq́; astronomis, quæ à Galileo Galileo patritio Florentino Patauini Gymnasij publico mathematico perspicilli nuper à se reperti beneficio sunt obseruata in Lunæ facie, fixis innumeris, Lacteo Circulo, stellis nebulosis, apprime verò in quatuor planetis circa Iovis stellam disparibus interuallis, atque periodis, celeritate mirabili circumuolutis; quos, nemini in hanc vsque diem cognitos, nouissimè author depræhendit primus; atque Medicea Sidera nuncupandos decrevit (Galilei, Galileo, 1610) | Latin | scanned book | 68 | 8.7 MB | https://spdf.joseluissaorin.com/commons/files/galilei-sidereus-nuncius-1610.spdf | | The Yellow Wall Paper (Gilman, Charlotte Perkins, 1901) | English | scanned book | 80 | 5.1 MB | https://spdf.joseluissaorin.com/commons/files/gilman-yellow-wall-paper-1901.spdf | | Les Fleurs du mal (Baudelaire, Charles, 1857) | French | scanned book | 264 | 12.8 MB | https://spdf.joseluissaorin.com/commons/files/baudelaire-fleurs-du-mal-1857.spdf | | Die Verwandlung (Kafka, Franz, 1917) | German | EPUB | 73 | 1.0 MB | https://spdf.joseluissaorin.com/commons/files/kafka-die-verwandlung-1917.spdf | | Le avventure di Pinocchio: storia di un burattino (Collodi, Carlo, 1902) | Italian | EPUB | 289 | 5.3 MB | https://spdf.joseluissaorin.com/commons/files/collodi-pinocchio-1902.spdf | | Reliquias de casa velha (Machado de Assis, Joaquim Maria, 1906) | Portuguese | EPUB | 240 | 2.2 MB | https://spdf.joseluissaorin.com/commons/files/machado-de-assis-reliquias-de-casa-velha-1906.spdf | | L'auca del senyor Esteve (Rusiñol, Santiago, 1907) | Catalan | EPUB | 265 | 5.7 MB | https://spdf.joseluissaorin.com/commons/files/rusinol-auca-del-senyor-esteve-1907.spdf | | Obras escogidas (Bécquer, Gustavo Adolfo, 1912) | Spanish | EPUB | 354 | 5.9 MB | https://spdf.joseluissaorin.com/commons/files/becquer-obras-escogidas-1912.spdf | | The Tell-Tale Heart (Poe, Edgar Allan, 1843) | English | audiobook | 21 | 296 KB | https://spdf.joseluissaorin.com/commons/files/poe-tell-tale-heart-1843.spdf | | El monte de las ánimas (Bécquer, Gustavo Adolfo, 1861) | Spanish | audiobook | 24 | 328 KB | https://spdf.joseluissaorin.com/commons/files/becquer-monte-de-las-animas-1861.spdf | | Menuet (Maupassant, Guy de, 1882) | French | audiobook | 14 | 244 KB | https://spdf.joseluissaorin.com/commons/files/maupassant-menuet-1882.spdf | The collection manifest: https://spdf.joseluissaorin.com/commons/commons.spdfl.json ### How each work was verified - **El garrote mas bien dado, y alcalde de Zalamea** (2026-10-07): Read page by page with the vision engine. Printed page numbers checked by eye against the page images on physical pages 1 ("Fol. 1"), 11, 16, 29 and 32 (the last, with the colophon "Año de 1746"): all match. All 32 folios were read on the page; none is inferred. The lines of the sample ("el honor / es patrimonio del alma") were found on the image of p. 11. - **Sidereus nuncius: Magna, longeqve admirabilia spectacula pandens, suspiciendaque proponens vnicuique, præsertim verò philosophis, atq́; astronomis, quæ à Galileo Galileo patritio Florentino Patauini Gymnasij publico mathematico perspicilli nuper à se reperti beneficio sunt obseruata in Lunæ facie, fixis innumeris, Lacteo Circulo, stellis nebulosis, apprime verò in quatuor planetis circa Iovis stellam disparibus interuallis, atque periodis, celeritate mirabili circumuolutis; quos, nemini in hanc vsque diem cognitos, nouissimè author depræhendit primus; atque Medicea Sidera nuncupandos decrevit** (2026-10-07): Foliated book: leaves numbered on the recto. Checked by eye against the page images: physical pages 7, 19, 35, 41 and 63 carry the printed leaf numbers 2, 8, 16, 17 and 28 (cited as fols. 2r, 8r, 16r, 17r, 28r); versos take the number of their leaf in brackets, e.g. [16v] on physical 36, where the text of 16r continues, and [28v] on physical 64, the FINIS page. The two unnumbered leaves inserted after leaf 16 (physical 37-40: Orion and Pleiades star maps, Nebulosa Orionis and Praesepe) have no folio, nor do the title leaf, binding and endpapers. No fragment mixes pages with and without a folio: the opening of 17r ("De Luna, de inerrantibus Stellis") cites fol. 17r. - **The Yellow Wall Paper** (2026-10-07): Read page by page with the vision engine. Printed page numbers checked by eye against the page images on physical pages 13 (p. 1), 40 (p. 28) and 67 (p. 55, the last page of the text): all match, and every page of the text (pp. 1-55) carries a folio read on the page. Front matter, the blank leaves after the text, the binding and the library slips (physical 1-12 and 68-80) have no folio. The title page reads "Charlotte Perkins Stetson"; the record uses the author's later name, Gilman. No fragment mixes pages with and without a folio: the opening of the story ("ancestral halls for the summer") cites pp. 1-3, its fragment starting at the top of p. 1. - **Les Fleurs du mal** (2026-10-07): Read page by page with the vision engine. Printed page numbers checked by eye against the page images on physical pages 14 (p. 6), 51 (p. 43), 55 (p. 47), 137 (p. 129) and 256 (p. 248): all match, with a constant offset of 8 between physical page and folio. Pages without a printed number (the first page of each poem, the table) carry the inferred folio in brackets, e.g. [11] for "Bénédiction" and [128] for "A une dame créole". Of the 102 entries of the book's own table, 96 were found on the page the table gives; of the other six, the table gives p. 43 for "Châtiment de l'orgueil", which begins on p. [44], the dedication (p. 1) has no folio, and four titles were not read as such although their pages are right. Physical 53 prints "44" by mistake and is cited as [45]. Four pages (physical 30, 53, 97 and 225) were refused by the vision engine as recitation of a known text and carry the old OCR layer instead, marked with low confidence: their folios are right, their text is poor. Front matter and endpapers have no folio. - **Die Verwandlung** (2026-10-07): Page numbers of the 1917 Kurt Wolff edition, taken from the page markers of the EPUB (Project Gutenberg's own page list leaves out 11 of them). All 71 pages (5-75) checked: each unit starts exactly at its page marker in the EPUB, with no page missing or repeated; p. 23 starts with "Leibe zu spüren bekommt" and p. 75 is the last page of the story. The Project Gutenberg header and licence carry no folio and are in fragments of their own; the last sentence ("ihren jungen Körper dehnte") cites pp. 74-75, the fragment covering the last paragraph. - **Le avventure di Pinocchio: storia di un burattino** (2026-10-07): Page numbers of the 1902 Bemporad edition, taken from the page markers of the EPUB (Project Gutenberg's own page list leaves out p. 268). All 287 page markers (pp. 5-300) checked: each unit starts exactly at its page marker in the EPUB, with no page missing or repeated. The Project Gutenberg header and licence carry no folio and are in fragments of their own; no fragment mixes pages with and without a folio. - **Reliquias de casa velha** (2026-10-07): Page numbers of the 1906 Garnier edition, taken from the page markers of the EPUB. All 238 page markers (I-III and 3-264) checked: each unit starts exactly at its page marker in the EPUB, with no page missing or repeated. The Project Gutenberg header and licence carry no folio and are in fragments of their own; no fragment mixes pages with and without a folio. - **L'auca del senyor Esteve** (2026-10-07): Page numbers of the first edition (Antoni López, [1907]), taken from the page markers of the EPUB (Project Gutenberg's own page list leaves out pp. 1, 3 and 204). All 263 page markers (pp. 1-280) checked: each unit starts exactly at its page marker in the EPUB, with no page missing or repeated. The Project Gutenberg header and licence carry no folio and are in fragments of their own; no fragment mixes pages with and without a folio. - **Obras escogidas** (2026-10-07): Page numbers of the 1912 Fernando Fé edition, taken from the page markers of the EPUB (Project Gutenberg's own page list leaves out pp. 18, 186, 209 and 250). All 352 page markers (pp. ii-xii and 1-353) checked: each unit starts exactly at its page marker in the EPUB, with no page missing or repeated. The Project Gutenberg header and licence carry no folio and are in fragments of their own; no fragment mixes pages with and without a folio. - **The Tell-Tale Heart** (2026-10-07): Checked without playing the audio. At 2:15, 10:00 and 16:30 the words the file places at that second were compared with an independent local transcription (whisper.cpp) of the same 30-second window: they are the same words, and the median difference between the start times of the matched words is at most 0.4 s (word times are interpolated within the transcriber's segments). Transcript against the Project Gutenberg text: word error rate 0.2 %. - **El monte de las ánimas** (2026-10-07): Checked without playing the audio. At 2:15, 10:10 and 18:20 the words the file places at that second were compared with an independent local transcription (whisper.cpp) of the same 30-second window: they are the same words, but the file's word times run about 1 s behind (median +0.9 to +1.3 s; they are interpolated within the transcriber's segments). Transcript against the text of the 1912 edition: word error rate 2.8 % (the reader used Wikisource's text). - **Menuet** (2026-10-07): Checked without playing the audio. At 2:15, 6:40 and 10:40 the words the file places at that second were compared with an independent local transcription (whisper.cpp) of the same 30-second window: they are the same words, and the median difference between the start times of the matched words is at most 0.2 s (word times are interpolated within the transcriber's segments). Transcript against the Project Gutenberg text: word error rate 1.4 %. ## How they were made and checked Every file was produced with `spdf build`, the reference producer. The `provenance` table of each file records which model read each page or each second, with what confidence and when; open any of them in the [validator](https://spdf.joseluissaorin.com/validator.md) to see it. **Validating a file is not enough**: a folio put on the wrong page passes the validator. So, before a work is published here, its printed folios (or, in a recording, its times) are checked by eye against the page images or the transcript, on several pages from the beginning, the middle and the end, and the table above says, for each work, which pages were checked and what was found. A work whose check fails is rebuilt, not published. The sources are verifiable public-domain works (Project Gutenberg, Internet Archive, LibriVox, Wikisource). The original bytes are named by their SHA-256 in every file, so anyone can check that a file was read from the source it claims. ## The manifest The whole collection is described by an `.spdfl.json` manifest: one entry per file, with its SHA-256, title, authors, year and download address. Point any SPDF tool at it to fetch or verify the whole set. ## Licence The works are in the public domain. The SPDF files (the reading: transcription, anchors, sections, vectors) are dedicated to the public domain under CC0 1.0. --- # Validator and inspector URL: https://spdf.joseluissaorin.com/validator > Drop a .spdf file to check it against the SPDF specification and see what is inside: record, units, fragments with anchors, figures, vector spaces. Everything runs in your browser; nothing is uploaded. The validator runs in the browser at `/validator`: drop a .spdf file and it is checked with spdf-format and SQLite in WebAssembly, without uploading it. From the command line: `npx spdf-format validate file.spdf`. Samples: SPDF in five pages (/muestras/spdf-in-five-pages.spdf), SPDF en cinco páginas (/muestras/spdf-en-cinco-paginas.spdf), El garrote mas bien dado, y alcalde de Zalamea (/commons/files/calderon-alcalde-de-zalamea-1746.spdf), Sidereus nuncius: Magna, longeqve admirabilia spectacula pandens, suspiciendaque proponens vnicuique, præsertim verò philosophis, atq́; astronomis, quæ à Galileo Galileo patritio Florentino Patauini Gymnasij publico mathematico perspicilli nuper à se reperti beneficio sunt obseruata in Lunæ facie, fixis innumeris, Lacteo Circulo, stellis nebulosis, apprime verò in quatuor planetis circa Iovis stellam disparibus interuallis, atque periodis, celeritate mirabili circumuolutis; quos, nemini in hanc vsque diem cognitos, nouissimè author depræhendit primus; atque Medicea Sidera nuncupandos decrevit (/commons/files/galilei-sidereus-nuncius-1610.spdf), The Yellow Wall Paper (/commons/files/gilman-yellow-wall-paper-1901.spdf), Les Fleurs du mal (/commons/files/baudelaire-fleurs-du-mal-1857.spdf). ## What it checks The validator runs the same checks, in the same order, as every conforming implementation, and reports the same codes. It uses `spdf-format`, the TypeScript implementation, with SQLite compiled to WebAssembly: **the file never leaves your computer**. | Code | Meaning | | --- | --- | | E001 | Not a SQLite database | | E002 | Unknown `application_id` or version | | E003 | A 5.0 file wrapped in gzip (warning: 5.0 files are distributed uncompressed) | | E010 · E011 | A required table or column is missing | | E012 | A required `spdf_meta` key is missing | | E013 | `documents` must hold exactly one row | | E020 | A trigger or a view is present | | E030 · E031 · E032 | A vector of the wrong length, an unknown space or an unknown dtype | | E040 · E041 · E042 | An invalid anchor, an unknown anchor type or characters out of range | | E050 · E051 | Invalid metadata JSON, or metadata that is not a CSL item | | E060 | A required extension this reader does not know | | E070 | The full-text index is out of sync with the fragments | | E080 · E081 · E082 | A blob, the content hash or the signature does not match | | E090 | Units are not numbered contiguously from 1 | | W100–W110 | Warnings: a semantic profile without vectors, a media profile without times, a newer minor version, a legacy file… | ## From the command line Every implementation validates too. With the TypeScript one: ```sh npx spdf-format validate darwin-origin.spdf ``` --- # Governance and RFCs URL: https://spdf.joseluissaorin.com/governance > Who maintains SPDF, how the specification changes (through public RFCs with conformance cases), how versions work, and the licences and patent commitment. How SPDF is maintained and how it changes: who decides, the RFC process that every normative change goes through, how versions are numbered and what they promise, and the registrations that will make SPDF files recognizable by archives and operating systems. ## Documents | Document | What it covers | |---|---| | [GOVERNANCE.md](GOVERNANCE.md) | The editor, the future technical committee, decisions, conflicts of interest, appeals, code of conduct, licences and the patent commitment. | | [RFC-PROCESS.md](RFC-PROCESS.md) | When an RFC is needed, its states and steps, discussion periods, and when an RFC counts as implemented. | | [VERSIONING.md](VERSIONING.md) | Major, minor and editorial versions, `user_version`, the compatibility promise, deprecation, and the separate versions of the conformance suite and the implementations. | | [RFCs](../spec/rfcs/) | The RFCs themselves, and the [template](../spec/rfcs/0000-template.md). | | [CONTRIBUTING.md](../CONTRIBUTING.md) | How to contribute to the specification, the conformance suite and the implementations. | | [CODE_OF_CONDUCT.md](../CODE_OF_CONDUCT.md) | Contributor Covenant 2.1. | | [SECURITY.md](../SECURITY.md) | How to report a vulnerability privately. | ## In short - SPDF is edited by **José Luis Saorín Ferrer**, who created it for Scholaris. A **technical committee** takes over disputed decisions once three independent implementations, maintained by at least two organizations, pass the conformance suite. - **Nothing changes in silence.** An accepted RFC lands with at least one conformance case; it is considered implemented when two independent implementations pass it. Discussion lasts at least 14 days, followed by a 7-day final comment period. - **Versions live in the file**: `PRAGMA user_version` is major × 100 + minor × 10 (5.0 is 500). Every 5.x reader reads every 5.y file; minor versions only add things readers can ignore; deprecated features are removed only in a major version, at least 24 months later. Every 5.x reader also reads the legacy 4.0 and 4.1 files. - **Licences**: the specification and documentation under CC BY 4.0, the code under `MIT OR Apache-2.0`, and a public commitment not to assert patents against implementations. ## The RFCs - [RFC 0001: SPDF 5.0](https://spdf.joseluissaorin.com/governance/rfcs/0001.md) (Accepted (2026-10-07)) - [RFC 0002: Conformance cases for exports and anchor resolution](https://spdf.joseluissaorin.com/governance/rfcs/0002.md) (Accepted (2026-10-07); normative text in SPEC §5.4 and §19; cases in) ## Registrations in preparation Drafts of the requests that will register SPDF with the bodies that identify and describe file formats. None has been submitted: each one starts with a note (in Spanish) saying who it is for, through which channel, and what is still missing. | Body | Request | Draft | |---|---|---| | IANA | Media type `application/vnd.spdf+sqlite3` (RFC 6838, vendor tree) | [iana-media-type.md](drafts/iana-media-type.md) | | IANA | Provisional registration of the `spdf` URI scheme (RFC 7595) | [uri-scheme-spdf.md](drafts/uri-scheme-spdf.md) | | The National Archives (UK), PRONOM | Format record and DROID signature for SPDF 5.0 | [pronom-submission.md](drafts/pronom-submission.md) | | Library of Congress | Format description for *Sustainability of Digital Formats* | [loc-sustainability-fdd.md](drafts/loc-sustainability-fdd.md) | | SQLite and file(1) | `application_id` 0x53504446 in SQLite's `magic.txt` and in libmagic | [sqlite-magic-entry.md](drafts/sqlite-magic-entry.md) | | W3C | Charter for a Community Group on citable processed documents | [w3c-community-group-charter.md](drafts/w3c-community-group-charter.md) | ## Contact . Security reports: see [SECURITY.md](../SECURITY.md). --- # Governance URL: https://spdf.joseluissaorin.com/governance/governance > Who decides what SPDF is, how, and how that changes as the format gains implementers and users. SPDF starts with a single editor and is designed to pass to a technical committee as soon as there are people outside… Who decides what SPDF is, how, and how that changes as the format gains implementers and users. SPDF starts with a single editor and is designed to pass to a technical committee as soon as there are people outside the original project to share it with. The key words MUST, SHOULD and MAY are to be read as described in BCP 14 (RFC 2119 and RFC 8174) when, and only when, they appear in capitals. ## Principles - **Nothing changes in silence.** Normative changes go through the [RFC process](RFC-PROCESS.md), in public, with recorded reasons. - **Conformance over authority.** What the format means is settled by the specification and its conformance cases, not by what any one implementation does, including the reference one. - **Open by licence.** The specification can be implemented by anyone, for any purpose, without asking (see [Licences and patents](#licences-and-patents)). - **Users outside software count.** Libraries, archives, scholars and readers in every language are part of the community, not only implementers. ## Phase 1: the editor SPDF was created by **José Luis Saorín Ferrer** for Scholaris and opened as a standard in October 2026. Until the technical committee exists, he is the **editor** and the **maintainer** of the specification and of this repository. The editor: - runs the RFC process: assigns numbers, opens and closes discussion periods, records decisions and their reasons; - decides on RFCs after public discussion, seeking consensus first, and writes down how every substantive objection was answered; - keeps the specification, its Spanish translation and the conformance suite consistent; - merges changes to each implementation's folder only after its maintainer's review, once each folder has a named maintainer; - represents the project in registrations (IANA, PRONOM, the Library of Congress, the SQLite and file(1) projects) and before standards bodies; - handles security reports as [`SECURITY.md`](../SECURITY.md) describes. Contact: . **Continuity.** If the editor cannot be reached for 90 days, the maintainers of the first-tier implementations that pass the conformance suite MAY convene an interim committee by the rules below, to keep the specification and the repository alive until the editor returns or a technical committee is formed. ## Phase 2: the technical committee ### When it is formed The editor calls for a technical committee within 90 days after both conditions hold: 1. **three independent implementations** (in the sense of [`RFC-PROCESS.md`](RFC-PROCESS.md#independent-implementations)) pass the whole conformance suite of the current version; and 2. they are maintained by **at least two organizations** independent of each other. The editor and the projects he leads count as one organization; a person maintaining an implementation on their own counts as their own organization. The editor MAY call for the committee earlier. ### Composition - Between 5 and 7 members. - The editor holds a seat while the role of editor exists. - At least two members are maintainers of implementations that pass the suite, and at least one member represents users of the format (a library, archive, publisher or research group) rather than an implementation. - **No organization holds more than one third of the seats** (rounded down, minimum one). If a member changes employer and the limit is exceeded, a seat is renewed early. - Terms are two years and staggered, so that about half the seats are renewed each year. ### How members are chosen - **First committee.** The editor opens a public call for nominations for 30 days (self-nominations welcome), publishes the candidates and their affiliations, and proposes a committee that meets the composition rules. It is confirmed unless a sustained objection, with reasons, is raised within 14 days; objections are resolved in public before the committee takes office. - **Later renewals.** Elections with approval voting. Voters are people with at least one merged contribution to the specification, the conformance suite or an implementation in the previous 24 months, plus the named maintainers of implementations. Ties are broken by the public random procedure of RFC 3797. - Vacancies are filled for the rest of the term by the same method, or by co-option confirmed as for the first committee when fewer than 6 months remain. ### The editor under the committee The editor keeps writing and maintaining the specification and running the RFC process, and is a member of the committee. Disputed decisions are taken by the committee. The committee may appoint additional editors, and may replace the editor by a two-thirds majority of its members. ## Making decisions 1. **Consensus first.** Decisions are sought by consensus: no sustained objection remains after discussion. Routine matters use lazy consensus: a proposal announced in public stands if nobody objects within 7 days. 2. **Voting when consensus fails.** Only when the discussion has been given a fair chance and progress requires a decision, the committee votes. A quorum is a majority of its members. Decisions take a simple majority of the members voting, except: - a new **major version** of the specification, changes to this document, [`RFC-PROCESS.md`](RFC-PROCESS.md) or [`VERSIONING.md`](VERSIONING.md), and replacing the editor need **two thirds of all members**; - the licences cannot be changed to anything less open than they are (see below). 3. **Records.** Every decision is recorded in public, in the RFC, issue or pull request it concerns, with the reasons and, for votes, who voted how. Meetings may be held, but their decisions are tentative until written down there, and anyone may object within 7 days of publication, giving technical reasons. 4. In phase 1 the editor takes decisions by the same standard: consensus sought first, reasons recorded, objections answered in writing. ## Conflicts of interest - Committee members and the editor publish their affiliations, employers and funding related to SPDF, and keep that information current. - A member MUST disclose any direct interest in a decision (for example, a proposal that favours or harms a product of their employer) and SHOULD abstain from voting on it. They may still take part in the discussion. - The editor's own products, among them Scholaris and SPDF Reader, are declared interests. While no committee exists, decisions that would give them an advantage over other implementations are explained in writing and remain open to appeal. ## Appeals - Anyone affected by a decision may appeal it in writing within **30 days** of its publication, giving the reasons and the outcome they seek. - **Phase 1:** the appeal goes to the editor, who reconsiders and answers in public within 30 days. If the appellant is not satisfied, the appeal stays on record and the first committee reviews it if the matter is still open. - **Phase 2:** the appeal goes to the whole committee, which answers within 30 days. Members who took the decision under appeal may speak, but the outcome needs a majority of the members who did not. The committee's answer is final. - Appeals do not suspend a decision unless the editor or the committee says so. - The licences allow anyone to fork the specification at any time. Forks are welcome to build on SPDF, but files and implementations that do not follow this specification must not be presented as conforming SPDF. ## Claims of conformance An implementation may describe itself as conforming to SPDF *version* for the kinds of case it claims only if it passes every case of those kinds in a suite that targets that version, and its report (`conformance.json`) is public. Partial support is welcome and should be described as such. ## Code of conduct Everyone who takes part follows the [Contributor Covenant 2.1](../CODE_OF_CONDUCT.md). - **Phase 1:** reports go to the editor at . Reports concerning the editor himself are recorded, and the reporter may ask for them to be reviewed by an independent mediator agreed with them, or by the first committee. - **Phase 2:** the committee names two of its members as conduct contacts. A person who is the subject of a report, or has a conflict of interest in it, takes no part in handling it. - Reports are kept confidential, as the code of conduct requires. ## Licences and patents - The **specification and the documentation** are licensed under [Creative Commons Attribution 4.0 International](https://creativecommons.org/licenses/by/4.0/) (CC BY 4.0). - The **code** in this repository (libraries, the reference producer, the reader, the conformance suite, the website) is licensed under the MIT licence or the Apache License 2.0, at the user's choice (`MIT OR Apache-2.0`). - Contributions are accepted under the same licence as the part of the repository they change (inbound = outbound), without a contributor licence agreement or a Developer Certificate of Origin sign-off for now (see [`CONTRIBUTING.md`](../CONTRIBUTING.md)). - **Patent commitment.** José Luis Saorín Ferrer, as author and editor, commits not to assert any patent claim that he owns or controls, now or in the future, against anyone for making, using, selling, offering, importing or distributing an implementation of the SPDF specification. Every contributor to the specification makes the same commitment, by the act of contributing, for the patent claims they own or control that would necessarily be infringed by implementing the text they contributed. The commitment is irrevocable. - No decision of the editor or of the committee may relicense the specification under terms less open than CC BY 4.0, or the code under terms less open than `MIT OR Apache-2.0`. Versions already published keep their licences in any case. ## Where the project lives - Repository: . It is private until the specification, the conformance suite and the first-tier libraries pass, and public afterwards. If a GitHub organization `spdf-format` is created, the repository moves there. - Website: . - The editor or the committee may later propose to continue the work in a standards venue (for example a W3C Community Group). Such a move is decided by RFC and keeps the licences and the patent commitment above. ## Changing this document This document changes by RFC, with the same discussion and final comment periods as the specification. Under the committee, the change needs two thirds of all members. --- # The RFC process URL: https://spdf.joseluissaorin.com/governance/rfc-process > Every normative change to SPDF goes through a public request for comments (RFC) and lands together with the conformance cases that check it. This document says when an RFC is needed, how it moves from draft to… Every normative change to SPDF goes through a public request for comments (RFC) and lands together with the conformance cases that check it. This document says when an RFC is needed, how it moves from draft to decision, and when it counts as implemented. The key words MUST, SHOULD and MAY are to be read as described in BCP 14 (RFC 2119 and RFC 8174) when, and only when, they appear in capitals. ## The rule **An accepted RFC lands with at least one conformance case; it is considered implemented when two independent implementations pass it.** Everything below exists to apply that rule fairly. ## When an RFC is needed An RFC is REQUIRED for any change to what a valid file is, what a reader, writer or validator must do, or what a function defined by the specification returns. In particular: - the schema, the canonical dump, `user_version` or `spdf_meta` keys; - anchor types, anchor members and the anchor URI; - the reference search, citation and export functions; - validation codes and their order; - integrity and signatures; - the legacy mapping; - the sidecar formats (`.spdfa.json`, `.spdfl.json`); - profiles, and the rules for extensions; - this process, [`VERSIONING.md`](VERSIONING.md) and [`GOVERNANCE.md`](GOVERNANCE.md). An RFC is NOT needed for: - **editorial changes**: typos, clearer wording, examples, diagrams, translations, as long as no implementation would have to change. Open a pull request labelled `editorial`. If anyone shows that an "editorial" change alters behaviour, it becomes an RFC; - **fixes to a wrong conformance case** that contradicts the specification: fixed in place and logged, as [`conformance/README.md`](../conformance/README.md) describes; - **vendor extensions**: tables named `x__` declared in the `extensions` table need no permission. An RFC is needed only to make an extension part of the specification; - **implementation changes** that do not change behaviour defined by the specification. When in doubt, open an issue and ask. ## States | State | Meaning | |---|---| | Draft | Written, not yet open for discussion. | | Discussion | Open for public comment, at least 14 days. | | Final comment period | Last call, 7 days, with a proposed disposition. | | Accepted | Decided in favour; merged with its normative text and at least one conformance case. | | Implemented | Two independent implementations pass all of its cases in CI. | | Rejected | Decided against, with reasons. | | Postponed | Good idea, wrong time; may be reopened. | | Withdrawn | Abandoned by its authors. | | Superseded | Replaced by a later RFC, which is named. | ## Steps 1. **Before writing (optional).** Open an issue to test the idea. It saves everyone time when the answer is "this already exists" or "this belongs in an extension". 2. **Draft.** Copy [`spec/rfcs/0000-template.md`](../spec/rfcs/0000-template.md) to `spec/rfcs/0000-short-title.md` and open a pull request. Fill in every section; the Conformance cases section may start as a sketch, but it MUST be complete before the final comment period. 3. **Discussion, at least 14 days.** The editor assigns the next free number, renames the file, sets the state to Discussion and announces it. Anyone may comment. Authors revise the text in the same pull request; each substantive revision is summarized in a comment so that late readers can follow. 4. **Final comment period, 7 days.** When the discussion has settled, the editor (or the technical committee, once it exists) announces a proposed disposition: accept, reject or postpone. A new substantive objection during this period returns the RFC to Discussion; the 7 days start again when it is resolved. 5. **Decision.** The editor, or the technical committee once it exists, decides by the rules in [`GOVERNANCE.md`](GOVERNANCE.md), and records in the RFC the date, the outcome and the reasons, including how each substantive objection was answered. 6. **Merge.** An accepted RFC is merged together with: - the normative text in `spec/SPEC.md` and its Spanish translation `spec/SPEC.es.md`; - any schema change in `spec/schema/`; - **at least one conformance case** in `conformance/`, with its line in `conformance/CHANGELOG.md` and the suite version bumped as [`VERSIONING.md`](VERSIONING.md) says. An RFC without a conformance case is not merged. 7. **Implemented.** When two independent implementations pass every case of the RFC in CI (their `conformance-` artifacts show it), the editor sets the state to Implemented and records which implementations and versions. A version of the specification is published as final only with Implemented RFCs. Accepted RFCs that are not yet implemented may appear in working drafts, marked as such. ## Independent implementations Two implementations are independent when neither wraps, binds or mechanically translates the other's code, and each was written from the specification. Bindings over the Rust core, including the C ABI and anything built on it, count together with the Rust implementation. Implementations may share the conformance suite and the reference oracle; that is what they are for. Independence of authorship is not required for this rule; it matters for forming the technical committee (see [`GOVERNANCE.md`](GOVERNANCE.md)). ## Where discussion happens - On the RFC's pull request, once the repository is public. Decisions reached in calls or meetings are tentative until they are written in the pull request. - Until the repository is public, the editor publishes RFCs in Discussion on the website and collects comments sent to ; comments are recorded in the pull request with their author's permission. - Contributions in English or Spanish are welcome. The normative text is written in English, with a faithful Spanish translation. ## Shortened periods The editor (or the committee) MAY shorten the discussion and final comment periods only to fix a security problem, or a defect that makes conforming implementations incompatible with each other. The reason is recorded in the RFC, and the change is open to an RFC that revisits it afterwards with the normal periods. ## Numbering and files - RFCs are numbered with four digits, in the order discussion opens. Numbers are never reused; a rejected or withdrawn RFC keeps its number and file. - `0000-template.md` is the template, not an RFC. - A Spanish version MAY sit next to the English one as `NNNN-short-title.es.md`; the English text prevails if they differ. - [RFC 0001](../spec/rfcs/0001-spdf-5.0.md), which defines SPDF 5.0, was accepted by the editor before this process existed, without its discussion periods. Every later RFC follows this document. --- # Versioning URL: https://spdf.joseluissaorin.com/governance/versioning > How the SPDF specification is numbered, how the number is written inside files, what each kind of version may change, and how the conformance suite and the libraries are versioned on their own. The normative rules… How the SPDF specification is numbered, how the number is written inside files, what each kind of version may change, and how the conformance suite and the libraries are versioned on their own. The normative rules are in [SPEC §23](../spec/SPEC.md#versioning); this document repeats them and adds the policy around them. The key words MUST, SHOULD and MAY are to be read as described in BCP 14 (RFC 2119 and RFC 8174) when, and only when, they appear in capitals. ## Specification versions The specification is numbered **MAJOR.MINOR**. Editorial corrections do not change that number; they are published as editorial releases with a third number. | Kind | Example | May change | Needs | |---|---|---|---| | Major | 5.0 → 6.0 | anything, including breaking changes | RFCs | | Minor | 5.0 → 5.1 | only OPTIONAL additions (see the compatibility promise) | RFCs | | Editorial | 5.0.0 → 5.0.1 | wording, examples, translations, corrections that change no behaviour | no RFC | An editorial release never changes what a valid file is or what an implementation must do. It does not change the version written in files, and conformance results do not depend on it. Releases of the specification are tagged in the repository as `spec-vMAJOR.MINOR.PATCH` (the first is `spec-v5.0.0`), separately from the tags of the libraries. Each published text states its maturity next to its number: - **Working draft**: stable enough to implement, open to change by RFC. SPDF 5.0 has been a working draft since 2026-10-07. - **Final**: every RFC it contains is Implemented (two independent implementations pass its cases) and the conformance suite covers every MUST that can be tested. A final version changes only through editorial releases; anything else waits for the next minor or major version. ## The version inside a file ```text PRAGMA user_version = MAJOR × 100 + MINOR × 10 ``` | Version | `user_version` | Bytes 60 to 63 | |---|---|---| | 4.0 (legacy) | 400 (or 0, see below) | `00 00 01 90` | | 4.1 (legacy) | 410 | `00 00 01 9A` | | 5.0 | 500 | `00 00 01 F4` | | 5.1 | 510 | `00 00 01 FE` | - The same version is written as text in `spdf_meta.spdf_version` (`"5.0"`), never with the editorial number. - Writers write the units digit as 0. Readers accept the whole range of their major (500 to 599 for 5.x), as the specification says. - A 5.x writer that uses nothing introduced after 5.0 SHOULD write 500, so that readers of 5.0 do not warn. - The encoding allows ten minor versions per major (x.0 to x.9). If a major ever needs more, the RFC that proposes the eleventh defines its encoding. - Legacy 4.x files could not always set `user_version` (some have 0); their version is in the `spdf` table, and the legacy rules of the specification ([SPEC §20](../spec/SPEC.md#legacy)) cover them. - The optional `version` parameter of the media type (`application/vnd.spdf+sqlite3; version=5.0`) is informative. The file header is authoritative. ## What readers do with versions - A reader of major M reads every minor of M, ignoring what it does not know. On a minor newer than its own it MAY warn (W105). - A reader refuses a major it does not know (E002), and SHOULD keep reading older majors. Every 5.x reader MUST read the legacy versions 4.0 and 4.1, and MAY import 3.0. ## The compatibility promise **A file that conforms to 5.0 is readable by every conforming reader of any later 5.x version, and its anchor URIs keep resolving. A 5.0 reader reads every 5.x file.** To keep that promise, a minor version: - MAY add OPTIONAL things only: tables, columns, `spdf_meta` keys, anchor members or anchor types, metadata members, and validation warnings or errors for things that were already forbidden; - MUST NOT remove or rename anything, change the meaning or the type of existing data, make something optional required, or change the canonical dump, the `content_sha256`, the anchor URIs, the citations or the reference search results of files that use nothing new; - states, in each RFC that adds something a reader of an earlier minor cannot interpret (a new anchor type, for example), what such a reader does with it: it reads the rest of the file and leaves the unknown part aside. The RFC also says how a validator of an earlier minor reports it, and a conformance case checks both. A major version may change or remove anything. Its RFC says how its readers treat files of the previous major, as 5.0 does for 4.x. ## Deprecation - A feature is deprecated by an RFC in a minor version, which gives the reason and the replacement. Deprecated features stay in the specification and readers keep reading them; writers SHOULD stop writing them. - A deprecated feature is removed no earlier than the next **major** version, and no earlier than **24 months** after the publication of the version that deprecated it. - The specification lists every deprecation with the version that deprecated it and the earliest version that may remove it. ## Things versioned on their own The specification, the conformance suite and the libraries have separate version numbers. None of them waits for the others to release. | What | Where the number lives | Scheme | |---|---|---| | Specification | `spec/SPEC.md`, `user_version`, `spdf_meta.spdf_version`; tags `spec-vX.Y.Z` | MAJOR.MINOR (+ editorial) | | Conformance suite | `conformance/manifest.json` (`suite_version`, and the `spdf_version` it targets) | semantic versioning | | Libraries | each package manifest; tags `X.Y.Z` (Go: `go/vX.Y.Z`) | one semantic version shared by all libraries | | SPDF Reader, `spdf build`, the website | their own manifests | their own | | Collection manifests | `spdf_library` member of `*.spdfl.json` (now `"1.0"`) | MAJOR.MINOR, with the same promise as the specification | | Extensions | `extensions.version` of each extension | chosen by its vendor | ### The conformance suite - **Major**: it targets a new major version of the specification, or a kind of case is retired. - **Minor**: new cases or new kinds of case. - **Patch**: a wrong case corrected in place, or a fix in the tools that changes no expectation. - While the suite is at 0.x, as now, minor releases may also correct expectations. - Every change is logged in `conformance/CHANGELOG.md`. Case ids are never renamed or reused. ### The libraries - The libraries follow a **common release train**: a semantic version tag without prefix (`0.1.0`, `0.2.0`…) marks a coordinated release of every library, with the same number in all of them. Go uses `go/vX.Y.Z` with the same number, as Go modules in a subdirectory require. - The libraries stay at 0.x until the 5.0 specification is final and there are two independent producers and at least three independent readers that pass the whole conformance suite. Then they release 1.0.0. - Each library states in its README which specification versions it supports, which product class and kinds of case it claims ([SPEC §21](../spec/SPEC.md#conformance)) and which suite version it passes; its runner reports its own version in `conformance.json`. - A library MUST NOT claim to support a specification version unless it passes every case of the kinds it claims in a suite that targets that version. - A new minor of the specification does not force a major release of the libraries: reading newer minors is already part of the compatibility promise. --- # Media type registration: application/vnd.spdf+sqlite3 URL: https://spdf.joseluissaorin.com/governance/drafts/iana-media-type > Borrador sin enviar. Solicitud de registro del tipo de medio application/vnd.spdf+sqlite3 en el árbol de fabricante (vnd.) del registro de tipos de medio de IANA, según la plantilla de la sección 5.6 de la RFC 6838,… > **Borrador sin enviar.** Solicitud de registro del tipo de medio > `application/vnd.spdf+sqlite3` en el árbol de fabricante (`vnd.`) del registro de tipos > de medio de IANA, según la plantilla de la sección 5.6 de la RFC 6838, con el sufijo > estructurado `+sqlite3` (registrado en IANA). Va a IANA por su formulario web > (), donde la revisa un experto designado. La > RFC 6838 recomienda, sin exigirlo, enviarla antes a la lista media-types@iana.org > para que la comente la comunidad; conviene hacerlo. > > Decidido el 7-10-2026 por el orquestador del proyecto: el tipo es > `application/vnd.spdf+sqlite3` y no `application/vnd.spdf`. El sufijo permite que las > herramientas genéricas de SQLite reconozcan el fichero, y SPDF ya cumple lo que pide el > registro del sufijo (un `application_id` propio en el desplazamiento 68 y su entrada en > `magic.txt`, véase `sqlite-magic-entry.md`). `SPEC.md` §24 ya define los > identificadores de fragmento que se describen abajo. > > Falta antes de enviarla: > > 1. Una versión estable de la especificación. `https://spdf.joseluissaorin.com/spec` > ya sirve `SPEC.md` (comprobado el 7-10-2026), pero es un borrador de trabajo que > cambia; conviene enlazar una versión fechada que no cambie (la de la etiqueta > `spec-v5.0.0`) y, a ser posible, enviarla cuando la 5.0 sea final. > 2. Confirmar quién figura como responsable del cambio (José Luis o, si se crea, la > organización `spdf-format` o el comité técnico). > 3. Por verificar: el texto exacto de las consideraciones de identificadores de > fragmento del registro del sufijo `+sqlite3`, para citarlo en el apartado > correspondiente. > > Comprobado el 7-10-2026: el nombre `vnd.spdf` no figura en el registro de IANA; la > plantilla sigue la RFC 6838; el registro de `application/vnd.sqlite3` sirvió de > referencia para las consideraciones de seguridad, que resumen las §2.4, §14 y §15 de > `SPEC.md`. --- **Type name:** application **Subtype name:** vnd.spdf+sqlite3 **Required parameters:** N/A **Optional parameters:** `version`: the version of the SPDF specification the file declares, as `MAJOR.MINOR` (syntax: `1*DIGIT "." 1*DIGIT`), for example `version=5.0`. The parameter is informative only. The authoritative version is stored in the file itself (the SQLite `user_version` header field and the `spdf_meta.spdf_version` row), and receivers MUST NOT rely on the parameter instead of the file. **Encoding considerations:** binary **Structured syntax suffix:** `+sqlite3`. The considerations of the `+sqlite3` suffix registration apply; an SPDF file is a SQLite 3 database that any SQLite tool can open read-only. The SPDF-specific rules below add to them. **Security considerations:** An SPDF file is a SQLite 3 database. The security considerations of `application/vnd.sqlite3` apply in full; in addition: 1. *Active content.* SQLite schemas can contain triggers and views, which are SQL code executed by the library on behalf of whoever uses the database, and virtual tables, which call module code. SPDF files MUST NOT contain triggers, views, or virtual tables other than the two FTS5 full-text tables the specification defines. Readers open files read-only (`SQLITE_OPEN_READONLY`, `PRAGMA query_only=1`), with `PRAGMA trusted_schema=OFF` and, where the SQLite binding exposes it, `SQLITE_DBCONFIG_DEFENSIVE`; they never load SQLite extensions, and they refuse files whose schema contains any of those objects. Files of the legacy versions 4.0 and 4.1 contain exactly three full-text synchronization triggers, which readers tolerate because a read-only connection never fires them. 2. *Malformed and hostile files.* The SQLite file format is complex, and crafted files can exercise bugs in the library. Following SQLite's own advice for untrusted databases, readers use a current SQLite release, disable memory-mapped I/O, enable `cell_size_check`, may run `quick_check` first, and bound the size of every value they read (the specification recommends 512 MiB) and the nesting depth of the JSON they parse (it recommends 64). An index that disagrees with its table can make queries return data that is not stored; validators check the full-text index on a private in-memory copy. Search terms reach the full-text engine only as quoted strings and SQL is always parameterized, so user input cannot inject query syntax. 3. *Compression.* The SPDF 5.x container is not compressed. Embedded objects (page images, the original document, audio or video) are stored in their own formats and carry those formats' considerations, including their own compression. Legacy 4.x files are SQLite databases wrapped in gzip; readers decompress them only up to a configurable limit (the specification's default is 4 GiB) to defeat decompression bombs. 4. *Embedded documents and images.* A file may embed the original it was read from (for example PDF, EPUB, HTML or SVG) and page images. They are untrusted input to their own decoders and may contain scripts or other active content. Readers MUST NOT execute embedded content, MUST NOT render SVG with scripts enabled, and SHOULD render embedded documents only in a sandbox. Blob keys are opaque strings: a reader that writes blobs to disk sanitizes them (no absolute paths, no `..`, no device names). 5. *Text and links.* Unit text is light Markdown; readers render it without raw HTML and escape it before inserting it into HTML. URLs may appear in the source reference, in image references, in the metadata and in web anchors. Readers MUST NOT fetch them automatically, because fetching discloses that the file was opened and can reach internal services; they fetch only on a user action, showing the address first. 6. *Language models.* Text read from a file may contain instructions aimed at language models. Applications that pass SPDF text to a model treat it as data, not as instructions. 7. *Privacy.* Besides the visible text, a file may hold the original document, page images, word-level timings and speaker names of recordings, and provenance records naming the tools and models that processed it and when. The content itself may be personal data (for example a recorded interview). Embedding vectors can be inverted to reconstruct much of the text they were computed from, so distributing the vectors of a text is close to distributing the text, and a writer that strips the text of a restricted document strips its vectors too. User annotations are kept outside the file, so sharing a document does not share its reader's notes. Writers compact files with `VACUUM` before distribution so that deleted data does not remain in free pages. 8. *Integrity and authenticity.* The format provides no confidentiality. It provides optional integrity: `content_sha256` is a SHA-256 of a canonical serialization of the content (RFC 8785), which covers embedded objects and vectors through their hashes but not the SQLite page layout, and `signature` is an Ed25519 signature (RFC 8032) of that hash. Verifiers recompute the hash rather than trust the stored value. Provenance, confidence values and metadata are claims made by the writer; a valid signature shows that the holder of the key vouched for them, and says nothing about whether the key should be trusted (key distribution is outside the specification). Unsigned files can be altered without detection. 9. *Anchor URIs.* SPDF anchor URIs name a document by the SHA-256 of its source. Sending such a URI to a resolver reveals which document and which passage a user is reading. **Interoperability considerations:** - SPDF files are ordinary SQLite 3 databases and can be opened by any SQLite tool. Full conformance requires FTS5 (part of the SQLite amalgamation since 3.9.0) with the `unicode61 remove_diacritics 2` tokenizer (SQLite 3.27.0 or later); safe opening requires `trusted_schema` (3.31.0 or later); the optional `trigram` index for Chinese, Japanese and Korean requires 3.34.0 or later. - Distributed files use the DELETE journal mode, so they are self-contained (no `-wal` or `-journal` companion file). - The version is stored as `PRAGMA user_version` = major × 100 + minor × 10. A reader of a major version reads every minor version of it; it refuses unknown major versions. - Files of the legacy versions 4.0 and 4.1, written by the Scholaris application before this registration, share the `.spdf` extension but are gzip-compressed SQLite databases (magic number `1F 8B`) without the SPDF `application_id`. Conforming SPDF readers read them. - Metadata is a CSL-JSON item; user annotations are kept outside the file as W3C Web Annotation documents (`*.spdfa.json`, served as `application/json` or `application/ld+json`); collections are JSON manifests (`*.spdfl.json`, `application/json`). Neither sidecar uses this media type. - A public conformance suite defines the expected results of reading, validating, searching, building anchor URIs and citing, and several independent implementations are tested against it. **Published specification:** SPDF: Semantic Processed Document Format, version 5.0. (Spanish translation: ). Source: . Licensed under CC BY 4.0. **Applications that use this media type:** SPDF Reader (desktop, mobile and web); the Scholaris citation application; the reference producer `spdf build`; the SPDF libraries for Rust, TypeScript, Python, Swift, Kotlin/JVM, Go, C#, PHP, Ruby, R and Julia, and their integrations with reference managers and document tools. **Fragment identifier considerations:** As RFC 6838 section 4.11 allows, this registration defines fragment identifier semantics specific to the type, in addition to those of the `+sqlite3` suffix. A fragment identifier on a URI that resolves to an `application/vnd.spdf+sqlite3` resource addresses a location in the single document the file holds. Its syntax is the parameter list of the SPDF anchor URI: `key=value` pairs joined by `&`, in the canonical order the specification defines, values percent-encoded as UTF-8. For example: ```text https://example.org/darwin.spdf#p=29&f=21&char=118,301 ``` `p` and `pe` are physical pages; `f` and `fe` printed folios; `t` a time range in seconds; `s` a section path; `para` a paragraph; `sl` a slide; `sh` and `rows` a sheet and its rows; `v` a verse or range of verses; `ref` a canonical reference (`scheme:ref`); `char` a character range; `xywh` a region. Where they overlap with existing standards the parameters use their syntax: `t=,` and `xywh=percent:,,,` as in W3C Media Fragments URI 1.0, and `char=,` as in RFC 5147, with positions counted in Unicode code points of the NFC-normalized text of the unit. Unknown keys are ignored. The same parameter list follows `#` in SPDF anchor URIs of the form `spdf:sha256-#…`, which name the document by the SHA-256 of its source instead of by location. **Additional information:** - Deprecated alias names for this type: N/A - Magic number(s): - offset 0, 16 bytes: `53 51 4C 69 74 65 20 66 6F 72 6D 61 74 20 33 00` (ASCII "SQLite format 3" followed by a NUL byte); - offset 68, 4 bytes: `53 50 44 46` (ASCII "SPDF"; the SQLite `application_id` 1397769286 = 0x53504446, stored big-endian); - offset 60, 4 bytes: the version as a big-endian integer, `00 00 01 F4` (500) for version 5.0, and in general from 500 to 599 for the 5.x versions. - File extension(s): `.spdf` - Macintosh file type code(s): none - Uniform Type Identifier (Apple platforms): `com.joseluissaorin.spdf`, conforming to `public.data` and `public.database`. **Person & email address to contact for further information:** José Luis Saorín Ferrer **Intended usage:** COMMON **Restrictions on usage:** N/A **Author:** José Luis Saorín Ferrer **Change controller:** José Luis Saorín Ferrer , editor of the SPDF specification. **Provisional registration? (standards tree only):** N/A --- # Format description: SPDF (Semantic Processed Document Format), version 5.0 URL: https://spdf.joseluissaorin.com/governance/drafts/loc-sustainability-fdd > Borrador sin enviar. Propuesta de ficha de formato (Format Description Document) para Sustainability of Digital Formats: Planning for Library of Congress Collections, de la Library of Congress. Sigue la estructura de… > **Borrador sin enviar.** Propuesta de ficha de formato (Format Description Document) > para *Sustainability of Digital Formats: Planning for Library of Congress > Collections*, de la Library of Congress. Sigue la estructura de las fichas > publicadas (comprobada el 7-10-2026 sobre la de SQLite 3, fdd000461, y las > explicaciones de términos del sitio). Las fichas las redacta y publica el equipo de > formatos de la Library of Congress; esto es una sugerencia para que la valoren, no un > texto que vayan a copiar tal cual. > > Vía: la página de contacto del sitio de formatos > (), que enlaza > la propia ficha de SQLite; el canal concreto (formulario o correo) está por verificar > al enviarla. > > Falta antes de enviarla: > > 1. Una versión final, o al menos fechada, de la especificación. Hoy la web la sirve > en `https://spdf.joseluissaorin.com/spec` como borrador de trabajo (comprobado el > 7-10-2026). > 2. Registro en IANA (`application/vnd.spdf+sqlite3`) y ficha en PRONOM, para poder citarlos > aquí; hoy figuran como pendientes. > 3. Más adopción. La ficha es honesta: hoy el formato lo usa sobre todo el proyecto de > su autor, y la Library of Congress puede decidir esperar. Conviene enviarla cuando > haya al menos un productor o un usuario institucional ajeno al proyecto. > > Por verificar: el nombre corto que asignarían (aquí `SPDF_5_0`), las facetas y los > identificadores de ficha de los formatos relacionados que no son SQLite, JSON ni > GeoPackage (TEI, ALTO, IIIF y W3C Web Annotation no aparecen en la lista de fichas > consultada). --- ## Format description properties | Property | Value | |---|---| | ID | to be assigned | | Short name | SPDF_5_0 (proposed) | | Content categories | text, dataset | | Format category | file-format | | Other facets | unitary, binary, structured | | Draft status | Preliminary (proposed) | ## Identification and description | | | |---|---| | Full name | SPDF (Semantic Processed Document Format), version 5.0 | | Description | SPDF is an open file format for documents that have been read once, by a PDF text layer, optical character recognition, a vision model, a speech recognizer or by hand, and stored so that every passage can be cited with its exact location in the source. A file is a SQLite 3 database (see SQLite_3) holding exactly one document: a bibliographic record as a CSL-JSON item; citable units in reading order (pages or leaves, time spans, slides, sections, spreadsheet rows) with their text in Unicode NFC and a JSON "anchor" giving the physical page and printed folio, leaf or column, verse, canonical reference, time range or slide; passages of about 150 to 300 words indexed for full-text search with the SQLite FTS5 module; and, optionally, section structure, figures with regions, embedding vectors from one or more models, provenance records, rights information, and the original file and page images as embedded binary objects. The SQLite header identifies the format: `application_id` "SPDF" at byte offset 68 and the version (500 for 5.0) at offset 60. A portable URI syntax (`spdf:sha256-…#p=29&f=21`) addresses any passage independently of the file's location. | | Production phase | Generally a middle- or final-state derivative: produced from an original (a printed book, a scan, a born-digital PDF, an EPUB, a recording, a slide deck, a web page) to make it searchable and citable. It does not replace the original, which it may embed or reference by its SHA-256 hash. | | Relationship to other formats | | | Subtype of | SQLite_3, SQLite, Version 3 (fdd000461) | | Has earlier version | SPDF 4.0 and 4.1 (no FDD): gzip-compressed SQLite databases with Spanish identifiers, internal to the Scholaris application; readers of 5.0 must read them | | May contain | JSON (fdd000381), in text columns (anchors, metadata, rights, provenance); embedded objects in their own formats (for example PNG or JPEG page images, the original PDF, EPUB, audio or video) | | Affinity to | GeoPackage_1_0 (fdd000419), another application format built on SQLite and identified by its `application_id`; TEI P5, ALTO and IIIF Presentation API 3.0, to which SPDF content can be exported; W3C Web Annotation, used by SPDF's annotation sidecar files (`.spdfa.json`); CSL-JSON, used for its metadata | ## Local use | | | |---|---| | LC experience or existing holdings | None known (for the Library of Congress to complete). | | LC preference | None (for the Library of Congress to complete). | ## Sustainability factors | Factor | Assessment | |---|---| | Disclosure | Fully documented, open specification, edited by José Luis Saorín Ferrer and published under the Creative Commons Attribution 4.0 licence, with a faithful Spanish translation. Version 5.0 is a working draft (first published 2026-10-07); changes go through a public request-for-comments process. | | Documentation | *SPDF: Semantic Processed Document Format, version 5.0* (specification, SQL schemas, JSON Schemas); a public conformance suite with expected results for reading, validating, searching, building anchor URIs and citing. | | Adoption | Very low as of October 2026. The format was created for the Scholaris citation application of the same author, whose earlier versions (4.0, 4.1) are internal to it. Libraries for version 5.0 in Rust, TypeScript, Python, Go, Swift, PHP and Ruby report passing the whole conformance suite; Kotlin, C#, R and Julia libraries, a reference producer and a free reader application are in development in the same project. No use by memory institutions or by unrelated producers is known yet. | | Licensing and patents | Specification under CC BY 4.0. Reference code under MIT or Apache License 2.0. The author and every contributor to the specification commit not to assert patents against implementations. The underlying SQLite format and library are in the public domain. No patents are known to apply. | | Transparency | High for the structure and text: any SQLite tool (for example the `sqlite3` command-line shell) can list the tables and read the text, which is UTF-8 in NFC, with JSON in text columns and light Markdown in unit text. Embedding vectors are little-endian binary arrays whose meaning depends on the model named in the file; they are derived data and can be recomputed. Embedded images and originals keep their own formats. Legacy 4.x files must be decompressed (gzip) before inspection. | | Self-documentation | Strong. The file records its format version and generator, a CSL-JSON bibliographic record with the source and confidence of each field, the SHA-256, media type and size of the original, rights (SPDX licence identifier, access level, holder), which reader produced the text of each unit and with what confidence, how each folio was obtained (read from the page or inferred), which model produced each set of vectors, and a provenance log of the processing stages. An optional content hash and Ed25519 signature allow checking the content and who vouched for it. | | Accessibility features | The format provides a text layer for scanned and image-only sources, reading order, section headings, footnotes separated from the main text, figure captions and descriptions, and time-aligned transcripts of recordings with speaker names and word timings, which support screen readers, search and captions. Tables can be written as Markdown tables in the unit text; there is no layout tagging comparable to tagged PDF. | | External dependencies | A SQLite library (widely available on all platforms) with the FTS5 module for search; no other software is required to read the text and metadata. Semantic search requires the embedding model named in the file to encode queries; the file remains fully readable without it. When the original is not embedded, it is referenced by URL or hash and must be preserved separately. | | Technical protection considerations | None. The format defines no encryption or access control; an encrypted SQLite database is not a valid SPDF file. The optional Ed25519 signature provides integrity and attribution, not protection. A `rights` object states the licence and the intended access level for information only. | ## Quality and functionality factors | Category | Factor | Assessment | |---|---|---| | Text | Normal rendering | Unit text in reading order, as light Markdown (CommonMark subset), suitable for linear reading, quotation, search and indexing. The literal spelling of the source is kept; a separate modernized-spelling layer is used only for search. | | Text | Integrity of document structure | Units (pages, leaves, time spans, slides, sections), a section hierarchy, paragraphs, verse lines, footnotes, running heads and footers kept apart from the text. | | Text | Integrity of layout and display | Not preserved by the text layer. Layout survives only in the optional page images or the embedded original. Regions of figures and passages can be recorded as fractions of the page image. | | Text | Support for mathematics, formulae, etc. | Not specified in version 5.0 beyond plain text and light Markdown. | | Text | Functionality beyond normal rendering | Exact citation: every passage carries its physical page and printed folio (roman, inferred, leaf or column), verse or canonical reference, with character offsets; a reference function produces short citations; anchor URIs; lexical, semantic and hybrid search; export to CSL-JSON, BibTeX, TEI, ALTO, IIIF and W3C Web Annotation. | | Still image | Normal rendering, clarity, color maintenance | Determined by the embedded page images or figures, which are stored in their own formats (for example PNG or JPEG) at whatever resolution the producer chose; SPDF adds regions and descriptions but no image encoding of its own. | | Sound | Normal rendering, fidelity, multiple channels | SPDF does not encode audio. It stores time-aligned transcripts (time spans with speakers and word timings in centiseconds) and may embed or reference the original recording, which keeps its own format and quality. | | Moving image | Normal rendering, clarity | As for sound: transcripts and time anchors, with the original video embedded or referenced. | | Dataset (metadata) | Normal functionality | Typed SQLite tables with a fixed schema; one document per file; JSON columns validated by JSON Schemas published with the specification. | | Dataset (metadata) | Support for software interfaces | Any SQLite interface; libraries in many programming languages tested against a common conformance suite; command-line tools and a validator. | | Dataset (metadata) | Data documentation (quality, provenance, etc.) | Per-unit reader and confidence, per-field metadata provenance, folio source (read or inferred), model and version of each vector space, processing log, and an optional signed content hash. | ## File type signifiers and format identifiers | Tag | Value | Note | |---|---|---| | Filename extension | spdf | Also used by the legacy versions 4.0 and 4.1, which are gzip-compressed. | | Internet Media Type | application/vnd.spdf+sqlite3 | Registration with IANA in preparation. | | Magic numbers | Hex: `53 51 4C 69 74 65 20 66 6F 72 6D 61 74 20 33 00` at offset 0 (ASCII "SQLite format 3" and NUL); `53 50 44 46` at offset 68 (ASCII "SPDF", the SQLite `application_id` 0x53504446); `00 00 01 F4` at offset 60 (`user_version` 500, version 5.0) | Legacy 4.x files begin with `1F 8B` (gzip) and cannot be told from other gzip files without decompressing. | | Other | Uniform Type Identifier `com.joseluissaorin.spdf`, conforming to `public.data` and `public.database` | Proposed in the specification. | | Pronom PUID | none yet | Submission in preparation. | | Wikidata Title ID | none yet | | ## Notes **General.** An SPDF file is meant to accompany its original, not to replace it: it records what was read from the original and where, so that citations stay exact even when the original is not at hand. The original's SHA-256 is always recorded, and anchor URIs use it, so the same passage can be addressed in every copy of the document. User annotations and collection lists are kept in separate JSON files (`.spdfa.json`, `.spdfl.json`), so the document file can stay unchanged. Distributed files must not contain triggers, views or foreign virtual tables, and readers open them read-only. **History.** SPDF was created by José Luis Saorín Ferrer for Scholaris, an application that inserts verified, page-exact citations into academic writing, under the name *Scholaris Processed Document Format*. Version 3.0 was a gzip-compressed SQLite database with `metadata` and `chunks` tables; versions 4.0 and 4.1 a gzip-compressed SQLite database with Spanish identifiers. All of them were internal to Scholaris. Version 5.0, published as a working draft on 7 October 2026 under the name *Semantic Processed Document Format*, is the first public version: uncompressed, with English identifiers, CSL-JSON metadata, defined text offsets, new anchor types, an anchor URI aligned with W3C Media Fragments and RFC 5147, optional signatures and a conformance suite. ## Format specifications - *SPDF: Semantic Processed Document Format, version 5.0*. José Luis Saorín Ferrer, 2026. . Source: . - SPDF conformance suite, version 0.2.0 (228 cases), in the same repository. ## Useful references - SQLite, *Database File Format*. - SQLite, *FTS5 Extension*. - Citation Style Language, CSL-JSON schema. - W3C, *Media Fragments URI 1.0 (basic)*, Recommendation, 2012. - W3C, *Web Annotation Data Model*, Recommendation, 2017. - RFC 5147, *URI Fragment Identifiers for the text/plain Media Type*. - RFC 8032, *Edwards-Curve Digital Signature Algorithm (EdDSA)*. - RFC 8785, *JSON Canonicalization Scheme (JCS)*. --- # PRONOM submission: SPDF (Semantic Processed Document Format) 5.0 URL: https://spdf.joseluissaorin.com/governance/drafts/pronom-submission > Borrador sin enviar. Propuesta de ficha nueva para PRONOM, el registro técnico de formatos de The National Archives (Reino Unido), que alimenta las herramientas de identificación DROID y FIDO. Los campos siguen la… > **Borrador sin enviar.** Propuesta de ficha nueva para PRONOM, el registro técnico de > formatos de The National Archives (Reino Unido), que alimenta las herramientas de > identificación DROID y FIDO. Los campos siguen la plantilla oficial de envío > («PRONOM Submission template», en Word y en hoja de cálculo) del repositorio > `digital-preservation/PRONOM_Research`. > > La vía actual es GitHub: un *pull request* con la investigación en la carpeta > `Submissions` de , o una > *issue* en ese repositorio. También aceptan correo a PRONOM@nationalarchives.gov.uk, > que es la vía para muestras que no deban publicarse. Lo que se sube a ese repositorio > se publica con licencia CC0 (muestras) y la descripción, con la Open Government > Licence. > > Falta antes de enviarla: > > 1. **Tipo de medio.** PRONOM solo admite tipos de medio registrados en IANA o que > figuren en la documentación oficial del formato. Lo ideal es enviar esto después > de registrar `application/vnd.spdf+sqlite3` (ver `iana-media-type.md`); si no, > hay que citar `SPEC.md` publicada como documentación oficial. > 2. **Muestras.** Preparar ficheros de ejemplo descargables y de dominio público: los de > `conformance/files/` sirven (Cervantes, Darwin, Hooke…), pero hay que publicarlos en > una URL estable o adjuntarlos al *pull request* (quedan bajo CC0). > 3. **Probar la firma con DROID** (por ejemplo con la utilidad de desarrollo de firmas > de Ross Spencer que enlaza la propia guía de PRONOM) sobre las siete muestras 5.0, > sobre los dos ficheros heredados y sobre un SQLite cualquiera, para comprobar que > no hay falsos positivos. > 4. Una versión estable de la especificación: la web ya la sirve en > `https://spdf.joseluissaorin.com/spec` (comprobado el 7-10-2026), pero como > borrador de trabajo; conviene citar una versión fechada que no cambie. > > Por verificar: si PRONOM prefiere una ficha por versión menor (5.0, 5.1…) o una ficha > «5.x» con la firma genérica; qué clasificación asignan (aquí se propone «Database», > con «Text (Structured)» como alternativa); y si quieren fichas aparte para las > versiones heredadas 4.0 y 4.1, que no se pueden identificar sin descomprimir. > Comprobado el 7-10-2026: la ficha de SQLite 3 es fmt/729 (tipo MIME > `application/x-sqlite3`, firma `53514C69746520666F726D6174203300` en el desplazamiento > 0); GZIP es x-fmt/266; los formatos basados en SQLite con `application_id` propio > (OGC GeoPackage, fmt/1700; Audacity 3, fmt/1826) declaran «Has priority over» respecto > de fmt/729 y usan la sintaxis de huecos `{n}` de DROID. --- ## Format | Field | Value | |---|---| | File format name | SPDF (Semantic Processed Document Format) | | Version | 5.0 | | Other names | Semantic Processed Document Format; Scholaris Processed Document Format (original name, versions 3.0 to 4.1); SPDF document | | PUID | to be assigned | | Format family | none | | Format type (classification) | Database (alternatively Text (Structured)) | | Extension(s) | spdf | | MIME / media type | application/vnd.spdf+sqlite3 (registration with IANA pending; defined in the official specification) | | Byte order | Big-endian (SQLite header integers); the format's own binary vector data is little-endian | | Disclosure | Open, fully documented; specification under CC BY 4.0 | | Developer | José Luis Saorín Ferrer, editor of the SPDF specification | | Support | The SPDF project, , contact jl@joseluissaorin.com | | Released | 5.0 working draft, 2026-10-07 | ## Description SPDF (Semantic Processed Document Format) is an open file format for documents that have been read once, by optical character recognition, a text layer, a speech recogniser or by hand, and stored so that every passage can be cited with its exact location: printed page or folio, leaf or column, verse, canonical reference, second of a recording, slide or spreadsheet rows. An SPDF file is a SQLite 3 database holding exactly one document: its bibliographic record as a CSL-JSON item, its citable units (pages, time spans, slides, sections, sheets) with their text in Unicode NFC, passages indexed for full-text search with the SQLite FTS5 module, optional section structure, figures, embedding vectors for semantic search, provenance records, and optionally the original file and page images as embedded binary objects. The SQLite header identifies the format: the `application_id` field at offset 68 holds the ASCII bytes "SPDF" and the `user_version` field at offset 60 holds the version (500 for version 5.0). The format was created by José Luis Saorín Ferrer for the Scholaris citation application, where it was called "Scholaris Processed Document Format". Versions 3.0 to 4.1 were internal to Scholaris and were stored compressed with gzip (version 3.0 as a gzip-compressed SQLite database with `metadata` and `chunks` tables; 4.0 and 4.1 as a gzip-compressed SQLite database with Spanish table names); those files also use the `.spdf` extension. Version 5.0 (October 2026) is the first public version: an uncompressed SQLite database with English identifiers, readable directly by any SQLite tool. Readers of version 5.0 are required to read the legacy versions 4.0 and 4.1 as well. Files may be accompanied by two kinds of JSON sidecar files, which are not SPDF files: `.spdfa.json` (user annotations in W3C Web Annotation format) and `.spdfl.json` (collection manifests). The format is used for scholarly reading, citation and search of books, articles, manuscripts, recordings, slides and web pages. ## Internal signatures ### Signature 1: SPDF 5.0 | Field | Value | |---|---| | Signature name | SPDF 5.0 | | Position type | Absolute from BOF | | Offset | 0 | | Max offset | 0 | | Value | `53514C69746520666F726D6174203300{44}000001F4{4}53504446` | Description: BOF, offset 0: 'SQLite format 3' followed by 0x00 (`53514C69746520666F726D6174203300`, 16 bytes), then a gap of 44 bytes, then at offset 60 the SQLite `user_version` 500 as a big-endian integer (`000001F4`), then a gap of 4 bytes, then at offset 68 the SQLite `application_id` 'SPDF' (`53504446`, 0x53504446 = 1397769286). The same signature as three byte sequences, if preferred: | Position type | Offset | Max offset | Value | |---|---|---|---| | Absolute from BOF | 0 | 0 | `53514C69746520666F726D6174203300` | | Absolute from BOF | 60 | 60 | `000001F4` | | Absolute from BOF | 68 | 68 | `53504446` | Checked against the seven version 5.0 files of the SPDF conformance suite: bytes 60 to 71 are `00 00 01 F4 00 00 00 00 53 50 44 46` in all of them. ### Signature 2 (optional): SPDF, any version from 5.0 If PRONOM prefers one record for all 5.x versions, or as a fallback for future minor versions, the `application_id` alone identifies the format, in the same way as the records for OGC GeoPackage (fmt/1700) and Audacity Project File 3.x (fmt/1826): | Position type | Offset | Max offset | Value | |---|---|---|---| | Absolute from BOF | 0 | 0 | `53514C69746520666F726D6174203300{52}53504446` | Future minor versions change only the `user_version` value: 5.1 is 510 (`000001FE`), 5.2 is 520 (`00000208`), and so on (major × 100 + minor × 10). ### Legacy versions 4.0 and 4.1 Legacy files begin with the gzip header (`1F8B08`) and are identified by DROID as GZIP Format (x-fmt/266). The SPDF database is only visible after decompression, where it has `application_id` 0 and `user_version` 410 (4.1), 400 or 0 (4.0), with tables named `spdf`, `documentos`, `unidades` and `fragmentos`. A `user_version` value on its own is not a reliable signature (other SQLite applications use values such as 400), so no signature is proposed for them. If PRONOM records them, version 5.0 would be related to them as "Is subsequent version of" SPDF 4.1; their signatures cannot collide with version 5.0, so no priority relationship is needed. ## External signature | Type | Value | |---|---| | File extension | spdf | ## Relationships | Relationship | Related format | |---|---| | Has priority over | SQLite Database File Format, version 3 (fmt/729) | | Is subtype of | SQLite Database File Format, version 3 (fmt/729) | ## Documentation - SPDF: Semantic Processed Document Format, version 5.0. José Luis Saorín Ferrer, 2026. . Licensed under CC BY 4.0. - SPDF conformance suite, version 0.1.0 (220 cases). - SQLite Database File Format. ## Samples for testing the signature Sample files built from short excerpts of public-domain works are part of the SPDF conformance suite in : - `conformance/files/*.spdf`: seven version 5.0 files (Cervantes, the *Lazarillo*, Bécquer, Hooke, Darwin, the *Analects*, NASA air-to-ground transmissions); - `conformance/legacy/*.spdf`: two legacy files, versions 4.1 (Garcilaso) and 4.0 (J. F. Kennedy), gzip-compressed; - `conformance/invalid/*.spdf`: files that each break one rule of the specification, useful as near misses (for example a version 5.0 database wrapped in gzip, or a SQLite database with an unknown `application_id`). ## Vendor and support José Luis Saorín Ferrer, editor of the SPDF specification. Contact: jl@joseluissaorin.com. ## Submitted by José Luis Saorín Ferrer, SPDF project. --- # application_id 0x53504446 ("SPDF") for SQLite's magic.txt and for file(1) URL: https://spdf.joseluissaorin.com/governance/drafts/sqlite-magic-entry > Borrador sin enviar. Dos propuestas para que las herramientas reconozcan un SPDF sin abrirlo: una línea en magic.txt, el fichero del árbol de fuentes de SQLite que lista los applicationid conocidos (la documentación… > **Borrador sin enviar.** Dos propuestas para que las herramientas reconozcan un SPDF > sin abrirlo: una línea en `magic.txt`, el fichero del árbol de fuentes de SQLite que > lista los `application_id` conocidos (la documentación del formato de fichero de > SQLite remite a él), y las reglas equivalentes en el fichero mágico de file(1) y > libmagic, que es lo que de verdad usan los sistemas operativos. > > Vías: > > 1. **SQLite.** El registro del sufijo `+sqlite3` en IANA pide añadir el valor a > `magic.txt` enviando un parche a la lista sqlite-users. Según tengo entendido esa > lista se cerró y la sustituyó el foro de SQLite (); por > verificar si basta un mensaje en el foro o si hay que escribir a los desarrolladores. > SQLite no acepta parches de código de terceros, pero `magic.txt` es un fichero de > texto que mantienen ellos: lo normal es pedir la línea y dejar que la añadan. > 2. **file(1) y libmagic.** El proyecto lo mantiene Christos Zoulas. Vías publicadas en > su README: el gestor de fallos y la lista > ; el código público está en GitHub (`file/file`), pero por > verificar si allí aceptan *pull requests* o solo las vías anteriores. Conviene > adjuntar un fichero de muestra (los de `conformance/files/` son de dominio público) y, > si lo aceptan, añadirlo también a su batería de pruebas (`file/file-tests`). > > Falta antes de enviarlas: > > - Que la especificación tenga una versión estable. La URL que citan los comentarios > de la regla (`https://spdf.joseluissaorin.com/spec`) ya responde (comprobado el > 7-10-2026), pero sirve un borrador de trabajo. > - Una URL pública para la muestra que cita el mensaje del foro. > - Decidir el nombre que se imprime: aquí «SPDF document», con el nombre completo en > `magic.txt` como pide la propuesta inicial; los demás valores del fichero usan > nombres cortos. > - El tipo de medio: libmagic suele usar tipos registrados o `x-` y puede pedir que > `application/vnd.spdf+sqlite3` esté ya en IANA (ver `iana-media-type.md`); si no, > la alternativa es dejar `application/vnd.sqlite3` y añadir solo el nombre. Por > verificar con el mantenedor. > > Probado el 7-10-2026 con file 5.41 (el de macOS) cargando las reglas con la variable > `MAGIC`: los siete ficheros 5.0 de la batería se reconocen, y un SQLite con otro > `application_id` (`invalid/E002-application-id.spdf`) no. Los diffs están hechos contra > las versiones de `magic.txt` (rama `master` del espejo de SQLite en GitHub) y de > `magic/Magdir/sql` (revisión 1.38) descargadas ese día. --- ## 1. SQLite `magic.txt` SQLite's documentation of the database header says that the `application_id` at offset 68 lets utilities such as file(1) identify application file formats, and that the list of assigned values is the `magic.txt` file in the SQLite source tree. SPDF uses `PRAGMA application_id = 1397769286` (0x53504446, ASCII "SPDF") and stores the specification version in `PRAGMA user_version` (major × 100 + minor × 10; 500 for version 5.0). ### Proposed lines ```text >68 belong =0x53504446 SPDF document (Semantic Processed Document Format) >>60 belong x \b, user_version %d - ``` The second line prints the format version, which SPDF keeps in `user_version`. Output with the current `magic.txt`: ```text quijote.spdf: SPDF document (Semantic Processed Document Format), user_version 500 - SQLite3 database ``` If a single line in the style of the other entries is preferred: ```text >68 belong =0x53504446 SPDF document - ``` which prints `SPDF document - SQLite3 database`. ### As a patch ```diff --- a/magic.txt +++ b/magic.txt @@ -30,4 +30,6 @@ >68 belong =0x45737269 Esri Spatially-Enabled Database - >68 belong =0x4d504258 MBTiles tileset - >68 belong =0x6a035744 TeXnicard card database +>68 belong =0x53504446 SPDF document (Semantic Processed Document Format) +>>60 belong x \b, user_version %d - >0 string =SQLite SQLite3 database ``` ### Message (forum post) > **Subject:** application_id 0x53504446 for SPDF, for magic.txt > > SPDF (Semantic Processed Document Format) is an open file format for documents that > have been read once (OCR, text layer or speech recognition) and stored so that every > passage can be cited with its exact page, folio, verse or time. Each file is a single > uncompressed SQLite 3 database. The specification (CC BY 4.0) is at > https://spdf.joseluissaorin.com/spec. > > SPDF files set `PRAGMA application_id = 1397769286` (0x53504446, "SPDF" in ASCII) and > keep the format version in `user_version` (500 for version 5.0). Could this value be > added to the list in magic.txt? Suggested lines: > > ``` > >68 belong =0x53504446 SPDF document (Semantic Processed Document Format) > >>60 belong x \b, user_version %d - > ``` > > A public-domain sample file is available at [URL of a sample]. Thank you for SQLite, > and for the application_id mechanism in particular. > > José Luis Saorín Ferrer, editor of the SPDF specification ## 2. file(1) and libmagic: `magic/Magdir/sql` In libmagic the SQLite 3 rules live in `magic/Magdir/sql`. Known application ids appear twice: once to set the media type and extensions, once to print the name. The generic rules that follow already print the user version, so no extra rule is needed for it. ### Patch ```diff --- a/magic/Magdir/sql +++ b/magic/Magdir/sql @@ -196,6 +196,13 @@ >>68 belong =0x4D504258 database !:mime application/vnd.sqlite3 !:ext mbtiles +# URL: https://spdf.joseluissaorin.com +# Reference: https://spdf.joseluissaorin.com/spec +# Note: SPDF, Semantic Processed Document Format, with application id 53504446h "SPDF" +# and the version as user version (500 for 5.0) +>>68 belong =0x53504446 database +!:mime application/vnd.spdf+sqlite3 +!:ext spdf >>68 default x database !:mime application/vnd.sqlite3 # no examples found with s3db sl3 suffix @@ -228,6 +235,7 @@ >>>68 belong =0x544f4952 (Riot Games patcher) >>>68 belong =0x544d5052 (Twinmotion project) >>>68 belong =0x43484557 (Chewing IME) +>>>68 belong =0x53504446 (SPDF document) # unknown application ID >>>68 default x >>>>68 belong !0 \b, application id %u ``` ### Result ```text $ file darwin.spdf darwin.spdf: SQLite 3.x database (SPDF document), user version 500 (0x1f4), last written using SQLite version … $ file --mime-type darwin.spdf darwin.spdf: application/vnd.spdf+sqlite3 $ file --extension darwin.spdf darwin.spdf: spdf ``` ### Message (bug tracker or mailing list) > **Subject:** magic: SQLite application id 0x53504446 (SPDF document) > > The attached patch to `magic/Magdir/sql` recognizes SPDF files (Semantic Processed > Document Format), SQLite 3 databases with `application_id` 0x53504446 ("SPDF"), and > gives them the media type `application/vnd.spdf+sqlite3` and the extension `spdf`. The format > keeps its version in `user_version`, which the existing rules already print. > Specification: https://spdf.joseluissaorin.com/spec. A public-domain sample is > attached; it can go into file-tests if useful. > > José Luis Saorín Ferrer, editor of the SPDF specification ## Legacy files SPDF 4.0 and 4.1 files (written by the Scholaris application before version 5.0) are SQLite databases compressed with gzip and have `application_id` 0. file(1) reports them as gzip data, and with `-z` as a SQLite 3.x database with user version 410 or 400 (or 0). No rule is proposed for them: a user version alone is not a reliable signature. --- # URI scheme registration (provisional): spdf URL: https://spdf.joseluissaorin.com/governance/drafts/uri-scheme-spdf > Borrador sin enviar. Solicitud de registro provisional del esquema de URI spdf en el registro de esquemas de URI de IANA, con la plantilla de la sección 7.4 de la RFC 7595. Los registros provisionales solo exigen que… > **Borrador sin enviar.** Solicitud de registro provisional del esquema de URI `spdf` > en el registro de esquemas de URI de IANA, con la plantilla de la sección 7.4 de la > RFC 7595. Los registros provisionales solo exigen que el nombre no esté ya registrado > y que la solicitud esté completa; no piden una especificación estable. La RFC 7595 > recomienda anunciar la propuesta en la lista uri-review@ietf.org para recibir > comentarios antes o a la vez que se envía a IANA. > > Vía: el formulario de IANA para esquemas de URI o un correo a iana@iana.org con esta > plantilla (por verificar cuál prefiere IANA hoy). > > Falta antes de enviarla: > > 1. Comprobar que `spdf` no figura ya en el registro de esquemas de URI (por > verificar el 7-10-2026 no se pudo; consultar > ). > 2. Enlazar una versión fechada de la especificación (etiqueta `spec-v5.0.0`) en lugar > del borrador de trabajo. > 3. Confirmar el responsable del cambio, como en el registro del tipo de medio > (`iana-media-type.md`). > > Cuando la 5.0 sea final y haya implementaciones independientes en uso, se puede pedir > el paso a registro permanente, que exige revisión de experto y una especificación > estable. --- **Scheme name:** spdf **Status:** Provisional **Applications/protocols that use this scheme name:** Applications that read, cite or annotate documents in the SPDF format (Semantic Processed Document Format): the SPDF libraries (Rust, TypeScript, Python, Swift, Kotlin/JVM, Go, C#, PHP, Ruby, R, Julia), the SPDF Reader, the Scholaris citation application and annotation sidecars (`*.spdfa.json`, W3C Web Annotation), which use `spdf` URIs as annotation targets. **Contact:** José Luis Saorín Ferrer **Change controller:** José Luis Saorín Ferrer , editor of the SPDF specification. **References:** SPDF: Semantic Processed Document Format, version 5.0, section 5 ("Anchor URI"). ; source: . **Scheme syntax:** ```abnf spdf-uri = "spdf:" docref [ "#" params ] docref = hash-ref / id-ref hash-ref = "sha256-" 64lhex ; lowercase hex SHA-256 of the source document id-ref = 1*( unreserved / pct-encoded ) ``` `params` is the list of anchor parameters defined in the specification (physical page, printed folio, time range, section path, paragraph, slide, sheet rows, verse, canonical reference, character range, region). The time and region parameters use the syntax of W3C Media Fragments URI 1.0 (`t=4160,4175.5`, `xywh=percent:10,20,30,10`) and the character range that of RFC 5147 (`char=118,301`). Example: ```text spdf:sha256-27ea8a4dd0bbf0a511246fe67f81c19084aa2da75c871fe7acde688fd182cb60#p=29&pe=30&f=Ir&fe=Iv&char=729,745 ``` **Scheme semantics:** An `spdf` URI names a document, independently of where copies of it are stored, by the SHA-256 of the bytes of its original (the scanned book, the recording, the PDF), and optionally a place in it: a page, a leaf, a time span, a verse, a passage. It is a name, not a locator: there is no network protocol and no resolution service. An application resolves it against the SPDF files it holds whose `source_sha256` matches. The document reference may instead be a document identifier local to a file; such URIs are only meaningful next to that file. **Encoding considerations:** Values are UTF-8 strings percent-encoded as in RFC 3986; every byte other than an unreserved character is percent-encoded with uppercase hexadecimal digits in the canonical form. Parsers also accept lowercase hexadecimal digits and unencoded non-ASCII characters (IRI form, RFC 3987). **Interoperability considerations:** The same parameter list is the fragment identifier of `application/vnd.spdf+sqlite3` resources (`https://example.org/x.spdf#p=5&f=1r`). A canonical form, the order of parameters and round-trip rules are defined by the specification and tested by its public conformance suite. **Security considerations:** An `spdf` URI reveals which document, and which passage of it, a user is reading: applications should not send such URIs to third parties without consent. Because the URI identifies a document by a hash, it cannot be used to retrieve the document, and an application must not fetch anything from the network to resolve it. Parsers must reject malformed percent-encoding and invalid UTF-8, and must not use decoded values as file paths or as query syntax. --- # Citable Processed Documents Community Group Charter URL: https://spdf.joseluissaorin.com/governance/drafts/w3c-community-group-charter > Borrador sin enviar. Propuesta de carta (charter) para un Community Group de W3C que se ocupe de SPDF y de su modelo de anclas. Sigue la plantilla oficial de cartas de Community Group (w3c/cg-charter, comprobada el… > **Borrador sin enviar.** Propuesta de carta (*charter*) para un Community Group de > W3C que se ocupe de SPDF y de su modelo de anclas. Sigue la plantilla oficial de > cartas de Community Group (`w3c/cg-charter`, comprobada el 7-10-2026); los apartados > de proceso, contribución, transparencia, elección de presidentes y enmiendas son su > texto, adaptado lo mínimo. > > Vía: se propone un grupo nuevo desde > con una cuenta de W3C (gratuita; no hace falta ser miembro de W3C ni pagar). Según el > proceso de los Community Groups, la propuesta está completa con un nombre que no use > otro grupo, una descripción del alcance y el apoyo de cinco personas; entonces W3C > anuncia el grupo. La carta se aprueba después, dentro del grupo. Por verificar si > W3C pide ahora redactar las cartas en el repositorio `w3c-cg/charter-proposals`. > > Falta antes de proponerlo: > > 1. **Nombre.** Aquí «Citable Processed Documents Community Group», que describe el > problema y no el producto; la alternativa es «SPDF Community Group». Comprobar que > no existe ya un grupo con ese nombre. > 2. **Cinco apoyos**, mejor de organizaciones distintas (implementadores, bibliotecas, > archivos, editores de TEI o IIIF), y un **segundo presidente** ajeno al proyecto. > 3. **Licencias.** Las contribuciones a las especificaciones del grupo quedan bajo el > CLA de W3C y los informes finales, bajo su acuerdo de especificación final (FSA). > La especificación SPDF actual es CC BY 4.0, que permite aportarla, pero conviene > confirmar con el Community Development Lead de W3C cómo conviven las dos licencias > y si el repositorio puede seguir con MIT OR Apache-2.0 para el código y la batería. > 4. Decidir por RFC, según `governance/GOVERNANCE.md`, si la especificación pasa a > desarrollarse en el grupo o el grupo solo la revisa; y que el repositorio sea > público. --- - **Status:** draft. This charter is a work in progress. To submit feedback, please use the issues of the repository where it is being developed: [repository URL]. - **This charter:** [URI] - **Previous charter:** none - **Start date:** [date the charter takes effect] - **Last modified:** 2026-10-07 ## Goals The mission of the Citable Processed Documents Community Group is to develop open, royalty-free specifications that let a document be **read once and cited forever**: a portable file format that keeps the text read from a book, scan, recording, slide deck, spreadsheet or web page together with the exact location of every passage in its source (printed page and folio, leaf, column, verse, canonical reference, time, slide), so that people, software and language-model agents can quote and cite it without approximation. The group starts from SPDF (Semantic Processed Document Format), version 5.0, an open specification with a conformance suite and implementations in several programming languages, and aims to make it a shared community specification, aligned with the W3C and other standards that already address annotation, fragments, bibliographic metadata and digital editions. ## Scope of work The group will work on: - **The SPDF file format**: the container (a single SQLite 3 database), its schema, the CSL-JSON metadata profile, text normalization and offsets, rights and provenance information, vector spaces for semantic search, integrity and signatures, profiles, extensions, and the reading of earlier versions. - **The anchor model and the anchor URI**: how a location in a document is described (pages and folios, leaves and columns, time ranges, sections and paragraphs, slides, spreadsheet rows, verses, canonical references, character ranges and regions) and how it is written as a URI, keeping the syntax of W3C Media Fragments and RFC 5147 where they overlap. - **Annotation and collection sidecars**: annotation files that use the W3C Web Annotation Data Model with an SPDF anchor selector, and collection manifests that list documents by hash. - **Reference behaviour that must be identical across implementations**: canonical serialization, validation, reference search, short citations and exports. - **A conformance test suite** for all of the above. - **Mappings** between SPDF anchors and the location models of TEI, IIIF, ALTO, CTS and W3C Web Annotation. Key use cases: verifiable citation in the humanities and social sciences; reading and citing early printed books, manuscripts and classical texts by folio, verse or canonical reference; citing oral history and recorded lectures to the second; archives and libraries publishing searchable, citable derivatives of their holdings; language-model agents that must quote a source and say exactly where. ## Out of scope - Optical character recognition, speech recognition, embedding models and ranking methods beyond the reference behaviour that makes implementations interoperable. - Specific products: readers, editors, producers or services. - Citation styles, which belong to the Citation Style Language project; bibliographic vocabularies beyond the CSL-JSON profile. - Changes to the Web Annotation, Media Fragments, IIIF, TEI, ALTO or CSL specifications themselves. The group may send them comments and requests. - Encryption, access control and digital rights management. - Protocols for storing, synchronizing or serving documents. ## Deliverables ### Specifications - **SPDF: Semantic Processed Document Format.** The file format, its anchor model and anchor URI, starting from version 5.0. Estimated: a first Community Group Report within 12 months of the group's start. - **SPDF annotation and collection sidecars.** The `.spdfa.json` profile of W3C Web Annotation with the SPDF anchor selector, and the `.spdfl.json` collection manifest. This may be published as part of the format specification or as a separate report. ### Non-normative reports The group may produce other Community Group Reports within the scope of this charter that are not Specifications, for instance: - Use cases and requirements for citable processed documents. - Mapping notes between SPDF anchors and TEI, IIIF, ALTO, CTS and CSL locators. - Implementation reports based on the conformance suite. ### Test suites and other software The group MAY produce test suites to support the Specifications. The SPDF conformance suite already exists and will be maintained by the group. Please see the GitHub LICENSE file for test suite contribution licensing information. ## Dependencies or liaisons - W3C *Web Annotation Data Model* and *Vocabulary* (Recommendations, 2017): annotation sidecars and selectors. - W3C *Media Fragments URI 1.0 (basic)* (Recommendation, 2012): time and spatial parameters of the anchor URI. - IETF: RFC 5147 (character ranges), RFC 8785 (JSON canonicalization), RFC 8032 (Ed25519), and IANA registrations for the media type `application/vnd.spdf+sqlite3` and the `spdf` URI scheme. - IIIF Consortium: Presentation API 3.0 and the region syntax of the Image API, for exports and mappings. - TEI Consortium: TEI P5, for exports and for the citation practices of digital editions. - Citation Style Language project: the CSL-JSON schema used for metadata. - Library of Congress: ALTO. - The CITE architecture: CTS URNs for canonical references. - SQLite: the file format and the `application_id` registry in its source tree. - Digital preservation registries: PRONOM (The National Archives, UK) and the Library of Congress *Sustainability of Digital Formats*. ## Community and Business Group Process The group operates under the Community and Business Group Process. Terms in this Charter that conflict with those of the Community and Business Group Process are void. As with other Community Groups, W3C seeks organizational licensing commitments under the W3C Community Contributor License Agreement (CLA). When people request to participate without representing their organization's legal interests, W3C will in general approve those requests for this group with the following understanding: W3C will seek and expect an organizational commitment under the CLA starting with the individual's first request to make a contribution to a group Deliverable. The section on Contribution Mechanics describes how W3C expects to monitor these contribution requests. The W3C Code of Conduct and the W3C Antitrust and competition policy apply to participation in this group. ## Work limited to charter scope The group will not publish Specifications on topics other than those listed under Specifications above. See below for how to modify the charter. ## Contribution mechanics Substantive Contributions to Specifications can only be made by Community Group Participants who have agreed to the W3C Community Contributor License Agreement (CLA). Reports other than Specifications published by this group should use the W3C Software and Document License where possible. Community Group participants agree to make all contributions in the GitHub repository the group is using for the particular document. This may be in the form of a pull request (preferred), by raising an issue, or by adding a comment to an existing issue. All GitHub repositories attached to the Community Group must contain a copy of the CONTRIBUTING and LICENSE files. ## Transparency The group will conduct all of its technical work in public. All technical work will occur in its GitHub repositories (and not in mailing list discussions). This is to ensure contributions can be tracked through a software tool. Meetings may be restricted to Community Group participants, but a public summary or minutes must be posted to a GitHub issue. ## Decision process This group will seek to make decisions where there is consensus. Normative changes to the Specifications follow a request-for-comments process: a proposal is discussed in public for at least 14 days, followed by a 7-day final comment period, and the Chairs assess consensus. **An accepted proposal lands with at least one conformance case; it is considered implemented when two independent implementations pass it.** A Specification is published as a final Community Group Report only with implemented proposals. Where consensus is not clear, the Chairs may issue a Call for Consensus to allow multi-day online feedback for a proposed course of action. It is expected that participants can earn Committer status through a history of valuable contributions, as is common in open source projects. After discussion and due consideration of different opinions, a decision should be publicly recorded as the resolution of a GitHub issue. If substantial disagreement remains (e.g., the group is divided) and the group needs to decide an Issue in order to continue to make progress, the Committers will choose an alternative that had substantial support (with a vote of Committers if necessary). Individuals who disagree with the choice are strongly encouraged to take ownership of their objection by taking ownership of an alternative fork. This is explicitly allowed (and preferred to blocking progress) to let implementation experience inform which spec is ultimately chosen by the group to move ahead with. Any decisions reached at any meeting are tentative and should be recorded in a GitHub Issue. Any group participant may object to a decision reached at an online or in-person meeting within 7 days of publication of the decision provided that they include clear technical reasons for their objection. The Chairs will facilitate discussion to try to resolve the objection according to this decision process. It is the Chairs' responsibility to ensure that the decision process is fair, respects the consensus of the CG, and does not unreasonably favor or discriminate against any group participant or their employer. ## Chair selection Participants in this group choose their Chair(s) and can replace their Chair(s) at any time using whatever means they prefer. However, if 5 participants, no two from the same organization, call for an election, the group must use the following process to replace any current Chair(s) with a new Chair, consulting the Community Development Lead on election operations (e.g., voting infrastructure and using RFC 3797). - Participants announce their candidacies. Participants have 14 days to announce their candidacies, but this period ends as soon as all participants have announced their intentions. If there is only one candidate, that person becomes the Chair. If there are two or more candidates, there is a vote. Otherwise, nothing changes. - Participants vote. Participants have 21 days to vote for a single candidate, but this period ends as soon as all participants have voted. The individual who receives the most votes, no two from the same organization, is elected chair. In case of a tie, RFC 3797 is used to break the tie. An elected Chair may appoint co-Chairs. Participants dissatisfied with the outcome of an election may ask the Community Development Lead to intervene. The Community Development Lead, after evaluating the election, may take any action including no action. Proposed initial Chairs: José Luis Saorín Ferrer (editor of SPDF), and a second Chair from another organization, to be identified before the group is proposed. ## Amendments to this charter The group can decide to work on a proposed amended charter, editing the text using the Decision Process described above. The decision on whether to adopt the amended charter is made by conducting a 30-day vote on the proposed new charter. The new charter, if approved, takes effect on either the proposed date in the charter itself, or 7 days after the result of the election is announced, whichever is later. A new charter must receive 2/3 of the votes cast in the approval vote to pass. The group may make simple corrections to the charter such as deliverable dates by the simpler group decision process rather than this charter amendment process. The group will use the amendment process for any substantive changes to the goals, scope, deliverables, decision process or rules for amending the charter. ## Licensing - Specifications: contributions under the W3C Community Contributor License Agreement (CLA); final reports under the W3C Community Final Specification Agreement (FSA). The SPDF specification brought to the group is already published under CC BY 4.0, and that publication remains available under its licence. - Other reports: the W3C Software and Document License where possible. - Test suites and software: as stated in the LICENSE file of the repository (currently MIT OR Apache-2.0). - The editor of SPDF and its contributors have committed not to assert patents against implementations of the specification; the CLA and the FSA add W3C's royalty-free patent commitments for contributions made in the group. --- # Contributing to SPDF URL: https://spdf.joseluissaorin.com/governance/contributing > SPDF is three things that change in different ways: a specification, a conformance suite that turns the specification into checkable cases, and many implementations that must agree with both. This guide says how to… SPDF is three things that change in different ways: a **specification**, a **conformance suite** that turns the specification into checkable cases, and many **implementations** that must agree with both. This guide says how to contribute to each. Everyone who takes part follows the [code of conduct](CODE_OF_CONDUCT.md). Security problems are reported privately, as [`SECURITY.md`](SECURITY.md) explains, never in a public issue. The repository is private until the specification, the conformance suite and the first-tier libraries pass; until then, contributions come from invited collaborators, and comments on the specification are welcome by email at . ## Contributing to the specification The normative text is [`spec/SPEC.md`](spec/SPEC.md), in English, with a faithful Spanish translation in [`spec/SPEC.es.md`](spec/SPEC.es.md). If they differ, the English text prevails. - **Editorial changes** (typos, clearer wording, examples, diagrams, translation fixes) that would not make any implementation change: open a pull request labelled `editorial`. Change both languages when the fix applies to both, or say in the pull request that the other language needs the same fix. - **Normative changes** (anything that changes what a valid file is, what a reader, writer or validator must do, or what a function returns) go through an RFC: copy [`spec/rfcs/0000-template.md`](spec/rfcs/0000-template.md) and follow [`governance/RFC-PROCESS.md`](governance/RFC-PROCESS.md). **An accepted RFC lands with at least one conformance case; it is considered implemented when two independent implementations pass it.** - **Questions and ambiguities** are contributions too. If two readings of the text are possible, open an issue with both, and an example file or case if you can. Do not settle it by copying what another implementation does. - Use the BCP 14 key words (MUST, SHOULD, MAY) only where they are meant, and in capitals. How versions change, and what a minor version may add, is in [`governance/VERSIONING.md`](governance/VERSIONING.md). ## Contributing to the conformance suite The suite in [`conformance/`](conformance/) is normative for behaviour. Its layout, case format and runner protocol are in [`conformance/README.md`](conformance/README.md). - **Only public-domain texts whose status anyone can verify.** Name the work, the edition and the reason it is in the public domain (for example the author's death date, or a government work). Layouts, folios, timings and hashes that are synthetic must say so in the document's CSL `note`. Never use personal data. - **Expectations come from an oracle, not from an implementation**: SQLite running the reference SQL, exact arithmetic, the reference oracle `tools/spdfref.py`, or cases written and reviewed by hand in `tools/manual/`. A case that only records what one library returns is not accepted. - After editing a source or a manual case, run `python3 conformance/tools/generar.py --sellar` and then `python3 conformance/tools/verificar.py` (Python 3.13, standard library only), and commit the regenerated files. - Add a line to [`conformance/CHANGELOG.md`](conformance/CHANGELOG.md) and bump the suite version as `governance/VERSIONING.md` says. - Case ids are never renamed or reused. A wrong case is fixed in place and the change is logged; a removed case keeps its id retired. - A new kind of case, or a case that changes what the specification requires, needs an RFC. ## Contributing to an implementation Each implementation lives in its own folder (`rust/`, `js/`, `python/`, `go/`, `swift/`, `kotlin/`, `dotnet/`, `php/`, `ruby/`, `r/`, `julia/`, `c/`) with its own CI workflow, `.github/workflows/.yml`. The same holds for `producer/`, `reader/`, `site/`, `integrations/` and `models/`. - **One folder per pull request.** Do not change another implementation's folder or workflow in the same pull request; if you find a bug in another implementation, open an issue or a separate pull request for it. - **Follow the specification, not another implementation**, including the reference one. When an implementation and the suite disagree, the suite wins until an RFC says otherwise. - **Run the conformance runner.** Every implementation ships a runner that discovers `conformance/cases/*.json` and prints the report described in `conformance/README.md`. A change must not turn a passing case into a failing one. CI uploads the report as the artifact `conformance-`. - **Keep the safety rules**: safe opening, size limits and the other requirements of [SPEC §2.4](spec/SPEC.md#container) and [§14](spec/SPEC.md#security) are not optional, even in tests and examples. - Follow the conventions of the language: its formatter, linter and test framework, as the folder's README describes. Add tests with the change. - New dependencies need a reason in the pull request. Prefer what the platform already provides; SQLite comes first. - Models are downloaded on demand and never committed or packaged. ## Pull requests - Keep them small and about one thing. Explain the why, not only the what. - Link the RFC, issue or conformance case the change relates to. - CI must pass, including the conformance runner of the folder you touched. - Before pushing, rebase on the current `main` (`git pull --rebase`). - Never commit secrets, credentials, API keys or personal data. ## Commit messages - Subject line: `: `, where `` is the top-level folder (`spec`, `conformance`, `governance`, `rust`, `js`, `python`, …) or `ci`, and the summary is in English or Spanish, at most 72 characters, without a final period. - A body, separated by a blank line, says why the change is needed when that is not obvious. - Add `Co-authored-by` lines for every co-author. A `Signed-off-by` line is welcome but not required. - Contributions prepared with the help of AI tools are welcome. The person who submits them is responsible for them and has reviewed them; say in the pull request which tools were used. ## Licensing of contributions Contributions are licensed under the licence of the part of the repository they change: | Part | Licence | |---|---| | Specification and documentation (`spec/`, `governance/`, prose in every folder) | [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/) | | Code (implementations, producer, reader, conformance suite, website, integrations) | MIT OR Apache-2.0, at the user's choice ([`LICENSE-MIT`](LICENSE-MIT), [`LICENSE-APACHE`](LICENSE-APACHE)) | The rule is **inbound = outbound**: by submitting a contribution you license it under the same licence as the part of the repository it changes, and you confirm that you wrote it or otherwise have the right to submit it under that licence. No contributor licence agreement and no Developer Certificate of Origin sign-off are required for now; if the project adopts one later, it will be announced through the RFC process and will not apply retroactively. Contributors to the specification also make the patent commitment described in [`governance/GOVERNANCE.md`](governance/GOVERNANCE.md#licences-and-patents): they will not assert patents they own or control against implementations of SPDF. ## Languages - The normative specification is written in English, with a faithful Spanish translation maintained alongside it. - Issues, pull requests, reviews and discussions may be in English or Spanish. - Documentation for people is published in English and, where possible, in Spanish. Spanish text is written with correct spelling, accents and `ñ` included; identifiers in code, schemas and JSON keys stay in English without accents. ## Reporting security problems Do not report vulnerabilities in public issues. Follow [`SECURITY.md`](SECURITY.md): email or, once the repository is public, use GitHub's private vulnerability reporting. ## En español Las contribuciones en español son igual de bienvenidas. La especificación normativa está en inglés y tiene una traducción española fiel en `spec/SPEC.es.md`; si se contradicen, manda el inglés. Las erratas y mejoras de redacción se proponen con un *pull request* marcado `editorial`; todo cambio normativo pasa por una RFC (`spec/rfcs/`, proceso en `governance/RFC-PROCESS.md`), y una RFC aceptada entra con al menos un caso de conformidad y se da por implementada cuando dos implementaciones independientes lo pasan. Los casos de la batería solo usan obras de dominio público comprobable. Cada implementación vive en su carpeta, con su propio flujo de CI, y un *pull request* toca una sola carpeta. Las contribuciones se aceptan con la misma licencia con la que se publican (*inbound = outbound*), sin DCO ni acuerdo de contribución; la especificación y la documentación se publican con CC BY 4.0 y el código con MIT OR Apache-2.0. Los fallos de seguridad se comunican en privado, como explica `SECURITY.md`. --- # Contributor Covenant Code of Conduct URL: https://spdf.joseluissaorin.com/governance/code-of-conduct > We as members, contributors, and leaders pledge to make participation in our community a harassment-free experience for everyone, regardless of age, body size, visible or invisible disability, ethnicity, sex… ## Our Pledge We as members, contributors, and leaders pledge to make participation in our community a harassment-free experience for everyone, regardless of age, body size, visible or invisible disability, ethnicity, sex characteristics, gender identity and expression, level of experience, education, socio-economic status, nationality, personal appearance, race, caste, color, religion, or sexual identity and orientation. We pledge to act and interact in ways that contribute to an open, welcoming, diverse, inclusive, and healthy community. ## Our Standards Examples of behavior that contributes to a positive environment for our community include: * Demonstrating empathy and kindness toward other people * Being respectful of differing opinions, viewpoints, and experiences * Giving and gracefully accepting constructive feedback * Accepting responsibility and apologizing to those affected by our mistakes, and learning from the experience * Focusing on what is best not just for us as individuals, but for the overall community Examples of unacceptable behavior include: * The use of sexualized language or imagery, and sexual attention or advances of any kind * Trolling, insulting or derogatory comments, and personal or political attacks * Public or private harassment * Publishing others' private information, such as a physical or email address, without their explicit permission * Other conduct which could reasonably be considered inappropriate in a professional setting ## Enforcement Responsibilities Community leaders are responsible for clarifying and enforcing our standards of acceptable behavior and will take appropriate and fair corrective action in response to any behavior that they deem inappropriate, threatening, offensive, or harmful. Community leaders have the right and responsibility to remove, edit, or reject comments, commits, code, wiki edits, issues, and other contributions that are not aligned to this Code of Conduct, and will communicate reasons for moderation decisions when appropriate. ## Scope This Code of Conduct applies within all community spaces, and also applies when an individual is officially representing the community in public spaces. Examples of representing our community include using an official e-mail address, posting via an official social media account, or acting as an appointed representative at an online or offline event. ## Enforcement Instances of abusive, harassing, or otherwise unacceptable behavior may be reported to the community leaders responsible for enforcement at . All complaints will be reviewed and investigated promptly and fairly. All community leaders are obligated to respect the privacy and security of the reporter of any incident. ## Enforcement Guidelines Community leaders will follow these Community Impact Guidelines in determining the consequences for any action they deem in violation of this Code of Conduct: ### 1. Correction **Community Impact**: Use of inappropriate language or other behavior deemed unprofessional or unwelcome in the community. **Consequence**: A private, written warning from community leaders, providing clarity around the nature of the violation and an explanation of why the behavior was inappropriate. A public apology may be requested. ### 2. Warning **Community Impact**: A violation through a single incident or series of actions. **Consequence**: A warning with consequences for continued behavior. No interaction with the people involved, including unsolicited interaction with those enforcing the Code of Conduct, for a specified period of time. This includes avoiding interactions in community spaces as well as external channels like social media. Violating these terms may lead to a temporary or permanent ban. ### 3. Temporary Ban **Community Impact**: A serious violation of community standards, including sustained inappropriate behavior. **Consequence**: A temporary ban from any sort of interaction or public communication with the community for a specified period of time. No public or private interaction with the people involved, including unsolicited interaction with those enforcing the Code of Conduct, is allowed during this period. Violating these terms may lead to a permanent ban. ### 4. Permanent Ban **Community Impact**: Demonstrating a pattern of violation of community standards, including sustained inappropriate behavior, harassment of an individual, or aggression toward or disparagement of classes of individuals. **Consequence**: A permanent ban from any sort of public interaction within the community. ## Attribution This Code of Conduct is adapted from the [Contributor Covenant][homepage], version 2.1, available at [https://www.contributor-covenant.org/version/2/1/code_of_conduct.html][v2.1]. Community Impact Guidelines were inspired by [Mozilla's code of conduct enforcement ladder][Mozilla CoC]. For answers to common questions about this code of conduct, see the FAQ at [https://www.contributor-covenant.org/faq][FAQ]. Translations are available at [https://www.contributor-covenant.org/translations][translations]. [homepage]: https://www.contributor-covenant.org [v2.1]: https://www.contributor-covenant.org/version/2/1/code_of_conduct.html [Mozilla CoC]: https://github.com/mozilla/diversity [FAQ]: https://www.contributor-covenant.org/faq [translations]: https://www.contributor-covenant.org/translations --- # Security policy URL: https://spdf.joseluissaorin.com/governance/security > SPDF files come from strangers. Opening one means parsing an untrusted database with a complex engine, and the specification asks every reader to do it safely (SPEC §2.4, §14 and §15). If you find a way around that,… SPDF files come from strangers. Opening one means parsing an untrusted database with a complex engine, and the specification asks every reader to do it safely ([SPEC §2.4](spec/SPEC.md#container), [§14](spec/SPEC.md#security) and [§15](spec/SPEC.md#privacy)). If you find a way around that, or any other vulnerability in the specification or in the software in this repository, please tell us privately first. ## Supported versions | Component | Supported | |---|---| | Specification | 5.0 (working draft), including the reading of legacy 4.0 and 4.1 files that it requires | | Libraries (Rust, TypeScript, Python, Go, Swift, Kotlin/JVM, C#, PHP, Ruby, R, Julia, C ABI) | the latest release of the common release train; while the libraries are at 0.x, only the latest minor release gets fixes | | Reference producer `spdf build` | the latest release | | SPDF Reader (desktop, mobile and web) | the latest release | | Website, browser validator and conformance suite | the current version on `main` | Earlier releases do not receive fixes; upgrading is the remedy. ## How to report Report vulnerabilities privately, by either of these channels: - **Email** to , with `[SPDF security]` at the start of the subject. If you need an encrypted channel, say so in a first message without details and we will agree on one. - **GitHub private vulnerability reporting** ("Report a vulnerability" in the Security tab of ), once the repository is public. Please do not open public issues, pull requests or discussions about a vulnerability before it is fixed and disclosed. A useful report includes: the component and version (or commit), the platform and the SQLite version, a description of the impact, and a minimal reproduction. A crafted `.spdf` file is the best reproduction; please build it from public-domain or synthetic content, never from personal data. ## What happens next | Step | Deadline | |---|---| | Acknowledgement of your report | within 72 hours | | First assessment (confirmed or not, severity, components affected) | within 14 days | | Coordinated public disclosure | within 90 days of the report | - We keep you informed while we work on a fix and agree the disclosure date with you. Disclosure may come earlier when a fix is released, and later only by mutual agreement, or if a fix needs a change to the specification that cannot be made safely in time. - If the vulnerability is being exploited, we may disclose sooner, with mitigations. - Fixes to the specification that cannot wait for the normal discussion periods follow the shortened procedure in [`governance/RFC-PROCESS.md`](governance/RFC-PROCESS.md). - Once the repository is public, advisories are published through GitHub Security Advisories, which can assign a CVE identifier. - We credit reporters in the advisory, unless you prefer to remain anonymous. ## In scope - **Files that escape safe opening**: an SPDF or legacy file that makes a conforming reader execute SQL from the file (through triggers, views, virtual tables or schema tricks), load an extension, write to the file or elsewhere on disk, read outside the file, or keep running without bound despite the limits the specification requires. - **Memory-safety and parsing bugs** in readers and validators: overflows, out-of-bounds reads, crashes or hangs triggered by a file, including in gzip decompression of legacy files, JSON columns, word timings, vector blobs (`f32`, `f16`, `i8`), anchor URI parsing and full-text queries. - **Signatures that validate when they should not**: a file whose `content_sha256` or Ed25519 signature verifies although its content differs from what was signed, two different contents with the same canonical dump, or a verifier that trusts a stored hash instead of recomputing it. - **Leaks through vectors or provenance**: an implementation or producer that claims to remove the text of a document but keeps vectors from which it can be recovered, or that writes paths, user names, keys or other personal data into `provenance` or `generator` against [SPEC §15](spec/SPEC.md#privacy). - **Content handling** in the reader, the website and the validator: script execution from unit text, metadata, SVG or embedded documents; automatic fetching of remote references; path traversal when exporting blobs; a browser validator or inspector that sends a file off the user's device. - **The specification itself**: a rule that, followed exactly, leaves conforming implementations unsafe. - **The supply chain of this repository**: CI workflows, release artifacts and published packages. ## Not a vulnerability - Bugs in SQLite itself: please report them to the SQLite project. Do tell us if the way SPDF readers use SQLite makes such a bug exploitable. - A file that is invalid but is refused or read safely. Validation errors are the expected behaviour. - Resource use within the configured limits. Very large files that stay within the limits a reader was given are not a denial of service; files that bypass the limits are. - The content of documents: wrong transcriptions, wrong folios or doubtful metadata are quality problems, and instructions written in a document's text are data. Report them as ordinary issues. - A valid signature made with a key you do not trust. The specification attributes content to a key; deciding which keys to trust is up to the user. - The Ed25519 test key in `conformance/`, which is public on purpose and signs only test files. - Vectors that reveal information about a text that the same file already contains. - Reports from automated scanners without a demonstrated impact, and missing security headers on static pages without a concrete attack. --- # RFC template URL: https://spdf.joseluissaorin.com/governance/rfcs/0000 > Copy this file to spec/rfcs/0000-short-title.md and open a pull request. The editor assigns the next free number when discussion opens and renames the file. Keep every section below: when one does not apply, write… - **Status**: Draft - **Authors**: Full Name (affiliation, if any) - **Created**: YYYY-MM-DD - **Discussion**: link to the pull request or to the public thread - **Specification**: the SPDF version this targets (for example 5.1) and whether the change is minor (additive) or major (breaking) - **Affects**: sections of `spec/SPEC.md`, schema files, JSON Schemas, kinds of conformance case - **Conformance cases**: ids of the cases added or changed - **Supersedes / superseded by**: RFC numbers, if any - **Decision**: filled in by the editor or the technical committee (date, outcome and reasons) Copy this file to `spec/rfcs/0000-short-title.md` and open a pull request. The editor assigns the next free number when discussion opens and renames the file. Keep every section below: when one does not apply, write "Not applicable" and one sentence saying why. Delete these instructions and the italic guidance under each heading before the final comment period. The process is described in [`governance/RFC-PROCESS.md`](../../governance/RFC-PROCESS.md); how versions change is in [`governance/VERSIONING.md`](../../governance/VERSIONING.md). ## Summary *One paragraph that someone who has never read the specification can follow: what changes and for whom.* ## Motivation *The problem, with real examples: a document that cannot be anchored today, a citation that comes out wrong, an implementation that cannot do something efficiently. Say who needs this (readers, producers, archives, users of a particular language or kind of source) and what happens if nothing changes.* ## Guide-level explanation *Explain the change as it would be taught to an implementer or a producer: new tables, columns, anchor members or functions, with a small worked example (a row, an anchor, a URI, a citation). No normative language here.* ## Normative changes *The exact changes, written so they can be merged into `spec/SPEC.md` as they are. Use the BCP 14 key words (RFC 2119 and RFC 8174: MUST, SHOULD, MAY) only where they are meant. For each change give the section, the current text and the new text. Cover, when they are affected:* - *the schema (`spec/schema/`) and the canonical dump;* - *anchors and the anchor URI (parameters, canonical order, parsing);* - *search, citation and export functions;* - *validation: new error or warning codes take the next free number in their range and their place in the check order;* - *integrity and signatures;* - *the legacy mapping (4.x to 5.x view);* - *sidecars (`.spdfa.json`, `.spdfl.json`);* - *the value of `PRAGMA user_version` and `spdf_meta.spdf_version` that files using the change must carry.* *The Spanish translation (`spec/SPEC.es.md`) is updated in the same pull request, or the editor updates it before the change is published.* ## Conformance cases *An accepted RFC lands with at least one conformance case; it is considered implemented when two independent implementations pass it. List each case: id, kind, what it checks, the source or file it uses, and how its expected output was obtained (SQLite running the reference SQL, exact arithmetic, hand-written and reviewed, the reference oracle), so that nobody has to trust a single implementation. Follow "Adding cases" in [`conformance/README.md`](../../conformance/README.md). Only public-domain texts whose status anyone can verify.* ## Backwards compatibility *Answer each question:* - *Do existing files stay valid, with the same canonical dump and the same `content_sha256`?* - *What does a reader of an earlier minor version of the same major do with a file that uses the change? It MUST still read it: see the compatibility promise in `governance/VERSIONING.md`. If it cannot, the change belongs in a new major version.* - *What must writers do differently, and when?* - *Is anything deprecated? Give the version that deprecates it and the earliest major version that may remove it (at least 24 months later).* - *Does the legacy 4.x mapping change?* ## Security and privacy *New ways for a file to make a reader do work, fetch something, execute something or reveal something: SQL objects, URLs, embedded content, sizes and limits, information about people, what vectors or provenance disclose, effects on hashes and signatures. If there are none, say why.* ## Alternatives *Other designs considered, including doing nothing, and why this one is better. Prior art in other formats and standards (TEI, IIIF, W3C Web Annotation, Media Fragments, CSL, EPUB, PDF) is welcome.* ## Unresolved questions *What must be settled before acceptance, and what is deliberately left for later.* ## Implementations *Filled in as implementations land. An RFC becomes Implemented when two independent implementations pass all of its cases in CI. Two implementations are independent when neither wraps or translates the other's code: bindings over the Rust core, including the C ABI, count as the Rust implementation.* | Implementation | Status | Version or commit | Cases passed | |---|---|---|---| | | | | | --- # RFC 0001: SPDF 5.0 URL: https://spdf.joseluissaorin.com/governance/rfcs/0001 > SPDF 5.0 is the first public version of the format. It keeps the idea of the Scholaris versions 4.0 and 4.1 (one document per SQLite file; every passage carries an anchor to the printed page, folio, verse, second or… - **Status**: Accepted (2026-10-07) - **Authors**: José Luis Saorín Ferrer (editor) - **Created**: 2026-10-07 - **Discussion**: review by the implementers of the libraries, recorded in the change log of `spec/CONTRACT.md` (drafts 0 to 1.2) - **Specification**: SPDF 5.0, a new major version. It replaces 4.1 as the current version; 4.0 and 4.1 become legacy versions that every reader still reads - **Affects**: the whole of `spec/SPEC.md`, `spec/schema/spdf-5.0.sql`, `spec/json-schema/`, `conformance/` - **Conformance cases**: the conformance suite 0.1.0 (220 cases) and 0.2.0 (228 cases, `cases_sha256` `2d28cb5169e426ce26fec1df6c7481e2e5a3b208f0f14f66411e51043aa442c3`) - **Supersedes / superseded by**: none - **Decision**: accepted by the editor on 2026-10-07. This RFC bootstraps the process and was accepted without the public discussion and final comment periods that apply to every later RFC (see [`governance/RFC-PROCESS.md`](../../governance/RFC-PROCESS.md)). It is published so that anyone can see what changed from 4.1 and why; any of its decisions can be revisited by a later RFC. ## Summary SPDF 5.0 is the first public version of the format. It keeps the idea of the Scholaris versions 4.0 and 4.1 (one document per SQLite file; every passage carries an anchor to the printed page, folio, verse, second or slide it comes from; several vector spaces may coexist) and turns it into an open standard: the container is plain uncompressed SQLite, the file identifies itself in its header, identifiers are in English, metadata is CSL-JSON, text offsets are defined, the anchor vocabulary grows, anchors get a portable URI aligned with W3C Media Fragments and RFC 5147, vectors can be stored in half precision or 8 bits, conformance profiles and an extension mechanism are defined, content can be hashed and signed, and distributed files may no longer contain code (triggers, views or foreign virtual tables). Every 5.0 reader still reads 4.0 and 4.1 files. The normative text is [`spec/SPEC.md`](../SPEC.md), whose Appendix A summarizes the changes from 4.1; section numbers below refer to it. ## Motivation SPDF 4.x was the internal export format of Scholaris. It worked for one application and one team, but it had properties that stand in the way of a format anyone can implement and archive: - **Compressed container.** A 4.x file is a SQLite database wrapped in gzip. A reader must decompress the whole file before reading a single row, so it cannot read by HTTP ranges or memory-map the file, and must defend itself against decompression bombs. Most of the bulk of a typical file (page images, an embedded original) is already in compressed formats and gains little from a second layer. - **No identification.** 4.x files carry `application_id` 0. The version is stored in a table (`spdf.spdf_version`) and only sometimes in `user_version`, which some platforms did not allow Scholaris to set (the 4.0 sample in the conformance suite has `user_version` 0). Tools such as file(1) or DROID cannot tell a 4.x file from any other gzip stream. - **Spanish identifiers.** Tables and columns are named in Spanish (`documentos`, `unidades`, `anio`, `ancla_fin`). That was natural inside Scholaris and is a barrier for everyone else. - **Bespoke metadata.** Document metadata used Scholaris's own JSON structure (`MetadatosDocumento`), which no reference manager understands. - **Undefined offsets.** Nothing said how to count a position inside a text, and the languages that implement SPDF count differently (UTF-16 code units in JavaScript, code points in Python, bytes in Rust and Go). - **Missing anchors.** Verse numbers, canonical references (Stephanus, Bekker, biblical, CTS) and leaf or column foliation, which are how poetry, classical texts, early printed books and manuscripts are cited, could not be expressed. - **No way to point at a passage from outside the file.** 4.x defined no URI for an anchor. - **Vectors in f32 only**, which makes files with several vector spaces large and is wasteful on phones. - **SQL code inside distributed files.** 4.x files carry three triggers that keep the full-text index in sync. Opening a file that contains SQL code written by someone else is a risk that a document format should not ask readers to take. - **No integrity, no profiles, no extension mechanism, no rights**: nothing to detect tampering, nothing to say what a minimal conforming file is, no way for a vendor to add data without forking the format, and nowhere to say under what terms a file may be shared. ## Guide-level explanation A 5.0 file is a SQLite 3 database that any SQLite tool can open. Its header says what it is: bytes 68 to 71 read `SPDF` (the `application_id`) and bytes 60 to 63 hold the version, 500. Inside there is exactly one document: - `documents`: one row with the CSL-JSON record of the work, the SHA-256 of the original file, its media type and size, and its rights; - `units`: the citable units in reading order (pages, time spans, slides, sections, sheets), each with its anchor and its text in NFC; - `fragments`: passages of about 150 to 300 words, each with the anchor of its start and, if it crosses units, of its end, indexed by FTS5 for lexical search; - optionally `sections`, `figures`, vector `spaces` and `vectors`, embedded `blobs` (page images, the original), `provenance` and `extensions`. An anchor is a small JSON object: ```json {"type":"page","physical":29,"printed":"21","source":"read","chars":[118,301]} ``` and the same location as a URI, which names the document by the hash of its original bytes so it survives renaming and copying: ```text spdf:sha256-3f2a…#p=29&f=21&char=118,301 ``` A citation is computed from the stored anchor and the CSL record, never generated: `(Darwin, 1859, p. 21)`. A producer may also store a SHA-256 of the canonical content and an Ed25519 signature over it, so that a reader can check that the content is the one the signer vouched for. ## Normative changes The changes against 4.1, by area, with the section of `spec/SPEC.md` that states each rule. ### Container and identification (§2, §24) 1. A 5.0 file is an **uncompressed** SQLite 3 database holding exactly one document; the header starts at byte 0. Page size 4096, journal mode DELETE and a final `VACUUM` are recommended. Collections are separate manifests (§17). 2. `PRAGMA application_id = 1397769286` (0x53504446, stored big-endian at offset 68, which reads `SPDF` in ASCII). 3. `PRAGMA user_version` = major × 100 + minor × 10 (5.0 is 500). Readers accept 500 to 599, may warn on a newer minor (W105), and refuse other majors (E002). 4. Extension `.spdf`; media type `application/vnd.spdf+sqlite3` (registration pending); Uniform Type Identifier `com.joseluissaorin.spdf`. 5. Files MUST NOT contain triggers, views, or virtual tables other than the FTS5 tables of the schema. Writers keep the full-text index in sync themselves. ### Safe opening (§2.4, §14) 6. Every reader opens files read-only, with `query_only` on, `trusted_schema` off, the defensive flag on and extension loading off; refuses triggers, views and foreign virtual tables (except the three legacy FTS triggers); and bounds the size of any single value (RECOMMENDED 512 MiB) and of decompressed gzip input (RECOMMENDED 4 GiB). Operations that write, such as the FTS5 `integrity-check`, run on a private copy. ### Schema with English identifiers (§3, §20) 7. Tables: `spdf` becomes `spdf_meta`, `documentos` `documents`, `unidades` `units`, `secciones` `sections`, `fragmentos` `fragments`, `figuras` `figures`, `espacios` `spaces`, `vectores` `vectors`, `procedencia` `provenance`; `blobs` keeps its name. Every column is renamed as §20.1 lists. 8. Removed: `documentos.estado` and `documentos.bibliotecas` (library membership belongs to collection manifests) and index-only columns. 9. Added: `documents.rights` (§16); `documents.source_ref`, nullable, replacing `original`; `spaces.dtype`, `spaces.truncated_from`, `spaces.task_prefixes`; `blobs.sha256`; `provenance.model`; the `extensions` table. 10. `spdf_meta` has the REQUIRED keys `spdf_version`, `profile`, `created`, `generator` and `document_id`; integrity keys are OPTIONAL. 11. `units.ord` is numbered from 1 and contiguous (4.x numbered units from 0). ### Metadata (§6) 12. `documents.metadata` is one CSL-JSON item plus an `spdf` extension object for what CSL cannot hold: provenance per field, a date range for undated works, the original language, ORCID identifiers. The record can be handed to Zotero, citeproc or Pandoc as it is. ### Text and offsets (§7) 13. All stored text is NFC. Positions inside a text are counted in **Unicode code points over the NFC text**, end exclusive, which every language can compute the same way. 14. The literal text is never modernized; the `search_text` column (introduced in 4.1) holds a modernized-spelling layer used only for search. ### Anchors (§4) 15. New anchor types: `verse` and `canonical` (schemes such as `stephanus`, `bekker`, `bible` or `cts`). 16. Page anchors gain `foliation`: `page` (default), `leaf` (`fol. 1r`) or `column` (`col. 45`). 17. Any anchor MAY carry `region` (fractions 0 to 1 of the unit image) and `chars` (code point range in the unit's NFC text). 18. Required members are fixed per type and validated (E040, E041, E042). ### Anchor URI (§5) 19. New URI form `spdf:#`, with an ABNF. The preferred docref is `sha256-` and the hex SHA-256 of the original, which survives renaming and copying. 20. Parameters have one canonical order and encoding, and `format(parse(uri))` reproduces the URI byte for byte. 21. Where SPDF overlaps with existing standards it uses their syntax: `t=` and `xywh=percent:` as in W3C Media Fragments URI 1.0, and `char=` as in RFC 5147. 22. Resolution rules say how a reader finds the unit a URI designates. ### Vectors (§9) 23. Vector components are little-endian `f32`, `f16` (IEEE binary16) or `i8` (value q/127); the space id records model, dimensions and dtype. 24. Spaces record Matryoshka truncation and the task prefixes used at encoding time; a compatibility rule says when one query vector serves several spaces; writers quantize with fixed rounding rules. 25. A file without vectors is valid. ### Search, citation and export (§8, §18, §19) 26. Reference algorithms that conformance tests: lexical search through FTS5 with the `unicode61 remove_diacritics 2` tokenizer (terms sent as written, without case folding), a route for Chinese, Japanese and Korean with an optional `trigram` index and a substring fallback, brute-force vector search, and hybrid search by reciprocal rank fusion with k = 10. Products MAY rank better. 27. A short citation function for Spanish and English, and exports to CSL-JSON and BibTeX (REQUIRED) and other formats. ### Profiles and extensions (§10, §11) 28. Profiles, declared in `spdf_meta.profile`: `core`, `semantic`, `media` and `full`. 29. Extensions are declared in the `extensions` table, with tables named `x__`. A reader that meets a **required** extension it does not know refuses the file (E060); optional ones are ignored. ### Integrity, signature and rights (§12, §13, §16) 30. A canonical JSON dump of the file (RFC 8785, with fixed rounding and ordering) is the conformance oracle and the basis of integrity. 31. `content_sha256` hashes that dump; `signature` is an Ed25519 signature (RFC 8032) over it. Because the dump covers blob and vector bytes through their hashes but not the SQLite page layout, the signature survives `VACUUM` and SQLite version changes. 32. `documents.rights` states the licence (SPDX), the access level and the holder. ### Validation (§22) 33. A deterministic check order and a closed list of error and warning codes. Conformance compares the sets of codes, not the messages. ### Sidecars (§17) 34. User annotations live outside the document, in `*.spdfa.json` files (W3C Web Annotation with an `SpdfAnchorSelector` and a `TextQuoteSelector`), so the document stays immutable and sharing it never shares its reader's notes. 35. Collections are `*.spdfl.json` manifests that list documents by hash. ### Legacy (§20) 36. Every reader reads 4.0 and 4.1 files: detects gzip, decompresses within the limit, tolerates exactly the triggers `fragmentos_ai`, `fragmentos_ad` and `fragmentos_au`, and presents the 5.0 view, with `"legacy": true` in the dump and warning W110. A gzip-wrapped 5.0 file is read but flagged (E003, a warning). 37. Version 3.0 MAY be supported through an importer. ## Conformance cases The conformance suite is the evidence for this RFC. Version 0.1.0 (2026-10-07) published 220 cases; version 0.2.0 (the same day) brought them to 228: `anchor_uri` 40, `cite` 79, `dump` 7, `legacy_dump` 2, `quantize` 6, `roundtrip` 7, `search_hybrid` 2, `search_lexical` 39, `search_vector` 5 and `validate` 41. They use seven 5.0 files built from public-domain texts, two authentic legacy files (4.0 and 4.1, gzip-wrapped, with FTS triggers), and invalid files that each break one rule. How each expectation is obtained (SQLite as the oracle for lexical search, exact arithmetic for vectors, hand-reviewed URIs and citations) is described in [`conformance/README.md`](../../conformance/README.md); the changes are in [`conformance/CHANGELOG.md`](../../conformance/CHANGELOG.md). ## Backwards compatibility - **5.0 is a new major version.** A 4.x reader cannot read 5.0 files: the schema is different and the container is no longer gzip. This is intended. - **5.0 readers read 4.x.** Every conforming 5.0 reader reads 4.0 and 4.1 files through the 5.0 view (§20). Nothing that a 4.x file holds about the document is lost; only the Scholaris shelf state (`estado`, `bibliotecas`) is dropped. - **Writers** produce 5.0 only. Scholaris keeps reading its 4.x files and exports 5.0 through the `spdf-format` library. - **Sidecars** are new and carry their own version (`spdf_library: "1.0"` in manifests). - **Nothing is deprecated** within 5.0. The obligation to read 4.x lasts for the whole 5.x line; a future major version decides by RFC whether its readers still read 4.x (see [`governance/VERSIONING.md`](../../governance/VERSIONING.md)). ## Security and privacy Sections 14 and 15 of the specification are new in 5.0. In short: - Files contain no triggers, views or foreign virtual tables, and readers refuse files that do, so opening a file runs no code written by its author. Readers open read-only, with `trusted_schema` off, defensive mode on and extensions disabled. - The 5.0 container is not compressed. Legacy gzip input is decompressed within a limit to defeat decompression bombs; values and JSON nesting are bounded. - Text is light Markdown rendered without raw HTML; remote references are never fetched automatically; blob keys are sanitized before anything is written to disk; text passed to language models is data, not instructions. - Vectors can leak the text they were computed from; provenance can leak details of the producer. Rights and confidentiality rules that apply to a text apply to its vectors. - `content_sha256` and the Ed25519 signature give integrity and attribution of the content, not confidentiality, and say nothing about whether the signer's key should be trusted. ## Alternatives - **Keep the gzip wrapper.** Rejected: it prevents HTTP range reads and memory mapping, forces full decompression and adds decompression-bomb handling, for little gain on already-compressed images and originals. Readers still accept it for legacy files. - **A ZIP package with JSON and a SQLite index inside** (as EPUB or OOXML do). Rejected: two levels of parsing, and the full-text index needs SQLite anyway; SQLite alone gives random access, an index and a single file. - **Pure JSON, Parquet or Arrow.** Rejected: no built-in full-text search, and either no random access (JSON) or no good fit for long text with anchors (columnar formats). - **Keep Spanish identifiers with English aliases.** Rejected: two names for every table and column would double the surface of every implementation. The specification keeps a faithful Spanish translation instead. - **UTF-16 code units, bytes or grapheme clusters for offsets.** Rejected: UTF-16 ties the format to JavaScript, bytes to one encoding, and grapheme clusters to a Unicode version. Code points over NFC are stable and cheap everywhere. - **IIIF-style regions (`pct:`) or bare fractions in `xywh`.** Rejected in favour of W3C Media Fragments (`percent:`); in Media Fragments bare numbers mean pixels, so bare fractions would have been misread. Mapping to IIIF Image API regions is trivial. - **Signing the file bytes, or wrapping the file in JWS or COSE.** Rejected: SQLite writes its own version into the header and `VACUUM` reorders pages, so byte signatures break without any change in content. A signature over the canonical content is stable. An envelope format can still be added later as an extension. - **Allowing triggers and views.** Rejected for safety; writers can keep the index in sync without them. ## Unresolved questions - The registration of `application/vnd.spdf+sqlite3` with IANA, and whether to use the registered `+sqlite3` structured syntax suffix instead (`application/vnd.spdf+sqlite3`). Drafts of this and other registrations are in `governance/drafts/`. - The provisional registration of the `spdf` URI scheme (RFC 7595), which §5.4 announces. - Whether the anchor URI parameters may also be used as the fragment identifier of a URL that points to a `.spdf` file (`https://example.org/x.spdf#p=29`), which the media type registration would like to say. - How a reader and a validator of an earlier minor version treat an anchor type, a dtype or a validation rule introduced by a later minor (today an unknown anchor type is error E041), so that the compatibility promise holds for validators too. - Markup for mathematics and other non-textual content inside unit text, which 5.0 does not specify. - A persistent identifier (DOI) for each published version of the specification. ## Implementations This RFC becomes Implemented when two independent implementations pass the whole conformance suite in CI, as their `conformance-` artifacts show. The table records what each implementation reported in its own commits on 2026-10-07; several independent implementations already report the whole suite, so the editor will mark the RFC Implemented once their CI artifacts confirm it. | Implementation | Folder | Reported on 2026-10-07 | |---|---|---| | Rust (reference, also the C ABI) | `rust/` | 228 of 228 (commit `38c75e5`) | | TypeScript (`spdf-format`, npm) | `js/` | 228 of 228, also in Chromium and Bun (commit `e063594`) | | Python (`spdf-format`, PyPI) | `python/` | 228 of 228 (commit `557af8f`) | | Go | `go/` | 228 of 228 (commit `9355f22`) | | Swift | `swift/` | 228 of 228 on macOS and the iOS simulator (commit `54d5525`) | | PHP and Ruby | `php/`, `ruby/` | 228 of 228 (commit `537202e`) | | Kotlin/JVM, C#, R | `kotlin/`, `dotnet/`, `r/` | in development | | Julia, C (over the Rust ABI) | `julia/`, `c/` | planned | | Reference producer `spdf build` (Python) | `producer/` | in development | | Scholaris (second producer) | external | reads 4.x; 5.0 export through `spdf-format` planned | --- # RFC 0002: Conformance cases for exports and anchor resolution URL: https://spdf.joseluissaorin.com/governance/rfcs/0002 > Add three kinds of conformance case: exportcsl and exportbibtex, which make the MUST of SPEC §19 testable, and resolve, which tests how an anchor URI is resolved against a file (SPEC §5.4). Fix the details those… - **Status:** Accepted (2026-10-07); normative text in SPEC §5.4 and §19; cases in conformance 0.4.0. Implemented when two independent implementations pass them. - **Author:** spec agent, for the editor (José Luis Saorín Ferrer) - **Created:** 2026-10-07 - **Affects:** SPEC §5.4, §19, §21; `conformance/` ## Summary Add three kinds of conformance case: `export_csl` and `export_bibtex`, which make the MUST of SPEC §19 testable, and `resolve`, which tests how an anchor URI is resolved against a file (SPEC §5.4). Fix the details those cases need: the BibTeX key algorithm, the field set and the resolution order. ## Motivation SPEC §19 says every implementation MUST export CSL-JSON and BibTeX, and SPEC §5.4 says how a reader finds the unit an anchor URI designates. Neither is covered by the suite 0.2.0, so twelve implementations can diverge silently: different BibTeX keys break the `\cite{}` commands of a user who switches tools, and different resolution sends a reader to a different page than the citation says. Both are the kind of disagreement SPDF exists to prevent. ## Guide-level explanation ```json {"id": "export-bibtex-quijote", "kind": "export_bibtex", "input": {"file": "files/quijote.spdf"}, "expect": {"entry_type": "book", "key": "cervantessaavedra1605", "fields": {"author": "Cervantes Saavedra, Miguel de", "title": "El ingenioso hidalgo don Quijote de la Mancha", "year": "1605", "publisher": "Juan de la Cuesta", "address": "Madrid", "language": "es", "note": "…"}}} ``` BibTeX is compared structurally (entry type, key, field map), not byte for byte, so implementations keep their own layout and escaping style, which BibTeX tools do not care about. ```json {"id": "resolve-quijote-leaf", "kind": "resolve", "input": {"file": "files/quijote.spdf", "uri": "spdf:sha256-fa38…4c75#f=1v"}, "expect": {"unit": "p6", "chars": null}} ``` ## Normative changes 1. SPEC §19, BibTeX key: take the `family` (or `literal`) of the first author, else the first word of the title; decompose with NFKD, drop every character that is not an ASCII letter, lowercase; if nothing is left, use `anon`; append the first year of `issued`, or `nd`. Collisions inside one export get the suffixes `a`, `b`, `c`… in the order of the documents. 2. SPEC §19, BibTeX fields: exactly the mapping already listed in §19; names as `Family, Given` joined with ` and `; a `literal` name wrapped in braces; `year` as a string; fields whose source is absent are omitted. 3. SPEC §5.4, resolution order: `p` (and `pe`) first; else `f` through `units.printed`, first unit in `ord` order; else `t` (the first unit with `t0 ≤ t < t1`, the last unit if `t` equals the document's end); else `s`/`para`, `sl`, `sh`/`rows`, `v`, `ref` against the units' anchors, then the fragments' anchors (giving their `unit`). The result is the unit id plus the `char` range and the `xywh` region, if any. A URI whose document reference does not match the file resolves to an error. ## Conformance cases About 20 cases on the existing corpus: `export_csl` and `export_bibtex` for every file in `files/` and `legacy/` (including the anonymous *Lazarillo*, the literal author of the NASA recording and the Chinese title of the *Analects*, which exercises the `anon` fallback), and `resolve` for every anchor type, for a folio printed twice and for a mismatched document reference. ## Backwards compatibility No file changes. Implementations that already export BibTeX may need to change their keys; that is the point. ## Security and privacy None beyond SPEC §14: exports copy metadata the file already exposes. ## Alternatives - Byte-exact BibTeX: rejected; layout differences are harmless and would make the cases brittle. - Leaving the key to each tool: rejected; stable keys across tools are what users need. - Better Bib(La)TeX keys (Better BibTeX style): possible later as an OPTIONAL profile. ## Unresolved questions - Should titles keep their capitalization protected with braces in the structural comparison, or should the comparison ignore braces? - `@online` (biblatex) versus `@misc` for web pages. - Whether `resolve` should also return the fragments that cover the resolved position. ## Implementations None yet. Accepting this RFC requires the cases in `conformance/` and two independent implementations passing them (`governance/RFC-PROCESS.md`). ## Decision (2026-10-07) Accepted, with these changes from the draft above; the normative text is SPEC §5.4 and §19, which prevail: - Keys: when the first author yields no ASCII letter, the first word of `title-short` is used before that of `title` (`lazarillo1554`, not `la1554`); `anon` stays the last fallback; negative years keep their sign. - BibTeX is compared as canonical text, line by line after trimming each line and dropping empty lines; protection braces count (they follow a deterministic rule). `@misc` is the default type; a biblatex profile with `@online` is left for a later RFC. - `locate` returns lists, not a single unit: all matching units (in `ord` order) and all matching fragments (in `n` order), plus `char` and `xywh`. The rule order is `p`, `f`, `t`, `sl`, `v`, `ref`, `s`, `sh`; `s` matches by path prefix unless `para` is given; `v` and `rows` match when their first value falls inside the anchor's range; a time equal to the end of the last timed unit matches it. Fragments are filtered by `char`. A reference to another document yields `document: false`. - Added in the same release: vector search over units and figures (result items carry `unit_id` or `figure_id`) and page-sequence checks for ALTO, TEI and IIIF exports (`export_structure`). --- # SPDF in Rust URL: https://spdf.joseluissaorin.com/docs/rust > How to install and use the Rust implementation of SPDF (spdf): open, validate, search and cite. Reference implementation; also provides the C ABI. - **Package**: `spdf` - **Install**: `cargo add spdf` - **Registry**: [crates.io](https://crates.io/crates/spdf) - **Tier**: first - **CI**: CI running - **Folder**: `rust/` *The README of this library is written in Spanish.* Espacio de trabajo Cargo con tres crates: | crate | qué es | |---|---| | [`crates/spdf`](crates/spdf) | la biblioteca: apertura segura, lectura 5.0 y legado 4.x, validación, volcado canónico, búsqueda léxica, vectorial e híbrida, anclas y URI, cita corta, CSL-JSON y BibTeX, escritor, conversión del legado, integridad y firma, lectura remota por rangos HTTP (`--features http`, experimental) | | [`crates/spdf-tools`](crates/spdf-tools) | la CLI `spdf` | | [`crates/spdf-ffi`](crates/spdf-ffi) | la ABI de C estable (`include/spdf.h`) | SQLite va empaquetado (`rusqlite` con `bundled`): la misma versión y el mismo FTS5 en macOS, Linux, Windows, iOS y Android. ## Uso rápido ```sh cd rust cargo build --release ./target/release/spdf validate ../conformance/files/quijote.spdf ./target/release/spdf search ../conformance/files/quijote.spdf --lexical hidalgo ./target/release/spdf conformance ../conformance > conformance.json ``` Desde otro crate del repositorio (por ejemplo el lector Tauri): ```toml spdf = { path = "../../rust/crates/spdf" } ``` ## Pruebas ```sh cargo test --workspace --all-features # unitarias, de propiedades, de API, remotas, conformidad y doctests cargo clippy --workspace --all-targets --all-features -- -D warnings cargo fmt --all -- --check cargo bench -p spdf --features http # véase RENDIMIENTO.md ``` La prueba `tests/conformance.rs` ejecuta toda la batería de `../conformance/cases`; el CI (`.github/workflows/rust.yml`) la corre en Linux, macOS y Windows y sube el informe como artefacto `conformance-rust`. ## Documentos - [`NOTAS.md`](NOTAS.md): decisiones propias y observaciones para el agente de la especificación. - [`RENDIMIENTO.md`](RENDIMIENTO.md): cifras. - [`PUBLICAR.md`](PUBLICAR.md): cómo publicar en crates.io. Licencia: MIT OR Apache-2.0. --- # SPDF in TypeScript URL: https://spdf.joseluissaorin.com/docs/js > How to install and use the TypeScript implementation of SPDF (spdf-format): open, validate, search and cite. Node (node:sqlite), Bun and the browser (sqlite-wasm). Powers the validator on this site. - **Package**: `spdf-format` - **Install**: `npm install spdf-format` - **Registry**: [npm](https://www.npmjs.com/package/spdf-format) - **Tier**: first - **CI**: CI running - **Folder**: `js/` Read, validate, search, cite and write **SPDF** files from JavaScript and TypeScript, in Node, Bun, Deno, Cloudflare Workers and the browser. SPDF (Semantic Processed Document Format) is an open format for documents that have been read once and can be cited forever: every passage carries its exact anchor (printed page or folio, second of a recording, slide, verse, canonical reference), so a citation can only print what the source says. A `.spdf` file is a plain SQLite 3 database with full-text indexes, optional embedding vectors and CSL-JSON metadata. Specification: [`spec/SPEC.md`](https://github.com/joseluissaorin/spdf/blob/main/spec/SPEC.md). This package is a native, independent implementation of SPDF 5.0 and of the legacy 4.0 and 4.1 formats (Scholaris). - **No required dependencies.** Node uses the built-in `node:sqlite`, Bun uses `bun:sqlite`; in the browser it uses the official SQLite WebAssembly build (`@sqlite.org/sqlite-wasm`, an optional peer dependency). Gzip, SHA-256 and Ed25519 come from the platform (`zlib`, `DecompressionStream`, WebCrypto). - **TypeScript first**: strict types, ESM only, declarations included. - **Conforming reader, semantic reader, writer and validator**, profiles `core`, `semantic` and `media`, with the ALTO, TEI and IIIF exports: it passes the whole SPDF conformance suite (341 cases of suite 0.4.1, every kind, nothing skipped) with `node:sqlite` (Node 22, 24 and 26), with `bun:sqlite`, and with `sqlite-wasm` in Node and in Chromium. - **Remote reading**: in the browser, `openRemote(url)` opens a file with HTTP Range requests and downloads only the pages a query touches (about 1 % of a 23 MiB book for a lexical search; figures below). ```sh npm install spdf-format npm install @sqlite.org/sqlite-wasm # only for the browser, Deno or Workers ``` ## Node Node 22.13 or newer (22.5–22.12 with `--experimental-sqlite`). On Node 24 and later the library works in memory (`serialize`/`deserialize`) and sets the defensive flag and the value-size limit; on Node 22 it goes through a temporary file. ```ts import { openSpdf, validate } from 'spdf-format'; const doc = await openSpdf('quijote.spdf'); // a path, bytes, or a Blob; 5.0 or legacy 4.x (gzip too) console.log(doc.version, doc.document.metadata.title); // '5.0' 'El ingenioso hidalgo…' for (const hit of await doc.searchLexical('lanza en astillero', { limit: 5 })) { console.log(doc.cite(hit.anchor, 'es', hit.anchor_end), hit.anchor_uri); // (Cervantes Saavedra, 1605, p. 23) spdf:sha256-3f2a…#p=29&f=23&char=0,159 } const page = await doc.unitByPrinted('23'); // the page whose printed folio is 23 await doc.close(); const report = await validate('quijote.spdf'); // { valid, version, profile, errors, warnings } ``` Lexical search follows the reference algorithm of the specification: words are OR-ed, quoted phrases (`"…"`, `“…”`, `«…»`, `„…“`) are AND-ed, and case and diacritics are folded by the FTS5 index (`unicode61 remove_diacritics 2`). Queries in Chinese, Japanese or Korean use the `trigram` index when the file has one, and a substring scan otherwise. ```ts const space = await doc.space('embeddinggemma-2@768'); // model, dims, dtype, task prefixes const q = await myModel.embed(`${space.task_prefixes?.query ?? ''}lanza en astillero`); await doc.searchVector(space.id, q, { limit: 5 }); // brute force, f32 / f16 / i8 await doc.searchHybrid('lanza en astillero', q, space.id, { limit: 5 }); // RRF, k = 10 ``` ## Browser ```ts import { configureBrowserEngine, openSpdf, openRemote, openBlob } from 'spdf-format'; // Only if your bundler moves sqlite3.wasm (esbuild, plain copies): tell the engine where it is. configureBrowserEngine({ wasmUrl: '/assets/sqlite3.wasm' }); const fromBytes = await openSpdf(await (await fetch('/quijote.spdf')).arrayBuffer()); const fromFile = await openBlob(fileInput.files[0]); // lazy inside a Worker (FileReaderSync) const remote = await openRemote('https://example.org/quijote.spdf'); // HTTP Range requests console.log(remote.source.stats()); // { requests, bytesFetched, chunksUsed, chunkSize } ``` - `openRemote` reads synchronously inside SQLite's VFS (synchronous `XMLHttpRequest`), so run it in a **Web Worker**; it also works on the main thread, where browsers warn about synchronous requests. Cross-origin servers must allow the `Range` request header and expose `Content-Range` (`Access-Control-Expose-Headers: Content-Range`). If the server ignores ranges, the file is downloaded whole (`fallbackToDownload: false` turns that into an error). Legacy gzip files are always downloaded whole. - `storage: 'opfs'` (or `'auto'`) keeps opened and written databases in the Origin Private File System (`opfs-sahpool` VFS, Workers only) instead of the WebAssembly heap: `configureBrowserEngine({ storage: 'auto' })`. - In Cloudflare Workers, where WebAssembly cannot be compiled at run time, initialize sqlite-wasm yourself with the precompiled module and pass it: `import { wasmEngine } from 'spdf-format/wasm'; openSpdf(bytes, { engine: wasmEngine({ sqlite3 }) })`. ### What a remote search costs Measured in Chromium 153 (headless) inside a Web Worker, on a 23.11 MiB SPDF of the whole *Don Quijote* (Project Gutenberg #2000: 1 382 pages, 2 723 fragments, one 768-dimension f32 space for fragments and pages), each operation on a freshly opened file, so the figures include opening it (`test/browser/run.mjs`): | Operation | Downloaded | Requests | Share of the file | |---|---:|---:|---:| | Open (header, schema, meta, document) | 32 KiB | 8 | 0.14 % | | Lexical search `Rocinante` (10 hits) | 180 KiB | 34 | 0.76 % | | Lexical search `molinos de viento` | 244 KiB | 44 | 1.03 % | | Lexical search `"Dulcinea del Toboso"` (phrase) | 216 KiB | 43 | 0.91 % | | Lexical search `hidalgo de la Mancha` | 268 KiB | 52 | 1.13 % | | The same lexical search again | 0 | 0 | cached | | `unit(700)` and its fragments | 72 KiB | 18 | 0.30 % | | Vector search, 768 × f32, brute force | 10.95 MiB | 811 | 47.4 % | | Hybrid search | 11.27 MiB | 869 | 48.8 % | Pages are fetched in 4 KiB chunks with an adaptive read-ahead (sequential misses double the next request up to 256 KiB) and an LRU cache of 64 MiB. Lexical results are ranked inside the FTS index and only the hits are read from `fragments`, which keeps a search at about 1 % of the file. Vector search is brute force by definition and reads every vector of the space: for remote use, ship an `i8` space (a quarter of the bytes) or query a server. ## Bun ```ts import { openSpdf } from 'spdf-format'; // resolves to bun:sqlite under Bun const doc = await openSpdf('quijote.spdf'); ``` Bun's binding has no switch for `SQLITE_DBCONFIG_DEFENSIVE` nor for value-size limits; read-only mode, `query_only` and `trusted_schema = OFF` still apply. ## Writing ```ts import { SpdfWriter, convertLegacy, generateSigningKey } from 'spdf-format'; const w = await SpdfWriter.create(); await w.setDocument({ id: 'quijote-1605', kind: 'scanned_pdf', mime: 'application/pdf', bytes: 123456, source_sha256: '3f2a…', // SHA-256 of the original metadata: { type: 'book', title: 'El ingenioso hidalgo don Quijote de la Mancha', author: [{ family: 'Cervantes Saavedra', given: 'Miguel de' }], issued: { 'date-parts': [[1605]] } }, }); await w.addUnits([{ id: 'p29', anchor: { type: 'page', physical: 29, printed: '21' }, text: '…', reader: 'pdf-text-layer' }]); await w.addFragments([{ id: 'f1', unit: 'p29', text: '…', anchor: { type: 'page', physical: 29, printed: '21', chars: [0, 159] } }]); const space = await w.addSpace({ provider: 'google', model: 'embeddinggemma-2', dims: 768, dtype: 'i8', modalities: ['text'] }); await w.addVectors(space, [{ target: 'fragment', id: 'f1', vector: embedding }]); // quantized to i8 const { privateKey } = await generateSigningKey(); const bytes = await w.finish({ signWith: privateKey }); // FTS rebuilt, VACUUM, content_sha256 + Ed25519 const v50 = await convertLegacy(legacyBytes); // a Scholaris 4.x file (gzip) as SPDF 5.0 const copy = await SpdfWriter.fromSpdf(doc); // a full 5.0 copy, e.g. to add a vector space ``` ## API at a glance | Area | Functions | |---|---| | Open | `openSpdf(input, opts)`, `openRemote(url)`, `openBlob(blob)` (browser), `validate(input)`, `dump(input)` | | Document | `doc.version`, `doc.legacy`, `doc.meta`, `doc.document`, `units()`, `unit(ord)`, `unitByPrinted()`, `sections()`, `fragments()`, `fragment(id)`, `figures()`, `spaces()`, `vectors()`, `blob(key)`, `blobs()`, `provenance()`, `extensions()`, `dump()`, `validate()` | | Search | `doc.searchLexical(q)`, `doc.searchVector(space, vec)`, `doc.searchHybrid(q, vec, space)` | | Anchors | `formatAnchorUri(docref, anchor, end)`, `parseAnchorUri(uri)`, `locatorToAnchor()`, `checkAnchor()`, `doc.locate(uriOrUrl)` (SPEC §5.4: units, fragments, `char`, `xywh`), `doc.resolve()` | | Citation | `cite(anchor, document, locale, end)`, `doc.cite(anchor, locale, end)` (es, en), `doc.citePassage(fragmentId, quote, locale)` (SPEC §18.2: cites the unit the quotation is in) | | Export | `toCslJson()`, `toCslJsonArray()`, `cslCitationItem()`, `toBibtex()` (keys `cervantessaavedra1605`, SPEC §19), `toAlto()`, `toTei()`, `toIiif()` | | Write | `SpdfWriter.create()`, `.fromSpdf()`, `.fromSource()`, `convertLegacy()`, `encodeVector()` | | Integrity | `doc.contentSha256()`, `doc.verifyIntegrity(publicKey?)`, `verifySignature()`, `generateSigningKey()` | | Engines | `nodeEngine()`, `bunEngine()`, `wasmEngine()`, `setDefaultEngine()`; the port is `SqlEngine`/`SqlConnection` | Entry points: `spdf-format` (picks Node, Bun or the browser by export condition), `spdf-format/node`, `spdf-format/bun`, `spdf-format/browser`, `spdf-format/core` (no engine; pass `{ engine }`), `spdf-format/wasm` and `spdf-format/cli`. Media type: `application/vnd.spdf+sqlite3` (`MEDIA_TYPE`); a `.spdf` URL may carry the anchor parameters as its fragment (`https://example.org/quijote.spdf#p=5&f=1r`), which `doc.locate()` accepts as well as `spdf:` URIs. ## Command line ```sh npx spdf-format validate quijote.spdf # exit status 1 if invalid npx spdf-format dump quijote.spdf --pretty # canonical JSON dump (JCS without --pretty) npx spdf-format search quijote.spdf molinos de viento --limit 5 npx spdf-format cite quijote.spdf f12 --locale en npx spdf-format cite quijote.spdf --bibtex npx spdf-format convert legacy-4.1.spdf out.spdf --hash npx spdf-format conformance path/to/spdf/conformance ``` ## Safety Files come from strangers. Every file is opened read-only, with `query_only`, `trusted_schema = OFF`, `mmap_size = 0`, `cell_size_check = ON`, the defensive flag where the binding has it, no extensions, a 512 MiB limit for any single value (`SQLITE_LIMIT_LENGTH` with `node:sqlite` on Node 24+ and with sqlite-wasm; checked by the library for blobs elsewhere) and a 4 GiB limit for gzip input. Files with triggers, views or virtual tables other than the FTS5 indexes are refused (`E020`), except the three FTS triggers of legacy 4.x files. Write operations of the validator (the FTS `integrity-check`) run on a private in-memory copy. ## Conformance `npm test` runs the unit tests and the whole shared suite with `node:sqlite` and `sqlite-wasm`; `npm run test:bun` runs it with `bun:sqlite`; `npm run test:browser` runs it in headless Chromium (Playwright, `--mute-audio`) together with the remote-reading, OPFS and Blob tests. CI uploads the runner report as the `conformance-js` artifact. ## License MIT OR Apache-2.0, at your option. --- # SPDF in Python URL: https://spdf.joseluissaorin.com/docs/python > How to install and use the Python implementation of SPDF (spdf-format): open, validate, search and cite. Standard library sqlite3 only; imported as spdf. - **Package**: `spdf-format` - **Install**: `pip install spdf-format` - **Registry**: [PyPI](https://pypi.org/project/spdf-format/) - **Tier**: first - **CI**: CI running - **Folder**: `python/` Read, validate, search, cite and write **SPDF** files from Python. SPDF (Semantic Processed Document Format) is an open format for documents that have been read once and can be cited forever: every passage carries its exact anchor (printed page or folio, second of a recording, slide, verse, canonical reference), so a citation can only print what the source says. A `.spdf` file is a plain SQLite 3 database with full-text indexes, optional embedding vectors and CSL-JSON metadata. This package is a native, independent implementation of SPDF 5.0 and of the legacy 4.0 and 4.1 formats. It is installed as `spdf-format` and imported as `spdf`. - Python 3.10 or newer, **standard library only** (`sqlite3`). - Optional extras: `numpy` (fast vector search), `crypto` (Ed25519 through `cryptography`; a pure-Python fallback is included), `pandas`, `arrow`. - Passes the whole SPDF conformance suite (0.4.1, 341 cases): reader, semantic reader, writer and validator, profiles `core`, `semantic` and `media`. ```sh pip install spdf-format # or: uv add spdf-format pip install "spdf-format[numpy]" # faster vector search ``` ## Quick start ```python import spdf with spdf.open("quijote.spdf") as f: # 5.0, or legacy 4.x (gzip-wrapped too) doc = f.document # spdf.Document, metadata is CSL-JSON print(doc.display_title, f.cite()) # El ingenioso hidalgo… (Cervantes Saavedra, 1605) for hit in f.search("lanza en astillero"): print(f.cite(hit.fragment), hit.anchor_uri) # (Cervantes Saavedra, 1605, p. 23) spdf:sha256-3f2a…#p=29&f=23&char=0,159 unit = f.unit_by_printed("23") # the page whose printed folio is 23 print(unit.text) ``` Lexical search follows the reference algorithm of the specification: words are OR-ed, quoted phrases (`"…"`, `“…”`, `«…»`, `„…“`) are AND-ed, and case and diacritics are folded by the FTS5 index (`unicode61 remove_diacritics 2`). Queries in Chinese, Japanese or Korean use the `trigram` index when the file has one, and a substring scan otherwise. ```python with spdf.open("darwin.spdf") as f: space = f.space("embeddinggemma-2@768") # model, dims, dtype, task prefixes qvec = my_model.encode(space.task_prefixes["query"] + "natural selection") f.search_vector(qvec, space=space.id, limit=5) f.search_hybrid("natural selection", qvec, space=space.id) # RRF, k = 10 f.search_vector(qvec, space=space.id, target="unit") # hits carry unit_id ``` Every hit has `id`, `target` (`fragment`, `unit` or `figure`), `score`, `via`, `anchor` and `anchor_uri`; `fragment_id`, `unit_id` and `figure_id` give the id for each target. ## Validate ```python report = spdf.validate("quijote.spdf") report.valid # True when there are no errors report.codes # {"W102"} … report.to_dict() # {"valid", "version", "profile", "errors", "warnings"} as in the spec ``` Validation follows the order of the specification and reports every code it finds: `E001` not SQLite, `E002` unknown version, `E003` gzip-wrapped 5.0 (warning), `E010`/`E011` missing table or column, `E012` missing metadata key, `E013` not exactly one document, `E020` trigger, view or foreign virtual table, `E030`–`E032` vectors and spaces, `E040`–`E042` anchors, `E050`/`E051` metadata, `E060` unknown required extension, `E070` FTS index out of sync, `E080` blob hash, `E081`/`E082` content hash and signature, `E090` unit order, and the warnings `W100`–`W110` (`W103`: a fragment crosses from one kind of `matter` to another, such as body text into a plate, or from a page with a printed folio to one without; writers should set `matter` on the units of paged documents and not let fragments cross). ## Write ```python with spdf.Writer("out.spdf") as w: w.add_document({ "id": "quijote", "kind": "pdf", "mime": "application/pdf", "bytes": 1203456, "source_sha256": "3f2a…", "source_ref": w.add_blob("original.pdf", "application/pdf", pdf_bytes), "metadata": {"type": "book", "title": "El ingenioso hidalgo don Quijote de la Mancha", "author": [{"family": "Cervantes Saavedra", "given": "Miguel de"}], "issued": {"date-parts": [[1605]]}, "language": "es"}, }) page = {"type": "page", "physical": 29, "printed": "23", "roman": False, "source": "read"} w.add_unit({"id": "u29", "ord": 1, "anchor": page, "text": "En un lugar de la Mancha…", "reader": "pdf-text-layer"}) w.add_fragment({"id": "f1", "unit": "u29", "ord": 1, "text": "En un lugar de la Mancha…", "anchor": {**page, "chars": [0, 24]}}) w.add_space({"id": "embeddinggemma-2@768", "provider": "google", "model": "embeddinggemma-2", "dims": 768, "modalities": ["text"]}) w.add_vector("fragment", "f1", "embeddinggemma-2@768", data=vector) # list, ndarray or bytes ``` The writer builds the file next to its destination, rebuilds the FTS index (distributed files carry no triggers), fills the required metadata keys (`spdf_version`, `profile`, `created`, `generator`, `document_id`), computes `content_sha256`, can sign it (`finalize(sign_key=…)`), compacts it with `VACUUM`, validates it and only then moves it into place. Text is normalized to NFC. `i8` and `f16` vectors are quantized as the specification says. `spdf.convert_legacy("old.spdf", "new.spdf")` converts a 4.x file to 5.0, and `spdf.write_source(full_dump, path)` rebuilds a file from a full dump (`f.full_dump()`). ## Anchors, URIs and citations ```python uri = spdf.make_uri("sha256-3f2a…", {"type": "page", "physical": 29, "printed": "21", "chars": [118, 301]}) # 'spdf:sha256-3f2a…#p=29&f=21&char=118,301' spdf.parse_uri(uri) # {"docref": "sha256-3f2a…", "locator": {"p": 29, "f": "21", "char": [118, 301]}} spdf.cite({"type": "page", "physical": 9, "printed": "1r", "foliation": "leaf"}, doc, locale="es") # '(Cervantes Saavedra, 1605, fol. 1r)' spdf.cite({"type": "time", "t0": 4160.0, "t1": 4175.5}, doc, locale="en") # '(Cortázar, 1959, 1:09:20)' ``` `f.locate(reference)` resolves an anchor URI, or the URL of a `.spdf` with an anchor fragment, against the file (SPEC §5.4): ```python f.locate("https://example.org/quijote.spdf#p=5&pe=6&char=101,278").to_dict() # {"document": True, "units": ["p5", "p6"], "fragments": ["q4"], "char": [101, 278], "xywh": None} ``` A quotation is cited by the unit it lies in, never by the start of the fragment that contains it (SPEC §18.2), so a passage on page 211 of a fragment that begins on an unnumbered plate cites page 211: ```python c = f.cite_passage("m4", "tube N N", locale="es") c.text # '(Hooke, 1665, p. 211)' c.uri # 'spdf:sha256-…#p=321&f=211&char=0,8' ``` Citations print only what the anchor says: inferred folios in brackets (`p. [21]`), unnumbered pages as `s. p.` / `n. pag.`, leaves and columns (`fol. 1r`, `col. 45`), `h:mm:ss` times, slides, sheets, verses and canonical references. Spanish uses `e` instead of `y` before the sound /i/ (`Gómez e Iglesias`). ## Bibliography and other exports | Export | API | CLI | |---|---|---| | CSL-JSON (Zotero, citeproc, Pandoc) | `f.to_csl_json()`, `spdf.bibliography.csl_citation_item()` | `spdf export -f csl` | | BibTeX | `f.to_bibtex()` | `spdf export -f bibtex` | | ALTO 4 XML (page units) | `f.to_alto()` | `spdf export -f alto` | | TEI P5 (minimal: header, `pb`, `p`, `lg`/`l`, `u`, `note`) | `f.to_tei()` | `spdf export -f tei` | | IIIF Presentation 3 manifest | `f.to_iiif(base_url)` | `spdf export -f iiif --base-url URL [--images-dir DIR]` | | JSON Lines of fragments | `spdf.interop.frames.fragment_records(f)` | `spdf export -f jsonl` | | pandas DataFrame | `f.to_pandas(vectors="space id")` | | | Arrow table | `f.to_arrow(vectors="space id")` | | | Canonical dump (JCS) | `f.dump()`, `f.dump_json()` | `spdf dump` | ALTO has no invented coordinates: SPDF stores text per unit, so blocks and lines carry none (they are optional in ALTO 4). The IIIF manifest paints each canvas with the page image, adds the unit text as a `supplementing` annotation, figure descriptions as `describing` annotations on `#xywh=percent:` regions, and sections as ranges; audio and video become one time-based canvas with a range per unit. ## Annotations and collections User annotations and libraries live outside the documents, as the specification's sidecar files: ```python from spdf import sidecars with spdf.open("quijote.spdf") as f: hit = f.search("lanza en astillero")[0] note = sidecars.annotation(f, hit.fragment, body="Origen del tópico.") # W3C Web Annotation sidecars.write_annotations("notas.spdfa.json", [note], label="Notas de lectura") lib = sidecars.library(["quijote.spdf", "lazarillo.spdf"], "Tesis: fuentes") sidecars.write_library("fuentes.spdfl.json", lib) ``` Each annotation targets the document by identity (`spdf:sha256-…`) with an `SpdfAnchorSelector` (the anchor URI) and a `TextQuoteSelector` (exact text with a little context), so it survives a re-reading that shifts offsets. ## Command line ```text spdf validate FILE… [--json] exit status 1 if a file is invalid spdf dump FILE [--pretty] canonical dump (RFC 8785) spdf info FILE | --env summary; --env shows the SQLite and FTS5 in use spdf search FILE QUERY [--vector JSON --space ID] [--mode lexical|vector|hybrid] [--json] spdf cite FILE [--fragment ID [--quote TEXT] | --unit ID | --uri URI] [--locale es|en] [--bibtex] spdf export FILE -f csl|bibtex|alto|tei|iiif|jsonl spdf convert OLD.spdf NEW.spdf legacy 4.x to 5.0 spdf sign FILE --key KEY / spdf verify FILE [--public-key ed25519:…] spdf conformance [DIR] run the conformance suite, print the report ``` Other packages can add subcommands through the `spdf.commands` entry point group: the entry point is a callable `register(subparsers)` that adds an `argparse` parser and sets `func` (a handler returning the exit status). The SPDF producer adds `spdf build` this way: ```toml [project.entry-points."spdf.commands"] build = "spdf_build.cli:register" ``` ## Safety Files are untrusted input. `spdf.open` checks the SQLite header before SQLite sees the file, decompresses gzip input to a private temporary file with a size limit (4 GiB by default), opens the database read-only by URI (`mode=ro`) with `query_only`, `trusted_schema=OFF`, `cell_size_check` and, on Python 3.12 or newer, `SQLITE_DBCONFIG_DEFENSIVE`; it never loads extensions, refuses triggers, views and virtual tables other than the format's FTS5 indexes (legacy files may keep their three FTS triggers), refuses files that require unknown extensions, and enforces a maximum blob size (512 MiB by default, `max_blob_size=`). ## FTS5 and Python builds Lexical search, writing and the FTS integrity check need SQLite's FTS5, which depends on how Python was built. It is present in the python.org installers, Homebrew, `uv python install` (python-build-standalone), conda and the usual Linux distributions; CI checks Linux, macOS and Windows with Python 3.10 to 3.13. `spdf info --env` tells you what you have. Without FTS5, reading, vector search, citations, exports and validation still work, and the operations that need it raise `spdf.Fts5UnavailableError` explaining what to do; on Linux, `pip install pysqlite3-binary` is picked up automatically as a fallback driver. ## Conformance The SPDF conformance suite lives in `conformance/` in the [repository](https://github.com/joseluissaorin/spdf). Run it with: ```sh spdf conformance path/to/conformance # prints {"impl", "version", "passed", "failed", "skipped"} ``` This implementation claims every kind of case: `dump`, `legacy_dump`, `roundtrip`, `validate`, `search_lexical`, `search_vector`, `search_hybrid`, `anchor_uri`, `cite`, `quantize`, `locate`, `cite_passage`, `export_csl`, `export_bibtex` and `export_structure` (checked on its own ALTO, TEI and IIIF output). CI publishes its report as the `conformance-python` artifact. ## Development ```sh cd python uv sync --group dev uv run pytest uv run ruff check src tests && uv run ruff format --check src tests uv run mypy ``` ## License Code: MIT OR Apache-2.0, at your option. The SPDF specification is CC BY 4.0. --- # SPDF in Swift URL: https://spdf.joseluissaorin.com/docs/swift > How to install and use the Swift implementation of SPDF (SPDF): open, validate, search and cite. Swift Package Manager, from the repository. Apple platforms and Linux. - **Package**: `SPDF` - **Install**: `.package(url: "https://github.com/joseluissaorin/spdf", from: "5.0.0")` - **Tier**: first - **CI**: CI running - **Folder**: `swift/` Native Swift implementation of **SPDF** (Semantic Processed Document Format): documents that have been read once and can be cited forever, because every passage carries its exact anchor (printed page, folio, second of a recording, slide, verse). - SwiftPM package `SPDF` for iOS 16+, macOS 13+, tvOS 16+, watchOS 9+, visionOS 1+ and Linux. - Uses the **system SQLite** (FTS5, `unicode61 remove_diacritics 2` and `trigram` are present on iOS and macOS; verified on the iOS simulator and on macOS). On Linux it links the distribution's `libsqlite3` (install `libsqlite3-dev` and `zlib1g-dev`). - No third-party code on Apple platforms (CryptoKit for SHA-256 and Ed25519, zlib for gzip). On Linux, [`swift-crypto`](https://github.com/apple/swift-crypto) provides the same CryptoKit API. - API with `async` variants, `Sendable` and `Codable` types; `SPDFFile` is thread-safe. - Conformance: passes the whole SPDF conformance suite (`../conformance`), every kind, on macOS and on the iOS simulator. ## Installation ```swift .package(url: "https://github.com/joseluissaorin/spdf", from: "0.1.0") // target dependency: .product(name: "SPDF", package: "spdf") ``` The repository root carries a thin `Package.swift` pointing into `swift/`, so the URL above is all SwiftPM needs; versions follow the monorepo-wide tags (`0.1.0`, …). ## Reading and searching ```swift import SPDF let file = try await SPDFFile.open(url) // 5.0, or a legacy 4.x .spdf (gzip) defer { file.close() } for hit in try await file.searchLexical("«lugar de la Mancha»", limit: 5) { let cite = try file.cite(hit.anchor!, end: hit.anchorEnd, locale: "es") print(hit.id, hit.score, hit.anchorURI, cite) // q4 1.889394 spdf:sha256-fa38…#p=5&pe=6&f=1r&fe=1v&char=101,278 (Cervantes Saavedra, 1605, fols. 1r-[1v]) } // Vector and hybrid search with your own query embedding (dot product when the // space is normalized, cosine otherwise). let vector: [Double] = embed(query) let semantic = try await file.searchVector(vector, space: "embeddinggemma-2@768", limit: 10) let hybrid = try await file.searchHybrid(query, vector: vector, space: "embeddinggemma-2@768", limit: 10) let units = try await file.units() // pages, time spans, slides… let bib = try file.exportBibTeX() // keys of SPEC §19: cervantessaavedra1605… let csl = try file.exportCSL() let alto = try file.exportALTO() // also exportTEI(), exportIIIF() // Cite a quotation by the unit it lies in (SPEC §18.2), not by its fragment's start. let cited = try file.citePassage(fragment: "m4", quote: "XXXIV.\ntube N N", locale: "es") // cited.text == "(Hooke, 1665, p. 211)", cited.uri == "spdf:sha256-ba9d…#p=319&pe=321&fe=211" // Resolve a reference (SPEC §5.4): an spdf: URI or a .spdf URL with a fragment. let where = try file.locate("https://example.org/quijote.spdf#p=7") // units ["p7"], fragments ["q5"] ``` Every method also has a synchronous form (`try file.searchLexical(…)`), handy in command-line tools and tests. ## Validation and canonical dump ```swift let report = SPDFValidator.validate(url) // ValidationResult: Codable print(report.valid, report.errorCodes, report.warningCodes) let json = try file.dumpJSON() // RFC 8785 canonical JSON let hash = try file.contentSHA256() // integrity hash (§8) ``` Opening is defensive: read-only, `query_only`, `trusted_schema=OFF`, `SQLITE_DBCONFIG_DEFENSIVE`, no extensions; files with triggers, views or foreign virtual tables (E020) or unknown required extensions (E060) are refused; strings and blobs are capped (512 MiB) and so is gunzipped input (4 GiB). ## Anchors and citations ```swift let anchor = Anchor(["type": "page", "physical": 29, "printed": "21", "source": "inferred"]) let uri = AnchorURI.format(docref: "sha256-3f2a…", anchor: anchor) // spdf:sha256-3f2a…#p=29&f=21 let parsed = try AnchorURI.parse(uri) // strict parser let text = Citation.cite(anchor, metadata: [ "type": "book", "title": "Arte nuevo de hacer comedias", "author": [["family": "Vega", "non-dropping-particle": "de", "given": "Lope"]], "issued": ["date-parts": [[1609]]], ], locale: "es") // (de Vega, 1609, p. [21]) ``` ## Writing ```swift let writer = try SPDFWriter(url: out, options: .init(generator: "my-app/1.0")) try writer.setDocument(SPDFDocumentInfo(id: "doc", kind: "pdf", metadata: ["type": "book", "title": "…"], sourceSHA256: sha, mime: "application/pdf", bytes: size)) let page = Anchor(["type": "page", "physical": 1, "printed": "1"]) try writer.add(SPDFUnit(id: "u1", anchor: page, text: "…", reader: "pdf-text-layer")) try writer.add(SPDFFragment(id: "f1", unit: "u1", text: "…", anchor: page)) try writer.add(VectorSpace(id: "embeddinggemma-2@768", provider: "local", model: "embeddinggemma-2", dims: 768)) try writer.addVector(target: .fragment, id: "f1", space: "embeddinggemma-2@768", values: embedding) // [Float] try writer.addBlob(key: "pages/0001.png", mime: "image/png", data: png) try writer.finish() // FTS rebuilt, content_sha256 written, VACUUM, atomic replace; no triggers ``` Every file the writer produces carries `spdf_meta.content_sha256`; pass `.init(signingKey: seed)` (a 32-byte Ed25519 seed) to sign it too (`signer`, `signature`, SPEC §8). Files signed this way verify with the Go, Rust, Python, C# and JavaScript implementations. `SPDFSeal.seal(url, signingKey:)` hashes and signs an existing file in place. `SPDFSource.write(_:to:)` builds a file from a full JSON dump (the format of `conformance/sources/`). ## Command line ```sh swift run spdf-swift validate file.spdf swift run spdf-swift dump file.spdf swift run spdf-swift search file.spdf "lugar de la Mancha" -n 5 swift run spdf-swift cite file.spdf q4 -locale en swift run spdf-swift export file.spdf bibtex swift run spdf-swift uri parse 'spdf:sha256-…#p=29&f=21' swift run spdf-swift build source.json out.spdf swift run spdf-swift seal out.spdf -key seed.hex swift run spdf-swift export file.spdf tei # also alto, iiif swift run spdf-swift conformance ../conformance -o conformance.json ``` ## Tests and conformance ```sh swift test # unit tests + the whole suite xcodebuild test -scheme SPDF-Package -destination 'platform=iOS Simulator,name=iPhone 17' ``` The runner (`SPDFConformance`) discovers the cases by listing `conformance/cases/*.json` and prints `{"impl","version","passed","failed","skipped"}`. Nothing is skipped: this is a full (reader and writer) implementation. ## License MIT OR Apache-2.0. --- # SPDF in Kotlin / JVM URL: https://spdf.joseluissaorin.com/docs/kotlin > How to install and use the Kotlin / JVM implementation of SPDF (io.github.joseluissaorin:spdf): open, validate, search and cite. Kotlin and Java on the JVM and Android. - **Package**: `io.github.joseluissaorin:spdf` - **Install**: `implementation("io.github.joseluissaorin:spdf:5.0.0")` - **Registry**: [Maven Central](https://central.sonatype.com/artifact/io.github.joseluissaorin/spdf) - **Tier**: first - **CI**: CI running - **Folder**: `kotlin/` Native Kotlin/JVM implementation of **SPDF** (Semantic Processed Document Format): documents that have been read once and can be cited forever, because every passage carries its exact anchor (printed page, folio, second of a recording, slide, verse). It is usable from Kotlin, Java and Android, and independent of the other implementations in this repository (it does not wrap the Rust ABI). - Maven coordinates: `io.github.joseluissaorin:spdf` (JVM), `io.github.joseluissaorin:spdf-android` (Android), both on top of `io.github.joseluissaorin:spdf-core`. - Java 17 or newer on the JVM; Android minSdk 23. Built with Kotlin 2.4 at language level 2.2, so apps on Kotlin 2.1 or newer can consume it. - Conformance: passes the whole SPDF conformance suite (`../conformance`, 309 cases in suite 0.4.0) with **both** SQLite adapters, on the JVM and on Android emulators (API 31 and 36), every kind (`dump`, `legacy_dump`, `roundtrip`, `validate`, `search_lexical`, `search_vector`, `search_hybrid`, `anchor_uri`, `cite`, `quantize`, `locate`, `export_csl`, `export_bibtex`, `export_structure`). Nothing is skipped: this is a full reader and writer. ## Modules | Artifact | What | SQLite | |---|---|---| | `spdf-core` | all the logic: safe opening, dump, validation, search, anchors, citation, export, writer, conformance runner | none: it talks to a small `SqlDriver` interface with typed values (`SqlValue`: NULL, INTEGER, REAL, TEXT, BLOB) | | `spdf` | `JdbcSqlDriver` and the `spdf` command-line tool | `org.xerial:sqlite-jdbc` (bundled SQLite with FTS5 and `trigram`) | | `spdf-android` | `AndroidxSqlDriver` | `androidx.sqlite` driver API with `BundledSQLiteDriver` (`androidx.sqlite:sqlite-bundled`): the app ships its own SQLite with FTS5 and `trigram`, which `android.database.sqlite` does not guarantee | Both adapters register themselves with `java.util.ServiceLoader`, so `SpdfFile.open(path)` picks the one on the classpath. You can always pass a driver explicitly (recommended on Android, where R8 may strip service files). Any other binding can be plugged in by implementing `SqlDriver` (three methods). ```kotlin // build.gradle.kts dependencies { implementation("io.github.joseluissaorin:spdf:0.1.0") // JVM // implementation("io.github.joseluissaorin:spdf-android:0.1.0") // Android } ``` ```xml io.github.joseluissaorin spdf 0.1.0 ``` ## What it does | | | |---|---| | Safe opening | read-only (`SQLITE_OPEN_READONLY`), `query_only`, `trusted_schema = OFF`, `mmap_size = 0`, `cell_size_check`, no extensions; files with triggers, views or foreign virtual tables refused (E020); unknown required extensions refused (E060); maximum value size (512 MiB) and maximum gunzipped size (4 GiB); WAL files are read through a copy with the header patched | | Versions | SPDF 5.0, and the legacy 4.0 / 4.1 files of Scholaris (Spanish schema, gzip-wrapped) through the 5.0 view, metadata mapped to CSL-JSON | | Dump | canonical JSON (RFC 8785 / JCS) of the whole file, `content_sha256` | | Validation | every code of SPEC §22 (E001–E090, W100–W110) in the reference order; FTS `integrity-check` on a private copy; Ed25519 signatures (platform provider, with a pure fallback for older Android) | | Search | lexical (FTS5 BM25, CJK route with `trigram` or substring), vector (`f32`, `f16`, `i8`; dot product or cosine), hybrid (reciprocal rank fusion, k = 10) | | Anchors | anchor ↔ URI (`spdf:sha256-…#p=29&f=21&char=118,301`), strict parser, canonical form | | Resolution | `locate(reference)`: an anchor URI, or the URL of a `.spdf` with the anchor as fragment (`https://…/quijote.spdf#p=5&f=1r`), resolved to units, fragments, `char` and `xywh` (SPEC §5.4) | | Citation | short author-date citation in Spanish and English | | Export | CSL-JSON (with the CSL `label`/`locator` of a cited passage) and BibTeX, one or several documents, with the key and field rules of SPEC §19 (same keys as every other implementation: `cervantessaavedra1605`, `lazarillo1554`, `anonnd`); ALTO 4, a minimal TEI and a IIIF Presentation 3 manifest (SPEC §19.4), with no invented coordinates or dimensions | | Writer | builds valid SPDF 5.0 files (FTS kept in sync, `VACUUM`, no triggers, atomic replace); f16/i8 quantization as the spec says; writes `content_sha256` by default and, given an Ed25519 key, `signer` and `signature` (SPEC §8) | | Typed reading | `document()`, `units()`, `fragments()`, `sections()`, `figures()`, `spaces()`, `provenance()`, `blob(key)` | ## Kotlin ```kotlin import io.github.joseluissaorin.spdf.* SpdfFile.open("quijote.spdf").use { f -> // also legacy .spdf (gzip) files for (hit in f.searchLexical("«lugar de la Mancha»", limit = 5)) { println("${hit.fragmentId} ${hit.score} ${hit.anchorUri}") println(f.cite(hit.anchor!!, hit.anchorEnd, locale = "es")) // q4 1.889394 spdf:sha256-fa38…#p=5&pe=6&f=1r&fe=1v&char=101,278 // (Cervantes Saavedra, 1605, fols. 1r-[1v]) } val query = DoubleArray(8) // your own query embedding val nearest = f.searchVector(query, space = "toy-embedding@8", limit = 5) val fused = f.searchHybrid("hidalgo", query, "toy-embedding@8", limit = 5) println(f.exportBibTeX()) // Resolution of a reference (SPEC §5.4) and structural exports (SPEC §19.4) val loc = f.locate("https://example.org/quijote.spdf#p=5&char=101,278") println("${loc.units} ${loc.fragments}") // [p5] [q4] val alto: String = f.exportAlto() val tei: String = f.exportTei() val manifest: String = f.exportIiif(base = "https://example.org/quijote") val citeproc = f.exportCslJson(Anchor.page(5, "1r", foliation = "leaf")) // with "label": "folio", "locator": "1r" } // Several documents in one bibliography (keys disambiguated with a, b, c…) val items = listOf("a.spdf", "b.spdf").map { p -> SpdfFile.open(p).use { it.metadata() } } println(Export.bibtex(items)) val csl = Json.compact(Export.cslItems(items)) // Validation and canonical dump val result = Validator.validate("file.spdf") println("${result.isValid} ${result.errorCodes} ${result.warningCodes}") SpdfFile.open("file.spdf").use { f -> val jcs: String = f.dumpJson() // RFC 8785 bytes val sum: String = f.contentSha256() // integrity hash of SPEC §18 } // Anchors and citations val a = Anchor.page(29, "21", source = "inferred") val uri = AnchorUri.format("sha256-3f2a…", a) // spdf:sha256-3f2a…#p=29&f=21 val parsed = AnchorUri.parse(uri) // strict: malformed URIs throw SpdfException val text = Citation.cite(a, null, mapOf( "type" to "book", "title" to "Arte nuevo de hacer comedias", "author" to listOf(mapOf("family" to "Vega", "non-dropping-particle" to "de", "given" to "Lope")), "issued" to mapOf("date-parts" to listOf(listOf(1609))), ), "es") // (de Vega, 1609, p. [21]) // Writing // content_sha256 is written by default; a 32-byte Ed25519 seed also signs the file. SpdfWriter.create(File("out.spdf"), WriterOptions(generator = "my-tool/1.0", signingKey = seed)).use { w -> w.setDocument(Document("doc", "pdf", "application/pdf", sha256Hex, mapOf("type" to "book", "title" to "…"))) w.addUnit(CitableUnit("u1", Anchor.page(1, "1"), "…", reader = "pdf-text-layer")) w.addFragment(Fragment("f1", "u1", "…", Anchor.page(1, "1"))) w.addSpace(Space("embeddinggemma-2@768", "local", "embeddinggemma-2", 768)) w.addVector("fragment", "f1", "embeddinggemma-2@768", embedding) // FloatArray or DoubleArray w.addBlob("pages/0001.png", "image/png", png) w.finish() // rebuilds FTS, VACUUM, atomic move; closing without finish() discards the file } ``` `Sources.write(source, file)` builds a file from a full JSON dump (the format of `conformance/sources/`). ## Java ```java import io.github.joseluissaorin.spdf.*; import io.github.joseluissaorin.spdf.sql.SqlDriver; import java.util.List; import java.util.Map; try (SpdfFile f = SpdfFile.open("quijote.spdf")) { List hits = f.searchLexical("lugar de la Mancha", 5); for (Hit h : hits) { System.out.println(h.getFragmentId() + " " + h.getAnchorUri()); System.out.println(f.cite(h.getAnchor(), h.getAnchorEnd(), "en")); } } ValidationResult r = Spdf.validate("quijote.spdf"); System.out.println(r.isValid() + " " + r.getErrorCodes()); try (SpdfWriter w = SpdfWriter.create("out.spdf")) { Document d = new Document("lazarillo", "pdf", "application/pdf", sha256Hex, Map.of("type", "book", "title", "La vida de Lazarillo de Tormes")); d.setTitle("Lazarillo de Tormes"); w.setDocument(d); w.addUnit(new CitableUnit("u1", Anchor.page(3, "A2r", "read", "leaf"), "Pues sepa Vuestra Merced", "pdf-text-layer")); Fragment frag = new Fragment("f1", "u1", "Pues sepa Vuestra Merced", Anchor.page(3, "A2r", "read", "leaf")); frag.setContext("Prólogo"); w.addFragment(frag); w.finish(); } AnchorUri.Parsed p = Spdf.parseUri("spdf:sha256-3f2a…#p=29&f=21"); SqlDriver driver = SqlDriver.defaultDriver(); // `default` is a Java keyword Location where = SpdfFile.open("quijote.spdf").locate("spdf:sha256-fa38…#f=1v"); String cite = Spdf.cite(Anchor.verse(12), null, Json.parseObject("{\"title\":\"Rimas\",\"author\":[{\"family\":\"Bécquer\"}]}"), "es"); ``` Every entry point is static for Java (`@JvmStatic`), optional parameters have overloads (`@JvmOverloads`), rows are classes with a constructor for the required members and setters for the rest, and nothing is `suspend`. `src/test/java` in the `spdf` module checks this. From Java the default adapter is `SqlDriver.defaultDriver()`. ## Android ```kotlin val driver = AndroidxSqlDriver() // BundledSQLiteDriver inside val options = OpenOptions(tempDir = context.cacheDir) // gunzipped and private copies go here SpdfFile.open(File(context.filesDir, "book.spdf"), driver, options).use { f -> val hits = f.searchLexical("golondrinas") } ``` `spdf-android` is a plain JVM library on purpose: it needs no Android SDK to build, and its tests run the androidx bundled driver on the host JVM (the artifact ships Linux, macOS and Windows natives too), including the whole conformance suite. Android apps consume it like any jar; Gradle resolves the Android variant of `androidx.sqlite:sqlite-bundled` (with the native library for each ABI) for them. The core avoids JDK APIs newer than Android API 23 (a test checks the bytecode). On a real Android runtime, the test-only module `spdf-android-device` (Android Gradle Plugin 9.4, Java instrumented test, never published) packages `../conformance` as test assets and runs the whole suite on a device or emulator. It is not part of the default build; enable it explicitly: ```sh # with an emulator or device attached (ANDROID_HOME pointing at the SDK) ANDROID_SERIAL=emulator-5584 ./gradlew -Pspdf.androidDevice=true :spdf-android-device:connectedAndroidTest ``` Results so far: 309/309 on Android 12 (API 31) and Android 16 (API 36) arm64 emulators. CI runs it on x86_64 emulators (API 31 and 35) in a separate job. ## Safety notes SPDF files come from strangers (SPEC §2.4, §14). What each adapter enforces: | | sqlite-jdbc | androidx.sqlite bundled | |---|---|---| | `SQLITE_OPEN_READONLY` | yes | yes | | `query_only`, `trusted_schema = OFF`, `mmap_size = 0`, `cell_size_check` | yes (set and checked by the core) | yes (set and checked by the core) | | extension loading | off | never enabled (no `addExtension`) | | `SQLITE_DBCONFIG_DEFENSIVE` | not exposed by sqlite-jdbc | not exposed by the driver API | | maximum value size | `SQLITE_LIMIT_LENGTH` | not exposed: no value can exceed the file size, so files larger than the limit are measured once when opened | `SpdfFile.safety` reports what was in force. Without the defensive flag, the read-only, query-only connection still refuses every write, and files with triggers, views or foreign virtual tables are refused before any query on user tables. The FTS `integrity-check` (a write) runs on a private temporary copy. ## Command line ```sh ./gradlew :spdf:installDist B=spdf/build/install/spdf/bin/spdf $B validate file.spdf $B dump file.spdf $B search file.spdf "lugar de la Mancha" -n 5 $B vsearch file.spdf toy-embedding@8 0,0.5,0.25,0.75,0.25,0,0.25,0 -n 3 $B hybrid file.spdf toy-embedding@8 0,0.5,0.25,0.75,0.25,0,0.25,0 selection -n 3 $B cite file.spdf q4 --locale en $B export file.spdf bibtex # also csl, alto, tei, iiif $B locate file.spdf 'spdf:sha256-…#f=1v' $B uri parse 'spdf:sha256-…#p=29&f=21' $B build source.json out.spdf $B conformance ../conformance -o conformance.json ``` ## Conformance ```sh ./gradlew build # unit tests; both adapters run the whole suite ./gradlew :spdf:conformance # runner report in build/conformance.json ./gradlew :spdf:run --args="conformance ../conformance -o conformance.json" ``` The runner discovers the cases by listing `conformance/cases/*.json` (`SPDF_CONFORMANCE_DIR` overrides the location) and prints `{"impl":"spdf-kotlin","version":"0.1.0","passed":[…], "failed":[…],"skipped":[…]}`; it exits non-zero if anything fails. ## Building Gradle 9.8 through the wrapper, Kotlin 2.4, any JDK 17 or newer (bytecode targets 17). `./gradlew publishAllPublicationsToBuildRepoRepository` writes exactly what would be published to `build/repo`; publishing to Maven Central is described in [`PUBLICAR.md`](PUBLICAR.md) (in Spanish). ## License MIT OR Apache-2.0. --- # SPDF in Go URL: https://spdf.joseluissaorin.com/docs/go > How to install and use the Go implementation of SPDF (github.com/joseluissaorin/spdf/go): open, validate, search and cite. Go module in the monorepo. - **Package**: `github.com/joseluissaorin/spdf/go` - **Install**: `go get github.com/joseluissaorin/spdf/go` - **Registry**: [pkg.go.dev](https://pkg.go.dev/github.com/joseluissaorin/spdf/go) - **Tier**: first - **CI**: CI running - **Folder**: `go/` Pure-Go implementation of **SPDF** (Semantic Processed Document Format): documents that have been read once and can be cited forever, because every passage carries its exact anchor (printed page, folio, second of a recording, slide, verse). - Module: `github.com/joseluissaorin/spdf/go` (package `spdf`) - SQLite without cgo ([`modernc.org/sqlite`](https://pkg.go.dev/modernc.org/sqlite), FTS5 and the `trigram` tokenizer included), so it cross-compiles to every Go target. - Go 1.25 or newer. - Conformance: passes the whole SPDF conformance suite (`../conformance`), every kind (`dump`, `legacy_dump`, `roundtrip`, `validate`, `search_*`, `anchor_uri`, `cite`). ```sh go get github.com/joseluissaorin/spdf/go ``` ## What it does | | | |---|---| | Safe opening | read-only, `query_only`, `trusted_schema=OFF`, `SQLITE_DBCONFIG_DEFENSIVE`, no extensions, files with triggers or views refused (E020), unknown required extensions refused (E060), max blob size (512 MiB) and max gunzipped size (4 GiB) | | Versions | SPDF 5.0, and the legacy 4.0 / 4.1 files of Scholaris (Spanish schema, gzip-wrapped) through the 5.0 view, metadata mapped to CSL-JSON | | Dump | canonical JSON (RFC 8785) of the whole file, `content_sha256` (§8) | | Validation | every code of the specification (E001–E090, W100–W110), Ed25519 signatures | | Search | lexical (FTS5 BM25, CJK route with `trigram` or substring), vector (`f32`, `f16`, `i8`), hybrid (reciprocal rank fusion, k = 10) | | Anchors | anchor ↔ URI (`spdf:sha256-…#p=29&f=21&char=118,301`), strict parser; resolution of a URI or of a `.spdf` URL with a fragment to units and fragments (`Locate`, SPEC §5.4) | | Citation | short author-date citation in Spanish and English; citation of a quotation by the unit it lies in (`CitePassage`, SPEC §18.2) | | Export | CSL-JSON (also citations with `label`/`locator`) and BibTeX, one or several documents with the keys of SPEC §19 (`cervantessaavedra1605`, `lazarillo1554`, collision suffixes); ALTO 4, minimal TEI and IIIF Presentation 3 (`ExportALTO`, `ExportTEI`, `ExportIIIF`; no invented coordinates or dimensions) | | Writer | builds valid SPDF 5.0 files (FTS kept in sync, `VACUUM`, no triggers) with `content_sha256` and, given a key, an Ed25519 signature (§8); `Seal` hashes and signs an existing file in place | ## Reading and searching ```go package main import ( "fmt" "log" spdf "github.com/joseluissaorin/spdf/go" ) func main() { f, err := spdf.Open("quijote.spdf", nil) // also legacy .spdf (gzip) files if err != nil { log.Fatal(err) } defer f.Close() hits, err := f.SearchLexical("«lugar de la Mancha»", 5) if err != nil { log.Fatal(err) } for _, h := range hits { cite, _ := f.Cite(h.Anchor, h.AnchorEnd, "es") fmt.Println(h.FragmentID, h.Score, h.AnchorURI, cite) // q4 1.889394 spdf:sha256-fa38…#p=5&pe=6&f=1r&fe=1v&char=101,278 (Cervantes Saavedra, 1605, fols. 1r-[1v]) } // Vector and hybrid search with your own query embedding. vec := make([]float64, 8) vhits, _ := f.SearchVector(vec, "toy-embedding@8", "fragment", 5) _ = vhits hy, _ := f.SearchHybrid("hidalgo", vec, "toy-embedding@8", 5) _ = hy bib, _ := f.ExportBibTeX() fmt.Print(bib) } ``` ## Validating and dumping ```go res := spdf.Validate("file.spdf", nil) fmt.Println(res.Valid, res.ErrorCodes(), res.WarningCodes()) f, _ := spdf.Open("file.spdf", nil) dump, _ := f.DumpJSON() // RFC 8785 bytes sum, _ := f.ContentSHA256() // integrity hash of §8 ``` ## Anchors and citations ```go a := spdf.Anchor{"type": "page", "physical": int64(29), "printed": "21", "source": "inferred"} uri := spdf.AnchorURI("sha256-3f2a…", a, nil) // spdf:sha256-3f2a…#p=29&f=21 docref, loc, err := spdf.ParseURI(uri) // strict: malformed URIs are errors text := spdf.Cite(a, nil, map[string]any{ "type": "book", "title": "Arte nuevo de hacer comedias", "author": []any{map[string]any{"family": "Vega", "non-dropping-particle": "de", "given": "Lope"}}, "issued": map[string]any{"date-parts": []any{[]any{int64(1609)}}}, }, "es") // (de Vega, 1609, p. [21]) ``` ## Citing a quotation ```go // A fragment of Micrographia runs from a plate into page 211: the quotation is cited // by the unit(s) it actually touches, not by the start anchor of its fragment. p, _ := f.CitePassage("m4", "XXXIV.\ntube N N", "es") fmt.Println(p.Text, p.URI) // (Hooke, 1665, p. 211) spdf:sha256-ba9d…#p=319&pe=321&fe=211 ``` ## Resolving references ```go r, _ := f.Locate("spdf:sha256-…#f=1v") // or "https://example.org/quijote.spdf#p=7" fmt.Println(r.Document, r.Units, r.Fragments, r.Char) // true [p6] [q4 q5] [] ``` ## Writing ```go w, err := spdf.Create("out.spdf", &spdf.WriterOptions{Generator: "my-tool/1.0"}) if err != nil { log.Fatal(err) } w.SetDocument(spdf.Document{ID: "doc", Kind: "pdf", Mime: "application/pdf", SourceSHA256: sha, Bytes: n, Metadata: map[string]any{"type": "book", "title": "…"}}) w.AddUnit(spdf.Unit{ID: "u1", Reader: "pdf-text-layer", Text: "…", Anchor: spdf.Anchor{"type": "page", "physical": int64(1), "printed": "1"}}) w.AddFragment(spdf.Fragment{ID: "f1", Unit: "u1", Text: "…", Anchor: spdf.Anchor{"type": "page", "physical": int64(1), "printed": "1"}}) w.AddSpace(spdf.SpaceDef{ID: "embeddinggemma-2@768", Provider: "local", Model: "embeddinggemma-2", Dims: 768}) w.AddVector("fragment", "f1", "embeddinggemma-2@768", embedding) // []float32, quantized for f16/i8 w.AddBlob("pages/0001.png", "image/png", png) if err := w.Close(); err != nil { // rebuilds FTS, VACUUM, atomic rename log.Fatal(err) } ``` Every file the Writer produces carries `spdf_meta.content_sha256`; pass `WriterOptions{SigningKey: ed25519.NewKeyFromSeed(seed)}` to sign it too (`signer`, `signature`). Files signed this way verify with the Rust, Python and JavaScript implementations (and tampering gives E081 or E082 in all of them). `spdf.Seal(path, key)` does the same for an existing file. `spdf.WriteSource(source, path)` builds a file from a full JSON dump (the format of `conformance/sources/`). ## Command line ```sh go install github.com/joseluissaorin/spdf/go/cmd/spdf@latest spdf validate file.spdf spdf dump file.spdf spdf search file.spdf "lugar de la Mancha" -n 5 spdf cite file.spdf q4 -locale en spdf export file.spdf bibtex spdf uri parse 'spdf:sha256-…#p=29&f=21' spdf build source.json out.spdf spdf seal out.spdf -key seed.hex # content_sha256 + Ed25519 signature spdf export file.spdf alto # also tei, iiif spdf conformance ../conformance -o conformance.json ``` ## Conformance ```sh go test ./... # unit tests and the whole suite go run ./cmd/spdf conformance ../conformance # the report of contract §11 ``` The runner discovers the cases by listing `conformance/cases/*.json` and prints `{"impl","version","passed","failed","skipped"}`. Nothing is skipped: this is a full (reader and writer) implementation. ## License MIT OR Apache-2.0. --- # SPDF in C# / .NET URL: https://spdf.joseluissaorin.com/docs/dotnet > How to install and use the C# / .NET implementation of SPDF (Spdf.Format): open, validate, search and cite. .NET with Microsoft.Data.Sqlite. - **Package**: `Spdf.Format` - **Install**: `dotnet add package Spdf.Format` - **Registry**: [NuGet](https://www.nuget.org/packages/Spdf.Format) - **Tier**: first - **CI**: CI failing - **Folder**: `dotnet/` Native C# implementation of **SPDF** (Semantic Processed Document Format): documents that have been read once and can be cited forever, because every passage carries its exact anchor (printed page, folio, second of a recording, slide, verse). - NuGet package: `Spdf.Format` · namespace `Spdf` · .NET 8 or newer. - SQLite through `Microsoft.Data.Sqlite.Core` with the `SQLitePCLRaw.bundle_e_sqlite3` native bundle (FTS5 and the `trigram` tokenizer included) on Windows, macOS and Linux. - An independent implementation: it does not wrap the Rust library or any other one. Ed25519 verification, RFC 8785 serialization and the exact rounding rules are written in C#. - Conformance: passes the whole SPDF conformance suite (`../conformance`), every kind (`dump`, `legacy_dump`, `roundtrip`, `quantize`, `validate`, `search_lexical`, `search_vector`, `search_hybrid`, `anchor_uri`, `cite`, `locate`, `export_csl`, `export_bibtex`, `export_structure`). Nothing is skipped. ```sh dotnet add package Spdf.Format ``` ## What it does | | | |---|---| | Safe opening | read-only, `query_only`, `trusted_schema=OFF`, `SQLITE_DBCONFIG_DEFENSIVE`, extension loading disabled, `mmap_size=0`, `cell_size_check=ON`; triggers, views and foreign virtual tables refused (E020), unknown required extensions refused (E060); max blob size (512 MiB) and max gunzipped size (4 GiB); files left in WAL mode are read from a private copy | | Versions | SPDF 5.0, and the legacy 4.0 / 4.1 files of Scholaris (Spanish schema, gzip-wrapped) through the 5.0 view, with metadata mapped to CSL-JSON | | Dump | canonical JSON (RFC 8785) of the whole file and `content_sha256` (§12, §13) | | Validation | every code of the specification (E001–E090, W100–W110), FTS integrity on an in-memory copy, Ed25519 signatures | | Search | lexical (FTS5 BM25, CJK route with `trigram` or substring), vector over fragments, units or figures (`f32`, `f16`, `i8`; dot product or cosine), hybrid (reciprocal rank fusion, k = 10) | | Anchors | anchor ↔ URI (`spdf:sha256-…#p=29&f=21&char=118,301`), strict parser, canonical form; resolution of `spdf:` URIs and `.spdf` URLs to units and fragments (§5.4) | | Citation | short author-date citation in Spanish and English | | Export | CSL-JSON (with the CSL `label`/`locator` of a citation) and BibTeX with the keys of §19.1 (`cervantessaavedra1605`, `lazarillo1554`, `anonnd`, collision suffixes `a`, `b`…); ALTO 4, minimal TEI and IIIF Presentation 3 (§19.4) | | Writer | builds valid SPDF 5.0 files (FTS kept in sync, `VACUUM`, no triggers, atomic replace), with `content_sha256` and an optional Ed25519 signature (§13) | ## Reading and searching ```csharp using Spdf; using var file = SpdfFile.Open("quijote.spdf"); // also legacy .spdf (gzip) files foreach (var hit in file.SearchLexical("«lugar de la Mancha»", limit: 5)) { Console.WriteLine($"{hit.FragmentId} {hit.Score:F6} {hit.AnchorUri}"); Console.WriteLine(file.Cite(hit.Anchor!, hit.AnchorEnd, "es")); // (Cervantes Saavedra, 1605, fols. 1r-[1v]) } // Vector and hybrid search with your own query embedding. double[] query = new double[8]; var nearest = file.SearchVector(query, space: "toy-embedding@8", target: "fragment", limit: 5); var pages = file.SearchVector(query, "toy-embedding@8", target: "unit"); // hits carry UnitId var fused = file.SearchHybrid("hidalgo", query, "toy-embedding@8", limit: 5); // Typed reading. Document doc = file.GetDocument(); IReadOnlyList pages = file.GetUnits(); IReadOnlyList fragments = file.GetFragments(); Blob? original = file.GetBlob("blob:original.pdf"); Console.Write(file.ExportBibTeX()); // @book{cervantessaavedra1605, … Console.WriteLine(file.ExportCslJson()); // [{"id":"cervantessaavedra1605", …}] BibTexEntry entry = file.ExportBibTeXEntry(); // entry type, key and fields // Several documents at once (keys disambiguated with a, b, c…), and a CSL citation item. string bib = BibliographyExport.BibTeX([file.GetMetadata(), other.GetMetadata()]); var cited = BibliographyExport.CslItems([file.GetMetadata()], hit.Anchor, hit.AnchorEnd); // + label, locator // Resolving a reference (§5.4): units, fragments, char range and region it designates. LocateResult where = file.Locate("spdf:sha256-…#p=5&pe=6&char=101,278"); LocateResult byUrl = file.Locate("https://example.org/quijote.spdf#f=1v"); // Structural exports (§19.4): ALTO 4, a minimal TEI and a IIIF Presentation 3 manifest. // No invented coordinates or dimensions; inferred folios bracketed (TEI, IIIF) or absent (ALTO). string alto = file.ExportAlto(); // Page per page unit, TextBlock / TextLine / String per word string tei = file.ExportTei(); // teiHeader, ,

, /, , string iiif = file.ExportIiif(new IiifOptions { Base = "https://example.org/quijote" }); var pageSequence = StructureExport.PageSequence(StructureFormat.Tei, tei); // read back from the XML ``` `SpdfFile.OpenAsync(path)` decompresses or copies asynchronously when the file needs it; `SpdfFile.Open(stream)` and `SpdfFile.Open(bytes)` read from memory. A `SpdfFile` is not thread-safe; open one per thread. ## Validating and dumping ```csharp ValidationResult result = SpdfValidator.Validate("file.spdf"); Console.WriteLine($"{result.Valid} {string.Join(",", result.ErrorCodes)} {string.Join(",", result.WarningCodes)}"); using var file = SpdfFile.Open("file.spdf"); string dump = file.DumpJson(); // RFC 8785 text string hash = file.ContentSha256(); // integrity hash of §13 ``` ## Anchors and citations ```csharp var anchor = Anchor.Page(29, "21", source: "inferred").WithChars(118, 301); string uri = AnchorUri.Format("sha256-3f2a…", anchor); // spdf:sha256-3f2a…#p=29&f=21&char=118,301 ParsedAnchorUri parsed = AnchorUri.Parse(uri); // FormatException if malformed var metadata = new Dictionary { ["type"] = "book", ["title"] = "Arte nuevo de hacer comedias", ["author"] = new List { new Dictionary { ["family"] = "Vega", ["non-dropping-particle"] = "de", ["given"] = "Lope" } }, ["issued"] = new Dictionary { ["date-parts"] = new List { new List { 1609L } } }, }; Console.WriteLine(Citation.Cite(anchor, null, metadata, "es")); // (de Vega, 1609, p. [21]) ``` JSON values (metadata, anchors, word timings) are plain trees: `null`, `bool`, `long`, `double`, `string`, `List` and `Dictionary`. `SpdfJson` parses and serializes them (`Canonical` is RFC 8785 with the six-decimal rounding of the specification; `Compact` keeps numbers as they are). ## Writing ```csharp using var w = SpdfWriter.Create("out.spdf", new SpdfWriterOptions { Generator = "my-tool/1.0" }); w.SetDocument(new Document { Id = "rimas", Kind = "pdf", Mime = "application/pdf", SourceSha256 = sha, Bytes = size, Metadata = new Dictionary { ["type"] = "book", ["title"] = "Rimas" }, }); w.AddUnit(new Unit { Id = "u1", Reader = "pdf-text-layer", Text = "…", Anchor = Anchor.Page(1, "1") }); w.AddFragment(new Fragment { Id = "f1", Unit = "u1", Text = "…", Anchor = Anchor.Page(1, "1").WithChars(0, 120) }); w.AddSpace(new Space { Id = "embeddinggemma-2@768:i8", Provider = "local", Model = "embeddinggemma-2", Dims = 768, DType = "i8" }); w.AddVector("fragment", "f1", "embeddinggemma-2@768:i8", embedding); // float[] or double[], quantized per §9.2 w.AddBlob("original.pdf", "application/pdf", bytes); w.Commit(); // document + spdf_meta, FTS rebuild, VACUUM, atomic move into place ``` Disposing a writer without `Commit()` discards the file. `SpdfSource.Write(source, path)` builds a file from a full JSON dump (the format of `conformance/sources/`), value for value. ### Integrity and signatures By default the writer stores `spdf_meta.content_sha256` (§13), computed on the finished file. Give it a 32-byte Ed25519 secret key (seed) to sign as well; it writes `signer` (`ed25519:` + base64 public key) and `signature` (base64 of the RFC 8032 signature over `spdf-content-sha256:` + the hex hash): ```csharp byte[] seed = LoadSeedFromYourKeyStore(); // 32 bytes; never commit it to a repository using var w = SpdfWriter.Create("signed.spdf", new SpdfWriterOptions { SigningKey = seed }); // … rows … w.Commit(); bool ok = SpdfValidator.Validate("signed.spdf").Valid; // recomputes the hash, verifies the signature ``` **The Ed25519 signer is not constant-time.** It is a portable BigInteger implementation (.NET 8 has no built-in Ed25519): signing handles the secret key with variable-time arithmetic, which can leak it through timing to anyone able to measure many signatures on the same machine. Sign only on a trusted machine, never in a shared or multi-tenant service. Verification uses public data only and is safe anywhere. `Ed25519.Sign`, `Ed25519.PublicKey`, `Ed25519.Verify` and `SpdfValidator.SignContentHash` are public for producers that sign outside the writer. ## Command line The repository includes a small CLI (`src/Spdf.Cli`, not published as a package): ```sh dotnet run --project src/Spdf.Cli -- validate file.spdf dotnet run --project src/Spdf.Cli -- dump file.spdf dotnet run --project src/Spdf.Cli -- search file.spdf "lugar de la Mancha" -n 5 dotnet run --project src/Spdf.Cli -- vsearch file.spdf toy-embedding@8 0.5,0.25,0.5,0.25,0,0.25,0.25,0.5 dotnet run --project src/Spdf.Cli -- cite file.spdf q4 --locale en dotnet run --project src/Spdf.Cli -- export file.spdf bibtex # also csl, alto, tei, iiif dotnet run --project src/Spdf.Cli -- sample demo.spdf --key-file seed.hex # small signed demo file dotnet run --project src/Spdf.Cli -- uri parse 'spdf:sha256-…#p=29&f=21' dotnet run --project src/Spdf.Cli -- build source.json out.spdf dotnet run --project src/Spdf.Cli -- conformance ../conformance -o conformance.json ``` ## Building and testing ```sh cd dotnet dotnet build dotnet test # unit tests and the whole conformance suite dotnet run --project src/Spdf.Cli -- conformance ../conformance -o conformance.json dotnet pack src/Spdf.Format -c Release -o artifacts ``` The test project finds the suite at `../conformance` relative to `dotnet/`; set `SPDF_CONFORMANCE_DIR` to use another copy. The runner discovers the cases by listing `conformance/cases/*.json` and prints `{"impl","version","passed","failed","skipped"}`. ## License MIT OR Apache-2.0. --- # SPDF in PHP URL: https://spdf.joseluissaorin.com/docs/php > How to install and use the PHP implementation of SPDF (joseluissaorin/spdf): open, validate, search and cite. PDO SQLite. - **Package**: `joseluissaorin/spdf` - **Install**: `composer require joseluissaorin/spdf` - **Registry**: [Packagist](https://packagist.org/packages/joseluissaorin/spdf) - **Tier**: second - **CI**: CI running - **Folder**: `php/` `joseluissaorin/spdf` reads, validates, searches, cites and writes **SPDF** files (Semantic Processed Document Format): documents that have been read once and can be cited forever. Every passage carries its exact anchor (printed page, folio, second of a recording, slide, verse), so a citation can only print what the source says. It is a native implementation of SPDF 5.0 over PDO SQLite. It also reads the legacy 4.0/4.1 files produced by Scholaris (gzip-wrapped, Spanish schema) through the 5.0 view. It is built for the PHP hosts where journals and libraries live: OJS, Omeka S, WordPress. ## Install ```sh composer require joseluissaorin/spdf ``` Requirements: PHP 8.1 or newer with `pdo_sqlite` (SQLite with FTS5; 3.44+ recommended), `intl`, `mbstring` and `zlib`. `sodium` (bundled with PHP) verifies signatures. ## Read, search and cite ```php use Spdf\Document; $doc = Document::open('lazarillo.spdf'); // read-only, safe opening echo $doc->title(), "\n"; // La vida de Lazarillo de Tormes… echo $doc->cite(['type' => 'image'], null, 'es'); // (Anónimo, 1554) foreach ($doc->searchLexical('"Antona Pérez" Tejares', 5) as $hit) { $f = $doc->fragment($hit['fragment_id']); echo $f['text'], ' ', $doc->cite($hit['anchor'], $f['anchor_end'], 'es'), "\n"; // hijo de Tomé González y de Antona Pérez… (Anónimo, 1554, p. [4]) echo $hit['anchor_uri'], "\n"; // spdf:sha256-3f2a…#p=10&f=4 } ``` Units, fragments, sections, figures, spaces, blobs and provenance are plain arrays with the 5.0 column names (`$doc->units()`, `$doc->fragments()`, `$doc->blob('blob:cover')`…). `$doc->metadata()` is the CSL-JSON item plus the `spdf` extension object. ### Vector and hybrid search ```php $query = $myEmbedder->embed('el ciego y el jarro de vino'); // list, same model as the space $doc->searchVector($query, 'embeddinggemma-2@768', 10); // f32, f16 and i8 spaces $doc->searchHybrid('ciego jarro', $query, 'embeddinggemma-2@768', 10); // RRF, k = 10 ``` These are the reference algorithms of the specification (§6): brute-force dot product (cosine when the space is not normalized) and reciprocal rank fusion over lists of depth `max(limit, 50)`. ### Anchors and URIs ```php use Spdf\AnchorUri; $uri = $doc->anchorUri($fragment['anchor'], $fragment['anchor_end']); $parsed = AnchorUri::parse('spdf:sha256-3f2a…#p=29&f=21&char=118,301'); // ['docref' => 'sha256-3f2a…', 'locator' => ['p' => 29, 'f' => '21', 'char' => [118, 301]]] AnchorUri::format($parsed['docref'], $parsed['locator']); // the same URI, byte for byte $doc->locate($uri); // {document, units, fragments, char, xywh} (SPEC §5.4) $doc->citePassage('f12', 'molinos de viento', 'es'); // {text, uri}: cites the unit the quote lies in (§18.2) ``` ### Bibliography ```php file_put_contents('lazarillo.json', $doc->cslJson()); // Zotero, Pandoc, citeproc file_put_contents('lazarillo.bib', $doc->bibtex()); // @book{lazarillo1554, ... ``` Keys and fields follow the specification (§19): the first author's name, or the first word of the short title, folded to ASCII and lowercased, plus the year (`cervantessaavedra1605`, `lazarillo1554`, `anonnd`); the CSL-JSON `id` is the same key. `Spdf\Bibliography::cslItems()` and `bibtexAll()` export several records and disambiguate colliding keys with `a`, `b`, `c`…; `cslItems([$meta], $anchor, $anchorEnd)` adds the CSL `label` and `locator` of a citation. ### ALTO, TEI and IIIF ```php file_put_contents('lazarillo.alto.xml', $doc->alto()); // ALTO 4, one Page per page unit file_put_contents('lazarillo.tei.xml', $doc->tei()); // TEI P5: pb, p, lg/l, u, note $manifest = $doc->iiif('https://revista.example.org/iiif/lazarillo'); // IIIF Presentation 3 ``` These are the optional exports of SPEC §19.4: printed folios only where the page carries them (`[iv]` marks an inferred folio in TEI and IIIF), sections as IIIF ranges, figures as `describing` annotations on their region, and no invented coordinates. ## Validate ```php $report = Spdf\Validator::validate('file.spdf'); // ['valid' => true, 'version' => '5.0', 'profile' => ['core', 'semantic'], // 'errors' => [], 'warnings' => []] ``` Error and warning codes are those of the specification (§12): `E001` not SQLite, `E020` trigger or view, `E070` FTS index out of sync, `E081` content hash mismatch… ## Write ```php use Spdf\Writer; $w = Writer::create('out.spdf', generator: 'my-journal/1.0', profile: 'core'); $w->document(['id' => 'art-12', 'kind' => 'pdf', 'source_sha256' => hash_file('sha256', 'art-12.pdf'), 'mime' => 'application/pdf', 'bytes' => filesize('art-12.pdf'), 'unit_count' => 1, 'metadata' => ['type' => 'article-journal', 'title' => 'Sobre el Lazarillo', 'author' => [['family' => 'Pérez', 'given' => 'Ana']], 'issued' => ['date-parts' => [[2026]]]]]); $w->unit(['id' => 'p1', 'ord' => 1, 'anchor' => ['type' => 'page', 'physical' => 1, 'printed' => '45'], 'text' => 'Texto de la página…', 'reader' => 'pdf-text-layer']); $w->fragment(['n' => 1, 'id' => 'f1', 'unit' => 'p1', 'ord' => 1, 'text' => 'Texto de la página…', 'anchor' => ['type' => 'page', 'physical' => 1, 'printed' => '45']]); $w->finish(contentHash: true); // FTS rebuilt, no triggers, VACUUM, atomic rename ``` ## Security Files are untrusted input. `Document::open()` opens them read-only with `PRAGMA query_only`, `trusted_schema=OFF`, never loads extensions, refuses triggers and views (except the three FTS triggers of legacy files), bounds blob sizes (512 MiB by default) and gzip inflation (4 GiB), and copies WAL-mode files instead of touching them. Limits are set with `new Spdf\Options(maxBlobBytes: …, maxInflatedBytes: …)`. PDO does not expose `SQLITE_DBCONFIG_DEFENSIVE`; the other measures cover what it guards in a read-only connection. ## In OJS, Omeka S and WordPress `examples/` holds three minimal integrations: - `show-and-cite.php`: a standalone page (title, whole-work citation, search, cited passages, BibTeX and CSL-JSON downloads). `SPDF_DIR=… php -S localhost:8080 examples/show-and-cite.php`. - `ojs/spdfViewer`: an OJS 3.4 generic plugin that renders `.spdf` galleys. - `omeka-s/SpdfViewer`: an Omeka S module with a file renderer for `application/vnd.spdf`. The plugin and the module are sketches to start from; they show the calls, not a finished product. ## Command line ```sh vendor/bin/spdf validate file.spdf vendor/bin/spdf dump file.spdf # canonical dump (RFC 8785) vendor/bin/spdf search file.spdf "molinos de viento" vendor/bin/spdf cite file.spdf f12 en vendor/bin/spdf bibtex file.spdf vendor/bin/spdf conformance ../conformance ``` ## Conformance `php bin/spdf conformance ../conformance` runs the shared suite of the repository and prints `{"impl":"joseluissaorin/spdf (PHP)","version":…,"passed":[…],"failed":[…],"skipped":[…]}`. CI runs it on PHP 8.1 to 8.4 and publishes the report as the `conformance-php` artifact. All kinds are claimed, `export_structure` included. ## License MIT OR Apache-2.0, at your option. The SPDF specification is CC BY 4.0. --- # SPDF in Ruby URL: https://spdf.joseluissaorin.com/docs/ruby > How to install and use the Ruby implementation of SPDF (spdf-format): open, validate, search and cite. On the sqlite3 gem. - **Package**: `spdf-format` - **Install**: `gem install spdf-format` - **Registry**: [RubyGems](https://rubygems.org/gems/spdf-format) - **Tier**: second - **CI**: CI running - **Folder**: `ruby/` `spdf-format` reads, validates, searches, cites and writes **SPDF** files (Semantic Processed Document Format): documents that have been read once and can be cited forever. Every passage carries its exact anchor (printed page, folio, second of a recording, slide, verse), so a citation can only print what the source says. Native implementation of SPDF 5.0 on the `sqlite3` gem. It also reads the legacy 4.0/4.1 files produced by Scholaris (gzip-wrapped, Spanish schema) through the 5.0 view. ## Install ```sh gem install spdf-format ``` ```ruby require "spdf" ``` Ruby 3.1 or newer. The `sqlite3` gem ships SQLite with FTS5; signatures are verified with the standard `openssl` library. ## Read, search and cite ```ruby Spdf::Document.open("lazarillo.spdf") do |doc| # read-only, safe opening puts doc.title puts doc.cite({ "type" => "image" }, locale: "es") # (Anónimo, 1554) doc.search_lexical('"Antona Pérez" Tejares', limit: 5).each do |hit| f = doc.fragment(hit["fragment_id"]) puts "#{f["text"]} #{doc.cite(hit["anchor"], f["anchor_end"], locale: "es")}" # hijo de Tomé González y de Antona Pérez… (Anónimo, 1554, p. [4]) puts hit["anchor_uri"] # spdf:sha256-3f2a…#p=10&f=4 end end ``` Rows are hashes with the 5.0 column names: `doc.units`, `doc.fragments`, `doc.sections`, `doc.figures`, `doc.spaces`, `doc.blob("blob:cover")`, `doc.metadata` (the CSL-JSON item plus the `spdf` extension object). ### Vector and hybrid search ```ruby query = embedder.embed("el ciego y el jarro de vino") # same model as the space doc.search_vector(query, space: "embeddinggemma-2@768", limit: 10) # f32, f16, i8 doc.search_hybrid("ciego jarro", query, space: "embeddinggemma-2@768") # RRF, k = 10 ``` ### Anchors, bibliography ```ruby Spdf::AnchorUri.parse("spdf:sha256-3f2a…#p=29&f=21&char=118,301") # {"docref" => "sha256-3f2a…", "locator" => {"p" => 29, "f" => "21", "char" => [118, 301]}} doc.locate("spdf:sha256-3f2a…#p=29") # {"document", "units", "fragments", "char", "xywh"} (SPEC §5.4) doc.cite_passage("f12", "molinos de viento") # {"text", "uri"}: cites the unit the quote lies in (§18.2) File.write("lazarillo.json", doc.csl_json) # Zotero, Pandoc, citeproc (id = BibTeX key) File.write("lazarillo.bib", doc.bibtex) # @book{lazarillo1554, ... (SPEC §19) Spdf::Bibliography.csl_items([meta], anchor, anchor_end) # adds CSL "label" and "locator" ``` ### ALTO, TEI and IIIF ```ruby File.write("lazarillo.alto.xml", doc.alto) # ALTO 4, one Page per page unit File.write("lazarillo.tei.xml", doc.tei) # TEI P5: pb, p, lg/l, u, note manifest = doc.iiif("https://biblioteca.example.org/iiif/lazarillo") # IIIF Presentation 3 ``` These are the optional exports of SPEC §19.4: printed folios only where the page carries them (`[iv]` marks an inferred folio in TEI and IIIF), sections as IIIF ranges, figures as `describing` annotations, and no invented coordinates. ## Validate ```ruby Spdf::Validator.validate("file.spdf") # {"valid" => true, "version" => "5.0", "profile" => ["core"], "errors" => [], "warnings" => []} ``` ## Write ```ruby Spdf::Writer.create("out.spdf", generator: "my-app/1.0") do |w| w.document("id" => "d1", "kind" => "pdf", "source_sha256" => Digest::SHA256.file("d1.pdf").hexdigest, "mime" => "application/pdf", "bytes" => File.size("d1.pdf"), "unit_count" => 1, "metadata" => { "type" => "book", "title" => "Lazarillo de Tormes", "issued" => { "date-parts" => [[1554]] } }) w.unit("id" => "p1", "ord" => 1, "anchor" => { "type" => "page", "physical" => 1, "printed" => "3" }, "text" => "Pues sepa Vuestra Merced…", "reader" => "pdf-text-layer") w.fragment("n" => 1, "id" => "f1", "unit" => "p1", "ord" => 1, "text" => "Pues sepa Vuestra Merced…", "anchor" => { "type" => "page", "physical" => 1, "printed" => "3" }) end ``` ## Security Files are untrusted input: they are opened read-only with `query_only` and `trusted_schema=OFF`, extensions are never loaded, triggers and views are refused (except the three FTS triggers of legacy files), blob sizes (512 MiB) and gzip inflation (4 GiB) are bounded, and WAL-mode files are copied. `Spdf::Document.open(path, max_blob_bytes: …)` changes the limits. The gem does not expose `SQLITE_DBCONFIG_DEFENSIVE`. ## Command line and conformance ```sh spdf validate file.spdf spdf dump file.spdf spdf search file.spdf "molinos de viento" spdf cite file.spdf f12 en spdf conformance path/to/spdf/conformance ``` `spdf conformance` runs every case of the shared suite and prints the report of the specification (§21). CI publishes it as the `conformance-ruby` artifact. All kinds are claimed, `export_structure` included. ## License MIT OR Apache-2.0, at your option. The SPDF specification is CC BY 4.0. --- # SPDF in R URL: https://spdf.joseluissaorin.com/docs/r > How to install and use the R implementation of SPDF (spdf): open, validate, search and cite. On RSQLite; fragments as data frames. - **Package**: `spdf` - **Install**: `remotes::install_github("joseluissaorin/spdf", subdir = "r")` - **Tier**: second - **CI**: CI running - **Folder**: `r/` Read, validate, search, cite and write **SPDF** files (Semantic Processed Document Format) from R. A SPDF file holds a document that has been read once and can be cited forever: every passage carries its exact anchor (printed page, folio, second of a recording, slide, verse), so a citation can only print what the source says. The package is a native implementation of SPDF 5.0 on RSQLite. It also reads the legacy 4.0/4.1 files produced by Scholaris (gzip-wrapped, Spanish schema) through the 5.0 view. Tables come back as tibbles. ## Install ```r # from CRAN, once published install.packages("spdf") # from the repository remotes::install_github("joseluissaorin/spdf", subdir = "r") ``` ## Read, search, cite ```r library(spdf) doc <- spdf_open(system.file("extdata", "quijote.spdf", package = "spdf")) spdf_info(doc) # title, authors, year, version, counts spdf_units(doc) # one row per citable unit (page, folio, time span...) fr <- spdf_fragments(doc) # searchable passages, anchors as list-columns hits <- spdf_search(doc, "hermoso") # finds the long-s "hermoso" through the modern layer hits$anchor_uri # spdf:sha256-27ea...#p=13&char=10,194 spdf_cite(spdf_metadata(doc), fr$anchor[[5]], fr$anchor_end[[5]], locale = "es") #> "(Cervantes Saavedra, 1608, fols. Ir-[Iv])" spdf_cite_passage(doc, "q5", "rozin, como tomaua la podadera.")$text #> "(Cervantes Saavedra, 1608, fol. [Iv])" the page the quotation is on spdf_locate(doc, hits$anchor_uri[1]) # list(document, units, fragments, char, xywh) cat(spdf_bibtex(doc)) # @book{cervantessaavedra1608, ... (also spdf_csl()) spdf_close(doc) ``` Vector and hybrid search take a query vector computed with the same model as the space: `spdf_search_vector(doc, v, "embeddinggemma-2@768")`, `spdf_search_hybrid(doc, "ciego jarro", v, "embeddinggemma-2@768")`. ## ALTO, TEI and IIIF ```r writeLines(spdf_alto(doc), "quijote.alto.xml") # ALTO 4, one Page per page unit writeLines(spdf_tei(doc), "quijote.tei.xml") # TEI P5: pb, p, lg/l, u, note writeLines(spdf_iiif_json(doc, "https://example.org/iiif/quijote"), "manifest.json") ``` ## Corpora ```r files <- list.files("corpus", pattern = "\\.spdf$", full.names = TRUE) spdf_corpus(files) # one row per document spdf_corpus_search(files, "\"molinos de viento\"") # one row per passage, with citation spdf_count_terms(files, c("honra", "fortuna")) # fragments and occurrences per work ``` The vignette `vignette("corpus", package = "spdf")` (in Spanish) walks through a digital-humanities workflow: searching a corpus and counting occurrences by work and year, with every number traceable to its page. ## Validate and write ```r spdf_validate("file.spdf") # list(valid, version, profile, errors, warnings) spdf_write("out.spdf", document = ..., units = ..., fragments = ...) ``` ## Security Files are untrusted input: `spdf_open()` connects read-only with `query_only` and `trusted_schema=OFF`, never loads extensions, refuses triggers, views and foreign virtual tables (except the three FTS triggers of legacy files), bounds blob sizes and gzip inflation, and copies WAL-mode files before opening them. ## Conformance `spdf_conformance("path/to/spdf/conformance")` runs the shared suite of the specification; `Rscript inst/scripts/conformance.R ../conformance` prints the JSON report. CI publishes it as the `conformance-r` artifact. All kinds are claimed, `export_structure` included. ## License MIT OR Apache-2.0, at your option. The SPDF specification is CC BY 4.0. The sample files in `inst/extdata` are short excerpts of public-domain works. --- # SPDF in Julia URL: https://spdf.joseluissaorin.com/docs/julia > How to install and use the Julia implementation of SPDF (SPDF.jl): open, validate, search and cite. On SQLite.jl. - **Package**: `SPDF.jl` - **Install**: `pkg> add SPDF` - **Tier**: second - **CI**: CI running - **Folder**: `julia/` Read, validate, search, cite and write **SPDF** files (Semantic Processed Document Format) from Julia. A SPDF file holds a document that has been read once and can be cited forever: every passage carries its exact anchor (printed page, folio, second of a recording, slide, verse), so a citation can only print what the source says. Native implementation of SPDF 5.0 on SQLite.jl. It also reads the legacy 4.0/4.1 files produced by Scholaris (gzip-wrapped, Spanish schema) through the 5.0 view. ## Install ```julia using Pkg Pkg.add("SPDF") # once registered; until then: Pkg.add(url = "https://github.com/joseluissaorin/spdf", subdir = "julia") ``` ## Read, search, cite ```julia using SPDF SPDF.open("quijote.spdf") do doc println(title(doc), " (", doc.version, ")") for hit in search_lexical(doc, "hermoso"; limit = 5) # the 1608 edition prints «hermoſo» f = SPDF.fragment(doc, hit.fragment_id) println(cite(doc, hit.anchor, f["anchor_end"]; locale = "es")) # (Cervantes Saavedra, 1608, s. p.) println(hit.anchor_uri) # spdf:sha256-27ea…#p=13&char=10,194 end # a quotation is cited by the page it lies in, not by the start of its fragment (§18.2) println(cite_passage(doc, "q5", "rozin, como tomaua la podadera.")["text"]) # (Cervantes Saavedra, 1608, fol. [Iv]) println(bibtex(doc)) end ``` `units(doc)`, `fragments(doc)`, `sections(doc)`, `figures(doc)`, `spaces(doc)`, `blobs(doc)` and `provenance(doc)` return vectors of `Dict`s with the 5.0 column names; `metadata(doc)` is the CSL-JSON item. `vectors(doc, space)` gives the decoded vectors (f32, f16 or i8) as `Dict(id => Vector{Float64})`. ```julia search_vector(doc, qvec, "embeddinggemma-2@768"; limit = 10) search_hybrid(doc, "ciego jarro", qvec, "embeddinggemma-2@768") # RRF, k = 10 parse_uri("spdf:sha256-3f2a…#p=29&f=21&char=118,301") locate(doc, "spdf:sha256-3f2a…#p=29") # Dict("document", "units", "fragments", "char", "xywh") csl_item(doc)["id"] # "cervantessaavedra1605", the BibTeX key (SPEC §19) validate("file.spdf") # Dict("valid" => true, "errors" => [], "warnings" => [], …) ``` ## ALTO, TEI and IIIF ```julia write("quijote.alto.xml", alto(doc)) # ALTO 4, one Page per page unit write("quijote.tei.xml", tei(doc)) # TEI P5: pb, p, lg/l, u, note manifest = iiif(doc, "https://example.org/iiif/quijote") # IIIF Presentation 3 (a Dict) ``` ## Write ```julia w = SPDF.Writer("out.spdf"; generator = "my-app/1.0", profile = "core") SPDF.document!(w, Dict("id" => "d1", "kind" => "pdf", "source_sha256" => bytes2hex(sha256(read("d1.pdf"))), "mime" => "application/pdf", "bytes" => filesize("d1.pdf"), "unit_count" => 1, "metadata" => Dict("type" => "book", "title" => "Lazarillo de Tormes", "issued" => Dict("date-parts" => [[1554]])))) SPDF.unit!(w, Dict("id" => "p1", "ord" => 1, "anchor" => Dict("type" => "page", "physical" => 1, "printed" => "3"), "text" => "Pues sepa Vuestra Merced…", "reader" => "pdf-text-layer")) SPDF.fragment!(w, Dict("n" => 1, "id" => "f1", "unit" => "p1", "ord" => 1, "text" => "Pues sepa Vuestra Merced…", "anchor" => Dict("type" => "page", "physical" => 1, "printed" => "3"))) SPDF.finish!(w) # FTS rebuilt, no triggers or views, VACUUM, atomic rename ``` ## Security Files are untrusted input: `SPDF.open` connects read-only (`mode=ro`), sets `query_only`, `trusted_schema=OFF` and `SQLITE_DBCONFIG_DEFENSIVE`, never loads extensions, refuses triggers, views and foreign virtual tables (except the three FTS triggers of legacy files), bounds blob sizes and gzip inflation, and copies WAL-mode files before opening them. Signatures are verified with a small pure-Julia Ed25519. ## Conformance `SPDF.conformance("path/to/spdf/conformance")` runs the shared suite; `Pkg.test()` runs it too when the package lives in the SPDF repository, and `julia --project=. bin/conformance.jl ../conformance` prints the JSON report. CI publishes it as the `conformance-julia` artifact. All kinds are claimed, `export_structure` included. ## License MIT OR Apache-2.0, at your option. The SPDF specification is CC BY 4.0. --- # SPDF in C URL: https://spdf.joseluissaorin.com/docs/c > How to install and use the C implementation of SPDF (libspdf): open, validate, search and cite. C ABI over the Rust core, for C, C++ and any FFI. - **Package**: `libspdf` - **Install**: `#include "spdf.h" /* link with -lspdf */` - **Tier**: second - **CI**: CI running - **Folder**: `c/` C and C++ access to **SPDF** files (Semantic Processed Document Format) through the C ABI of the Rust reference implementation (`rust/crates/spdf-ffi`, header `spdf.h`). Everything the other implementations do is here: safe opening of SPDF 5.0 and legacy 4.x files, validation, the canonical dump, the reference lexical, vector and hybrid searches, anchor URIs (format, parse, locate), short citations, CSL-JSON and BibTeX, and writing files from a dump. SQLite is bundled inside the library. This folder adds: - `CMakeLists.txt`: builds the Rust library with cargo and exposes the CMake target `spdf::spdf` (static by default, `-DSPDF_SHARED=ON` for the shared library, `-DSPDF_PREBUILT_DIR=…` to use a prebuilt `libspdf_ffi`); - `include/spdf.hpp`: a header-only C++17 wrapper (RAII `spdf::Document`, exceptions, `std::string` results); - `examples/`: `search.c`, `validate.c` and `search.cpp`; - `tests/conformance.c`: the conformance runner, which drives every case through the ABI; - `src/sjson.c`: a small JSON reader used by the examples and the runner (not part of the ABI); - `spdf.pc.in`: a pkg-config file for installations. ## Build ```sh cmake -S c -B build -DCMAKE_BUILD_TYPE=Release cmake --build build ctest --test-dir build --output-on-failure ``` Requirements: a C11 and C++17 compiler, CMake 3.16+, and Rust (cargo) unless `SPDF_PREBUILT_DIR` points to a built library. Static linking pulls in `-lpthread -ldl -lm` on Linux and `-framework Security -framework CoreFoundation` on macOS; the CMake target adds them. ## C ```c #include "spdf.h" SpdfDoc *doc = NULL; if (spdf_open("quijote.spdf", NULL, &doc) != SPDF_OK) { /* read-only, safe opening */ fprintf(stderr, "%s\n", spdf_last_error()); /* {"status","code","message"} */ return 1; } char *hits = NULL; if (spdf_search_lexical(doc, "hermoso", 10, &hits) == SPDF_OK) { /* the 1608 edition prints «hermoſo» */ puts(hits); /* [{"fragment_id":"q1","score":…,"via":["lexical"],"anchor":{…},"anchor_uri":"spdf:sha256-27ea…#p=13&char=10,194"}] */ spdf_string_free(hits); } char *cite = NULL; /* a quotation is cited by the page it lies in (SPEC §18.2) */ spdf_cite_passage(doc, "q5", "rozin, como tomaua la podadera.", "es", &cite); puts(cite); /* {"text":"(Cervantes Saavedra, 1608, fol. [Iv])","uri":"spdf:sha256-27ea…#p=30&f=Iv&char=130,161",…} */ spdf_string_free(cite); spdf_close(doc); ``` Every function returns an `int` status (`SPDF_OK` = 0) and writes its result through an out parameter; strings returned by the library are freed with `spdf_string_free`, byte buffers with `spdf_bytes_free`. Complex values are JSON with the shapes of the specification. ## C++ ```cpp #include "spdf.hpp" spdf::Document doc("quijote.spdf"); std::string hits = doc.search("hidalgo", 5); // JSON array std::string where = doc.locate("spdf:sha256-…#p=5"); // {"document","units","fragments","char","xywh"} std::cout << doc.bibtex(); // @book{cervantessaavedra1608, … std::string tei = doc.tei(); // also alto() and iiif(base_url) std::string bib = spdf::export_bibtex({&doc, &other}); // several documents, keys disambiguated ``` Errors throw `spdf::Error` (with `status()` and the JSON of `spdf_last_error()`). ## Conformance `build/spdf_conformance ../conformance` runs every case through the ABI and prints the report of the specification; `ctest` runs it. CI publishes it as the `conformance-c` artifact. All kinds are claimed: `roundtrip` through `spdf_write_from_dump`, `quantize` through `spdf_quantize`, `locate` through `spdf_locate`, the exports through `spdf_export_csl_multi`, `spdf_export_bibtex_multi` and `spdf_export_structure`, and `cite_passage` through `spdf_cite_passage`. ## License MIT OR Apache-2.0, at your option. The SPDF specification is CC BY 4.0. --- # MCP server: spdf-mcp URL: https://spdf.joseluissaorin.com/integrations/mcp > Any agent searches a folder of SPDF files and cites with the exact folio. A [Model Context Protocol](https://modelcontextprotocol.io) server for [SPDF](https://spdf.joseluissaorin.com) files. Point it at a folder of `.spdf` documents and any agent (Claude, ChatGPT, Cursor, Zed, your own) can list them, search them, read a passage, look at the figures and **cite with the exact printed folio, without being able to invent one**. It is a thin layer over [`spdf-format`](../../js), the official TypeScript implementation: the text comes from the file, the citation comes from the stored anchor, and a page that does not exist is an error, never an approximation. ## Run it ```sh npx spdf-mcp ~/Library/SPDF # stdio (what desktop clients use) npx spdf-mcp ~/Library/SPDF --locale es # citations in Spanish by default npx spdf-mcp ~/Library/SPDF --http 8765 # Streamable HTTP on http://127.0.0.1:8765/mcp ``` Options: `--http `, `--host ` (default `127.0.0.1`), `--locale en|es`, `--no-recursive`. Several folders or files can be given. Unreadable or unsafe files (for example one with a view or a trigger) are skipped and reported by `list_documents`. ### Claude Code ```sh claude mcp add spdf -- npx spdf-mcp ~/Library/SPDF ``` ### Any client with a JSON configuration (Claude Desktop, Cursor, Zed…) ```json { "mcpServers": { "spdf": { "command": "npx", "args": ["spdf-mcp", "/path/to/library"] } } } ``` ## Tools All tools are read-only. Each returns readable JSON as text and the same data as structured content. | Tool | Arguments | Returns | | --- | --- | --- | | `list_documents` | `filter?`, `refresh?` | Every document: `doc` reference (`sha256-…` of the original), title, authors, year, kind, language, units, fragments, figures, vector spaces; and the files that were skipped | | `search` | `query`, `docs?`, `limit?` (1–50), `vector?` + `space?`, `locale?` | Passages with their literal `text`, `citation`, `anchor_uri`, section, context and score. Lexical search uses the SPDF reference algorithm (accent-insensitive, `"phrases"`); with a query vector in a space the files carry it is hybrid (reciprocal rank fusion, k = 10). Results from several files are merged by rank | | `read_passage` | `doc` + one of `fragment_id`, `folio`, `page`, `time`; or `uri`; `around?`; `locale?` | The literal text, citation, anchor URI, who read it and with what confidence, warnings, and optionally the neighbouring fragments | | `cite` | same as `read_passage`, plus `reference?` | `citation`, `anchor_uri` and `quote` together; with `reference` also CSL-JSON and BibTeX | | `list_figures` | `doc?`, `figure_id?`, `include_image?`, `locale?` | Figures, plates and frames with caption, description, citation, anchor URI and region; the image itself when asked for one figure | | `get_metadata` | `doc` | CSL-JSON, BibTeX, rights, SHA-256 of the original and provenance | `doc` accepts the full reference from `list_documents`, a unique prefix of its hash, the document id or the file name. ### What the citations look like | Passage | `cite` returns | | --- | --- | | Physical page 2, printed folio 1 | `(Saorín Ferrer, 2026, p. 1)` | | A plate whose folio was inferred | `(Saorín Ferrer, 2026, p. [3])` plus a warning to keep the brackets | | A cover with no printed folio | `(Saorín Ferrer, 2026, n. pag.)` (`s. p.` in Spanish) plus a warning | | A folio that does not exist | an error: *"… has no page with printed folio 21. Printed folios: 1, 2, 3, 4, 5. Do not cite it."* | The server also sends the model a short set of instructions (the MCP `instructions` field) with the rules for citing without inventing; they are the same as on [SPDF for agents](https://spdf.joseluissaorin.com/agents). ## As a library ```js import { Library, createServer } from 'spdf-mcp'; const lib = await Library.open(['./library'], { locale: 'en' }); const server = createServer(lib); // an McpServer from @modelcontextprotocol/sdk await server.connect(myTransport); ``` ## Tests ```sh cd js && npm ci && npm run build # the official library, once cd integrations/spdf-mcp && npm ci && npm test ``` The tests use the MCP SDK itself as the client, three ways: in memory (every tool and every error path), over **stdio** against the compiled binary, and over **Streamable HTTP** (including the DNS-rebinding guard). They run against the sample files in [`../fixtures`](../fixtures). ## Security - Files are opened read-only, with `trusted_schema=OFF`, defensive mode and no extensions, and files with triggers or views are refused (`spdf-format` does this; the server never writes). - The HTTP transport listens on `127.0.0.1` by default, is stateless, accepts only `POST /mcp` and rejects requests whose `Host` is not the one it listens on. There is no authentication: do not expose it to a network you do not trust. ## Licence MIT OR Apache-2.0. --- # LangChain.js loader URL: https://spdf.joseluissaorin.com/integrations/langchain-js > Passages as LangChain documents with citation and anchor URI. A [LangChain.js](https://js.langchain.com) document loader for [SPDF](https://spdf.joseluissaorin.com) files. Every passage becomes a `Document` with its **literal text** and, in the metadata, its **citation with the exact printed folio** (or second, slide, verse) and its **anchor URI**, so a retrieval-augmented answer can cite the page a reader will find on paper instead of a chunk number. It sits on [`spdf-format`](../../js), the official TypeScript implementation: the citation is computed from the anchor stored in the file, never generated. ## Install ```sh npm install spdf-langchain @langchain/core ``` ## Use ```js import { SpdfLoader } from 'spdf-langchain'; const docs = await new SpdfLoader('darwin-origin.spdf').load(); docs[0].pageContent; // the literal passage docs[0].metadata.citation; // '(Darwin, 1859, p. 21)' docs[0].metadata.anchor_uri; // 'spdf:sha256-…#p=29&f=21&char=118,301' // A folder (recursive), Spanish citations, one document per page, skipping broken files: const pages = await new SpdfLoader('library/', { locale: 'es', granularity: 'unit', skipInvalid: true }).load(); // Streaming: for await (const d of new SpdfLoader('library/').lazyLoad()) console.log(d.metadata.citation); ``` When you answer from retrieved documents, quote `pageContent` and cite with `metadata.citation`; keep `metadata.anchor_uri` next to the claim. ## Options | Option | Default | Meaning | | --- | --- | --- | | `granularity` | `'fragment'` | `'fragment'` (passages of 150 to 300 words, the unit SPDF searches and cites) or `'unit'` (whole pages, time spans, slides) | | `locale` | `'en'` | Language of `citation`: `'en'` or `'es'` | | `recursive` | `true` | Descend into subfolders | | `skipInvalid` | `false` | Skip files that cannot be opened safely (with a warning on stderr) instead of failing | | `embeddingsFrom` | none | A vector space stored in the files (for example `all-MiniLM-L6-v2@384`): its vector goes to `metadata.embedding`, so you can index without re-embedding when your query model is the same | ## Metadata All values are scalars (string, number or boolean) and keys whose value would be null are left out, so every vector store accepts them (Chroma, for one, rejects nulls). The keys match the Python loaders. | Key | Example | | --- | --- | | `citation` | `(Saorín Ferrer, 2026, p. 1)`; `p. [3]` when the folio was inferred; `n. pag.` (`s. p.`) when the page has none | | `anchor_uri` | `spdf:sha256-50d9…#p=2&f=1&char=15,307` | | `printed_folio`, `physical_page`, `folio_inferred` | `'1'`, `2`, `false` (page anchors; no `printed_folio` key when the page has none) | | `t0`, `t1`, `speaker` | seconds (time anchors) | | `slide`, `line_from`, `line_to` | slides and verses | | `title`, `authors`, `year`, `language`, `kind` | from the CSL record | | `section`, `context` | `'I. Anchors'`, one line situating the passage | | `fragment_id` or `unit_id`, `ord`, `spdf_doc_id`, `docref`, `source`, `spdf_version`, `anchor_type` | identifiers (`spdf_doc_id`, not `doc_id`: vector stores and parent-document retrievers overwrite `doc_id`) | | `anchor`, `anchor_end` | the full anchors as JSON strings | ## Tests ```sh cd js && npm ci && npm run build # the official library, once cd integrations/langchain-js && npm ci && npm test ``` They run against [`../fixtures`](../fixtures) and include a LangChain retriever over an in-memory vector store that returns documents with their citation intact. ## Licence MIT OR Apache-2.0. --- # LlamaIndex.TS reader URL: https://spdf.joseluissaorin.com/integrations/llamaindex-js > Passages as LlamaIndex documents, with the stored vectors if you want them. A [LlamaIndex.TS](https://ts.llamaindex.ai) reader for [SPDF](https://spdf.joseluissaorin.com) files. Every passage becomes a `Document` with its **literal text** and, in the metadata, its **citation with the exact printed folio** (or second, slide, verse) and its **anchor URI**, so a retrieval-augmented answer can cite the page a reader will find on paper instead of a chunk number. It can also hand over the **vectors already stored in the file**, so an index can be built without embedding anything again. It sits on [`spdf-format`](../../js), the official TypeScript implementation: the citation is computed from the anchor stored in the file, never generated. ## Install ```sh npm install spdf-llamaindex @llamaindex/core ``` ## Use ```js import { VectorStoreIndex } from 'llamaindex'; import { SpdfReader } from 'spdf-llamaindex'; const docs = await new SpdfReader({ locale: 'en' }).loadData('library/'); docs[0].metadata.citation; // '(Darwin, 1859, p. 21)' docs[0].metadata.anchor_uri; // 'spdf:sha256-…#p=29&f=21&char=118,301' const index = await VectorStoreIndex.fromDocuments(docs); const nodes = await index.asRetriever({ similarityTopK: 5 }).retrieve('natural selection'); nodes.map((n) => n.node.metadata.citation); ``` The `citation` is visible to the LLM (`MetadataMode.LLM`), so a query engine's answer can quote it; the anchor JSON, hashes and identifiers are excluded from the text that gets embedded (`EXCLUDED_EMBED_METADATA`, `EXCLUDED_LLM_METADATA`). With `SimpleDirectoryReader`, register it for the extension: `fileExtToReader: { spdf: new SpdfReader() }` (it implements `loadDataAsContent`). ### Reusing the stored vectors ```js const docs = await new SpdfReader({ embeddingsFrom: 'all-MiniLM-L6-v2@384' }).loadData('library/'); // docs[i].embedding is set from the file: index with an embed model of the same space. ``` ## Options | Option | Default | Meaning | | --- | --- | --- | | `granularity` | `'fragment'` | `'fragment'` (passages of 150 to 300 words) or `'unit'` (whole pages, time spans, slides) | | `locale` | `'en'` | Language of `citation`: `'en'` or `'es'` | | `recursive` | `true` | Descend into subfolders | | `skipInvalid` | `false` | Skip files that cannot be opened safely instead of failing | | `embeddingsFrom` | none | A vector space stored in the files: sets `Document.embedding` | ## Metadata Scalars only (string, number or boolean; keys with a null value are left out): `citation`, `anchor_uri`, `printed_folio`, `physical_page`, `folio_inferred`, `t0`/`t1`/`speaker`, `slide`, `line_from`/`line_to`, `title`, `authors`, `year`, `language`, `kind`, `section`, `context`, `fragment_id` or `unit_id`, `ord`, `spdf_doc_id`, `docref`, `source`, `spdf_version`, and the full `anchor`/`anchor_end` as JSON strings. See the table in [`spdf-langchain`](../langchain-js/README.md#metadata). ## Tests ```sh cd js && npm ci && npm run build # the official library, once cd integrations/llamaindex-js && npm ci && npm test ``` They run against [`../fixtures`](../fixtures) and build a `VectorStoreIndex` with a deterministic toy embedding (no network, no keys) to check that retrieved nodes keep their citation. ## Licence MIT OR Apache-2.0. --- # LangChain loader (Python) URL: https://spdf.joseluissaorin.com/integrations/langchain-python > The same loader for LangChain in Python. A [LangChain](https://www.langchain.com/) document loader for **SPDF** files (Semantic Processed Document Format), so that retrieval-augmented answers cite the printed page instead of a chunk number. An SPDF file is a SQLite database holding one document that has already been read: every passage (fragment) carries its exact anchor (physical page and printed folio, second of a recording, slide, verse…). `SpdfLoader` turns each passage into a LangChain `Document` whose metadata holds a ready-made short citation such as `(Saorín Ferrer, 2026, p. 1)` and a portable anchor URI such as `spdf:sha256-…#p=2&f=1&char=15,307`, both computed by the official library [`spdf-format`](../../python). ## Install ```bash pip install spdf-langchain ``` Requires Python 3.10 or later, `langchain-core>=0.3` and `spdf-format` (standard library only). From a checkout of the repository: ```bash pip install -e python/ -e "integrations/langchain-python[test]" ``` ## Usage ```python from spdf_langchain import SpdfLoader loader = SpdfLoader("spdf-in-five-pages.spdf", locale="en") # a file, a folder or a list of both docs = loader.load() # or: for doc in loader.lazy_load(): ... docs[0].page_content # the literal passage, exactly as in the source docs[0].metadata["citation"] # '(Saorín Ferrer, 2026, p. 1)' docs[0].metadata["anchor_uri"] # 'spdf:sha256-50d9…5f4c#p=2&f=1&char=15,307' docs[0].id # 'sha256-50d9…5f4c:f2-1' (stable across runs) ``` - A **folder** is searched recursively for `*.spdf` files (hidden files and folders are skipped); a **list** may mix files and folders. `lazy_load()` yields the documents one by one, one file open at a time; `load()`, `aload()` and `alazy_load()` come from `BaseLoader`. - Unsafe or invalid files are refused by `spdf-format`: loading one raises its error (`spdf.UnsafeFileError`, `spdf.NotSpdfError`…, all subclasses of `spdf.SpdfError`) with the validation code (`E020`…) and the file path in the message. A missing path raises `FileNotFoundError`. - Legacy SPDF 4.0 and 4.1 files (also gzip-wrapped) read like 5.0 ones. - Vector stores in recent `langchain-core` versions take the ids from `Document.id`; with older ones, pass them yourself: `store.add_documents(docs, ids=[d.id for d in docs])`. ### Options | Option | Default | Meaning | | --- | --- | --- | | `granularity` | `"fragment"` | `"fragment"`: one document per passage (about 150 to 300 words). `"unit"`: one per page, time span, slide… | | `locale` | `"en"` | Locale of `citation`: `"en"` or `"es"` (`"es-ES"` works; others fall back to English). | | `with_vectors` | `None` | Id of a vector space stored in the files (`"all-MiniLM-L6-v2@384"`). Puts the stored vector in `metadata["vector"]` (a list of floats) and the space id in `metadata["vector_space"]`. Off by default, because most vector stores expect flat metadata. A file without that space raises `spdf.SpdfError`. | ## Metadata Values are flat scalars (`str`, `int`, `float`, `bool`), so every vector store accepts them (the only exception is `vector`, and only if you ask for it). **A key whose value would be null is left out** (Chroma and others reject `None`): the cover of a book has no `printed_folio` key, a page has no `t0`. | Key | Type | Example (first passage of the English fixture) | Notes | | --- | --- | --- | --- | | `source` | str | `fixtures/spdf-in-five-pages.spdf` | Path the file was read from. | | `spdf_version` | str | `5.0` | `4.0` or `4.1` for legacy files. | | `spdf_doc_id` | str | `spdf-in-five-pages` | The document id inside the file. Not called `doc_id`, which LangChain's multi-vector and parent-document retrievers (and LlamaIndex vector stores) use for their own ids. | | `docref` | str | `sha256-50d94244…5f4c` | Document reference used by anchor URIs (SHA-256 of the original). | | `title` | str | `SPDF in five pages` | | | `authors` | str | `Saorín Ferrer` | | | `year` | int | `2026` | | | `language` | str | `en` | BCP 47. | | `kind` | str | `pdf` | `pdf`, `epub`, `audio`, `video`… | | `fragment_id` | str | `f2-1` | Fragment granularity only. | | `unit_id` | str | `u2` | The unit (page…) where the passage starts. | | `anchor_type` | str | `page` | `page`, `time`, `section`, `slide`, `sheet`, `web`, `image`, `verse`, `canonical`. | | `physical_page` | int | `2` | Page anchors: position of the page in the file. | | `printed_folio` | str | `1` | The folio as printed (`"xiv"`, `"1r"`). | | `folio_inferred` | bool | `false` | True when the folio was deduced, not read; the citation prints it in brackets, `p. [3]`. | | `section` | str | `I. Anchors` | Heading path joined with `" / "`. In unit granularity, the sections that share the unit are joined with `" \| "`. | | `context` | str | `SPDF in five pages, I. Anchors` | One line that situates the passage (fragment granularity). | | `anchor` | str | `{"chars":[15,307],"confidence":1,"physical":2,"printed":"1","source":"read","type":"page"}` | The start anchor as canonical JSON; `json.loads` it for the full object. | | `anchor_end` | str | | End anchor (JSON) when the passage crosses into another unit. | | `t0`, `t1` | float | | Seconds, for time anchors (recordings). | | `anchor_uri` | str | `spdf:sha256-50d94244…5f4c#p=2&f=1&char=15,307` | Resolve it with `spdf.open(path).locate(uri)`; parse it with `spdf.parse_uri`. | | `citation` | str | `(Saorín Ferrer, 2026, p. 1)` | Short author-date citation in the chosen locale. | | `vector`, `vector_space` | list, str | | Only with `with_vectors`. | `page_content` is always the literal passage (`fragments.text` or `units.text`), never the modernised-spelling search layer, which SPDF forbids quoting. ## End-to-end example (no API key) Retrieval with a toy embedding and an in-memory vector store, then a prompt in which every passage carries its citation; pipe the prompt into any chat model. ```python import hashlib import math import re from langchain_core.embeddings import Embeddings from langchain_core.prompts import ChatPromptTemplate from langchain_core.vectorstores import InMemoryVectorStore from spdf_langchain import SpdfLoader class HashingEmbeddings(Embeddings): """A toy bag-of-words embedding: no model to download, no API key.""" def __init__(self, dim: int = 512) -> None: self.dim = dim def _vec(self, text: str) -> list[float]: v = [0.0] * self.dim for word in re.findall(r"\w+", text.lower()): v[int(hashlib.md5(word.encode()).hexdigest(), 16) % self.dim] += 1.0 norm = math.sqrt(sum(x * x for x in v)) or 1.0 return [x / norm for x in v] def embed_documents(self, texts: list[str]) -> list[list[float]]: return [self._vec(t) for t in texts] def embed_query(self, text: str) -> list[float]: return self._vec(text) docs = SpdfLoader("integrations/fixtures/spdf-in-five-pages.spdf", locale="en").load() store = InMemoryVectorStore.from_documents(docs, embedding=HashingEmbeddings()) # needs numpy question = "How is a plate without a printed folio cited?" hits = store.similarity_search(question, k=2) for d in hits: print(d.metadata["citation"], d.metadata["anchor_uri"]) prompt = ChatPromptTemplate.from_messages([ ("system", "Answer from the passages only. After each claim, copy the citation of its passage."), ("human", "{context}\n\nQuestion: {question}"), ]) context = "\n\n".join(f"{d.page_content} {d.metadata['citation']}" for d in hits) messages = prompt.invoke({"context": context, "question": question}) # answer = chat_model.invoke(messages) ``` Output: ```text (Saorín Ferrer, 2026, p. [3]) spdf:sha256-50d94244…5f4c#p=4&f=3&char=0,133 (Saorín Ferrer, 2026, p. 2) spdf:sha256-50d94244…5f4c#p=3&f=2&char=19,258 ``` The plate carries no printed number; its folio is inferred, so the citation prints it in brackets. The context the model receives reads: ```text Plate I. A page with its folio and a manicule pointing at a passage. This plate carries no printed number; its folio, 3, is inferred. (Saorín Ferrer, 2026, p. [3]) Every unit records who read it: … the citation puts it in brackets. (Saorín Ferrer, 2026, p. 2) ``` ## Reusing the vectors stored in the file SPDF files may ship vectors (`f.spaces()` in `spdf-format` lists them, with model, size and any task prefixes). When your embedding model is the one that produced a space, load the vectors instead of embedding every passage again, for example with FAISS: ```python from langchain_community.vectorstores import FAISS # pip install langchain-community faiss-cpu from langchain_huggingface import HuggingFaceEmbeddings # pip install langchain-huggingface from spdf_langchain import SpdfLoader docs = SpdfLoader("library/", with_vectors="all-MiniLM-L6-v2@384").load() pairs = [(d.page_content, d.metadata.pop("vector")) for d in docs] # keep the metadata flat store = FAISS.from_embeddings( pairs, HuggingFaceEmbeddings(model_name="sentence-transformers/all-MiniLM-L6-v2"), # embeds queries only metadatas=[d.metadata for d in docs], ids=[d.id for d in docs], ) print(store.similarity_search("inferred folio", k=1)[0].metadata["citation"]) ``` ## Development ```bash cd integrations/langchain-python uv venv && uv pip install -e ../../python -e ".[test]" .venv/bin/python -m pytest ``` The tests use the shared fixtures in `integrations/fixtures/`. ## License MIT OR Apache-2.0, at your option. --- # LlamaIndex reader (Python) URL: https://spdf.joseluissaorin.com/integrations/llamaindex-python > The same reader for LlamaIndex in Python. A [LlamaIndex](https://www.llamaindex.ai/) reader for **SPDF** files (Semantic Processed Document Format), so that retrieval-augmented answers cite the printed page instead of a chunk number. An SPDF file is a SQLite database holding one document that has already been read: every passage (fragment) carries its exact anchor (physical page and printed folio, second of a recording, slide, verse…). `SpdfReader` turns each passage into a LlamaIndex `Document` whose metadata holds a ready-made short citation such as `(Saorín Ferrer, 2026, p. 1)` and a portable anchor URI such as `spdf:sha256-…#p=2&f=1&char=15,307`, both computed by the official library [`spdf-format`](../../python). ## Install ```bash pip install spdf-llamaindex ``` Requires Python 3.10 or later, `llama-index-core>=0.12` and `spdf-format` (standard library only). From a checkout of the repository: ```bash pip install -e python/ -e "integrations/llamaindex-python[test]" ``` ## Usage ```python from spdf_llamaindex import SpdfReader reader = SpdfReader(locale="en") docs = reader.load_data("spdf-in-five-pages.spdf") # a file, a folder or a list of both docs[0].text # the literal passage, exactly as in the source docs[0].metadata["citation"] # '(Saorín Ferrer, 2026, p. 1)' docs[0].metadata["anchor_uri"] # 'spdf:sha256-50d9…5f4c#p=2&f=1&char=15,307' docs[0].id_ # 'sha256-50d9…5f4c:f2-1' (stable across runs) ``` - A **folder** is searched recursively for `*.spdf` files (hidden files and folders are skipped); a **list** may mix files and folders. `lazy_load_data()` yields the documents one by one, one file open at a time. - With `SimpleDirectoryReader`, register the reader for the extension: `SimpleDirectoryReader("library/", file_extractor={".spdf": SpdfReader()})`. An `fsspec` filesystem passed as `fs=` is honoured. - `extra_info={...}` adds metadata to every document (and wins over the reader's keys). - Unsafe or invalid files are refused by `spdf-format`: loading one raises its error (`spdf.UnsafeFileError`, `spdf.NotSpdfError`…, all subclasses of `spdf.SpdfError`) with the validation code (`E020`…) and the file path in the message. A missing path raises `FileNotFoundError`. - Legacy SPDF 4.0 and 4.1 files (also gzip-wrapped) read like 5.0 ones. ### Options | Option | Default | Meaning | | --- | --- | --- | | `granularity` | `"fragment"` | `"fragment"`: one document per passage (about 150 to 300 words). `"unit"`: one per page, time span, slide… | | `locale` | `"en"` | Locale of `citation`: `"en"` or `"es"` (`"es-ES"` works; others fall back to English). | | `include_embeddings` | `None` | Id of a vector space stored in the files (`"all-MiniLM-L6-v2@384"`). Sets `Document.embedding` from the stored vectors, so that an index whose embedding model matches that space does not embed the passages again. A file without that space raises `spdf.SpdfError`; a passage without a stored vector keeps `embedding=None` and is embedded by the index. | | `excluded_embed_metadata_keys` | all but `title`, `section` | Keys kept out of the text that is embedded. | | `excluded_llm_metadata_keys` | all but `title`, `authors`, `year`, `section`, `citation` | Keys hidden from the LLM. By default the LLM sees the `citation` line next to each passage and can copy it into its answer. | ## Metadata Values are flat scalars (`str`, `int`, `float`, `bool`), so every vector store accepts them. **A key whose value would be null is left out** (Chroma and others reject `None`): the cover of a book has no `printed_folio` key, a page has no `t0`. | Key | Type | Example (first passage of the English fixture) | Notes | | --- | --- | --- | --- | | `source` | str | `fixtures/spdf-in-five-pages.spdf` | Path the file was read from. | | `spdf_version` | str | `5.0` | `4.0` or `4.1` for legacy files. | | `spdf_doc_id` | str | `spdf-in-five-pages` | The document id inside the file. Not called `doc_id`: LlamaIndex vector stores overwrite `doc_id`, `document_id` and `ref_doc_id` with the node's reference document id. | | `docref` | str | `sha256-50d94244…5f4c` | Document reference used by anchor URIs (SHA-256 of the original). | | `title` | str | `SPDF in five pages` | | | `authors` | str | `Saorín Ferrer` | | | `year` | int | `2026` | | | `language` | str | `en` | BCP 47. | | `kind` | str | `pdf` | `pdf`, `epub`, `audio`, `video`… | | `fragment_id` | str | `f2-1` | Fragment granularity only. | | `unit_id` | str | `u2` | The unit (page…) where the passage starts. | | `anchor_type` | str | `page` | `page`, `time`, `section`, `slide`, `sheet`, `web`, `image`, `verse`, `canonical`. | | `physical_page` | int | `2` | Page anchors: position of the page in the file. | | `printed_folio` | str | `1` | The folio as printed (`"xiv"`, `"1r"`). | | `folio_inferred` | bool | `false` | True when the folio was deduced, not read; the citation prints it in brackets, `p. [3]`. | | `section` | str | `I. Anchors` | Heading path joined with `" / "`. In unit granularity, the sections that share the unit are joined with `" \| "`. | | `context` | str | `SPDF in five pages, I. Anchors` | One line that situates the passage (fragment granularity). | | `anchor` | str | `{"chars":[15,307],"confidence":1,"physical":2,"printed":"1","source":"read","type":"page"}` | The start anchor as canonical JSON; `json.loads` it for the full object. | | `anchor_end` | str | | End anchor (JSON) when the passage crosses into another unit. | | `t0`, `t1` | float | | Seconds, for time anchors (recordings). | | `anchor_uri` | str | `spdf:sha256-50d94244…5f4c#p=2&f=1&char=15,307` | Resolve it with `spdf.open(path).locate(uri)`; parse it with `spdf.parse_uri`. | | `citation` | str | `(Saorín Ferrer, 2026, p. 1)` | Short author-date citation in the chosen locale. | | `vector_space` | str | `all-MiniLM-L6-v2@384` | Only when `include_embeddings` attached a vector. | The document text is always the literal passage (`fragments.text` or `units.text`), never the modernised-spelling search layer, which SPDF forbids quoting. ## End-to-end example (no API key) A complete retrieval-augmented query with a toy embedding and LlamaIndex's `MockLLM`; swap them for your models. The sources of the answer carry their citations. ```python import hashlib import math import re from llama_index.core import VectorStoreIndex from llama_index.core.embeddings import BaseEmbedding from llama_index.core.llms import MockLLM from spdf_llamaindex import SpdfReader class HashingEmbedding(BaseEmbedding): """A toy bag-of-words embedding: no model to download, no API key.""" dim: int = 512 def _vec(self, text: str) -> list[float]: v = [0.0] * self.dim for word in re.findall(r"\w+", text.lower()): v[int(hashlib.md5(word.encode()).hexdigest(), 16) % self.dim] += 1.0 norm = math.sqrt(sum(x * x for x in v)) or 1.0 return [x / norm for x in v] def _get_text_embedding(self, text: str) -> list[float]: return self._vec(text) def _get_query_embedding(self, query: str) -> list[float]: return self._vec(query) async def _aget_query_embedding(self, query: str) -> list[float]: return self._vec(query) docs = SpdfReader(locale="en").load_data("integrations/fixtures/spdf-in-five-pages.spdf") index = VectorStoreIndex(docs, embed_model=HashingEmbedding()) engine = index.as_query_engine(llm=MockLLM(), similarity_top_k=2) response = engine.query("How is a plate without a printed folio cited?") for source in response.source_nodes: print(source.node.metadata["citation"], source.node.metadata["anchor_uri"]) ``` Output: ```text (Saorín Ferrer, 2026, p. [3]) spdf:sha256-50d94244…5f4c#p=4&f=3&char=0,133 (Saorín Ferrer, 2026, p. 2) spdf:sha256-50d94244…5f4c#p=3&f=2&char=19,258 ``` The plate carries no printed number; its folio is inferred, so the citation prints it in brackets. What the LLM receives for each passage is: ```text title: SPDF in five pages authors: Saorín Ferrer year: 2026 section: III. Read once, query many citation: (Saorín Ferrer, 2026, p. [3]) Plate I. A page with its folio and a manicule pointing at a passage. … ``` Fragments are already passage-sized, so the example passes the documents straight to `VectorStoreIndex(docs, …)`. `VectorStoreIndex.from_documents(docs, …)` also works (the splitter copies the metadata to every node), but it creates new nodes without the stored embeddings. ## Reusing the vectors stored in the file SPDF files may ship vectors (`f.spaces()` in `spdf-format` lists them, with model, size and any task prefixes). When your embedding model is the one that produced a space, load the vectors instead of embedding every passage again: ```python from llama_index.core import VectorStoreIndex from llama_index.embeddings.huggingface import HuggingFaceEmbedding # pip install llama-index-embeddings-huggingface from spdf_llamaindex import SpdfReader docs = SpdfReader(include_embeddings="all-MiniLM-L6-v2@384").load_data("library/") embed = HuggingFaceEmbedding(model_name="sentence-transformers/all-MiniLM-L6-v2") # local, embeds queries only index = VectorStoreIndex(docs, embed_model=embed) # not from_documents: keep the stored vectors print(index.as_retriever().retrieve("inferred folio")[0].node.metadata["citation"]) ``` ## Development ```bash cd integrations/llamaindex-python uv venv && uv pip install -e ../../python -e ".[test]" .venv/bin/python -m pytest ``` The tests use the shared fixtures in `integrations/fixtures/`. ## License MIT OR Apache-2.0, at your option. --- # Zotero 7 and 8 plugin URL: https://spdf.joseluissaorin.com/integrations/zotero > Import an SPDF as an item, attach it, copy a citation with the folio. A Zotero 7 and Zotero 8 plugin for [SPDF](https://spdf.joseluissaorin.com) (Semantic Processed Document Format) files: documents that were read once and keep, for every passage, its exact anchor (printed folio, physical page, second of a recording, slide…). With the plugin, Zotero can: - **Import SPDF as Item…** (Tools menu and item context menu): pick an `.spdf` file and get a Zotero item built from the file's own CSL-JSON metadata, in the selected library and collection, with the `.spdf` file attached. - **Attach SPDF…** (item context menu, one regular item selected): attach an `.spdf` file to an item you already have. - **Copy Citation with Folio…** (item context menu, an item with an SPDF attachment, or the attachment itself, or a sibling attachment such as the PDF): type a printed folio, a physical page or an anchor URI, and the clipboard receives the short citation and, on the next line, the anchor URI: ``` (Saorín Ferrer, 2026, p. [3]) spdf:sha256-50d942445564fe24effe743701f0c9a16fde414e6eb1f0ce2095b866d0555f4c#p=4&f=3 ``` The citation is always computed from the anchor stored in the file, never from what was typed, so it can only print what the source says: an inferred folio is printed in brackets (`p. [3]`), a page without folio is `s. p.` in Spanish and `n. pag.` in English, and a folio or page that the document does not have is reported as missing and nothing is copied. The menus and dialogs are in English and Spanish (Fluent, `locale/en-US` and `locale/es-ES`). Citations follow Zotero's interface language: Spanish when it is Spanish, English otherwise (SPEC §18). ## Install 1. Download `spdf-zotero-.xpi` (or build it, see below). 2. In Zotero: **Tools → Plugins**, then the gear menu → **Install Plugin From File…**, and choose the `.xpi`. The plugin needs Zotero 7 or Zotero 8 (`strict_min_version` 6.999, `strict_max_version` 8.*). It installs nothing else and downloads nothing. ## Use **Import SPDF as Item…** creates the item with `Zotero.Utilities.Item.itemFromCSLJSON` from `documents.metadata` without its `spdf` extension (exactly what `spdf-format`'s `toCslJson` exports). Legacy Scholaris files (SPDF 4.0 and 4.1, usually gzip-wrapped, with Spanish table names) are read too; their metadata is mapped to CSL-JSON by `spdf-format` (`mapLegacyMetadata`). The file is then copied into Zotero's storage with `Zotero.Attachments.importFromFile` (with the media type `spdf-format` declares, `MEDIA_TYPE`; the specification fixes `application/vnd.spdf+sqlite3`), and the new item is selected. Attachments are recognised as SPDF by that type, by the older `application/vnd.spdf` or by the `.spdf` extension. Both Import and Attach add one line to the item's **Extra** field: ``` SPDF: sha256-<64 hex digits> ``` That is the document reference of the file (its `source_sha256`), the same one anchor URIs carry, so a URI found in a manuscript can be matched to the item later. Other lines of Extra are kept; the line is not repeated. **Copy Citation with Folio…** accepts: | You type | Meaning | | --- | --- | | `145`, `xiv`, `1r`, `p. 145`, `pág. 12` | a printed folio, as printed (`XIV` also finds `xiv`) | | `[21]` | the same, with the brackets of an inferred folio | | `145-146`, `pp. 2-[3]` | a range of folios | | `p=12`, `#12`, `p=12-13` | a physical page (position in the original), or a range | | `f=A-3` | a printed folio, explicitly (for folios that contain a dash) | | `spdf:sha256-…#p=29&f=21` | a full anchor URI; `char=` and `xywh=` are kept | | `spdf:sha256-…` | the whole document: `(Saorín Ferrer, 2026)` | An anchor URI is resolved as SPEC §5.4 says: `p` first, then `f`, then `t` (time), then slides, sheets, verses, canonical references and sections. If the item has several SPDF attachments, a URI picks the one whose document it names, and a folio asks which attachment to use. If two pages carry the same folio, the plugin asks which one. What the plugin refuses: files that are not SQLite, of an unknown version, with views or triggers, or with an unknown required extension (SPEC §2.4). It does not run the full validator when importing (see Design); use the web validator or `npx spdf-format validate file.spdf` to audit a file. ## Design - **The format is not reimplemented.** Everything SPDF-specific (opening safely, legacy 4.x mapping, anchors and anchor URIs, the short citation, CSL-JSON export) comes from the official TypeScript library, `spdf-format`, whose engine-less core (`spdf-format/core`) is bundled into the plugin by esbuild. - **SQLite is Zotero's own.** `src/engine.ts` is a `spdf-format` `SqlEngine` over `Sqlite.sys.mjs` (`resource://gre/modules/Sqlite.sys.mjs`), the asynchronous mozStorage wrapper Zotero 7 (Firefox 115) and Zotero 8 (Firefox 140) ship. The whole `SpdfDocument` API (units, fragments, `cite`, `anchorUri`, search, `validate`, `dump`) therefore works inside Zotero with no second SQLite in the package. - **Read only.** Files are opened with `Sqlite.openConnection({ path, readOnly: true })` (`SQLITE_OPEN_READONLY`); mozStorage keeps extension loading off; the core then sets `PRAGMA query_only = 1` and `PRAGMA trusted_schema = OFF`, and refuses views and triggers. Only temporary copies (a decompressed legacy file, the private copy used by the FTS5 integrity check) are written, in Zotero's temp directory, and they are deleted when the connection closes. - **Column names.** A `mozIStorageRow` can be read by index or by a known name but cannot list its columns, while the core expects rows keyed by column name. `src/sql.ts` infers the names the way SQLite assigns them (the `AS` alias, the column of a bare reference, the expression text otherwise, fixed names for PRAGMAs), and the engine checks every inferred name against the first row with `getResultByName`, so a wrong guess is an error, never mislabelled data. The tests check the inference against SQLite's own names and run every statement the core issues through it. - **Gzip.** Legacy files are gunzipped with `DecompressionStream`, taken from Zotero's main window because the plugin sandbox does not have the Compression Streams API, with the 4 GiB output limit of SPEC §2.3. - **Compartments.** Bytes from `IOUtils`, mozStorage or the main window are copied into typed arrays of the plugin's own realm, because the core tests `instanceof Uint8Array`. The sandbox also lacks `structuredClone`, which the core uses to copy metadata; the bundle gets a JSON-based fallback (`src/shims/structured-clone.ts`). - **One file.** `bootstrap.js` is the whole plugin: the bundle followed by the bootstrap hooks. The plugin never loads a second script from its own `jar:` URL, and the file is pure ASCII so its decoding never depends on the script loader. - **Menus.** On Zotero 8 the entries are registered with `Zotero.MenuManager` (targets `main/menubar/tools` and `main/library/item`); on Zotero 7, which has no menu API, they are added to `menu_ToolsPopup` and `zotero-itemmenu` in each main window, as Zotero's sample plugin does. Labels come from Fluent in both cases. - **No full validation on import.** The validator's content-hash step re-reads every blob (page images, the original PDF), and mozStorage returns BLOBs as JavaScript arrays of numbers, which is slow and memory-hungry for a large book. Firefox's SQLite may also lack FTS5, which the integrity check needs. A reference manager only needs the metadata and the anchors, and the safety checks a reader must make are made. ## Build and test ```sh cd integrations/zotero npm install # spdf-format comes from ../../js (file: dependency) npm run build # dist/addon/ and dist/spdf-zotero-.xpi npm test # typecheck, build, then all tests ``` `spdf-format` must be built first (`cd js && npm run build`), since the plugin bundles `js/dist`. The build is reproducible: same sources, same `.xpi` bytes. For development, Zotero can load the unpacked plugin: in a **test profile**, create a text file named `spdf@joseluissaorin.com` in the profile's `extensions` directory whose only line is the absolute path of `dist/addon/`, then start Zotero with `-purgecaches`. ### How it is tested There is no Zotero in the test run. Everything that does not need Zotero runs in Node with `node:sqlite`, against the shared fixtures (`integrations/fixtures`) and the legacy conformance files (`conformance/legacy`). That corpus is rebuilt from time to time, so the tests find legacy files by the kind of anchor they hold (pages, times, sections) and take the expected citations from the official Node engine of `spdf-format`, checked against the rules of SPEC §18: - `test/engine.test.ts`: the adapter runs over a stand-in for `Sqlite.sys.mjs` built on `node:sqlite` whose rows behave like `mozIStorageRow` (values by index, no column names, BLOBs as arrays of octets, one statement per call, the same parameter binding rules). Through it, every fixture and legacy file gives the same canonical dump, units, fragments, citations, search results and validation report as the official Node engine of `spdf-format`. Read-only opening, temp-file cleanup, gzip limits and refusals (E001, E020) are checked too. - `test/sql.test.ts`: column-name inference compared with SQLite's own names, statement splitting, placeholder renaming. - `test/locate.test.ts`: folio, page, range and URI lookup and the citations, for example `(Saorín Ferrer, 2026, p. 1)` for physical page 2, `p. [3]` for the inferred plate, `s. p.` / `n. pag.` for the cover; in legacy files every printed folio, physical page, time (`h:mm:ss` from one hour on) and section paragraph; and "not found" for folios, pages, times and paragraphs that do not exist. - `test/commands.test.ts`: the three commands against a fake `Zotero` that records items, CSL-JSON, attachments, Extra and the clipboard. - `test/plugin.test.ts`: startup, both menu paths (DOM for Zotero 7, `MenuManager` for Zotero 8), windows opening and closing, shutdown, and the real host code (`Services.prompt`, `FilePicker`, `IOUtils`, `PathUtils`, `Localization`) over fakes. - `test/l10n.test.ts`: English and Spanish have the same messages, every id used in the code exists, Spanish has its accents and « » quotes. - `test/xpi.test.ts`: the `.xpi` unzips, the manifest declares `spdf@joseluissaorin.com`, 6.999 to 8.*, and `bootstrap.js` itself runs startup, import, citation and shutdown in a `vm` context holding only the globals of Zotero's plugin sandbox (no `window`, `console`, `DecompressionStream` or `structuredClone`). ### What still has to be checked by hand None of this has run inside a real Zotero yet. Before a release, in a test profile of **Zotero 7** and of **Zotero 8**: 1. The plugin installs from the `.xpi`, shows in Tools → Plugins, and can be disabled, enabled and removed without errors in the Error Console. 2. The three entries appear with their labels (English and Spanish UI); in Zotero 7 in the Tools menu and the item context menu (DOM path); in Zotero 8 through `Zotero.MenuManager`, and Attach / Copy citation appear only when they apply. 3. Import: the file picker filters `.spdf`; the item gets the right type, creators, date and Extra line; the file is copied into storage; it lands in the selected collection; a read-only group library is refused. 4. `Sqlite.openConnection({ readOnly: true })` opens files in Zotero's storage and in arbitrary folders, and the column-name check passes on real `mozIStorageRow`s (the fake models them from the IDL and Zotero's own use of them). 5. A legacy gzip-wrapped 4.x file imports (`DecompressionStream` from the main window, temp file in Zotero's `tmp` directory, removed afterwards). 6. Copy Citation: `Services.prompt` dialogs, the clipboard content, the progress notice; the "not found" warnings. 7. Large files (hundreds of MB, thousands of pages) open and cite in reasonable time. 8. The `update_url` in the manifest (`https://spdf.joseluissaorin.com/zotero/updates.json`) is served by the website, or is removed; until then Zotero simply finds no updates. ## License MIT OR Apache-2.0, like the rest of the SPDF code. The specification is CC BY 4.0. --- # Pandoc filter URL: https://spdf.joseluissaorin.com/integrations/pandoc > SPDF anchors in Markdown become citations with the printed folio, in any CSL style. `spdf.lua` is a [Pandoc](https://pandoc.org) Lua filter that lets you cite a place in an [SPDF](../../spec/SPEC.md) file from Markdown and get a real citation, formatted by citeproc in any CSL style, with the folio **printed in the source**. An SPDF file is a document that has already been read: every unit knows its physical position in the file and the folio printed on the page (or its second, slide, verse…). You write the anchor; the filter looks it up in your SPDF files, adds the work to the bibliography and hands citeproc the locator a reader will find on paper: ```markdown --- spdf-library: ~/Library/SPDF --- Physical page 2 carries the printed folio 1 [@spdf:sha256-50d942445564fe24effe743701f0c9a16fde414e6eb1f0ce2095b866d0555f4c#p=2]. An inferred folio goes in brackets [@spdf:sha256-50d942445564fe24effe743701f0c9a16fde414e6eb1f0ce2095b866d0555f4c#p=4]. The cover has no folio [@spdf:sha256-50d942445564fe24effe743701f0c9a16fde414e6eb1f0ce2095b866d0555f4c#p=1]. # References ``` ```console $ pandoc paper.md --lua-filter spdf.lua --citeproc -t plain --wrap=none [WARNING] Scripting warning at spdf.lua line 46 column 1: spdf: [@spdf:sha256-50d942445564fe24effe743701f0c9a16fde414e6eb1f0ce2095b866d0555f4c#p=1]: physical page 1 of spdf-in-five-pages.spdf has no printed folio; cited as unnumbered (n. pag.) Physical page 2 carries the printed folio 1 (Saorín Ferrer 2026, 1). An inferred folio goes in brackets (Saorín Ferrer 2026, [3]). The cover has no folio (Saorín Ferrer 2026, n. pag.). References Saorín Ferrer, José Luis. 2026. SPDF in Five Pages. Spdf.joseluissaorin.com. https://spdf.joseluissaorin.com/validator. ``` Every output in this file is real, produced with Pandoc 3.9 and its default style, Chicago author-date, mostly from the sample booklet [`integrations/fixtures/spdf-in-five-pages.spdf`](../fixtures/) (six page units: a cover without folio, then printed folios 1, 2, an inferred [3], 4 and 5); the same cases are checked by the tests in [`test/`](test/). ## Why Pandoc, and not Calibre Academic writing in Markdown already goes through Pandoc and citeproc: that is the step where an author's sources become citations in the style a journal or a university asks for. Turning a position in a file (the 29th page of a PDF) into the folio a reader finds on paper (p. 21, or p. [21] when the folio was inferred, or *n. pag.* when there is none) belongs exactly there, just before the style formats it. Calibre is a reading library, and reading SPDF files is what the SPDF Reader is for; a Calibre plugin would be one more place to read, and would not help anyone cite. ## Requirements and installation - **Pandoc 3.1.1 or later** (the filter uses `pandoc.json`; tested with 3.9). - **The `sqlite3` command-line tool**, 3.33 or later (JSON output); 3.37 or later is recommended, because the filter then runs it in safe mode. macOS ships it; on Debian or Ubuntu `apt install sqlite3`, on Fedora `dnf install sqlite`. No compiled Lua module is needed. - For legacy SPDF 4.x files, which are gzip-compressed: `gzip`, `head` and a POSIX `sh` (all standard on macOS and Linux). Copy [`spdf.lua`](spdf.lua) next to your document, or into the `filters` folder of your Pandoc user data directory (`~/.local/share/pandoc/filters/` on macOS and Linux), where `--lua-filter spdf.lua` finds it from anywhere. Run it **before** citeproc: ```sh pandoc paper.md --lua-filter spdf.lua --citeproc -o paper.pdf ``` or with a defaults file (`pandoc -d spdf.yaml paper.md -o paper.pdf`): ```yaml # spdf.yaml filters: - spdf.lua - citeproc metadata: spdf-library: ~/Library/SPDF ``` ## Writing citations The key of the citation is an SPDF anchor URI ([SPEC §5](../../spec/SPEC.md#anchor-uri)): `spdf:` followed by the document reference and, after `#`, the anchor parameters. - The document reference is `sha256-` and the 64 hexadecimal digits of the document's `source_sha256` (recommended: it is the same in every copy of the file), or the document id, percent-encoded. To read both from a file: `sqlite3 -readonly book.spdf "SELECT source_sha256, id FROM documents"`. - The parameters say where: `p=` the physical page (the position in the file), `f=` the printed folio, `pe=` / `fe=` the end of a range, `t=` seconds (`t=4160`, `t=12,24.5`, `t=1:09:20`), `s=` and `para=` a section path and paragraph, `sl=` a slide, `sh=` and `rows=` a sheet, `v=` verse lines, `ref=` a canonical reference (`ref=stephanus:514a`). `char=` and `xywh=` narrow a unit and do not change the citation. Everything Pandoc offers around a citation keeps working: prefixes, suffixes, several citations in one bracket, suppressing the author, in-text citations. To keep them short, the examples from here on cite the booklet by its document id, `spdf-in-five-pages`; `sha256-50d942445564fe24effe743701f0c9a16fde414e6eb1f0ce2095b866d0555f4c` gives the same output. ```markdown Prefix and suffix are kept [see @spdf:spdf-in-five-pages#p=5, emphasis added]. Two citations in one bracket [@spdf:spdf-in-five-pages#p=2; compare @spdf:spdf-in-five-pages#p=6, for the summary]. The author can be suppressed [-@spdf:spdf-in-five-pages#p=3]. A range [@spdf:spdf-in-five-pages#p=2&pe=3], or by folios [@spdf:spdf-in-five-pages#f=4&fe=5]. @spdf:spdf-in-five-pages#p=2 says so. As @spdf:spdf-in-five-pages#p=4, argues, inferred folios are bracketed. ``` ```text Prefix and suffix are kept (see Saorín Ferrer 2026, 4, emphasis added). Two citations in one bracket (Saorín Ferrer 2026, 1; compare Saorín Ferrer 2026, 5, for the summary). The author can be suppressed (2026, 2). A range (Saorín Ferrer 2026, 1–2), or by folios (Saorín Ferrer 2026, 4–5). Saorín Ferrer (2026, 1) says so. As Saorín Ferrer (2026, [3]), argues, inferred folios are bracketed. ``` ### How Pandoc reads these keys In Pandoc's Markdown a citation key starts with a letter, a digit or `_` and may contain letters, digits, `_` and the internal punctuation `: . # $ % & - + ? < > ~ /`. The equals sign is not among them, so **the key stops at the first parameter name** and the rest of the anchor lands in the citation's suffix, or, for an in-text citation, in the text that follows. These are the Pandoc 3.9 ASTs (`pandoc -t native`) the filter is built on: | You write | `citationId` | Where the rest goes | |---|---|---| | `[@spdf:sha256-…#p=29]` | `spdf:sha256-…#p` | suffix `=29` | | `[see @spdf:sha256-…#f=21, emphasis added]` | `spdf:sha256-…#f` | suffix `=21, emphasis added`, prefix `see` | | `[@spdf:sha256-…#p=2&f=1]` | `spdf:sha256-…#p` | suffix `=2&f=1` | | `@spdf:my-doc#p=3 says` | `spdf:my-doc#p` (in-text) | the next text element, `=3` | | `[@{spdf:sha256-…#p=2&pe=3}]` | the whole URI | nothing (braced form) | | `@{spdf:sha256-…#p=2} [emphasis added]` | the whole URI (in-text) | suffix `emphasis added` | | `[@spdf:sha256-…]` | `spdf:sha256-…` | the whole document, no locator | The filter puts the pieces back together, so all of these work as written. A few forms do not reach the filter as citations, and need another spelling: | Instead of | Write | Why | |---|---|---| | `@spdf:…#p=2 [emphasis added]` | `@{spdf:…#p=2} [emphasis added]` | Pandoc attaches a bracketed suffix only directly after the key, and here `=2` comes between them. | | `[@{spdf:my doc#p=2}]` | `[@{spdf:my%20doc#p=2}]` | A space ends the key even inside braces. Percent-encode it, as the URI grammar asks anyway; the filter warns when it meets such text. | | `@spdf:…#f=xiv.` meaning folio `xiv.` | `[@{spdf:…#f=xiv.}]` or `f=xiv%2E` | Prose punctuation stuck to the end (`. , : ; ! ?`, closing quotes, an ellipsis) is taken as punctuation of your sentence; in the braced form every character counts. | **CommonMark and GFM.** `commonmark`, `commonmark_x` and `gfm` have no citation syntax in Pandoc 3.9: `[@spdf:…#p=2]` arrives as plain text. The filter finds those pieces and reads them again with Pandoc's own Markdown reader, so the same syntax works with `--from=commonmark_x`; the prefix and suffix of such a citation must be plain text. ## What the citation prints The filter resolves the anchor against the file and writes the locator into the suffix in a form citeproc recognises as a CSL locator (`, {p. [3]}`), so the **style** decides how to print it: Chicago author-date writes `[3]`, a style that shows labels writes `p. [3]`. The rules are those of the specification ([SPEC §18](../../spec/SPEC.md#citation)): | Anchor | Example | Chicago author-date | |---|---|---| | page, folio read | `#p=2` | `(Saorín Ferrer 2026, 1)` | | page, folio inferred | `#p=4` | `(Saorín Ferrer 2026, [3])` | | page without folio | `#p=1` | `(Saorín Ferrer 2026, n. pag.)` and a warning; `s. p.` in Spanish | | folio given directly | `#f=3` | `(Saorín Ferrer 2026, [3])`: checked against the units, bracketed if inferred | | range | `#p=3&pe=4`, `#f=4&fe=5` | `(Saorín Ferrer 2026, 2–[3])`, `(Saorín Ferrer 2026, 4–5)` | | range from an unnumbered page | `#p=14&pe=29` in the 1608 *Quixote* | `fol. Ir` and a warning: the unnumbered end is left out | | leaf (`foliation: leaf`), roman | `#p=29`, `#p=30`, `#p=29&pe=30` | `(Cervantes Saavedra [1605] 1608, fol. Ir)`, `fol. [Iv]`, `fols. Ir–[Iv]` | | time (ground elapsed time) | `#t=369959`, `#t=369966,369976` | `102:45:59`, `102:46:06-102:46:16` | | section | `#s=學而第一¶=1` | `(孔子, n.d., § 學而第一, para. 1)` | | section with a printed page | `#s=XXI¶=1&f=159` | `(Bécquer [1871] 1885, 159)` | | verse | `#v=2`, `#v=1-3` | `v. 2`, `vv. 1–3` | | slide | `#sl=2` | `slide 2` (`diap. 2` in Spanish) | | sheet | `#sh=Data&rows=4-9` | `Data, rows 4-9` | | canonical | `#ref=stephanus:514a`, `#ref=analects:1.2` | `514a`, `(孔子, n.d., 1.2)` | | whole document | no parameters | `(Saorín Ferrer 2026)` | When both `p` and `f` are given, `p` decides and a disagreeing `f` is reported. A folio printed on several pages (`f=1` in front matter and body) takes the first and warns; add `p=` to choose. A page without a printed folio is **never** cited by its position in the file: that number does not exist on paper. An end without a printed folio never takes part in a range ([SPEC §18.1](../../spec/SPEC.md#citation)): the folio of the other end is cited alone, and the page is unnumbered only when neither end has a folio. Verses, sections, paragraphs, slides, sheets and canonical references are looked up in the anchors of the units and of the fragments ([SPEC §5.4](../../spec/SPEC.md#anchor-uri)), so a reference kept on a fragment (`analects:1.2`, a line of a poem) is found, and one the file does not anchor is reported. Pandoc only reads a locator label in the terms of the locale citeproc is using, with no English fallback: a German document needs `S.` for a page, a Spanish one `f.` for a folio. The filter asks citeproc for those terms itself, once per run, so labels work in every CSL locale. With the test style [`test/styles/labels.csl`](test/styles/labels.csl), which prints the label citeproc recognised, a document with `lang: de-DE` gives `(Saorín Ferrer, 2026, [page] S. 1)` and `(Cervantes Saavedra, 1608, [folio] Fol. [Iv])`, and one with `lang: es-ES` gives `[page] p. 1` and `[folio] f. [Iv]`. Locators that CSL has no label for (an unnumbered page, a time, a slide, a sheet, a canonical reference, a section) are written after an empty locator, `{}, 1:09:20`, so that citeproc does not mistake `1:09:20` or `514a` for a page number. In Spanish: ```markdown --- lang: es --- La portada no lleva folio [@spdf:sha256-6abda0640aaed500ee9673212fab17b83ef4326f5ade42d6c6cabec3f8996df7#p=1]. Según @spdf:spdf-en-cinco-paginas#p=3, el formato guarda el folio impreso. ``` ```text La portada no lleva folio (Saorín Ferrer 2026, s. p.). Según Saorín Ferrer (2026, 2), el formato guarda el folio impreso. ``` ## Options Set them in the document's YAML metadata, in a defaults file, or with `-M`. | Field | Default | Meaning | |---|---|---| | `spdf-library` | `.` and the folder of the input file | A folder, or a list of folders and files. Folders are searched for `*.spdf` (not recursively). Relative paths are taken from the working directory, then from the input file's folder; `~/` is your home folder. | | `spdf-locale` | from `lang` | `en` or `es`: the language of the few words the filter writes itself (`n. pag.` / `s. p.`, `slide` / `diap.`, `para.` / `párr.`, `rows` / `filas`). Any `lang` other than Spanish gives English. | | `spdf-links` | `false` | `true` adds the canonical anchor URI of each resolved citation as a link in a footnote (inside a footnote, in parentheses after the citation). | | `spdf-sqlite3` | `sqlite3` | The `sqlite3` program to run. | With `spdf-links: true`: ```text A citation gets a note with its anchor (Saorín Ferrer 2026, [3])[1]. A range links to its canonical anchor (Saorín Ferrer 2026, 2–[3])[2]. Two citations get one note (Saorín Ferrer 2026, 1; 2026, 5)[3]. Inside a footnote the anchor goes in parentheses.[4] [1] spdf:sha256-50d942445564fe24effe743701f0c9a16fde414e6eb1f0ce2095b866d0555f4c#p=4&f=3 [2] spdf:sha256-50d942445564fe24effe743701f0c9a16fde414e6eb1f0ce2095b866d0555f4c#p=3&pe=4&f=2&fe=3 [3] spdf:sha256-50d942445564fe24effe743701f0c9a16fde414e6eb1f0ce2095b866d0555f4c#p=2&f=1; spdf:sha256-50d942445564fe24effe743701f0c9a16fde414e6eb1f0ce2095b866d0555f4c#p=6&f=5 [4] As shown in Saorín Ferrer (2026, 2) (spdf:sha256-50d942445564fe24effe743701f0c9a16fde414e6eb1f0ce2095b866d0555f4c#p=3&f=2). ``` The link is the canonical URI of what was found ([SPEC §5.2](../../spec/SPEC.md#anchor-uri)), whatever form you wrote: a citation by folio links to its physical page as well. ## References For each cited document the filter takes the CSL-JSON item stored in the file (without its `spdf` extension object), reads it with Pandoc's own CSL JSON reader, and adds it to the document's `references` metadata under the key `spdf-` followed by the first 12 hex digits of the document's `source_sha256` (`spdf-50d942445564`). The citation's key is rewritten to it, so citeproc, `link-citations`, `nocite` and every CSL style see an ordinary reference. Nothing is duplicated. If a reference with that key already exists, it is kept. If your `references` or `bibliography` files already hold the same work (same DOI, same ISBN, or same title, year and first author), the citation uses **your** key and your entry: ```text The English booklet is already in the references as saorin2026, so the citation uses that key (Saorín Ferrer 2026b, 1) and the bibliography lists it once, next to a citation written by hand (Saorín Ferrer 2026b, 4). The Spanish booklet is in the bibliography file as cinco (Saorín Ferrer 2026a, 2). ``` Without `--citeproc`, Pandoc's Markdown writer shows what the filter did, which is also a way to hand a resolved manuscript to someone without SPDF files (`-t markdown -s` also writes the references). For ```markdown A [see @spdf:spdf-in-five-pages#p=4, emphasis added]. B @spdf:spdf-in-five-pages#p=1 says. C [@spdf:spdf-in-five-pages#p=2] ``` `pandoc --lua-filter spdf.lua -t markdown` writes: ```markdown A [see @spdf-50d942445564, {p. \[3\]}, emphasis added]. B @spdf-50d942445564 [{}, n. pag.] says. C [@spdf-50d942445564, {p. 1}] ``` ## Warnings and unresolved citations An anchor that cannot be resolved is never dropped or guessed. The filter warns on stderr and leaves the citation visibly marked: its key becomes the anchor URI you wrote, which citeproc prints in bold with a question mark and reports again: ```text [WARNING] Scripting warning at spdf.lua line 46 column 1: spdf: [@spdf:spdf-in-five-pages#p=99]: page p=99 is not in spdf-in-five-pages.spdf; the citation is left unresolved [WARNING] Citeproc: citation spdf:spdf-in-five-pages#p=99 not found A page the file does not have (spdf:spdf-in-five-pages#p=99?). ``` This happens for an unknown document, a page, folio, time, verse, slide, sheet, section or canonical reference the file does not have, a malformed parameter (`p=0`), a truncated hash, or a time cited in a paged document. A citation next to it in the same bracket is still resolved. An unknown parameter name (`pg=9`) is reported and ignored, as the specification asks. Files that cannot be used are skipped with the reason: ```text spdf: skipping roto.spdf: it contains view x_rotura; SPDF files must not carry triggers, views or foreign virtual tables (E020) spdf: skipping E001-not-sqlite.spdf: it is not a SQLite database (E001) spdf: skipping E002-application-id.spdf: it is not an SPDF file (unknown application_id, E002) spdf: skipping E013-two-documents.spdf: its documents table does not hold exactly one row (E013) spdf: library path not found: no-such-folder ``` The warnings go through Pandoc's own log (Pandoc prefixes them with the line of the filter that emitted them; the message is what follows `spdf:`). So `--quiet` silences them, and `--fail-if-warnings` makes Pandoc exit with status 3: use it in a build that must not publish an unresolved citation. If `sqlite3` is missing and the document cites SPDF anchors, Pandoc stops with: ```text spdf.lua: the sqlite3 command-line tool was not found (looked for 'sqlite3'). Install it (macOS ships it; Debian/Ubuntu: apt install sqlite3; Fedora: dnf install sqlite) or point the metadata field spdf-sqlite3 at it. ``` A document without SPDF citations never calls `sqlite3`. ## Legacy SPDF 4.0 and 4.1 files Files written by Scholaris before SPDF 5.0 are SQLite databases wrapped in gzip, with Spanish table and column names. The filter recognises the gzip magic bytes, decompresses the file into a temporary folder (refusing more than 4 GiB), and reads it through the 5.0 view of [SPEC §20](../../spec/SPEC.md#legacy): `documentos`, `unidades`, `huella`, the anchor members (`fisica`, `impresa`, `origen: deducido`…) and the `MetadatosDocumento` object, mapped to CSL (title and subtitle, authors, editors, dates, publisher, place…). ```text Garcilaso, a folio inferred from its neighbours (Garcilaso de la Vega [1543] 1919, [7]), a folio read on the page (Garcilaso de la Vega [1543] 1919, 159), by document id and folio (Garcilaso de la Vega [1543] 1919, 159), a range (Garcilaso de la Vega [1543] 1919, [7]–159), and a page the file does not have (spdf:garcilaso#p=10?). Kennedy, by paragraph (Kennedy 1962, para. 15) (Kennedy 1962, para. 16), and a paragraph the excerpt does not have (spdf:kennedy-rice#para=40?). Apollo 11, a recording in ground elapsed time: a moment (National Aeronautics and Space Administration 1969, 102:46:16), a span (National Aeronautics and Space Administration 1969, 102:46:18-102:46:23), and a time outside the excerpt (spdf:apolo11-tierra#t=99?). ``` ## Security SPDF files come from strangers, so the filter follows the safe opening of [SPEC §2.4](../../spec/SPEC.md#container) as far as the `sqlite3` tool allows. Each query runs in a separate `sqlite3` process opened with `-readonly`, `-safe` (3.37 or later: no `load_extension`, no file or shell commands), `-batch`, `-bail` and `-init /dev/null` (your `~/.sqliterc` is not read), with `.dbconfig defensive on`, `PRAGMA trusted_schema = OFF`, `query_only = 1`, `mmap_size = 0`, `cell_size_check = ON` and a 512 MiB limit on any value. A file with a trigger, a view or a virtual table other than the FTS5 index is refused before any of its tables is read (legacy files may keep their three FTS triggers, which cannot fire on a read-only connection). The SQL is fixed text in the filter; nothing from a file or from your document is ever put into a query, executed, or fetched from the network. ## Tests ```sh make test # or: test/run.sh [case ...] ``` The runner needs `pandoc`, `sqlite3` and `gzip`. Each case in `test/cases/` is run with `--lua-filter spdf.lua --citeproc -t plain` and compared with `test/expected/` three ways: the text, the warnings, and the citations and references the filter produced (`-t native` through [`test/inspect.lua`](test/inspect.lua)). The cases cover printed folios from `p=`, checks of `f=`, inferred folios, the cover without folio, ranges, leaves, times, sections, verses, slides, sheets, canonical references, unknown documents, pages and parameters, prefixes and suffixes, several citations in one bracket, in-text citations and their punctuation, the braced form, CommonMark input, a Spanish document, locator labels in German and Spanish, merging with existing references and bibliography files, anchor links, invalid files (`roto.spdf` and the `conformance/invalid` files: none may crash the filter), legacy gzip files, and a missing `sqlite3`. They read the shared fixtures in [`integrations/fixtures/`](../fixtures/) and some files of the conformance suite (`conformance/files`, `conformance/legacy`, `conformance/invalid`), and build `test/build/mixed.spdf` from [`test/fixtures/mixed.sql`](test/fixtures/mixed.sql) for the anchor types those lack. Those files are rebuilt by other parts of the repository, so the cases name them by placeholders (`HASH_EN`, `HASH_QUIJOTE`, `HASH_EN_12` for the citekey…) that the runner fills with each file's current `source_sha256` and puts back in the outputs; the table is `HASHED` in `test/run.sh`. `test/run.sh --update` rewrites the expected outputs; read the diff before keeping it. ## Limitations - The filter's own words exist in English and Spanish only. Locator labels follow any CSL locale, but the probe uses the locale's terms: a style whose own `` section redefines the `page`, `folio`, `column` or `verse` terms may not see the label. - Times have no CSL label (Pandoc's locator parser does not know CSL 1.0.2 `timestamp`), so they are plain text in the suffix, `m:ss` below one hour and `h:mm:ss` above, as in SPEC §18; a `t=a,b` pair prints as a range. - `char=` and `xywh=` are kept in links but not printed: a citation locates a unit. - Citations in metadata fields (title, abstract) are not resolved. - Library folders are not searched recursively, `.spdfl.json` collection manifests are not read, and remote files are never fetched. - On Windows, `sqlite3.exe` must be on `PATH`; without a POSIX `sh`, legacy gzip files are decompressed in memory and without the 4 GiB cap. - Opening costs two short `sqlite3` runs per library file and one more per cited document; with a library of thousands of files, list the ones you cite. ## License MIT OR Apache-2.0, like the rest of the code in this repository. The test fixtures in `test/fixtures/` and the test style in `test/styles/` are dedicated to the public domain (CC0 1.0). # ===== CASTELLANO ===== --- # Especificación de SPDF 5.0 URL: https://spdf.joseluissaorin.com/es/especificacion > La especificación normativa del formato SPDF 5.0: contenedor, esquema, anclas y su URI, ficha CSL-JSON, búsqueda, vectores, integridad, validación, cita y conformidad. - **Estado:** borrador de trabajo, 2026-10-07. Lo bastante estable para implementarse; los cambios pasan por el proceso de RFC (`spec/rfcs/`) y se registran en `spec/CONTRACT.md` hasta que la 5.0 sea definitiva. - **Editor:** José Luis Saorín Ferrer. - **Esta versión:** `spec/SPEC.es.md` en . - **Original inglés:** [`SPEC.md`](SPEC.md). Este documento es la traducción española fiel de `SPEC.md`; en caso de discrepancia prevalece el texto inglés. - **Licencia:** esta especificación se publica bajo [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/). El código del repositorio es MIT OR Apache-2.0. Quienes contribuyen se comprometen a no hacer valer patentes contra las implementaciones. ## Resumen SPDF es un formato de fichero abierto y portátil para documentos que se han **leído una vez y pueden citarse para siempre**. Un fichero `.spdf` contiene el texto de un documento (un libro impreso, un escaneo, una grabación, una presentación de diapositivas, una hoja de cálculo, una página web) como un conjunto de unidades citables y de fragmentos buscables, y cada fragmento lleva un **ancla** exacta: la página o la hoja impresas, el segundo de una grabación, la diapositiva, el verso, la referencia canónica. Una cita producida a partir de un fichero SPDF solo puede imprimir lo que dice la fuente. El contenedor es una base de datos SQLite 3 sin más, los metadatos son un elemento CSL-JSON, el índice de texto completo usa tokenizadores que vienen con cualquier SQLite, y pueden convivir vectores de embedding opcionales de varios modelos. Cualquier lenguaje con SQLite puede leer SPDF sin bibliotecas especiales. ## Estado de este documento Esta es la primera versión pública del formato (las versiones anteriores, de la 3.0 a la 4.1, eran internas de Scholaris y se tratan como legado en [§20](#legacy)). La batería de conformidad de `conformance/` forma parte de la especificación: cuando este texto y un caso de conformidad discrepan, la discrepancia es un error que hay que resolver mediante el proceso de RFC; mientras no se resuelva, las implementaciones siguen el caso de conformidad. ## Índice - [Prefacio](#preface) - [1. Convenciones y terminología](#terminology) - [2. Contenedor](#container) - [3. Esquema](#schema) - [4. Anclas](#anchors) - [5. URI de ancla](#anchor-uri) - [6. Metadatos](#metadata) - [7. Texto, normalización y desplazamientos](#text-normalization) - [8. Búsqueda de referencia](#search) - [9. Espacios vectoriales](#vectors) - [10. Perfiles](#profiles) - [11. Extensiones](#extensions) - [12. Volcado canónico](#dump) - [13. Integridad y firmas](#integrity) - [14. Consideraciones de seguridad](#security) - [15. Consideraciones de privacidad](#privacy) - [16. Derechos](#rights) - [17. Anotaciones y colecciones](#annotations) - [18. Cita breve](#citation) - [19. Exportaciones](#exports) - [20. Formatos legados](#legacy) - [21. Conformidad](#conformance) - [22. Validación](#validation) - [23. Versionado y compatibilidad](#versioning) - [24. Tipo de medio e identificación de ficheros](#media-type) - [25. Internacionalización](#i18n) - [Referencias](#references) - [Apéndice A. Cambios respecto a SPDF 4.1](#changes) ## Prefacio SPDF nació dentro de Scholaris, una aplicación escrita por José Luis Saorín Ferrer para insertar en la escritura académica citas verificadas y exactas en la página. Scholaris necesitaba leer una fuente una sola vez (con la capa de texto de un PDF, un modelo de visión o un reconocedor de voz), conservar lo que había leído y responder, años después, la única pregunta que una cita tiene que responder con honradez: *¿dónde dice exactamente esto la fuente?* La respuesta tenía que sobrevivir a que el fichero original se moviera, a que se sustituyera el modelo de lectura y a que se reescribiera el motor de búsqueda. El resultado fue un fichero por documento, el *Scholaris Processed Document Format*, que pasó por una versión de JSON y SQLite comprimida con gzip (3.0) y por un esquema SQLite con nombres en español (4.0 y 4.1). La versión 5.0 es la primera pensada para todo el mundo. Conserva lo que la experiencia demostró acertado y abandona lo que la ataba a un solo programa: los identificadores están en inglés, el contenedor no va comprimido para que pueda proyectarse en memoria y leerse por rangos HTTP, los metadatos son CSL-JSON sin más para que Zotero, citeproc y Pandoc los entiendan, y cada número de un caso de conformidad procede de un oráculo que cualquiera puede volver a ejecutar. El nombre pasó a ser *Semantic Processed Document Format*; las siglas no cambiaron. Cinco principios guían cada decisión de esta especificación: 1. **Las anclas, primero.** Cada fragmento sabe exactamente de dónde procede: página física y folio impreso, hoja y cara, segundo, diapositiva, verso, referencia canónica. Nada que no pueda anclarse es citable. 2. **Procedencia.** Un fichero dice quién leyó cada unidad y con qué confianza, qué modelo produjo cada vector y cómo se obtuvo cada campo de los metadatos. Los datos derivados pueden recalcularse a partir del original más las unidades. 3. **Leer una vez, consultar muchas.** Leer un documento es caro (modelos de visión, reconocimiento de voz, corrección humana); consultarlo tiene que ser barato, posible sin conexión y posible desde cualquier lenguaje con SQLite. 4. **Portabilidad.** Un fichero, un documento, ningún servidor, ninguna dependencia propietaria, ninguna capa de compresión que deshacer, ningún código dentro del fichero. Lectores escritos en muchos lenguajes superan la misma batería de conformidad. 5. **Cita honrada.** Una cita imprime solo lo que dice un ancla. Un folio que se infirió se imprime entre corchetes; una página sin numerar se cita como sin numerar; la grafía modernizada que se usa para buscar nunca se cita textualmente. ## 1. Convenciones y terminología Las palabras clave «DEBE», «NO DEBE», «OBLIGATORIO», «DEBERÁ», «NO DEBERÁ», «DEBERÍA», «NO DEBERÍA», «RECOMENDADO», «NO RECOMENDADO», «PUEDE» y «OPCIONAL» de este documento, así como sus formas de plural y de femenino, se interpretan como sus equivalentes ingleses descritos en BCP 14 [RFC 2119] [RFC 8174] cuando, y solo cuando, aparecen en mayúsculas, como aquí. La correspondencia es la siguiente: | español | inglés (BCP 14) | |---|---| | DEBE, DEBEN | MUST | | NO DEBE, NO DEBEN | MUST NOT | | OBLIGATORIO, OBLIGATORIA, OBLIGATORIOS, OBLIGATORIAS | REQUIRED | | DEBERÁ, DEBERÁN | SHALL | | NO DEBERÁ, NO DEBERÁN | SHALL NOT | | DEBERÍA, DEBERÍAN | SHOULD | | NO DEBERÍA, NO DEBERÍAN | SHOULD NOT | | RECOMENDADO, RECOMENDADA, RECOMENDADOS, RECOMENDADAS | RECOMMENDED | | NO RECOMENDADO, NO RECOMENDADA, NO RECOMENDADOS, NO RECOMENDADAS | NOT RECOMMENDED | | PUEDE, PUEDEN | MAY | | OPCIONAL, OPCIONALES | OPTIONAL | El ABNF sigue [RFC 5234]. El JSON sigue [RFC 8259]; «objeto JSON», «array», «cadena» y «número» tienen el significado que les da RFC 8259. El SQL sigue el dialecto de SQLite. - **Documento**: la obra que describe un fichero SPDF (una por fichero). - **Original**: los bytes a partir de los cuales se leyó el documento (PDF, conjunto de imágenes, audio, EPUB…). - **Unidad**: una división citable del documento: una página u hoja, un intervalo de tiempo, una diapositiva, una sección, un rango de una hoja de cálculo. Las unidades están ordenadas y se numeran desde 1. - **Fragmento**: un pasaje buscable y citable de unas 150 a 300 palabras, con el ancla de su inicio y, si atraviesa unidades, la de su fin. - **Ancla**: un objeto JSON que localiza una unidad, un fragmento o una figura en el documento ([§4](#anchors)). - **URI de ancla**: la forma textual de un ancla, `spdf:#` ([§5](#anchor-uri)). - **Espacio**: un espacio vectorial, es decir, el modelo, las dimensiones y la codificación que produjeron un conjunto de vectores de embedding ([§9](#vectors)). - **Lector**: software que abre ficheros SPDF y expone su contenido. **Escritor**: software que crea ficheros SPDF. **Validador**: software que comprueba si los ficheros se ajustan a esta especificación. **Productor**: un escritor que además lee originales (OCR, reconocimiento de voz, vectores). - **Punto de código**: un valor escalar Unicode. Las longitudes y los desplazamientos de esta especificación cuentan puntos de código, nunca bytes ni unidades de código UTF-16. - **NFC**: la forma de normalización C de Unicode [UAX #15]. - **JCS**: el esquema de canonicalización de JSON (JSON Canonicalization Scheme) [RFC 8785]. ## 2. Contenedor ### 2.1 Fichero Un fichero SPDF 5.0 es un fichero de base de datos SQLite 3 [SQLITE-FORMAT] que contiene exactamente un documento. NO DEBE ir envuelto en ninguna capa de compresión ni de archivo: la cabecera de la base de datos DEBE empezar en el byte 0. Los escritores DEBEN establecer: - `PRAGMA application_id = 1397769286` (0x53504446 en hexadecimal). SQLite lo guarda en orden big-endian en el desplazamiento 68 de la cabecera, de modo que los bytes 68 a 71 dicen «SPDF» en ASCII. - `PRAGMA user_version = 500`. El valor codifica la versión de la especificación como mayor × 100 + menor × 10 (5.0 → 500, 5.1 → 510). - La fila `spdf_version` de `spdf_meta` con el valor `"5.0"` ([§3.2](#schema)). Los escritores DEBERÍAN usar un tamaño de página de 4096 bytes y el diario de reversión (rollback journal) en modo `DELETE` (nunca dejar un fichero `-wal` o `-journal` junto a un fichero distribuido), y ejecutar `VACUUM` tras la última escritura para que el fichero no tenga páginas libres. Los escritores NO DEBERÍAN usar `auto_vacuum`. Un fichero NO DEBE contener disparadores (triggers) ni vistas, y NO DEBE contener tablas virtuales distintas de las tablas FTS5 definidas en [§3](#schema). Los escritores mantienen por sí mismos sincronizado el índice de texto completo (por ejemplo, con `INSERT INTO fragments_fts(fragments_fts) VALUES('rebuild')` antes de `VACUUM`). ### 2.2 Nombre y tipo La extensión de fichero es `.spdf`. El tipo de medio es `application/vnd.spdf+sqlite3` ([§24](#media-type)). Un fichero contiene un documento; las bibliotecas de documentos se describen mediante un manifiesto de colección aparte ([§17](#annotations)). ### 2.3 Entrada comprimida con gzip Los ficheros legados 4.x son bases de datos SQLite envueltas en gzip [RFC 1952] ([§20](#legacy)). Por ello, los lectores DEBEN aceptar un fichero que empiece por los bytes mágicos de gzip `1F 8B`, descomprimirlo (en memoria o en un fichero temporal) con un límite configurable del tamaño descomprimido (valor por defecto RECOMENDADO: 4 GiB) y continuar con el resultado. Un fichero 5.0 envuelto en gzip es legible pero no conforme: los validadores lo notifican como `E003` en la lista de avisos ([§22](#validation)). ### 2.4 Apertura segura Los ficheros SPDF vienen de desconocidos. Todo lector DEBE abrirlos como sigue, y DEBE rechazar el fichero si su enlace (binding) con SQLite no permite cumplir alguno de los pasos: 1. Abrir la base de datos en solo lectura (`SQLITE_OPEN_READONLY`, o el parámetro de URI `mode=ro`). Nunca abrir un fichero distribuido en lectura y escritura en su ubicación. 2. `PRAGMA query_only = 1` y `PRAGMA trusted_schema = OFF`. 3. Activar `SQLITE_DBCONFIG_DEFENSIVE` donde el enlace lo ofrezca, y mantener desactivada la carga de extensiones (`sqlite3_enable_load_extension(db, 0)`; nunca llamar a `load_extension`). 4. Leer `sqlite_master` y rechazar el fichero si contiene un disparador, una vista o una tabla virtual distinta de `fragments_fts` y `fragments_fts_trigram` declaradas `USING fts5`. A los ficheros legados 4.x se les permiten exactamente los tres disparadores `fragmentos_ai`, `fragmentos_ad` y `fragmentos_au` ([§20](#legacy)), que nunca se activan en una conexión de solo lectura. 5. Imponer un tamaño máximo configurable a cada valor BLOB o TEXT que se lea (valor por defecto RECOMENDADO: 512 MiB), por ejemplo con `sqlite3_limit(db, SQLITE_LIMIT_LENGTH, …)`. Los lectores DEBERÍAN además desactivar la E/S proyectada en memoria (`PRAGMA mmap_size = 0`) y activar `PRAGMA cell_size_check = ON` para ficheros de fuentes no fiables, y PUEDEN ejecutar `PRAGMA quick_check` antes de usarlos. Las operaciones que necesitan escribir, como la orden `integrity-check` de FTS5, DEBEN ejecutarse sobre una copia privada (por ejemplo, una copia en memoria hecha con la API de copia de seguridad), nunca sobre el fichero. [§14](#security) explica las amenazas. ## 3. Esquema ### 3.1 Visión general El esquema normativo es el script SQL [`schema/spdf-5.0.sql`](schema/spdf-5.0.sql), que se reproduce íntegro a continuación. Todas sus tablas son OBLIGATORIAS, aunque estén vacías; solo `fragments_fts_trigram` es OPCIONAL. Los nombres, los tipos y las restricciones de las columnas DEBEN ser los que figuran en él. Los escritores NO DEBEN añadir columnas a estas tablas; los datos que no encajan van a tablas de extensión ([§11](#extensions)). Los lectores DEBEN ignorar las columnas que no conocen (una versión menor posterior puede añadir columnas OPCIONALES, [§23](#versioning)). El JSON guardado en columnas TEXT DEBE ser JSON válido [RFC 8259] codificado en UTF-8; los escritores PUEDEN serializarlo de cualquier forma (el volcado canónico lo vuelve a serializar, [§12](#dump)). Las marcas de tiempo son cadenas ISO 8601 / RFC 3339 en UTC con el sufijo `Z`. Los identificadores (columnas `id`) son cadenas no vacías que elige el escritor; son opacos, distinguen entre mayúsculas y minúsculas y son estables durante toda la vida del fichero. ```sql PRAGMA application_id = 1397769286; -- 0x53504446, "SPDF" PRAGMA user_version = 500; CREATE TABLE spdf_meta (key TEXT PRIMARY KEY, value TEXT NOT NULL); CREATE TABLE documents ( id TEXT PRIMARY KEY, kind TEXT NOT NULL, metadata TEXT NOT NULL, source_sha256 TEXT NOT NULL, source_ref TEXT, mime TEXT NOT NULL, bytes INTEGER NOT NULL, unit_count INTEGER NOT NULL, duration REAL, created TEXT NOT NULL, updated TEXT NOT NULL, title TEXT, authors TEXT, year INTEGER, language TEXT, rights TEXT); CREATE TABLE units ( id TEXT PRIMARY KEY, document TEXT NOT NULL REFERENCES documents(id), ord INTEGER NOT NULL, anchor TEXT NOT NULL, text TEXT NOT NULL DEFAULT '', notes TEXT, header TEXT, footer TEXT, image TEXT, thumbnail TEXT, reader TEXT NOT NULL, confidence REAL NOT NULL DEFAULT 1, printed TEXT, t0 REAL, t1 REAL, words TEXT); CREATE INDEX units_doc ON units(document, ord); CREATE INDEX units_printed ON units(document, printed); CREATE TABLE sections ( id TEXT PRIMARY KEY, document TEXT NOT NULL, parent TEXT, level INTEGER NOT NULL, title TEXT NOT NULL, unit_from TEXT NOT NULL, unit_to TEXT, summary TEXT); CREATE TABLE fragments ( n INTEGER PRIMARY KEY, id TEXT NOT NULL UNIQUE, document TEXT NOT NULL, unit TEXT NOT NULL, ord INTEGER NOT NULL, text TEXT NOT NULL, context TEXT NOT NULL DEFAULT '', section TEXT, anchor TEXT NOT NULL, anchor_end TEXT, search_text TEXT); CREATE INDEX fragments_doc ON fragments(document, ord); CREATE INDEX fragments_unit ON fragments(unit); CREATE VIRTUAL TABLE fragments_fts USING fts5( text, context, section, search_text, content='fragments', content_rowid='n', tokenize='unicode61 remove_diacritics 2'); -- OPTIONAL: -- CREATE VIRTUAL TABLE fragments_fts_trigram USING fts5( -- text, content='fragments', content_rowid='n', tokenize='trigram'); CREATE TABLE figures ( id TEXT PRIMARY KEY, document TEXT NOT NULL, unit TEXT NOT NULL, image TEXT NOT NULL, caption TEXT, description TEXT, anchor TEXT NOT NULL); CREATE TABLE spaces ( id TEXT PRIMARY KEY, provider TEXT NOT NULL, model TEXT NOT NULL, version TEXT, dims INTEGER NOT NULL, dtype TEXT NOT NULL DEFAULT 'f32', normalized INTEGER NOT NULL DEFAULT 1, truncated_from INTEGER, modalities TEXT NOT NULL, task_prefixes TEXT, created TEXT); CREATE TABLE vectors ( target TEXT NOT NULL, id TEXT NOT NULL, space TEXT NOT NULL REFERENCES spaces(id), document TEXT NOT NULL, data BLOB NOT NULL, PRIMARY KEY (target, id, space)); CREATE TABLE blobs (key TEXT PRIMARY KEY, mime TEXT NOT NULL, sha256 TEXT NOT NULL, data BLOB NOT NULL); CREATE TABLE provenance ( document TEXT NOT NULL, stage TEXT NOT NULL, provider TEXT, model TEXT, detail TEXT, ms INTEGER, at TEXT NOT NULL); CREATE TABLE extensions (name TEXT PRIMARY KEY, version TEXT NOT NULL, required INTEGER NOT NULL DEFAULT 0); ``` ### 3.2 `spdf_meta` Pares clave/valor sobre el fichero. Claves OBLIGATORIAS: | clave | valor | |---|---| | `spdf_version` | `"5.0"` | | `profile` | nombres de perfil separados por espacios, un subconjunto de `core semantic media full` ([§10](#profiles)); siempre incluye `core` | | `created` | momento de creación del fichero (UTC) | | `generator` | `name/version` del escritor, p. ej. `spdf-producer/0.3.1` | | `document_id` | igual a `documents.id` | Claves OPCIONALES: `content_sha256`, `signature`, `signer` ([§13](#integrity)) y `license_note` (texto libre para personas). Versiones posteriores o extensiones (con el prefijo `x__`) PUEDEN añadir otras claves; los lectores DEBEN ignorar las claves que no conocen. ### 3.3 `documents` Exactamente una fila. - `kind`: uno de `pdf` (PDF con una capa de texto utilizable), `scanned_pdf` (PDF leído por visión), `photos` (un conjunto de fotografías de páginas), `image` (una sola imagen), `audio`, `video`, `document` (DOCX, ODT, RTF, HTML, Markdown, texto plano), `epub`, `slides`, `sheet`, `web`. Los lectores DEBEN aceptar tipos desconocidos y tratarlos como `document`. - `metadata`: el elemento CSL-JSON con la extensión `spdf` ([§6](#metadata)). - `source_sha256`: SHA-256 en hexadecimal en minúsculas de los bytes del original. Identifica el documento a través de sus copias y es la referencia de documento preferida en las URI de ancla. - `source_ref`: dónde está el original: `blob:` cuando va dentro del fichero, una URL absoluta, o NULL. - `mime`, `bytes`: tipo de medio y tamaño en bytes del original. - `unit_count`: número de filas de `units` (una discrepancia es el aviso W102). - `duration`: segundos, para audio y vídeo; NULL en los demás casos. - `created`, `updated`: cuándo se creó el registro del documento y cuándo se modificó por última vez. - `title`, `authors`, `year`, `language`: copias desnormalizadas para filtrar sin analizar el JSON: el `title` de CSL; los apellidos (o los nombres literales) de la lista `author` de CSL unidos con `"; "`; el primer año de `issued`; el `language` de CSL. Cuando están presentes, DEBEN coincidir con `metadata`. - `rights`: objeto JSON de derechos ([§16](#rights)) o NULL. ### 3.4 `units` Una fila por unidad citable, `ord` = 1, 2, 3… sin huecos (E090), en orden de lectura. - `anchor`: el ancla de la unidad ([§4](#anchors)). - `text`: el texto completo de la unidad tal como se leyó, en NFC y en Markdown ligero ([§7](#text-normalization)). Cadena vacía para las unidades sin texto (una página en blanco, una fotografía). - `notes`: array JSON de cadenas (notas al pie separadas del cuerpo) o NULL. - `header`, `footer`: cabeceras y pies de página recurrentes, que se mantienen fuera de `text`, o NULL. - `image`, `thumbnail`: `blob:` o URL de la imagen de la unidad (página, fotograma, diapositiva) y de su miniatura, o NULL. - `reader`: lo que produjo `text` (`pdf-text-layer`, `tesseract-5`, `gemma-4-e4b`, `whisper-large-v3-turbo`, `human`…). `confidence`: de 0 a 1. - `printed`: el folio impreso de una unidad de tipo página, copiado de su ancla, para que los lectores puedan «ir a la página 145» con un índice. - `t0`, `t1`: inicio y fin en segundos, copiados de un ancla de tiempo; NULL en los demás casos. - `words`: tiempos de las palabras para audio y vídeo ([§7.4](#text-normalization)) o NULL. ### 3.5 `sections` El árbol de encabezados. `level` empieza en 1; `parent` es el id de la sección que la contiene o NULL; `unit_from` y `unit_to` son los ids de la primera y de la última unidad (`unit_to` es NULL cuando la sección termina con el documento); `summary` es un texto OPCIONAL en la lengua del documento. ### 3.6 `fragments` - `n`: un entero positivo, único y estable: es el rowid que usa el índice FTS5 (un rowid implícito puede cambiar con `VACUUM`). - `unit`: id de la unidad donde empieza el fragmento. `ord`: orden de lectura dentro del documento (creciente con la posición del fragmento en el texto). - `text`: el pasaje literal, en NFC, exactamente como en la fuente (nunca modernizado). - `context`: una línea que sitúa el fragmento en la obra («Capítulo III: la lucha por la existencia»), que usa la búsqueda; cadena vacía si no la hay. - `section`: array JSON de cadenas, la ruta de encabezados, o NULL. - `anchor`: ancla del inicio del fragmento. `anchor_end`: ancla de su fin cuando pasa a otra unidad; NULL en los demás casos. - `search_text`: la capa de grafía modernizada ([§25.3](#i18n)): texto que se usa SOLO para la búsqueda (`aſsi` → `así`, `V. M.` → `vuestra merced`). Cadena vacía cuando no aporta nada; NULL cuando no se ha calculado. NO DEBE mostrarse como el texto de la fuente ni citarse textualmente. ### 3.7 `fragments_fts` y `fragments_fts_trigram` `fragments_fts` es un índice FTS5 de contenido externo sobre `fragments` con las columnas `text`, `context`, `section` y `search_text`, en este orden, y el tokenizador `unicode61 remove_diacritics 2`, que ofrece cualquier SQLite con FTS5. DEBE estar sincronizado con `fragments` (E070). `fragments_fts_trigram` es OPCIONAL, indexa solo `text` con el tokenizador `trigram` (SQLite 3.34 o posterior) y DEBERÍA estar presente cuando el documento está mayoritariamente en chino, japonés o coreano. ### 3.8 `figures` Figuras, láminas, tablas en forma de imagen, fotografías dentro de una página. `image` es el `blob:` de una imagen recortada, o la imagen de la unidad junto con una `region` en el ancla. `caption` es el pie de figura impreso, si lo hay; `description` es una descripción en la lengua del documento (para la accesibilidad y la búsqueda). `anchor` lleva normalmente una `region`. ### 3.9 `spaces` y `vectors` Véase [§9](#vectors). `vectors.target` es `fragment`, `unit` o `figure`, y `vectors.id` es el id de esa fila; `data` es el vector en orden little-endian. ### 3.10 `blobs` Contenido binario incluido en el fichero: el original, imágenes de página, figuras recortadas, miniaturas. `key` es una cadena opaca (por convención, con forma de ruta: `pages/0001.png`), `mime` su tipo de medio, `sha256` el SHA-256 en hexadecimal en minúsculas de `data` (E080). Las demás tablas se refieren a un blob como `blob:`. ### 3.11 `provenance` Una fila por paso de producción: `stage` (`reading`, `transcription`, `folios`, `metadata`, `embedding`, `figures`…), `provider`, `model`, `detail` (objeto JSON o NULL), `ms` (duración en milisegundos) y `at` (marca de tiempo UTC). Véase [§15](#privacy) para saber qué no registrar. ### 3.12 `extensions` Véase [§11](#extensions). ## 4. Anclas ### 4.1 Generalidades Un ancla es un objeto JSON con un miembro `type` de tipo cadena. Los tipos que define esta versión y sus miembros son: | tipo | miembros OBLIGATORIOS | miembros OPCIONALES | |---|---|---| | `page` | `physical` (entero ≥ 1), `printed` (cadena o null) | `roman` (booleano), `foliation` (`page`, `leaf`, `column`; por defecto, `page`), `source` (`read`, `inferred`, `epub`, `none`), `confidence` (0–1) | | `time` | `t0`, `t1` (segundos, 0 ≤ t0 ≤ t1) | `speaker` (cadena) | | `section` | `path` (array de cadenas) | `paragraph` (entero ≥ 1), `printed` (cadena) | | `slide` | `n` (entero ≥ 1) | | | `sheet` | `sheet` (cadena), `row_from`, `row_to` (enteros) | | | `web` | `url` (cadena) | `path`, `paragraph`, `accessed` (fecha ISO) | | `image` | | | | `verse` | `line_from` (entero) | `line_to` (entero), `printed` (cadena) | | `canonical` | `scheme` (cadena), `ref` (cadena) | | Toda ancla PUEDE llevar además: - `region`: `{"x", "y", "w", "h"}`, números entre 0 y 1, fracciones de la anchura y de la altura de la imagen de la unidad, con el origen en la esquina superior izquierda; - `chars`: `[start, end]`, desplazamientos en puntos de código dentro del `text` en NFC de la unidad del ancla, `0 ≤ start ≤ end ≤ length`, con el fin excluido (E042); - `matter`: la clase de materia de la unidad: `body` (el texto de la obra), `front` (preliminares: portada, índice, licencias, dedicatoria, prólogo de una edición), `back` (índices, colofón, apéndices de una edición), `plate` (una lámina o desplegable fuera de las páginas de texto), `cover` (cubierta), `library` (exlibris, sellos, páginas de la biblioteca o del digitalizador, licencias de una edición digital) o `blank` (en blanco). Si falta, vale `body`; los lectores DEBEN tratar como `body` los valores que no conozcan. Los escritores DEBERÍAN indicarla en las unidades de los documentos paginados siempre que no sea `body`. En esta especificación, un «entero» es un número JSON de valor entero: `10` y `10.0` son el mismo valor JSON y ambos son enteros. Un ancla cuyo JSON no es válido, o a la que le falta un miembro OBLIGATORIO o lo tiene con un tipo erróneo, no es válida (E040); un `type` desconocido es E041. Los lectores DEBEN conservar los miembros que no conocen cuando copian anclas. ### 4.2 Páginas, folios y hojas `physical` es la posición de la página en el original, contando desde 1 (el índice de página del PDF, el número de la foto). `printed` es el folio exactamente como está impreso en la página («23», «xiv», «A-3», «1r»), o null cuando la página no lleva número. - `roman: true` marca los folios en números romanos (preliminares). - `foliation` describe qué cuentan los números impresos: `page` (cada página numerada), `leaf` (cada hoja numerada, con las caras `r`ecto y `v`erso, impresas como `"1r"`, `"1v"`) o `column` (columnas numeradas, como en algunos diccionarios y en los primeros libros impresos). - `source` indica cómo se obtuvo `printed`: `read` (visto en la página), `inferred` (deducido de las páginas vecinas, p. ej. un verso sin numerar), `epub` (de una lista de páginas EPUB), `none` (sin folio; `printed` es null). - Un folio inferido se cita entre corchetes, `p. [21]`; una página sin folio se cita como sin numerar ([§18](#citation)). Un productor NO DEBE inventar folios: si ninguna prueba respalda un número, `printed` es null y `source` es `none`. ### 4.3 Tiempo, secciones, versos y referencias canónicas Las anclas de tiempo localizan grabaciones en segundos desde el inicio del original; `speaker` nombra a quien habla. Las anclas de sección localizan texto sin paginar (EPUB, DOCX, HTML) por ruta de encabezados y número de párrafo, y PUEDEN añadir la página impresa equivalente cuando la edición ofrece una lista de páginas. Las anclas de verso cuentan versos (`line_from`, `line_to`), tal como los numeran las ediciones impresas. Las anclas canónicas usan un sistema de cita independiente de cualquier edición: `stephanus` (Platón), `bekker` (Aristóteles), `bible` (libro capítulo:versículo), `cts` (un URN de CTS [CTS]) o cualquier otro esquema documentado; los esquemas se escriben en ASCII en minúsculas. ### 4.4 Inicio y fin El `anchor` de un fragmento localiza su inicio; `anchor_end`, cuando está presente, localiza su fin y tiene el mismo `type`. La cita del fragmento entero imprime entonces un rango (`pp. 145-146`). - La **unidad final** de un fragmento es la primera unidad posterior a su unidad inicial (en orden de `ord`) cuya ancla es igual a `anchor_end` una vez quitados `chars` y `region` de ambas. - `chars` en `anchor` da la parte del fragmento que está en la unidad inicial, y `chars` en `anchor_end`, la parte que está en la unidad final (normalmente `[0, b]`). Los escritores DEBERÍAN indicar los dos en los fragmentos que cruzan unidades, para que los lectores sepan de qué unidad viene cada parte del pasaje. - Los escritores NO DEBERÍAN dejar que un fragmento cruce de una unidad de una clase de `matter` a otra (del texto de la obra a una lámina, una cubierta, una página de la biblioteca o una licencia), ni de una página con folio impreso a otra sin él: la cita de un fragmento así mezclaría localizadores de naturaleza distinta. Los validadores notifican esos fragmentos con W103. ## 5. URI de ancla ### 5.1 Sintaxis Una URI de ancla nombra un lugar de un documento con independencia de cualquier fichero: ``` spdf:sha256-3f2a…c9#p=29&f=21&char=118,301 ``` La referencia de documento es `sha256-` seguido de los 64 dígitos hexadecimales en minúsculas de `documents.source_sha256` (RECOMENDADA: es la misma para todas las copias del documento), o el id del documento con codificación porcentual. El fragmento es una lista de parámetros. Los parámetros reutilizan la sintaxis de W3C Media Fragments [MEDIA-FRAGMENTS] para el tiempo (`t=`) y el espacio (`xywh=`) y la sintaxis de RFC 5147 [RFC 5147] para los rangos de caracteres (`char=`), de modo que las herramientas que conocen esos estándares pueden interpretarlos. La forma canónica se define mediante este ABNF [RFC 5234]: ```abnf spdf-uri = "spdf:" docref [ "#" params ] docref = hash-ref / id-ref hash-ref = "sha256-" 64lhex lhex = DIGIT / %x61-66 ; 0-9 a-f id-ref = 1*vchar ; percent-encoded document id params = param *( "&" param ) param = p / pe / f / fe / t / s / para / sl / sh / rows / v / ref / char / xywh p = "p=" posint ; physical page pe = "pe=" posint ; physical end page f = "f=" value ; printed folio fe = "fe=" value ; printed end folio t = "t=" number [ "," number ] ; seconds, Media Fragments npt s = "s=" value *( "/" value ) ; section path para = "para=" uint ; paragraph sl = "sl=" posint ; slide sh = "sh=" value ; sheet name rows = "rows=" uint "-" uint ; sheet rows v = "v=" uint [ "-" uint ] ; verse lines ref = "ref=" value ":" value ; canonical scheme ":" reference char = "char=" uint "," uint ; code points, RFC 5147 style xywh = "xywh=percent:" number "," number "," number "," number value = *vchar vchar = unreserved / pct-encoded unreserved = ALPHA / DIGIT / "-" / "." / "_" / "~" pct-encoded = "%" HEXDIG HEXDIG ; uppercase in the canonical form posint = %x31-39 *DIGIT uint = "0" / posint number = uint [ "." 1*DIGIT ] ``` En la forma canónica, los parámetros aparecen como mucho una vez y en el orden de la regla `param` anterior (`p`, `pe`, `f`, `fe`, `t`, `s`, `para`, `sl`, `sh`, `rows`, `v`, `ref`, `char`, `xywh`); los valores son cadenas UTF-8 en las que todo byte que no sea un carácter no reservado se codifica en porcentaje con dígitos hexadecimales en mayúsculas; en `s`, los separadores entre los elementos de la ruta son barras `/` literales y una `/` dentro de un elemento es `%2F`; en `ref`, el primer `:` literal separa el esquema de la referencia, y los dos puntos que haya dentro de ellos son `%3A`. Los números usan la forma decimal más corta de ECMAScript (`4160`, `4175.5`, `0.125`), nunca un exponente. ### 5.2 De un ancla a una URI El formateo hace corresponder un ancla (y, opcionalmente, un ancla de fin) con parámetros: | ancla | parámetros | |---|---| | `page` | `p` = `physical`; `f` = `printed` si no es null; con una página de fin: `pe` = su `physical` si es distinto, `fe` = su `printed` si no es null y es distinto de `printed` | | `time` | `t` = `t0` y después `t1` (o el `t1` del ancla de fin) | | `section`, `web` | `s` = `path` si no está vacío; `para` = `paragraph`; `f` = `printed`; `fe` como para las páginas | | `slide` | `sl` = `n` | | `sheet` | `sh` = `sheet`; `rows` = `row_from`-`row_to` | | `verse` | `v` = `line_from`, o `line_from`-`line_to` cuando `line_to` está presente y es distinto; `f` = `printed` | | `canonical` | `ref` = `scheme`:`ref` | | `image` | ninguno | | cualquiera | `char` = `chars`; `xywh` = `region` × 100, como `percent:` | Los valores de `t` se redondean a 6 decimales. Los valores de `xywh` son las fracciones × 100 redondeadas a 4 decimales (`0.125` → `12.5`, `0.333333` → `33.3333`). Una URI sin parámetros (`spdf:`) designa el documento entero. ### 5.3 Análisis sintáctico El análisis devuelve la referencia de documento y un objeto **localizador** con un miembro por cada parámetro presente: `p`, `pe`, `para`, `sl` (enteros); `f`, `fe`, `sh` (cadenas); `t` (array de uno o dos números); `s` (array de cadenas); `rows` (dos enteros); `v` (uno o dos enteros); `ref` (objeto con `scheme` y `ref`); `char` (dos enteros); `xywh` (cuatro fracciones: los valores porcentuales divididos entre 100 y redondeados a 6 decimales). Los analizadores DEBEN aceptar la codificación porcentual con dígitos hexadecimales en minúsculas, los parámetros en cualquier orden, los caracteres no ASCII sin codificar (forma IRI [RFC 3987]), el prefijo `npt:` y las formas de reloj `h:mm:ss[.f]` y `mm:ss[.f]` en `t`. Los analizadores DEBEN ignorar los parámetros cuyo nombre no conocen. Los analizadores DEBEN rechazar: un esquema distinto de `spdf:`; una referencia de documento vacía; un parámetro repetido; números mal formados; `p`, `pe` o `sl` iguales a 0; un rango de `char` o de `t` cuyo fin precede a su inicio; `xywh` sin la unidad `percent:` (las coordenadas en píxeles no pueden resolverse sin la imagen); una codificación porcentual que no se decodifica como UTF-8 válido. Formatear un localizador analizado DEBE devolver la URI canónica byte a byte. La batería de conformidad comprueba el formateo, el análisis y la ida y vuelta para cada tipo de ancla. ### 5.4 Resolución `locate(file, reference)` resuelve una URI de ancla, o la URL de un recurso SPDF con identificador de fragmento ([§24](#media-type)), contra un fichero, y devuelve: ```json {"document": true, "units": ["p5", "p6"], "fragments": ["q4"], "char": [101, 278], "xywh": null} ``` 1. **Referencia.** Una URI `spdf:` se analiza como en el [§5.3](#anchor-uri); `document` es verdadero cuando su referencia de documento es `sha256-` seguido del `source_sha256` del fichero, o el id de documento del fichero. Cualquier otra referencia (una URL `https:`, una ruta de fichero) designa el propio fichero: `document` es verdadero y el texto que sigue a su primer `#`, si lo hay, se analiza como la lista de parámetros del §5.3. Cuando `document` es falso, `units` y `fragments` están vacíos (las implementaciones PUEDEN notificarlo como un error; los ejecutores de la batería traducen ese error a `document: false`). 2. **Regla.** El primer parámetro presente en el orden `p`, `f`, `t`, `sl`, `v`, `ref`, `s`, `sh` elige el predicado de abajo. Sin ninguno de ellos (sin fragmento, o solo con `char` y `xywh`), la referencia designa el documento entero y `units` y `fragments` están vacíos. 3. **Predicado** sobre un ancla (los miembros ausentes del ancla nunca coinciden): - `p`: un ancla `page` con `p ≤ physical ≤ pe` (`pe` vale `p` por defecto); - `f`: `printed` igual a `f` (en las unidades, la columna `units.printed`); - `t`: un ancla `time` con `t0 ≤ t < t1`, donde `t` es el primer valor del parámetro; la última unidad con ancla de tiempo (en orden de `ord`) también coincide cuando `t` es igual a su `t1`; - `sl`: un ancla `slide` con `n = sl`; - `v`: un ancla `verse` con `line_from ≤ v ≤ line_to` (`line_to` vale `line_from` por defecto), donde `v` es el primer valor del parámetro; - `ref`: un ancla `canonical` con el mismo `scheme` y el mismo `ref`; - `s`: un ancla `section` o `web` cuya `path` empieza por los elementos de `s`; cuando está presente `para`, la `path` debe ser igual a `s` y `paragraph` igual a `para`; - `sh`: un ancla `sheet` con `sheet = sh` y, cuando está presente `rows`, `row_from ≤ a ≤ row_to` para su primer valor `a`. 4. **Coincidencias.** `units` son los id de las unidades cuya ancla coincide, en orden de `ord`. `fragments` son los id de los fragmentos cuya `anchor` de inicio o cuya `anchor_end` coincide, en orden de `n` (un fragmento que termina en una página se encuentra desde esa página). Cuando no coincide ninguna unidad pero sí algunos fragmentos, `units` son las unidades iniciales distintas de esos fragmentos, en orden de `ord`. 5. **Caracteres.** `char` se refiere al texto de la primera unidad de `units`. Cuando está presente `char` = `[c, d]`, un fragmento se conserva en `fragments` solo si su unidad inicial es esa unidad y su `anchor` tiene `chars` = `[a, b]` que se solapan con el rango, o si su unidad final ([§4.4](#anchors)) es esa unidad y su `anchor_end` tiene `chars` que se solapan con él; `[a, b]` se solapa con `[c, d]` cuando `a < d` y `c < b` (si `c < d`), o cuando `a ≤ c < b` (si `c = d`). 6. `char` y `xywh` se copian del localizador, o son null. Pueden coincidir varias unidades (dos páginas con el folio impreso «1», un número de verso repetido en dos poemas): `locate` las devuelve todas y el lector deja elegir al usuario; `p` siempre deshace la ambigüedad entre páginas, y por eso las URI formateadas lo llevan. Se prevé el registro provisional del esquema de URI `spdf` [RFC 7595]; la solicitud está redactada en `governance/drafts/uri-scheme-spdf.md`. ## 6. Metadatos ### 6.1 Elemento CSL-JSON `documents.metadata` es un elemento CSL-JSON [CSL-JSON] que describe el documento tal como ha de citarse: como mínimo, `type` (un tipo CSL como `book`, `article-journal`, `chapter`, `thesis`, `speech`, `interview`, `broadcast`, `motion_picture`, `webpage`, `dataset`, `graphic`) y `title` (E051 si falta alguno de los dos). Miembros habituales: `author`, `editor`, `translator`, `interviewer` (arrays de nombres `{family, given}` o `{literal}`, con las partículas CSL `non-dropping-particle` y `dropping-particle` cuando hace falta), `issued` (`{"date-parts": [[year, month, day]]}`), `original-date`, `title-short`, `original-title`, `container-title`, `collection-title`, `publisher`, `publisher-place`, `volume`, `issue`, `page`, `edition`, `DOI`, `ISBN`, `ISSN`, `URL`, `accessed`, `language` (BCP 47), `abstract`, `note`. El miembro `id` es OPCIONAL dentro del fichero; las exportaciones lo establecen ([§19](#exports)). Los escritores NO DEBEN inventar metadatos. Un campo que no puede respaldarse con el original o con una fuente externa citada se omite. ### 6.2 El objeto de extensión `spdf` El miembro `spdf` del elemento contiene lo que CSL no puede expresar. Todos sus miembros son OPCIONALES: ```json "spdf": { "provenance": {"title": {"source": "title-page", "confidence": 0.99}, "issued": {"source": "colophon", "confidence": 0.95}}, "undated": {"from": 1600, "to": 1610, "basis": "printer active years"}, "original_language": "fr", "subtitle": "con anotaciones de Fernando de Herrera", "orcid": {"Foucault, Michel": "0000-0000-0000-0000"} } ``` - `provenance`: para cada campo CSL, de dónde procede el valor (`reading`, `title-page`, `colophon`, `crossref`, `openalex`, `wikidata`, `user`, `epub`, `pdf`, …) y una confianza entre 0 y 1. - `undated`: para obras sin fecha impresa, un rango verosímil (`from`, `to`, años, negativos para los años antes de Cristo) y las pruebas en que se basa (`basis`). NO DEBE copiarse en `issued`: una cita imprime «s. f.» / «n.d.» ([§18](#citation)). - `original_language`: etiqueta BCP 47 de la lengua original de una traducción. - `subtitle`: el subtítulo cuando el `title` de CSL es «Título: Subtítulo». - `orcid`: identificadores ORCID por nombre («Apellidos, Nombre»). Las extensiones PUEDEN añadir otros miembros con el prefijo `x__`. ## 7. Texto, normalización y desplazamientos ### 7.1 Codificación y normalización Todo el texto está en UTF-8 y en NFC. Los escritores DEBEN normalizar a NFC antes de guardar y antes de calcular desplazamientos. Los escritores NO DEBEN guardar U+0000, sustitutos (surrogates) desemparejados ni no-caracteres, y NO DEBERÍAN guardar otros caracteres de control salvo U+0009 (tabulador) y U+000A (salto de línea). Las líneas terminan solo con U+000A. ### 7.2 Desplazamientos Los desplazamientos de `chars` ([§4.1](#anchors)) y todas las longitudes de esta especificación cuentan puntos de código del texto en NFC. Las implementaciones cuyas cadenas son UTF-16 (JavaScript, Java, C#, `NSString` de Swift) DEBEN convertir: un carácter fuera del plano multilingüe básico cuenta como un punto de código, pero como dos unidades UTF-16. ### 7.3 Markdown ligero `units.text` PUEDE usar este subconjunto de CommonMark [COMMONMARK]: párrafos separados por una línea en blanco; encabezados de `#` a `######`; `*emphasis*` y `**strong**`; listas con `-` y con `1.`; citas con `>`; tablas al estilo de GitHub; marcas de nota al pie `[^1]` cuyo texto va a `notes`. Los lectores NO DEBEN interpretar el HTML en bruto de `text`; lo muestran como texto. Los desplazamientos cuentan los caracteres guardados, marcado incluido. Los fragmentos DEBERÍAN conservar el marcado de su unidad de origen para que `fragments.text` sea una subcadena del `text` de la unidad siempre que el fragmento no atraviese unidades. Los turnos de palabra de las transcripciones empiezan con la etiqueta `**Name:**` seguida de un espacio (`**Neil Armstrong:** Houston, Tranquility Base here.`). ### 7.4 Tiempos de las palabras `units.words` es el objeto JSON `{"v": 1, "t0": , "cs": [start, duration, start, duration, …]}`: un par de enteros por palabra, en centésimas de segundo desde `t0` (que es igual al `t0` de la unidad). Las palabras son las secuencias maximales de caracteres que no son espacio en blanco del `text` de la unidad, una vez eliminadas las etiquetas de hablante (`**Name:**`); por tanto, `cs` contiene exactamente el doble de enteros que palabras hay. Los lectores lo usan para resaltar la palabra que se está pronunciando y para convertir un rango de `char` en un rango de tiempo. ## 8. Búsqueda de referencia La búsqueda de referencia define lo que comprueba la conformidad: resultados que toda implementación devuelve de forma idéntica a partir del mismo fichero. Los productos PUEDEN ordenar mejor (palabras vacías, expansión de consultas, reordenación, filtros); aun así, DEBEN ofrecer el comportamiento de referencia para superar la batería, y DEBERÍAN señalar la diferencia en su documentación. Un elemento de resultado es `{"fragment_id", "score", "via", "anchor", "anchor_uri"}`, donde `via` enumera los métodos que han contribuido (`"lexical"`, `"vector"`) en ese orden, `anchor` es el ancla del fragmento y `anchor_uri` es la URI formateada a partir de `anchor` y `anchor_end` con la referencia de documento `sha256-`. ### 8.1 Búsqueda léxica Dadas una cadena de consulta y un límite: 1. **Normalizar**: `q` = NFC(consulta). 2. **Frases**: recorrer `q` de izquierda a derecha. Una comilla de apertura `"` (U+0022), `“` (U+201C), `«` (U+00AB) o `„` (U+201E) abre una frase que cierra, respectivamente, la siguiente comilla `"`, `”` (U+201D), `»` (U+00BB) o, en el caso de `„`, `“` o `”`. El texto entre las comillas es la frase. Una comilla de apertura sin comilla de cierre se trata como un separador. 3. **Palabras**: secuencias maximales de caracteres cuya categoría general Unicode es letra (L), marca (M) o número (N). Un término de frase son las palabras de la frase unidas con un espacio; las frases sin palabras se descartan. 4. **Términos**: si hay al menos un término de frase, los términos son los términos de frase (las palabras sueltas fuera de las comillas se descartan) y el operador es `AND`. En caso contrario, los términos son las palabras de `q` y el operador es `OR`. Los términos duplicados se eliminan, conservando el primero, y se comparan por la clave `lower(remove_Mn(NFD(term)))` (minúsculas por defecto de Unicode, tras eliminar las marcas que no ocupan espacio); la clave se usa solo para detectar duplicados. No se eliminan palabras vacías. Sin términos, el resultado está vacío. 5. **Cadena MATCH**: cada término, **tal como está escrito** (sin plegado de mayúsculas ni descomposición), es una cadena FTS5: `"` + el término con cada `"` duplicada + `"`; las cadenas se unen con ` AND ` o ` OR `. El tokenizador pliega por sí mismo las mayúsculas y los diacríticos; plegar la consulta de antemano rompería coincidencias (`Straße`, `fin`). 6. **Consulta**: ```sql SELECT f.n, f.id, bm25(fragments_fts, 1.0, 0.5, 0.5, 1.0) AS r FROM fragments_fts JOIN fragments f ON f.n = fragments_fts.rowid WHERE fragments_fts MATCH ?1 ORDER BY r, f.n LIMIT ?2 ``` La puntuación es −r. Como `search_text` es la cuarta columna indexada, una consulta en grafía moderna encuentra la grafía antigua sin ningún paso especial. Como `n` es el rowid del índice, las implementaciones PUEDEN ordenar solo dentro del índice (`SELECT rowid, bm25(…) FROM fragments_fts WHERE fragments_fts MATCH ?1 ORDER BY 2, 1 LIMIT ?2`) y buscar los ids de fragmento únicamente de las filas devueltas; el resultado es idéntico. Los lectores que obtienen los ficheros por rangos HTTP DEBERÍAN hacerlo así, ya que evita leer todos los fragmentos que coinciden. 7. **Ruta CJK**: si `q` contiene un punto de código de alguno de los rangos U+2E80–U+2FDF, U+3040–U+30FF, U+3100–U+312F, U+3130–U+318F, U+31A0–U+31FF, U+3400–U+4DBF, U+4E00–U+9FFF, U+A960–U+A97F, U+AC00–U+D7AF, U+F900–U+FAFF, U+FF66–U+FF9F o U+20000–U+3FFFF, el paso 6 se sustituye por lo siguiente: - si existe `fragments_fts_trigram` y todos los términos tienen al menos 3 puntos de código, la misma cadena MATCH se ejecuta contra `fragments_fts_trigram`, ordenada por `bm25(fragments_fts_trigram)` y después por `n`; la puntuación es −bm25; - en caso contrario (no hay índice de trigramas, o hay un término de menos de 3 puntos de código, que un índice de trigramas no puede encontrar), se ejecuta sobre `fragments.text` la **alternativa de reserva por subcadena**: para cada fragmento, `hits` = el número de términos `t` con `instr(text, t) > 0`; se devuelven los fragmentos con `hits` ≥ 1 (con el operador `OR`) o con `hits` = número de términos (con `AND`), ordenados por `hits` de forma descendente y después por `n`; la puntuación es `hits`. ### 8.2 Búsqueda vectorial Dados un espacio, un objetivo (`fragment` por defecto, o `unit`, `figure`), un vector de consulta de `dims` números y un límite: se compara la consulta con todos los vectores de ese espacio y de ese objetivo por fuerza bruta. Cada componente se convierte en un número IEEE 754 binary64 (f32 y f16 de forma exacta; i8 como q/127). La consulta se usa tal como se da, sin normalizar. Cuando el espacio tiene `normalized = 1`, la puntuación es el producto escalar; en caso contrario, es la similitud del coseno. Los resultados se ordenan por puntuación descendente y después por el `n` del fragmento, el `ord` de la unidad o el `id` de la figura. Un elemento de resultado cuyo objetivo es `unit` o `figure` lleva `unit_id` o `figure_id` en lugar de `fragment_id`, y la URI de ancla del ancla propia de la unidad o de la figura. Los productos PUEDEN usar índices aproximados; la referencia es exhaustiva. ### 8.3 Búsqueda híbrida Se ejecutan la búsqueda léxica y la búsqueda vectorial (objetivo `fragment`), cada una con profundidad `max(limit, 50)`, y se fusionan mediante la fusión por rango recíproco (reciprocal rank fusion) [RRF] con k = 10: puntuación = Σ 1/(10 + rango) sobre las listas que contienen el fragmento, con el rango empezando en 1. Se ordena por puntuación descendente y después por `n`; se conservan `limit` resultados. La constante 10 se midió en Scholaris: el clásico 60 aplana las listas cortas y buenas. ### 8.4 Comparación en la conformidad El orden de los resultados DEBE coincidir exactamente; las puntuaciones DEBEN coincidir con una tolerancia absoluta de 1e-6. Las puntuaciones léxicas son las del propio `bm25()` de SQLite, que es el oráculo. ## 9. Espacios vectoriales ### 9.1 Espacios Una fila de `spaces` describe cómo se produjo un conjunto de vectores: - `id`: `@` para vectores `f32` y `@:` en los demás casos (`embeddinggemma-2@768`, `embeddinggemma-2@256:i8`). PUEDE seguir un sufijo `+` para separar vectores del mismo modelo calculados a partir de entradas distintas (el Scholaris legado usa `+contexto`). - `provider` (quién ejecutó el modelo: `local`, `google`, `inferbox`…), `model`, `version`. - `dims`: número de componentes. - `dtype`: `f32` (IEEE 754 binary32), `f16` (binary16) o `i8` (byte con signo; el valor es q/127). Cualquier otro valor no es válido (E032). - `normalized`: 1 si todos los vectores guardados tienen norma euclídea unitaria (antes de la cuantización). - `truncated_from`: para el truncamiento Matryoshka [MRL], la dimensión original (`768` para un vector recortado a 256); NULL en los demás casos. Los vectores truncados DEBERÍAN renormalizarse antes de guardarse, con `normalized = 1`. - `modalities`: array JSON de las modalidades de entrada que acepta el modelo (`text`, `image`, `audio`, `video`, `pdf`). - `task_prefixes`: objeto JSON con los prefijos o las instrucciones que se usaron al codificar, `{"document": "…", "query": "…"}`, para que un lector pueda codificar las consultas de la misma manera; NULL si no los hay. Un fichero PUEDE contener varios espacios; un fichero sin espacios es válido (perfil `core`). ### 9.2 Vectores `vectors.data` es el vector como `dims` valores little-endian del `dtype` del espacio, de modo que su longitud es `dims` × 4, 2 o 1 bytes (E030). Todo vector se refiere a un espacio de `spaces` (E031). Los escritores cuantizan como sigue: f32 → f16 con el redondeo IEEE al par más cercano; f32 → i8 con `q = clamp(round_half_away_from_zero(v × 127), −127, 127)`. El valor −128 no se usa. Un valor que no cabe en el dtype (un f32 finito por encima de 65504 que se redondearía a infinito en f16, un valor no finito) es un error para el escritor, y nunca se guarda en silencio. ### 9.3 Compatibilidad entre espacios y cuantizaciones Dos espacios son **compatibles**, y un mismo vector de consulta sirve para ambos, cuando `provider`, `model`, `version`, `dims`, `normalized`, `truncated_from` y `task_prefixes` son iguales; `dtype` puede ser distinto. Por tanto, un lector que dispone de un modelo PUEDE buscar en un espacio `f32` y en su copia `i8` con la misma consulta. Los espacios que difieren en cualquier otro campo no son comparables: los lectores NO DEBEN mezclar puntuaciones de espacios incompatibles, y NO DEBEN comparar vectores de dimensiones distintas. Un espacio Matryoshka (`truncated_from` = 768, `dims` = 256) es compatible con una consulta solo si la consulta se truncó a las mismas dimensiones y se renormalizó. ## 10. Perfiles `spdf_meta.profile` declara qué promesas hace un fichero. Los perfiles son etiquetas acumulativas; un fichero enumera todos los perfiles que cumple. | perfil | requisitos | |---|---| | `core` | OBLIGATORIO en todo fichero. Todas las tablas de [§3](#schema); al menos una unidad; todas las unidades, los fragmentos y las figuras con ancla; texto en NFC; índice FTS sincronizado. | | `semantic` | Al menos un espacio, y vectores para todos los fragmentos en al menos un espacio. Un fichero `semantic` sin vectores provoca W100. | | `media` | `kind` es `audio` o `video`; las unidades llevan anclas `time` y `t0`/`t1`; `duration` tiene valor; `words` DEBERÍA estar presente. Un fichero `media` sin anclas de tiempo provoca W101. | | `full` | `semantic` y, para audio y vídeo, `media`; para los tipos paginados, imágenes de página (`units.image`) y figuras cuando el original las tiene. | Los lectores NO DEBEN rechazar un fichero por su perfil; los perfiles dicen a los lectores qué esperar y a los validadores qué comprobar. ## 11. Extensiones Los datos que esta especificación no define van a **tablas de extensión** llamadas `x__` (letras ASCII en minúsculas, dígitos y `_`; `` es un nombre que controla el autor, p. ej. `x_scholaris_claims`). Cada extensión en uso se declara en la tabla `extensions` con su `name` (`_` o el prefijo de la tabla), una `version` y `required`: - `required = 0`: los lectores que no conocen la extensión la ignoran. - `required = 1`: el fichero no puede entenderse sin ella; un lector que no la conoce DEBE rechazar el fichero con E060. Las extensiones NO DEBEN cambiar el significado de las tablas del núcleo, NO DEBEN añadirles columnas y NO DEBERÍAN ser obligatorias. Las tablas de extensión no forman parte del volcado canónico. Una extensión que resulta útil a varias implementaciones pasa a formar parte del núcleo mediante el proceso de RFC ([§23](#versioning)). ## 12. Volcado canónico El volcado canónico es una vista JSON de un fichero que todas las implementaciones producen de forma idéntica. Es el oráculo de la batería de conformidad y la entrada del hash de integridad. ```jsonc {"spdf_version": "5.0", // legacy files: "4.0"/"4.1" and "legacy": true "meta": {"": "", …}, // every spdf_meta row "fts": {"tokenizer": "unicode61 remove_diacritics 2", "trigram": false}, "document": {"id", "kind", "metadata", "source_sha256", "source_ref", "mime", "bytes", "unit_count", "duration", "created", "updated", "title", "authors", "year", "language", "rights"}, "units": [{"id", "ord", "anchor", "text", "notes", "header", "footer", "image", "thumbnail", "reader", "confidence", "printed", "t0", "t1", "words"}], "sections": [{"id", "parent", "level", "title", "unit_from", "unit_to", "summary"}], "fragments": [{"n", "id", "unit", "ord", "text", "context", "section", "anchor", "anchor_end", "search_text"}], "figures": [{"id", "unit", "image", "caption", "description", "anchor"}], "spaces": [{"id", "provider", "model", "version", "dims", "dtype", "normalized", "truncated_from", "modalities", "task_prefixes", "created"}], "vectors": {"": {"count": 6, "sha256": ""}}, "blobs": [{"key", "mime", "bytes", "sha256"}], "provenance": [{"stage", "provider", "model", "detail", "ms", "at"}], "extensions": [{"name", "version", "required"}]} ``` Reglas: 1. Todos los miembros enumerados están presentes. El NULL de SQL se convierte en `null`; INTEGER, en un entero JSON; REAL, en un número JSON; TEXT, en una cadena. Las columnas que contienen JSON (`metadata`, `rights`, `anchor`, `anchor_end`, `notes`, `words`, `section`, `modalities`, `task_prefixes`, `detail`) se analizan y se incrustan como valores JSON. Se omite la columna `document` de las tablas hijas. Los booleanos guardados como enteros (`normalized`, `required`) siguen siendo enteros. 2. Todo número que no sea entero, incluidos los que están dentro del JSON analizado, se redondea a 6 decimales (redondeo de la mitad al par sobre su valor binario exacto); un resultado de −0 se convierte en 0. 3. Orden: `units` por `ord`; `fragments` por `n`; `sections`, `figures` y `spaces` por `id`; `blobs` por `key`; `extensions` por `name` (orden de puntos de código, que es la intercalación `BINARY` de SQLite sobre UTF-8); `provenance` por los bytes UTF-8 de la serialización JCS de cada entrada. `fts.trigram` es verdadero si y solo si existe `fragments_fts_trigram`; `fts.tokenizer` es el valor de la opción `tokenize` de `fragments_fts` tal como está declarado, sin sus comillas y con cada secuencia de espacios en blanco reducida a un solo espacio (`unicode61 remove_diacritics 2`; `unicode61`, el valor por defecto de FTS5, si no está presente). 4. `vectors` tiene un miembro por cada `vectors.space` distinto: `count` es el número de filas y `sha256` el SHA-256 en hexadecimal de sus blobs `data` concatenados por orden de `target` y después de `id`. 5. `blobs[].bytes` y `blobs[].sha256` se calculan a partir de `data`, no se copian de la columna `sha256`. 6. La serialización, siempre que importen los bytes (cálculo de hashes), es JCS [RFC 8785]: sin espacios en blanco, con los miembros de objeto ordenados por las unidades de código UTF-16 de sus nombres, los números en la forma de ECMAScript (`1`, no `1.0`; `0.000001`, no `1e-6`) y las cadenas en UTF-8 con solo `"`, `\` y U+0000–U+001F escapados. Para los ficheros legados, el volcado es la vista 5.0 definida en [§20](#legacy), con el `spdf_version` legado y `"legacy": true`. ## 13. Integridad y firmas `spdf_meta.content_sha256` (OPCIONAL) es el SHA-256 en hexadecimal en minúsculas de la serialización JCS del volcado canónico del que se han eliminado los miembros `meta.content_sha256`, `meta.signature` y `meta.signer`. Cubre todo el contenido salvo las tablas de extensión y es independiente de la disposición de páginas de SQLite, de modo que dos escritores que guardan el mismo contenido producen el mismo hash. `spdf_meta.signature` (OPCIONAL, requiere `content_sha256` y `signer`) es la codificación base64 estándar, con relleno, de una firma Ed25519 [RFC 8032] sobre los bytes ASCII de la cadena `spdf-content-sha256:` seguida del `content_sha256` en hexadecimal. `spdf_meta.signer` es `ed25519:` seguido de la codificación base64 estándar de la clave pública de 32 bytes. Los validadores que encuentren `content_sha256` DEBEN recalcularlo (E081 si no coincide) y, si hay una firma, DEBEN verificarla (E082 si la verificación falla). Una firma válida prueba que el titular de la clave produjo este contenido; no dice nada sobre si la clave es de confianza. Los lectores DEBERÍAN mostrar quién firmó (la clave, o un nombre que el usuario le haya asociado) y NO DEBEN presentar una clave desconocida como de confianza. ## 14. Consideraciones de seguridad Un fichero SPDF es una base de datos escrita por otra persona. Abrirlo es analizar una entrada no fiable con un motor complejo. Las amenazas, y las reglas de esta especificación que les dan respuesta: - **Código en el esquema.** Los disparadores, las vistas y las tablas virtuales pueden ejecutar SQL o llamar a módulos cuando se usa la base de datos. Los ficheros NO DEBEN contenerlos ([§2.1](#container)); los lectores DEBEN rechazarlos, abrir en solo lectura con `query_only`, `trusted_schema = OFF` y el indicador defensivo, y no cargar nunca extensiones ([§2.4](#container)). Los disparadores FTS legados se toleran solo porque nunca se activan en una conexión de solo lectura. - **Bases de datos mal formadas.** SQLite es robusto frente a ficheros corruptos, pero recomienda precauciones adicionales con los que no son fiables [SQLITE-SECURITY]: desactivar la E/S proyectada en memoria, activar `cell_size_check`, fijar límites de longitud, considerar `quick_check`. - **Bombas de descompresión.** La entrada gzip (legado) DEBE descomprimirse con un límite de tamaño ([§2.3](#container)). - **Valores desmesurados.** Los blobs, los textos y los valores JSON DEBEN estar acotados; los analizadores de JSON DEBERÍAN limitar la profundidad de anidamiento (valor RECOMENDADO: 64). - **Inyección en las consultas.** El texto del usuario nunca llega a FTS5 como sintaxis: cada término es una cadena FTS5 entre comillas ([§8.1](#search)). El SQL siempre va parametrizado. Las implementaciones PUEDEN limitar el número de términos (valor RECOMENDADO: 64) para acotar el coste de la consulta. - **Rutas.** Las claves de los blobs son cadenas opacas, no nombres de fichero. Un lector que extrae blobs a disco DEBE sanearlas (sin rutas absolutas, sin `..`, sin nombres de dispositivo). - **Referencias remotas.** `source_ref`, `image`, `thumbnail`, `URL` y las anclas `web` pueden apuntar a la red. Los lectores NO DEBEN descargarlas automáticamente: la descarga revela que se abrió el fichero y puede alcanzar servicios internos. Solo hay que descargarlas ante una acción del usuario, y mostrando antes la dirección. - **Contenido activo.** `text` es Markdown ligero; los lectores NO DEBEN representar el HTML en bruto que contenga, y DEBEN escapar el texto antes de insertarlo en HTML. Las imágenes de los blobs son una entrada no fiable para los decodificadores de imágenes; el SVG NO DEBE representarse con los scripts activados. - **Procedencia falsificada.** La procedencia, la confianza y los metadatos son afirmaciones del escritor. Solo una firma con una clave de confianza ([§13](#integrity)) los atribuye. - **Entradas de los modelos.** El texto leído de un fichero puede contener instrucciones dirigidas a modelos de lenguaje («ignora las instrucciones anteriores…»). Las aplicaciones que pasan texto SPDF a un modelo DEBEN tratarlo como datos, no como instrucciones. ## 15. Consideraciones de privacidad - **Los vectores pueden filtrar el texto.** Los vectores de embedding pueden invertirse: hay ataques publicados que reconstruyen la mayor parte de una entrada breve a partir de su vector [VEC2TEXT]. Distribuir los vectores de un texto es casi como distribuir el texto. Las reglas de derechos y de confidencialidad que se aplican al texto se aplican a sus vectores ([§16](#rights)); un escritor al que se pide eliminar el texto de un documento restringido DEBE eliminar también sus vectores. - **La procedencia puede delatar al productor.** Los escritores NO DEBERÍAN registrar en `provenance.detail` ni en `generator` rutas de ficheros locales, nombres de usuario, nombres de máquina, identificadores de cuenta, claves de API ni instrucciones (prompts) que contengan datos personales. - **Personas en los documentos.** Las entrevistas y las grabaciones nombran a los hablantes y pueden contener datos personales. Los productores DEBERÍAN permitir a los usuarios eliminar o seudonimizar los nombres de `speaker`, y los lectores NO DEBERÍAN indexar nombres de hablantes en servicios compartidos sin consentimiento. - **Las anotaciones son personales.** Las anotaciones del usuario viven fuera del fichero, en ficheros auxiliares `.spdfa.json` ([§17](#annotations)), para que compartir un documento nunca comparta las notas de quien lo lee. - **La apertura es observable** solo si un lector descarga referencias remotas; véase [§14](#security). ## 16. Derechos `documents.rights` es NULL o un objeto JSON: ```json {"license": "CC-BY-4.0", "access": "open", "holder": "Universidad de La Laguna", "note": "Text and images under CC BY 4.0; page scans courtesy of the library."} ``` - `license`: un identificador o una expresión de licencia SPDX [SPDX] (`CC-BY-4.0`, `CC0-1.0`), o la URL de una licencia o de una declaración de derechos (para el dominio público, `https://creativecommons.org/publicdomain/mark/1.0/`; para las declaraciones de derechos, `http://rightsstatements.org/vocab/InC/1.0/`). - `access`: `open` (cualquiera puede recibir el fichero), `restricted` (solo el público que permita el titular: una clase, una biblioteca) o `private` (copia personal). - `holder`: el titular de los derechos, o null. `note`: texto libre. SPDF no concede derechos. Un fichero hecho a partir de una obra protegida por derechos de autor es una copia de esa obra, vectores incluidos ([§15](#privacy)). Los productores DEBERÍAN rellenar `rights` cuando conocen los derechos y DEBERÍAN poner `access` en `private` por defecto cuando no los conocen, y los lectores DEBERÍAN mostrar `rights` antes de compartir un fichero. La condición de dominio público depende de la jurisdicción; `note` es el lugar para indicar de cuál. ## 17. Anotaciones y colecciones ### 17.1 Anotaciones en `.spdfa.json` Las anotaciones del usuario (subrayados, notas, etiquetas) se guardan fuera del documento, en un fichero con la extensión `.spdfa.json`, como una `AnnotationCollection` de W3C Web Annotation [WEB-ANNOTATION] en JSON-LD: ```json {"@context": "http://www.w3.org/ns/anno.jsonld", "type": "AnnotationCollection", "spdf_annotations": "1.0", "label": "Notas de lectura", "first": {"type": "AnnotationPage", "items": [ {"id": "urn:uuid:7b0c…", "type": "Annotation", "motivation": "commenting", "created": "2026-10-07T09:00:00Z", "body": {"type": "TextualBody", "value": "Origen del tópico.", "format": "text/plain", "language": "es"}, "target": {"source": "spdf:sha256-3f2a…c9", "selector": [ {"type": "SpdfAnchorSelector", "value": "spdf:sha256-3f2a…c9#p=29&f=21&char=118,301"}, {"type": "TextQuoteSelector", "exact": "En un lugar de la Mancha", "prefix": "", "suffix": ", de cuyo nombre"}]}}]}} ``` `target.source` es la URI de ancla sin fragmento. El `SpdfAnchorSelector` lleva la URI de ancla completa; el `TextQuoteSelector` [WEB-ANNOTATION] permite que la anotación sobreviva a una relectura que altere los desplazamientos. Los lectores que no pueden resolver el ancla DEBERÍAN recurrir a la cita textual como alternativa de reserva. El miembro `spdf_annotations` indica la versión de este perfil. PUEDEN añadirse selectores de otros tipos (un `FragmentSelector` conforme a Media Fragments para el tiempo y la región) para las herramientas que no conocen SPDF. ### 17.2 Colecciones en `.spdfl.json` Una biblioteca es un manifiesto, no un contenedor: ```json {"spdf_library": "1.0", "name": "Tesis: fuentes", "created": "2026-10-07T00:00:00Z", "items": [{"sha256": "3f2a…c9", "title": "El ingenioso hidalgo…", "authors": "Cervantes Saavedra", "year": 1605, "url": "https://example.org/quijote.spdf", "file_sha256": "…"}]} ``` `sha256` es el `source_sha256` del documento (la identidad que usan las URI de ancla); `url` y `file_sha256` (SHA-256 de los bytes del fichero `.spdf`) son OPCIONALES y permiten a un lector descargar y comprobar una copia. Los elementos están en el orden en que los ordenó el usuario. Los esquemas JSON (JSON Schema) de ambos ficheros auxiliares están en [`json-schema/`](json-schema/). ## 18. Cita breve `cite(anchor, anchor_end, metadata, locale)` produce una cita autor-fecha entre paréntesis, la forma que comparten la mayoría de los estilos, para que todas las implementaciones impriman el mismo localizador. Las bibliografías completas y los demás estilos se producen a partir del elemento CSL-JSON con un procesador CSL ([§19](#exports)). ### 18.1 Cita de un ancla ``` ( names ", " year [ ", " locator ] ")" ``` Se definen las configuraciones regionales (locale) `es` y `en`; una configuración regional se asigna por su subetiqueta de lengua principal (`es-ES` → `es`), y cualquier otra recurre a `en` como alternativa de reserva. **Nombres**, tomados del `author` de CSL. El nombre de una persona es `literal` si está presente; si no, la `non-dropping-particle`, un espacio y `family`; si no, `given`. Con un autor, ese nombre; con dos, `A y B` (es) o `A and B` (en), donde el español escribe `e` en lugar de `y` cuando el segundo nombre empieza por el sonido /i/ (`i`, `í`, `hi` o `hí` no seguidos de vocal: `Gómez e Iglesias`, `Gómez e Hidalgo`, pero `Gómez y Hierro`); con tres o más, `A et al.` en ambas configuraciones regionales. Sin autores, el `title-short` o, en su defecto, el `title` hasta sus primeros dos puntos, sin espacios en los extremos. **Año**: el primer año de `issued`; los años negativos se escriben como `375 a. C.` (es) o `375 BC` (en). Sin año, `s. f.` (es) o `n.d.` (en). El rango de `spdf.undated` no se imprime en una cita breve. **Localizador**: | ancla | es | en | |---|---|---| | página, folio leído | `p. 145` | `p. 145` | | página, folio romano | `p. xiv` | `p. xiv` | | página, folio inferido | `p. [21]` | `p. [21]` | | página sin folio | `s. p.` | `n. pag.` | | rango de páginas (fin con otro folio) | `pp. 145-146`, `pp. 20-[21]` | igual | | hoja / rango de hojas | `fol. 1r`, `fol. [2v]`, `fols. 1r-[1v]` | igual | | columna / rango de columnas | `col. 45`, `cols. 45-46` | igual | | tiempo (t0, segundos redondeados hacia abajo) | `1:09:20`, `0:42` | igual | | rango de tiempo (t1 del ancla de fin) | `0:12-0:24` | igual | | sección o web con `printed` | como una página | como una página | | sección o web | `§ 3.2 El panóptico, párr. 4` | `§ 3.2 El panóptico, para. 4` | | diapositiva | `diap. 3` | `slide 3` | | hoja de cálculo | `Datos, filas 4-9`, `Datos, fila 4` | `Datos, rows 4-9`, `Datos, row 4` | | verso | `v. 1234`, `vv. 1234-1240` | igual | | canónica | `514a` | igual | | imagen | (sin localizador) | (sin localizador) | Los tiempos se escriben `h:mm:ss` a partir de una hora y `m:ss` por debajo (las horas no se reinician: tiempo de misión `109:24:48`). Un rango solo se imprime cuando los dos extremos tienen folio impreso y los folios son distintos; los corchetes marcan por separado cada extremo inferido. Un extremo sin folio impreso nunca forma parte de un rango: la cita imprime solo el folio del otro extremo (`p. 211`, nunca `pp. s. p.-211`), y `s. p.` / `n. pag.` solo cuando ninguno de los dos lo tiene. Las etiquetas salen siempre de la foliación del extremo que tiene folio impreso: una página sin numerar seguida de la hoja Ir da `fol. Ir`; un rango cuyos dos extremos impresos tienen foliaciones distintas etiqueta cada extremo (`p. xiv-fol. 1r`); las anclas `section` y `web` cuentan como páginas. El localizador se omite cuando quedaría vacío, lo que da `(Hooke, 1665)`. ### 18.2 Cita de un pasaje Una cita DEBE localizar el pasaje que reproduce, no el fragmento que da la casualidad de contenerlo. `cite_passage(fragment, quote, locale)` cita una cita textual tomada de un fragmento: 1. Se divide el fragmento en sus partes: el texto de la unidad inicial entre los dos valores de `anchor.chars` y, en un fragmento que cruza unidades, el texto de la unidad final ([§4.4](#anchors)) entre los dos valores de `anchor_end.chars` (el texto entero de la unidad cuando falta `chars`). 2. Si la cita textual está en la parte inicial, se cita el ancla de la unidad inicial, con `chars` igual a la posición de la cita en esa unidad. Si no, y está en la parte final, se cita solo el ancla de la unidad final, con sus `chars`. Si no, y abarca las dos partes, se cita el rango que va del ancla de la unidad inicial al de la unidad final, sin `chars`, con la regla de rangos de arriba (un extremo sin folio no cuenta). 3. El resultado es la cita corta del §18 y la URI de ancla del ancla o rango citado ([§5](#anchor-uri)). Los lectores y las herramientas de cita NO DEBEN citar un pasaje con la `anchor` de inicio de su fragmento cuando el pasaje no está en la unidad inicial: una cita de la segunda página de un fragmento que empieza en una lámina sin numerar cita el folio de la segunda página. ## 19. Exportaciones Las implementaciones DEBEN exportar CSL-JSON y BibTeX tal como definen los §19.1 a §19.3, y PUEDEN exportar los demás formatos del §19.4. Las exportaciones nunca inventan datos: los campos ausentes del fichero están ausentes de la exportación. Una exportación recibe uno o varios documentos, en orden. ### 19.1 Claves Cada documento exportado recibe una clave, que se usa como `id` de CSL y como clave de BibTeX: 1. Se toma el primer nombre de la lista `author` de CSL: su `family`, si no su `literal`, si no su `given`. Se pliega: se descompone con NFKD, se conservan solo las letras ASCII `A`–`Z` y `a`–`z` y se pasa a minúsculas. (`Cervantes Saavedra` → `cervantessaavedra`.) 2. Si queda vacía (sin autor, o sin ninguna letra ASCII en el nombre), se pliega del mismo modo la primera palabra, separada por espacios, de `title-short`, o de `title` cuando no hay `title-short`. (`Lazarillo de Tormes` → `lazarillo`.) 3. Si sigue vacía, se usa `anon`. 4. Se añade el primer año de `issued` en decimal (los años negativos conservan el signo), o `nd` cuando no lo hay: `cervantessaavedra1605`, `anonnd`. 5. Cuando la misma clave aparece más de una vez en una exportación, cada aparición recibe un sufijo en el orden de la exportación: `a`, `b`, … `z`, `aa`, `ab`… ### 19.2 CSL-JSON La exportación CSL-JSON es una matriz JSON con un elemento por documento: el elemento `metadata` sin su miembro `spdf`, con `id` igual a la clave. Se compara como JSON. La cita de un pasaje añade al elemento la `label` y el `locator` de CSL de un ancla y de un ancla final opcional, para que un procesador de CSL pueda imprimirla en cualquier estilo: | ancla | `label` | `locator` | |---|---|---| | `page`, foliación `page` / `leaf` / `column` | `page` / `folio` / `column` | el folio como en el [§18](#citation): `145`, `[21]`, `xiv`, `1r`, rangos `145-146`, `1r-[1v]`; sin `label` ni `locator` cuando `printed` es null | | `section` o `web` con `printed` | `page` | como en las páginas | | `section` o `web` con `paragraph` | `paragraph` | el número de párrafo | | otras `section` o `web` con ruta | `section` | el último elemento de la ruta | | `time` | `timestamp` | `1:09:20`, rangos `0:12-0:24` (como en el §18) | | `verse` | `verse` | `1234` o `1234-1240` | | `canonical` | `section` | el `ref` | | `sheet` | `line` | `4` o `4-9` | | `slide`, `image` | ninguna | ninguno (CSL no tiene localizador de diapositiva; la cita corta del §18 la imprime) | ### 19.3 BibTeX La exportación BibTeX es texto con una entrada por documento, en el orden de la exportación, separadas por una línea vacía: ```bibtex @book{cervantessaavedra1605, author = {Cervantes Saavedra, Miguel de}, title = {{El} ingenioso hidalgo don {Quijote} de la {Mancha}}, year = {1605}, publisher = {Juan de la Cuesta}, address = {Madrid}, language = {es} } ``` - **Tipo de entrada** según el `type` de CSL: `book` → `book`; `article-journal`, `article-magazine`, `article-newspaper` → `article`; `chapter` → `incollection`; `paper-conference` → `inproceedings`; `thesis` → `phdthesis`; `report` → `techreport`; cualquier otro → `misc`. - **Campos**, en este orden, cada uno solo cuando su origen está presente y no vacío: `author` (`author` de CSL), `editor` (`editor`), `title`, `year` (primer año de `issued`), `journal` en las entradas `article` o, si no, `booktitle` (`container-title`), `publisher`, `address` (`publisher-place`), `series` (`collection-title`), `volume`, `number` (`issue`), `pages` (`page`), `edition`, `doi` (`DOI`), `isbn` (`ISBN`), `url` (`URL`), `language`, `note`. - **Valores**: se escriben `{…}` en UTF-8. En todos los valores, `\` pasa a ser `\textbackslash{}`, `{` pasa a ser `\{` y `}` pasa a ser `\}`; no se escapa nada más. - **Nombres**: un nombre `literal` se escribe entre llaves, `{National Aeronautics and Space Administration}`; si no, el apellido (precedido de la `non-dropping-particle` y un espacio, si la hay) y el nombre de pila (`given`) se escriben `Apellido, Nombre`, o entre llaves cuando solo existe uno de los dos. Los nombres se unen con ` and `. - **Mayúsculas**: en `title` y en `journal`/`booktitle`, toda palabra separada por espacios que contiene una letra mayúscula (categoría general Unicode Lu) se encierra entre llaves, después de escaparla, para que los estilos no la pasen a minúsculas: `{El} ingenioso hidalgo don {Quijote}`. - **Comparación**: dos exportaciones son iguales cuando, después de quitar los espacios iniciales y finales de cada línea y de eliminar las líneas vacías, sus líneas son idénticas. ### 19.4 Otros formatos - **ALTO** [ALTO] (PUEDE): ALTO 4, un `Page` por cada unidad de página, con `PHYSICAL_IMG_NR` = `physical` y `PRINTED_IMG_NR` = `printed` solo cuando `printed` no es null y su `source` no es `inferred` (ALTO registra números impresos, y un folio deducido no está impreso); un `TextBlock` por párrafo y un `TextLine` por línea; coordenadas solo cuando el productor las tiene (desde una extensión), nunca inventadas. - **TEI** [TEI] (PUEDE, mínima): `teiHeader` a partir de los metadatos (`titleStmt`, `publicationStmt` con los derechos, `sourceDesc` con los campos de CSL), y un `body` con un `` antes de cada unidad de página, cuyo `n` es el folio tal como se cita en el §18 sin su etiqueta (`n="ii"`, `n="[iv]"`, `n="1r"`; sin `n` en las páginas sin numerar) y cuyo `facs` es la imagen de la unidad, si la hay; `

` para los párrafos, ``/`` para el verso, `` para los turnos de palabra y `` para las notas. - **IIIF Presentation 3** [IIIF] (PUEDE): un `Manifest` con un `Canvas` por unidad, en orden de `ord`; la `label` de un lienzo de página es `{"none": [n]}`, con `n` como el `n` de TEI, y los lienzos de las páginas sin numerar no llevan `label`; la imagen de la unidad es la anotación de pintado (*painting*) y el texto, una anotación `supplementing`; el audio y el vídeo son un único lienzo temporal con `duration` y un `Range` por unidad o sección; las secciones pasan a ser `structures`; las anclas con región pasan a ser destinos `#xywh=percent:`. - **Web Annotation** (PUEDE): citas y resultados de búsqueda como anotaciones con los selectores del [§17.1](#annotations). La batería de conformidad comprueba, en los documentos paginados, la secuencia de páginas de estas exportaciones: los pares `PHYSICAL_IMG_NR`/`PRINTED_IMG_NR` de ALTO, el `n` de cada `pb` de TEI y la `label` de cada lienzo de página de IIIF, en orden. ## 20. Formatos legados ### 20.1 SPDF 4.0 y 4.1 Los ficheros de Scholaris 4.x DEBEN ser legibles para todos los lectores. Son bases de datos SQLite, normalmente envueltas en gzip, con identificadores en español. Detección, tras descomprimir: una tabla `spdf` (`clave`, `valor`) cuya fila `spdf_version` empieza por `4.`, o un `user_version` 400 o 410 junto con una tabla `documentos`. `application_id` es 0. El esquema se reproduce literalmente en [`schema/spdf-4.1.sql`](schema/spdf-4.1.sql) y [`schema/spdf-4.0.sql`](schema/spdf-4.0.sql) (a la 4.0 le faltan `unidades.palabras` y `fragmentos.texto_busqueda`). Los ficheros legados contienen los disparadores `fragmentos_ai`, `fragmentos_ad` y `fragmentos_au`, que [§2.4](#container) tolera. Los lectores presentan los ficheros legados a través de la **vista 5.0**: - **Tablas**: `spdf` → `spdf_meta` (`clave` → `key`, `valor` → `value`), `documentos` → `documents`, `unidades` → `units`, `secciones` → `sections`, `fragmentos` → `fragments`, `figuras` → `figures`, `espacios` → `spaces`, `vectores` → `vectors`, `blobs` → `blobs`, `procedencia` → `provenance`; sin extensiones. - **Columnas**: `tipo` → `kind`, `metadatos` → `metadata`, `huella` → `source_sha256`, `original` → `source_ref`, `unidades` → `unit_count`, `duracion` → `duration`, `creado` → `created`, `actualizado` → `updated`, `titulo` → `title`, `autores` → `authors`, `anio` → `year`, `idioma` → `language`; `orden` → `ord`, `ancla` → `anchor`, `texto` → `text`, `notas` → `notes`, `cabecera` → `header`, `pie` → `footer`, `imagen` → `image`, `miniatura` → `thumbnail`, `lector` → `reader`, `confianza` → `confidence`, `impresa` → `printed`, `palabras` → `words`; `padre` → `parent`, `nivel` → `level`, `unidad_desde` → `unit_from`, `unidad_hasta` → `unit_to`, `resumen` → `summary`; `unidad` → `unit`, `contexto` → `context`, `seccion` → `section`, `ancla_fin` → `anchor_end`, `texto_busqueda` → `search_text`; en las figuras, `pie` → `caption`, `descripcion` → `description`; `proveedor` → `provider`, `modelo` → `model`, `normalizado` → `normalized`, `modalidades` → `modalities` (`texto` → `text`, `imagen` → `image`); `objetivo` → `target` (`fragmento` → `fragment`, `unidad` → `unit`, `figura` → `figure`), `espacio` → `space`, `valores` → `data`; `clave` → `key`, `datos` → `data`; `fase` → `stage`, `detalle` → `detail`, `cuando` → `at`. - **Tipos de documento**: `pdf`, `pdf_escaneado` → `scanned_pdf`, `fotos` → `photos`, `imagen` → `image`, `audio`, `video`, `documento` → `document`, `epub`, `presentacion` → `slides`, `hoja` → `sheet`, `web`. - **Anclas**: `tipo` → `type` (`pagina` → `page`, `tiempo` → `time`, `seccion` → `section`, `diapositiva` → `slide`, `hoja` → `sheet`, `web`, `imagen` → `image`), `fisica` → `physical`, `impresa` → `printed`, `romana` → `roman`, `origen` → `source` (`leido` → `read`, `deducido` → `inferred`, `epub`, `ninguno` → `none`), `confianza` → `confidence`, `hablante` → `speaker`, `ruta` → `path`, `parrafo` → `paragraph`, `n`, `hoja` → `sheet`, `filaDesde` → `row_from`, `filaHasta` → `row_to`, `consultada` → `accessed`, `region`. Los miembros desconocidos se conservan tal cual. - **Metadatos** (`MetadatosDocumento` → CSL-JSON): `titulo` → `title`, o `"titulo: subtitulo"` con `title-short` = `titulo` y `spdf.subtitle` = `subtitulo`; `tituloOriginal` → `original-title`; `autores`, `editores`, `traductores`, `entrevistadores` (`{nombre, apellidos, orcid}`) → `author`, `editor`, `translator`, `interviewer` (`{family: apellidos, given: nombre}`, omitiendo las partes vacías; el ORCID va a `spdf.orcid` con la clave `"apellidos, nombre"`); `fecha` → `issued` con todas sus partes de fecha cuando su año es igual a `anio` o no hay `anio`, y, si no, `anio` → `issued`; `anioOriginal` → `original-date`; `editorial` → `publisher`; `lugar` → `publisher-place`; `revista` o, en su defecto, `contenedor` → `container-title`; `coleccion` → `collection-title`; `volumen` → `volume`; `numero` → `issue`; `paginas` → `page`; `edicion` → `edition`; `doi` → `DOI`; `isbn` → `ISBN`; `url` → `URL`; `idioma` → `language`; `resumen` → `abstract`; `idiomaOriginal` → `spdf.original_language`; `sinFecha` `{desde, hasta, fundamento}` → `spdf.undated` `{from, to, basis}`; `procedencia` → `spdf.provenance`, con los nombres de campo convertidos como arriba y `fuente` → `source` (`lectura` → `reading`, `usuario` → `user`, `colofon` → `colophon`, `impresores` → `printers`, los demás sin cambios), `confianza` → `confidence`. `tipoCSL` → `type`; sin él, el tipo es `article-journal` cuando está presente `revista` y, si no, se decide según el tipo de documento: `audio` y `presentacion` → `speech`, `video` → `motion_picture`, `web` → `webpage`, `hoja` → `dataset`, `imagen` y `fotos` → `graphic`, cualquier otro → `book`. Se omiten las cadenas vacías, los nulos y los arrays vacíos. - **Otras reglas**: las claves de `spdf_meta` `creado` → `created` y `generador` → `generator`, las demás sin cambios; se descartan `documentos.estado` y `documentos.bibliotecas`; `rights` es null; los espacios reciben `dtype` `f32`, `truncated_from` y `task_prefixes` null y `created` a partir de `creado`; `provenance.model` es null; se calculan los hashes de los blobs; las unidades se renumeran `ord` = 1, 2, 3… por orden de (`orden`, `id`), porque la 4.x numera las unidades desde 0; `fragments.ord` conserva `orden`. Referencias en `original`, `imagen` y `miniatura`: una cadena vacía se convierte en null (en las figuras sigue siendo una cadena vacía), un valor igual a una clave de `blobs` se convierte en `blob:` y cualquier otro valor se conserva como referencia opaca. El ancla 4.x no tiene `foliation`; los folios por hojas no existían en la 4.x. ### 20.2 SPDF 3.0 y anteriores Las versiones v1 a v3 de Scholaris escribían bases de datos SQLite envueltas en gzip con las tablas `metadata` (key, value, incluida `schema_version`) y `chunks`, entre otras. Los lectores PUEDEN importarlas; importar es una conversión con pérdidas (los enlaces intermodales, las escenas y algunos vectores no tienen cabida en la 5.0) y el importador DEBERÍA informar de lo que ha descartado. La versión 3.0 está documentada con carácter histórico en el repositorio de Scholaris; esta especificación no la define. ## 21. Conformidad ### 21.1 Clases de producto - Un **lector conforme** abre los ficheros de forma segura ([§2.4](#container)), lee ficheros 5.0 y ficheros legados 4.x, produce el volcado canónico ([§12](#dump)), analiza y formatea URI de ancla ([§5](#anchor-uri)), produce citas breves ([§18](#citation)), ejecuta la búsqueda léxica de referencia ([§8.1](#search)) y exporta CSL-JSON y BibTeX ([§19](#exports)). Un **lector semántico** ejecuta además las búsquedas vectorial e híbrida de referencia. - Un **escritor conforme** produce ficheros que validan sin errores ni avisos para los perfiles que declaran y cuyo volcado canónico es igual al volcado que se dio al escritor (ida y vuelta). - Un **validador conforme** notifica exactamente los códigos de [§22](#validation) para los casos de validación de la batería. ### 21.2 Niveles Una implementación declara su clase y los perfiles que cubre, por ejemplo «lector y escritor, perfiles core y semantic». Su declaración se respalda con la batería de conformidad: supera todos los casos de los tipos que exige su clase (`dump`, `legacy_dump`, `anchor_uri`, `cite`, `cite_passage`, `search_lexical`, `validate`, `locate`, `export_csl`, `export_bibtex` para los lectores; además, `search_vector` y `search_hybrid` para los lectores semánticos; además, `roundtrip` y `quantize` para los escritores; y `export_structure` para las implementaciones que exportan ALTO, TEI o IIIF), con la versión de la batería con la que se probó. PUEDEN existir implementaciones parciales, pero NO DEBEN llamarse conformes. ### 21.3 La batería La batería (`conformance/` en el repositorio) es normativa en cuanto al comportamiento. Su protocolo está en `conformance/README.md`: formato de los casos, informe del ejecutor (runner) y convención de integración continua. Cada publicación de la batería tiene una versión y un manifiesto con el número de casos y su hash. ## 22. Validación ### 22.1 Procedimiento Un validador comprueba un fichero en este orden; un paso marcado con *fin* termina la validación: 1. Si el fichero empieza por `1F 8B`, descomprimirlo ([§2.3](#container)). 2. Si el resultado no es una base de datos SQLite: E001, *fin*. 3. Determinar la versión: `application_id` 1397769286 con `user_version` 500–599 es 5.x; la detección de legado de [§20.1](#legacy) es 4.x; cualquier otra cosa: E002, *fin*. Para 4.x, notificar W110 y comprobar solo que existen las tablas `spdf`, `documentos`, `unidades`, `fragmentos` y `fragmentos_fts` (E010 por cada una) y que no hay ningún disparador ni vista aparte de los tres disparadores tolerados (E020); *fin*. 4. Un fichero 5.x envuelto en gzip: E003 en los avisos. Una versión menor superior a 0: W105; en ese fichero, los tipos de ancla desconocidos (E041) y los dtype desconocidos (E032) se notifican en los avisos en lugar de en los errores, porque una versión menor posterior puede definirlos. 5. Disparadores, vistas y tablas virtuales ajenas: E020 por cada uno. 6. Tablas obligatorias (E010 por cada una) y columnas obligatorias (E011 por cada una). 7. Claves de `spdf_meta` (E012 por cada una). 8. `documents` contiene exactamente una fila (E013); `metadata` y `rights` son JSON válido (E050); `metadata` tiene `type` y `title` de tipo cadena (E051). 9. Extensiones obligatorias que el validador no conoce (E060). 10. `units.ord` es 1…N (E090); `unit_count` es igual a N (W102). 11. Anclas de unidades, fragmentos (inicio y fin) y figuras (E040, E041, E042); fragmentos que cruzan de una clase de `matter` a otra o una frontera de folio (W103, [§4.4](#anchors)). 12. Espacios: `dtype` (E032). Vectores: espacio conocido (E031), longitud (E030). 13. Índice FTS sincronizado: ejecutar `INSERT INTO fragments_fts(fragments_fts, rank) VALUES('integrity-check', 1)` (y lo mismo sobre `fragments_fts_trigram`) en una copia privada; un error es E070. 14. Blobs: el `sha256` guardado es igual al calculado (E080). 15. Si hasta aquí no ha habido errores y `content_sha256` está presente: recalcularlo (E081); si coincide y `signature` está presente, verificarla (E082). 16. Avisos de perfil: W100, W101. El resultado es un objeto JSON (esquema en [`json-schema/validation-result.schema.json`](json-schema/validation-result.schema.json)): ```json {"valid": false, "version": "5.0", "profile": ["core"], "errors": [{"code": "E090", "message": "units.ord is not 1..N", "where": "units"}], "warnings": []} ``` `valid` es verdadero si y solo si `errors` está vacío. `version` es null cuando se desconoce. Los mensajes son texto libre; la conformidad compara los conjuntos de códigos. ### 22.2 Códigos | código | significado | |---|---| | E001 | no es una base de datos SQLite (o gzip defectuoso) | | E002 | `application_id` o versión desconocidos | | E003 | fichero 5.0 envuelto en gzip (se notifica como aviso) | | E010 | falta una tabla obligatoria | | E011 | falta una columna obligatoria | | E012 | falta una clave obligatoria de `spdf_meta` | | E013 | `documents` no contiene exactamente una fila | | E020 | hay un disparador, una vista o una tabla virtual ajena | | E030 | longitud del vector ≠ dims × tamaño del dtype | | E031 | el vector se refiere a un espacio desconocido | | E032 | dtype desconocido | | E040 | ancla no válida (JSON incorrecto, miembro obligatorio ausente o de tipo erróneo) | | E041 | tipo de ancla desconocido | | E042 | `chars` fuera de rango | | E050 | JSON de metadatos o de derechos no válido | | E051 | metadatos sin `type` y `title` de tipo cadena | | E060 | extensión obligatoria desconocida | | E070 | índice FTS desincronizado | | E080 | el sha256 del blob no coincide | | E081 | `content_sha256` no coincide | | E082 | la firma no se verifica | | E090 | `units.ord` no es contiguo desde 1 | | W100 | perfil `semantic` sin vectores | | W101 | perfil `media` sin anclas de tiempo | | W102 | `unit_count` ≠ número de unidades | | W103 | un fragmento cruza entre unidades de distinta `matter`, o entre una página con folio impreso y otra sin él | | W105 | versión menor más reciente que la del validador | | W110 | fichero legado 4.x | Los códigos nunca se reutilizan con otro significado. Los códigos nuevos se añaden en versiones menores. ## 23. Versionado y compatibilidad La especificación se versiona como MAYOR.MENOR; las correcciones editoriales no cambian la versión. `user_version` la codifica ([§2.1](#container)). - Una versión **menor** (5.1, 5.2…) solo añade cosas OPCIONALES: tablas, columnas, claves de `spdf_meta`, miembros o tipos de ancla, miembros de metadatos, avisos o errores de validación para cosas que ya estaban prohibidas. Un lector 5.0 lee todos los ficheros 5.x e ignora lo que no conoce; PUEDE emitir un aviso (W105). Los validadores notifican los tipos de ancla y los dtype de una versión menor más reciente como avisos, no como errores ([§22.1](#validation)). Un escritor 5.x que no usa nada nuevo DEBERÍA escribir `user_version` 500. - Una versión **mayor** (6.0) puede cambiar o eliminar cosas. Los lectores DEBEN rechazar las versiones mayores que no conocen (E002) y DEBERÍAN seguir leyendo las versiones mayores anteriores (como la 5.0 lee la 4.x). - **Obsolescencia**: una funcionalidad se declara obsoleta en una versión menor, con el motivo y su sustituto, y se elimina no antes de la siguiente versión mayor y al menos 24 meses después. - **Promesa**: un fichero conforme con la 5.0 será legible para todo lector conforme de cualquier versión 5.x posterior, y sus URI de ancla seguirán resolviéndose. - La batería de conformidad y cada biblioteca tienen sus propios números de versión; el manifiesto de la batería indica qué versión de la especificación prueba. Los cambios se proponen y se deciden mediante el proceso de RFC de `spec/rfcs/` y `governance/`. ## 24. Tipo de medio e identificación de ficheros - Tipo de medio: `application/vnd.spdf+sqlite3` (registro ante la IANA en preparación; plantilla en `governance/drafts/iana-media-type.md`). El sufijo estructurado `+sqlite3` indica a las herramientas genéricas que el fichero es una base de datos SQLite 3. Parámetro OPCIONAL `version` (`"5.0"`). Codificación: binaria. Los ficheros de legado 4.x son datos gzip y no tienen un tipo registrado propio. - Identificadores de fragmento: en un recurso de tipo `application/vnd.spdf+sqlite3`, el identificador de fragmento es la regla `params` del [§5.1](#anchor-uri), con el significado que tiene en una URI de ancla para el documento de ese recurso: `https://example.org/quijote.spdf#p=5&f=1r`. - Extensión: `.spdf`. Ficheros auxiliares: `.spdfa.json` y `.spdfl.json`, servidos como `application/json` (o `application/ld+json` para las anotaciones). - Números mágicos: los bytes 0–15 son `53 51 4C 69 74 65 20 66 6F 72 6D 61 74 20 33 00` («SQLite format 3» y un NUL); los bytes 68–71 son `53 50 44 46` («SPDF»); los bytes 60–63 contienen `user_version` en big-endian (`00 00 01 F4` para la 5.0). Los ficheros legados 4.x empiezan por `1F 8B` y no pueden distinguirse de otros ficheros gzip sin descomprimirlos. - Uniform Type Identifier (plataformas de Apple): `com.joseluissaorin.spdf`, conforme a `public.data` y `public.database`, hasta que se acuerde un identificador neutral respecto al fabricante. ## 25. Internacionalización ### 25.1 Lenguas y sistemas de escritura Las etiquetas de lengua son BCP 47 [BCP 47]: `es`, `en-GB`, `la`, `grc` (griego antiguo), `lzh` (chino literario), `ar`. `documents.language` es la lengua principal; los fragmentos en otras lenguas no necesitan etiquetarse en esta versión. El texto se guarda en orden lógico, sea cual sea su dirección. ### 25.2 Texto de derecha a izquierda El árabe, el hebreo, el siríaco y las demás escrituras de derecha a izquierda se guardan en orden lógico, sin caracteres de control bidireccional salvo los presentes en la fuente. Los lectores los muestran con el algoritmo bidireccional de Unicode [UAX #9] y DEBERÍAN aislar las cadenas que aporta el usuario (`dir="auto"`). Las URI de ancla codifican ese texto en porcentaje, de modo que son neutras respecto a la dirección; cuando se muestra una forma IRI, los lectores DEBERÍAN aislarla. Los desplazamientos de `chars` cuentan puntos de código en orden lógico. ### 25.3 Textos antiguos El `text` de una unidad o de un fragmento es el texto de la fuente, nunca modernizado: la s larga (`ſ`), las alternancias `u`/`v` e `i`/`j`, las abreviaturas y las tildes se mantienen como están impresas. La capa `search_text` lleva una forma modernizada que solo usa la búsqueda. El tokenizador `unicode61` con `remove_diacritics 2` ya pliega las mayúsculas, los diacríticos latinos y la `ſ`; no pliega los acentos ni los espíritus del griego, las ligaduras como `æ` y `œ`, ni la `ß`, de modo que los productores DEBERÍAN poner en `search_text` las formas plegadas que necesiten (para el griego politónico, el texto sin diacríticos). Las citas reproducen `text`, nunca `search_text`. ### 25.4 Chino, japonés y coreano El tokenizador `unicode61` trata una secuencia de caracteres han como un único token. Los ficheros cuyo texto es mayoritariamente CJK DEBERÍAN incluir `fragments_fts_trigram`; la búsqueda de referencia encuentra entonces las subcadenas de tres o más caracteres por trigramas y las más cortas por subcadena ([§8.1](#search)). Los productores PUEDEN añadir a `search_text` una forma segmentada (palabras separadas por espacios). ### 25.5 Números y folios Los folios impresos se guardan tal como están impresos, en cualquier escritura (`"xiv"`, `"٣٤"`, `"三"`). Los lectores NO DEBEN convertirlos para citar; PUEDEN ofrecer conversiones para la navegación. ## Referencias ### Normativas - [BCP 47] Phillips, A., Davis, M., «Tags for Identifying Languages», BCP 47, RFC 5646. - [COMMONMARK] CommonMark Spec, versión 0.31.2, . - [CSL-JSON] Citation Style Language, esquema CSL-JSON, . - [MEDIA-FRAGMENTS] W3C, «Media Fragments URI 1.0 (basic)», Recomendación, 2012. - [RFC 1952] Deutsch, P., «GZIP file format specification version 4.3». - [RFC 2119] Bradner, S., «Key words for use in RFCs to Indicate Requirement Levels». - [RFC 3986] Berners-Lee, T., et al., «Uniform Resource Identifier (URI): Generic Syntax». - [RFC 3987] Duerst, M., Suignard, M., «Internationalized Resource Identifiers (IRIs)». - [RFC 5147] Wilde, E., Duerst, M., «URI Fragment Identifiers for the text/plain Media Type». - [RFC 5234] Crocker, D., Overell, P., «Augmented BNF for Syntax Specifications: ABNF». - [RFC 8032] Josefsson, S., Liusvaara, I., «Edwards-Curve Digital Signature Algorithm (EdDSA)». - [RFC 8174] Leiba, B., «Ambiguity of Uppercase vs Lowercase in RFC 2119 Key Words». - [RFC 8259] Bray, T., «The JavaScript Object Notation (JSON) Data Interchange Format». - [RFC 8785] Rundgren, A., et al., «JSON Canonicalization Scheme (JCS)». - [SQLITE-FORMAT] SQLite, «Database File Format», . - [SQLITE-FTS5] SQLite, «SQLite FTS5 Extension», . - [UAX #15] Unicode Standard Annex #15, «Unicode Normalization Forms». - [WEB-ANNOTATION] W3C, «Web Annotation Data Model», Recomendación, 2017. ### Informativas - [ALTO] Library of Congress, «ALTO: Technical Metadata for Layout and Text Objects», versión 4. - [CTS] Protocolo y esquema de URN «Canonical Text Services», . - [IIIF] IIIF Consortium, «IIIF Presentation API 3.0». - [MRL] Kusupati, A., et al., «Matryoshka Representation Learning», NeurIPS 2022. - [RFC 6838] Freed, N., Klensin, J., Hansen, T., «Media Type Specifications and Registration Procedures». - [RFC 7595] Thaler, D., et al., «Guidelines and Registration Procedures for URI Schemes». - [RRF] Cormack, G. V., Clarke, C. L. A., Büttcher, S., «Reciprocal Rank Fusion outperforms Condorcet and individual rank learning methods», SIGIR 2009. - [SPDX] SPDX License List, . - [SQLITE-SECURITY] SQLite, «Defense Against The Dark Arts», . - [TEI] TEI Consortium, «TEI P5: Guidelines for Electronic Text Encoding and Interchange». - [UAX #9] Unicode Standard Annex #9, «Unicode Bidirectional Algorithm». - [VEC2TEXT] Morris, J. X., et al., «Text Embeddings Reveal (Almost) As Much As Text», EMNLP 2023. ## Apéndice A. Cambios respecto a SPDF 4.1 - Contenedor sin comprimir con `application_id` y `user_version`; gzip solo para el legado. - Identificadores en inglés; los ficheros legados se leen a través de la vista 5.0. - Metadatos como elemento CSL-JSON con el objeto de extensión `spdf`. - Nuevos tipos de ancla `verse` y `canonical`; `foliation` (hojas y columnas); `chars` y `region` en cualquier ancla. - URI de ancla con ABNF, alineada con W3C Media Fragments y RFC 5147. - `spaces.dtype` (`f32`, `f16`, `i8`), `truncated_from`, `task_prefixes`; compatibilidad entre espacios. - Perfiles, extensiones, `rights`, hashes de los blobs, `model` en la procedencia. - Ni disparadores ni vistas en los ficheros distribuidos; procedimiento de apertura segura. - Volcado canónico, hash de integridad y firmas Ed25519. - Unidades numeradas desde 1. - Eliminados: `documentos.estado` y `documentos.bibliotecas` (la pertenencia a bibliotecas corresponde a los manifiestos de colección). --- # SPDF: documentos leídos una vez, citables siempre URL: https://spdf.joseluissaorin.com/es > SPDF es un formato de fichero abierto para documentos que ya se han leído. Cada pasaje lleva su ancla exacta (página impresa, folio, segundo, diapositiva, verso), así que una cita solo puede imprimir lo que dice la fuente. # Se lee una vez. Se cita siempre. [Leer la especificación](https://spdf.joseluissaorin.com/es/especificacion.md) · [Validar un fichero](https://spdf.joseluissaorin.com/es/validador.md) · [Abrir el lector](https://spdf.joseluissaorin.com/reader/) URI de ancla: `spdf:sha256-a156ce5ac9858e29e74bdc90c424fa1b77621bde7204d43faf18c2e726243ab0#p=41&f=17r&char=0,1424` → (Galilei, 1610, fol. 17r) Página física 41 del fichero, folio impreso 17r, caracteres 0 a 1424 de esa página, en *Sidereus nuncius*. La cita se calcula a partir del ancla guardada al leer; no se adivina nada. [Ábrelo en el validador](https://spdf.joseluissaorin.com/es/validador#url=/commons/files/galilei-sidereus-nuncius-1610.spdf). ## Por qué un formato Leer bien un documento es lento y caro: reconocer el texto, transcribir, encontrar los folios impresos, dividir en secciones, calcular vectores. SPDF guarda el resultado para que nadie tenga que hacerlo dos veces, y para que todo lo que se cite a partir de él se pueda comprobar. ### I. Anclas Cada pasaje sabe dónde está. Los fragmentos se guardan con el sitio del que salieron: página física y folio impreso (romano, deducido o por hojas), segundo y tiempos de cada palabra en audio y vídeo, diapositiva, rango de filas, verso o una referencia canónica como Stephanus 514a. Las anclas se escriben como URI portátiles. `{"type":"page","physical":29,"printed":"21"}` ### II. Procedencia Cada campo dice quién lo escribió. Cada unidad registra qué lector produjo su texto (la capa de texto de un PDF, un modelo de visión, un reconocedor de voz) y con qué confianza; la ficha registra de dónde salió cada campo (colofón, portada, catálogo). Los folios deducidos se citan entre corchetes. `reader: gemma-4-e4b · confidence: 0.97` ### III. Una lectura, muchas consultas Lo caro se hace una sola vez. El reconocimiento de texto, la transcripción, las secciones y los vectores se pagan al producir el fichero. Después responde a consultas léxicas, semánticas e híbridas sin conexión, incluso en un teléfono, con el índice de texto completo del propio SQLite y vectores de varios modelos a la vez. `fts5 unicode61 · f32 | f16 | i8 · RRF k = 10` ### IV. Portabilidad Un fichero, cualquier lenguaje, ningún servidor. Un .spdf es una base de datos SQLite 3 corriente: ni contenedor propio ni cuenta. Se puede proyectar en memoria o leer por rangos HTTP, y lo abren doce implementaciones independientes, todas probadas con la misma batería de conformidad. `un documento = un fichero` ### V. Honestidad de la cita Una cita solo puede imprimir lo que dice la fuente. Las citas cortas y la bibliografía (CSL-JSON, BibTeX) se derivan del ancla guardada y de la ficha CSL; no se generan. Los agentes tienen la misma garantía con el servidor MCP: buscan, citan el pasaje y dan el folio exacto, y no pueden inventárselo. `(Saorín Ferrer, 2026, p. [3])` ## Dentro de un .spdf Un único fichero SQLite, sin comprimir para que se pueda leer por rangos, sin disparadores ni vistas. Los lectores lo abren en solo lectura, en modo defensivo, y nunca cargan extensiones. El esquema es lo bastante pequeño para aprenderlo en una tarde. | Tabla | Qué guarda | | --- | --- | | `spdf_meta` | versión, perfil, programa que lo generó, identificador | | `documents` | una ficha CSL-JSON, con la procedencia de cada campo | | `units` | las unidades citables: páginas, tramos de tiempo, diapositivas, hojas | | `fragments` | pasajes de 150 a 300 palabras con sus anclas | | `fragments_fts` | índice FTS5, insensible a las tildes | | `sections` | el árbol de encabezados | | `figures` | figuras, láminas y fotogramas, con región y descripción | | `spaces` | espacios vectoriales, declarados como modelo@dims | | `vectors` | f32, f16 o i8 en little-endian | | `blobs` | el original y las imágenes, con su SHA-256 | | `provenance` | qué produjo qué, con qué modelo y cuándo | | `extensions` | tablas x_proveedor_nombre, obligatorias u opcionales | ### Perfiles - `core`: Texto y anclas. Basta para buscar y citar. - `semantic`: Lo anterior más vectores de uno o varios modelos. - `media`: Lo básico más audio y vídeo con tiempos por palabra. - `full`: Todo lo anterior. ## Doce implementaciones, una batería La de Rust es la de referencia y además ofrece una ABI de C. Las demás son nativas e independientes: cada una abre, valida, vuelca, busca, escribe anclas, cita y construye ficheros, y todas pasan los mismos casos de conformidad en cada commit. - **Rust** (`spdf`): CI en marcha. https://spdf.joseluissaorin.com/es/documentacion/rust.md - **TypeScript** (`spdf-format`): CI en marcha. https://spdf.joseluissaorin.com/es/documentacion/js.md - **Python** (`spdf-format`): CI en marcha. https://spdf.joseluissaorin.com/es/documentacion/python.md - **Swift** (`SPDF`): CI en marcha. https://spdf.joseluissaorin.com/es/documentacion/swift.md - **Kotlin / JVM** (`io.github.joseluissaorin:spdf`): CI en marcha. https://spdf.joseluissaorin.com/es/documentacion/kotlin.md - **Go** (`github.com/joseluissaorin/spdf/go`): CI en marcha. https://spdf.joseluissaorin.com/es/documentacion/go.md - **C# / .NET** (`Spdf.Format`): CI en rojo. https://spdf.joseluissaorin.com/es/documentacion/dotnet.md - **PHP** (`joseluissaorin/spdf`): CI en marcha. https://spdf.joseluissaorin.com/es/documentacion/php.md - **Ruby** (`spdf-format`): CI en marcha. https://spdf.joseluissaorin.com/es/documentacion/ruby.md - **R** (`spdf`): CI en marcha. https://spdf.joseluissaorin.com/es/documentacion/r.md - **Julia** (`SPDF.jl`): CI en marcha. https://spdf.joseluissaorin.com/es/documentacion/julia.md - **C** (`libspdf`): CI en marcha. https://spdf.joseluissaorin.com/es/documentacion/c.md ## Por dónde empezar ### Validar Suelta un .spdf en el validador: lo comprueba contra la especificación y te enseña lo que hay dentro, en tu navegador, sin subir nada. https://spdf.joseluissaorin.com/es/validador.md ### Leer El Lector SPDF abre, busca y cita ficheros SPDF en macOS, Windows, Linux, iOS, Android y la web, con modelos locales y sin cuenta. https://spdf.joseluissaorin.com/es/descargas.md ### Construir spdf build convierte un PDF, un escaneado, un EPUB o una grabación en un SPDF con modelos locales o con tu propia clave. O empieza por SPDF Commons, una pequeña colección de obras de dominio público. https://spdf.joseluissaorin.com/es/commons.md ### Para agentes Cada hoja de esta web tiene un gemelo en Markdown (la misma dirección terminada en .md), /llms.txt los reúne y /llms-full.txt trae la especificación entera. El servidor spdf-mcp permite a cualquier agente buscar en una carpeta de ficheros SPDF y citar con el folio exacto. https://spdf.joseluissaorin.com/es/agentes.md --- # Implementaciones URL: https://spdf.joseluissaorin.com/es/implementaciones > Doce implementaciones nativas e independientes de SPDF 5.0, todas comprobadas con la misma batería de conformidad en cada commit, y lo que dice hoy el CI de cada una. SPDF no es una biblioteca con envoltorios para otros lenguajes. Es una especificación con varias **implementaciones nativas e independientes**, escrita cada una con las costumbres de su lenguaje y comprobadas todas con los mismos casos de conformidad. La de Rust es la de referencia y además ofrece una ABI de C para quien prefiera no tratar con SQLite directamente. La tabla se rehace a partir de la integración continua del repositorio cada vez que se publica esta web. Aquí no hay nada escrito a mano: si un run está en rojo, sale en rojo. | Implementación | Paquete | Nivel | CI en main | Último run | | --- | --- | --- | --- | --- | | Rust | `spdf` | primer nivel | CI en marcha | 2026-10-07 | | TypeScript | `spdf-format` | primer nivel | CI en marcha | 2026-10-07 | | Python | `spdf-format` | primer nivel | CI en marcha | 2026-10-07 | | Swift | `SPDF` | primer nivel | CI en marcha | 2026-10-07 | | Kotlin / JVM | `io.github.joseluissaorin:spdf` | primer nivel | CI en marcha | 2026-10-07 | | Go | `github.com/joseluissaorin/spdf/go` | primer nivel | CI en marcha | 2026-10-07 | | C# / .NET | `Spdf.Format` | primer nivel | CI en rojo | 2026-10-07 | | PHP | `joseluissaorin/spdf` | segundo nivel | CI en marcha | 2026-10-07 | | Ruby | `spdf-format` | segundo nivel | CI en marcha | 2026-10-07 | | R | `spdf` | segundo nivel | CI en marcha | 2026-10-07 | | Julia | `SPDF.jl` | segundo nivel | CI en marcha | 2026-10-07 | | C | `libspdf` | segundo nivel | CI en marcha | 2026-10-07 | | Productor de referencia (spdf build) | `producer/` | — | CI en verde · 16/16 casos | 2026-10-07 | | Lector SPDF | `reader/` | — | CI en marcha | 2026-10-07 | | Batería de conformidad | `conformance/` | — | CI en marcha | 2026-10-07 | | Integraciones | `integrations/` | — | CI en marcha | 2026-10-07 | | Esta web | `site/` | — | CI en marcha | 2026-10-07 | Comprobado el 2026-10-07. JSON: https://spdf.joseluissaorin.com/status.json ## Clases de producto Lo que declara cada implementación, según las clases de producto de la [especificación](https://spdf.joseluissaorin.com/es/especificacion.md#conformance) (lector, lector semántico, escritor, validador), tal como lo registra el editor de la especificación en el README del repositorio. La columna de CI de arriba es lo que comprueban hoy las máquinas; esta tabla es lo declarado. | Implementación | Carpeta | Lector | Lector semántico | Escritor | Validador | ALTO / TEI / IIIF | Batería 0.3.0 | Batería 0.4.0 | |---|---|---|---|---|---|---|---|---| | Rust (referencia) | `rust` | sí | sí | sí | sí | sin probar | 229/229 | pendiente | | TypeScript | `js` | sí | sí | sí | sí | sin probar | 229/229 | 309/309 | | Python | `python` | sí | sí | sí | sí | sin probar | 229/229 | pendiente | | Swift | `swift` | sí | sí | sí | sí | sin probar | 229/229 | pendiente | | Kotlin / JVM | `kotlin` | sí | sí | sí | sí | sin probar | 229/229 | pendiente | | Go | `go` | sí | sí | sí | sí | sin probar | 229/229 | pendiente | | C# | `dotnet` | sí | sí | sí | sí | sin probar | 229/229 | pendiente | | PHP | `php` | sí | sí | sí | sí | sin probar | 229/229 | 309/309 | | Ruby | `ruby` | sí | sí | sí | sí | sin probar | 229/229 | 309/309 | | R | `r` | sí | sí | sí | sí | sin probar | 229/229 | 309/309 | | Julia | `julia` | sí | sí | sí | sí | sin probar | 229/229 | 309/309 | | C (Rust C ABI) | `c` | sí | sí | sí | sí | sin probar | 229/229 | 309/309 | | Productor `spdf build` | `producer` | n/a | n/a | lo que produce valida | n/a | n/a | pruebas propias 16/16 | pendiente | ## Lo que hace cada implementación Todas las implementaciones, en todos los lenguajes, hacen las mismas ocho cosas, y la batería de conformidad las comprueba todas: 1. **Abren con seguridad**: en solo lectura, con `query_only`, `trusted_schema=OFF` y el modo defensivo donde el enlace con SQLite lo permite, sin cargar nunca extensiones, rechazando los ficheros con disparadores o vistas y con límites al tamaño de los blobs y de la descompresión. 2. **Validan** un fichero e informan con los códigos de error y de aviso de la especificación ([§ Validación](https://spdf.joseluissaorin.com/es/especificacion.md#validation)). 3. **Leen** SPDF 5.0 y los ficheros heredados 4.0 y 4.1 de Scholaris, que suelen venir envueltos en gzip y usan identificadores en castellano. 4. **Vuelcan** un fichero a JSON canónico (RFC 8785), el oráculo con el que se compara a todas las demás. 5. **Buscan**: léxica (FTS5), vectorial (fuerza bruta, en f32, f16 o i8) e híbrida (fusión por rango recíproco con k = 10). 6. **Escriben y leen URI de ancla**, byte a byte, en los dos sentidos. 7. **Citan**: citas cortas de autor y año en castellano y en inglés, y bibliografía en CSL-JSON y BibTeX. 8. **Construyen**: crean un SPDF válido desde cero y reconstruyen un fichero a partir de su volcado. ## Cómo se comprueba la conformidad La batería está en `conformance/`: volcados de origen, ficheros generados, ficheros heredados, ficheros rotos a propósito y un caso en JSON por cada comprobación. Un caso tiene un `id`, un tipo (`dump`, `validate`, `search_lexical`, `search_vector`, `search_hybrid`, `anchor_uri`, `cite`, `legacy_dump`, `roundtrip`), una entrada y el resultado esperado. Cada implementación trae un ejecutor que pasa todos los casos e imprime una línea de JSON: ```json {"impl": "rust", "version": "5.0.0", "passed": ["dump-001", "…"], "failed": [], "skipped": []} ``` El CI de cada implementación falla si falla un solo caso y sube ese JSON como el artefacto `conformance-`, que es lo que cuenta la tabla de arriba. ## Niveles - **Primer nivel**: Rust, TypeScript, Python, Swift, Kotlin/JVM, Go y C#. Se publican a la vez que cada versión de la especificación. - **Segundo nivel**: PHP, Ruby, R, Julia y C. La misma batería; se publican cuando están listas. ## Productores Leer un documento es cosa del productor, no de la biblioteca. Hay dos productores independientes, y que se entiendan entre sí es condición para declarar estable el formato: - `spdf build`, el productor de referencia, en Python, con modelos locales (EmbeddingGemma 2, Gemma 4, Whisper) o con tu propia clave de API. - [Scholaris](https://scholaris.joseluissaorin.com), en TypeScript, donde nació el formato. --- # Documentación URL: https://spdf.joseluissaorin.com/es/documentacion > Cómo abrir, validar, buscar y citar ficheros SPDF en Rust, TypeScript, Python, Swift, Kotlin, Go, C#, PHP, Ruby, R, Julia y C. Todas las implementaciones ofrecen las mismas operaciones con los nombres y las costumbres de su lenguaje. Elige el tuyo: cada hoja trae la línea de instalación, un primer ejemplo y el README completo de la biblioteca (en inglés, como el código). - [Rust](https://spdf.joseluissaorin.com/es/documentacion/rust.md): `cargo add spdf`. Implementación de referencia; da además la ABI de C. - [TypeScript](https://spdf.joseluissaorin.com/es/documentacion/js.md): `npm install spdf-format`. Node (node:sqlite), Bun y el navegador (sqlite-wasm). Es la que mueve el validador de esta web. - [Python](https://spdf.joseluissaorin.com/es/documentacion/python.md): `pip install spdf-format`. Solo sqlite3 de la biblioteca estándar; se importa como spdf. - [Swift](https://spdf.joseluissaorin.com/es/documentacion/swift.md): `.package(url: "https://github.com/joseluissaorin/spdf", from: "5.0.0")`. Swift Package Manager, desde el repositorio. Plataformas de Apple y Linux. - [Kotlin / JVM](https://spdf.joseluissaorin.com/es/documentacion/kotlin.md): `implementation("io.github.joseluissaorin:spdf:5.0.0")`. Kotlin y Java sobre la JVM y Android. - [Go](https://spdf.joseluissaorin.com/es/documentacion/go.md): `go get github.com/joseluissaorin/spdf/go`. Módulo de Go dentro del monorepo. - [C# / .NET](https://spdf.joseluissaorin.com/es/documentacion/dotnet.md): `dotnet add package Spdf.Format`. .NET con Microsoft.Data.Sqlite. - [PHP](https://spdf.joseluissaorin.com/es/documentacion/php.md): `composer require joseluissaorin/spdf`. PDO SQLite. - [Ruby](https://spdf.joseluissaorin.com/es/documentacion/ruby.md): `gem install spdf-format`. Sobre la gema sqlite3. - [R](https://spdf.joseluissaorin.com/es/documentacion/r.md): `remotes::install_github("joseluissaorin/spdf", subdir = "r")`. Sobre RSQLite; los fragmentos como data frames. - [Julia](https://spdf.joseluissaorin.com/es/documentacion/julia.md): `pkg> add SPDF`. Sobre SQLite.jl. - [C](https://spdf.joseluissaorin.com/es/documentacion/c.md): `#include "spdf.h" /* link with -lspdf */`. ABI de C sobre el núcleo de Rust, para C, C++ y cualquier FFI. ## Las mismas operaciones en todas | Operación | Qué devuelve | | --- | --- | | abrir | Un acceso de solo lectura a un fichero 5.0 o heredado 4.x, abierto con seguridad | | validar | `{valid, version, profile, errors, warnings}` con los códigos de la especificación | | volcar | El JSON canónico del fichero (RFC 8785) | | buscar (léxica, vectorial, híbrida) | Resultados `{fragment_id, score, via, anchor, anchor_uri}` | | URI de ancla | Escribir y leer URI `spdf:`, byte a byte | | citar | `(Apellido, Año, localizador)` en castellano o en inglés | | exportar | CSL-JSON y BibTeX | | construir | Un SPDF nuevo y válido a partir de tus propios datos | ## Sin ninguna biblioteca Un fichero SPDF es una base de datos SQLite. Cualquier herramienta que hable SQLite lo lee; solo hay que tener cuidado con las reglas de seguridad y con las anclas. ```sh sqlite3 -readonly darwin-origin.spdf "SELECT key, value FROM spdf_meta" sqlite3 -readonly darwin-origin.spdf \ "SELECT f.id, u.printed, substr(f.text, 1, 80) FROM fragments_fts JOIN fragments f ON f.n = fragments_fts.rowid JOIN units u ON u.id = f.unit WHERE fragments_fts MATCH 'selection' LIMIT 5" ``` Los ficheros heredados 4.x de Scholaris vienen envueltos en gzip: antes, `gzip -dc viejo.spdf > viejo.sqlite`. --- # Cómo citar URL: https://spdf.joseluissaorin.com/es/citar > Cómo citar la especificación de SPDF en tu trabajo, y cómo cita SPDF los documentos que guarda, con el folio, el segundo o el verso exactos. ## Citar la especificación Si SPDF te sirve en tu investigación, cita la propia especificación, con la versión que hayas usado. Cita recomendada: > Saorín Ferrer, José Luis. 2026. *SPDF: Semantic Processed Document Format. Specification, version 5.0.* https://spdf.joseluissaorin.com/spec ```bibtex @techreport{spdf-5.0, author = {Saor{\'\i}n Ferrer, Jos{\'e} Luis}, title = {{SPDF}: Semantic Processed Document Format. Specification, version 5.0}, year = {2026}, institution = {spdf.joseluissaorin.com}, url = {https://spdf.joseluissaorin.com/spec}, note = {CC BY 4.0} } ``` ```json [ { "id": "spdf-5.0", "type": "report", "title": "SPDF: Semantic Processed Document Format. Specification, version 5.0", "author": [ { "family": "Saorín Ferrer", "given": "José Luis" } ], "issued": { "date-parts": [ [ 2026 ] ] }, "URL": "https://spdf.joseluissaorin.com/spec", "version": "5.0", "language": "en" } ] ``` ```yaml cff-version: 1.2.0 message: "Si usas SPDF, cita su especificación así." title: "SPDF: Semantic Processed Document Format. Specification" version: "5.0" date-released: 2026-10-07 authors: - family-names: "Saorín Ferrer" given-names: "José Luis" url: "https://spdf.joseluissaorin.com/spec" license: CC-BY-4.0 ``` Cada versión publicada de la especificación tendrá su DOI; mientras tanto, cita la dirección y la versión. ## Cómo cita SPDF lo que guarda La razón de ser del formato es que una cita **se calcula a partir del ancla guardada; no se genera**. Todas las implementaciones tienen la misma función `cite`, comprobada por la batería de conformidad, que recibe un ancla, la ficha CSL del documento y una lengua: | Ancla | En castellano | En inglés | | --- | --- | --- | | página, folio impreso leído en la página | `(Darwin, 1859, p. 21)` | `(Darwin, 1859, p. 21)` | | página, folio deducido de las vecinas | `(Darwin, 1859, p. [21])` | `(Darwin, 1859, p. [21])` | | página, folio en romanos | `(Woolf, 1929, p. xiv)` | `(Woolf, 1929, p. xiv)` | | foliación por hojas | `(Cervantes, 1605, fol. 1r)` | `(Cervantes, 1605, fol. 1r)` | | una página sin número impreso | `(Darwin, 1859, s. p.)` | `(Darwin, 1859, n. pag.)` | | intervalo | `(Darwin, 1859, pp. 21-22)` | `(Darwin, 1859, pp. 21-22)` | | momento de una grabación | `(Cortázar, 1977, 1:09:20)` | `(Cortázar, 1977, 1:09:20)` | | diapositiva | `(Gould, 2024, diap. 3)` | `(Gould, 2024, slide 3)` | | verso | `(Milton, 1667, vv. 234-240)` | `(Milton, 1667, vv. 234-240)` | | referencia canónica (año CSL −375) | `(Plato, 375 a. C., 514a)` | `(Plato, 375 BC, 514a)` | Dos autores se unen con *y* en castellano (con *e* ante el sonido /i/, como pide la norma) y con *and* en inglés; tres o más pasan a *et al.* Una página sin folio impreso nunca se cita con su posición en el fichero disfrazada de número de página. Las referencias bibliográficas completas se exportan en **CSL-JSON** (siempre) y en **BibTeX**, así que se les puede aplicar cualquier estilo CSL (Chicago, APA, MLA, ISO 690…) con citeproc, Zotero o Pandoc. ## URI de ancla Cada pasaje se puede señalar con una URI portátil que sobrevive a que el fichero cambie de nombre o se copie, porque nombra el documento por el SHA-256 de los bytes del original: ```text spdf:sha256-3f2a9c…#p=29&f=21&char=118,301 ``` Los parámetros siguen W3C Media Fragments y la RFC 5147 donde coinciden: `p` página física, `f` folio impreso, `t` segundos, `s` ruta de secciones, `sl` diapositiva, `v` verso, `ref` referencia canónica, `char` intervalo de caracteres, `xywh` región en porcentaje. La gramática completa está en la [especificación](https://spdf.joseluissaorin.com/es/especificacion.md#anchor-uri). ## Citar desde tus herramientas de escritura - **Pandoc**: escribe `[@spdf:sha256-3f2a9c…#p=29]` en Markdown y el [filtro de Pandoc](https://spdf.joseluissaorin.com/es/integraciones.md#pandoc) lo convierte en una cita de verdad con el folio impreso, en cualquier estilo CSL. - **Zotero**: el [complemento de Zotero](https://spdf.joseluissaorin.com/es/integraciones.md#zotero) importa la ficha CSL de un SPDF como ítem, adjunta el fichero y copia una cita con el folio. - **Agentes**: el [servidor MCP](https://spdf.joseluissaorin.com/es/integraciones.md#mcp) da a cualquier agente una herramienta `cite` que devuelve a la vez la cita, la URI de ancla y el texto citado, así que no puede citar una página que no diga lo que afirma. --- # SPDF para agentes URL: https://spdf.joseluissaorin.com/es/agentes > Cómo leen esta web y usan los ficheros SPDF los modelos de lenguaje y los agentes: gemelos en Markdown, llms.txt, el servidor MCP y las reglas para citar sin inventar. Esta web está escrita para que la lean igual las personas que las máquinas. Todo lo que aquí lee una persona, un agente lo puede pedir como texto plano. ## Leer esta web - Cada hoja tiene un **gemelo en Markdown**: la misma dirección terminada en `.md` (la portada es `/es.md`). Las hojas también responden en Markdown si se piden con `Accept: text/markdown`, y por defecto a `curl` y a `wget`. - [`/llms.txt`](https://spdf.joseluissaorin.com/llms.txt) enumera todas las hojas con una línea de descripción, en inglés y en castellano. - [`/llms-full.txt`](https://spdf.joseluissaorin.com/llms-full.txt) trae **la especificación entera** y todas las hojas de la web en un solo fichero. - [`/status.json`](https://spdf.joseluissaorin.com/status.json) da el estado del CI y los casos de conformidad de cada implementación, en JSON. - [`/sitemap.xml`](https://spdf.joseluissaorin.com/sitemap.xml) enumera todas las hojas con sus alternativas de lengua. `robots.txt` da la bienvenida a los buscadores y a los rastreadores de IA, también para entrenar. - Las hojas llevan JSON-LD de schema.org: la especificación como `TechArticle`, las implementaciones como `SoftwareSourceCode` y SPDF Commons como `Dataset`. ## Usar ficheros SPDF desde un agente El [servidor MCP](https://spdf.joseluissaorin.com/es/integraciones.md#mcp) `spdf-mcp` apunta a una carpeta de ficheros `.spdf` y da a cualquier cliente MCP (Claude, ChatGPT, Cursor, Zed, tu propio agente) estas herramientas: | Herramienta | Qué hace | | --- | --- | | `list_documents` | Los documentos de la carpeta, con título, autores, año, tipo y número de unidades | | `search` | Búsqueda léxica (o híbrida, si los ficheros llevan vectores y se da un vector de consulta) en todos los documentos, con anclas | | `read_passage` | El texto literal de un fragmento, de una unidad (página, tramo de tiempo, diapositiva) o de un intervalo, por id, folio impreso o URI de ancla | | `cite` | La cita corta con el folio o el segundo exactos, la URI de ancla y el texto citado, en castellano o en inglés | | `list_figures` | Figuras, láminas y fotogramas con pie, descripción y ancla; si se pide, la imagen | | `get_metadata` | La ficha CSL-JSON y el BibTeX de un documento | ```sh npx spdf-mcp ~/Biblioteca/SPDF # stdio npx spdf-mcp ~/Biblioteca/SPDF --http 8765 # Streamable HTTP, opcional ``` ## Reglas para citar sin inventar 1. **El texto, del fichero; la cita, del ancla.** Toma el texto de un pasaje de `read_passage` o del resultado de la búsqueda, y su cita de `cite`. No escribas nunca un número de página por tu cuenta. 2. **Folio impreso, no posición.** Una página tiene una posición física en el fichero y, casi siempre, un folio impreso. Se cita el folio impreso, y `cite` ya lo hace. Si una página no tiene folio impreso, la cita dice `s. p.` (`n. pag.` en inglés): no lo sustituyas por la posición. 3. **Los corchetes significan deducido.** `p. [21]` quiere decir que el folio se dedujo de las páginas vecinas, no que se leyó en la página. Conserva los corchetes. 4. **Guarda la URI de ancla.** Ponla junto a la afirmación (en una nota, un enlace o un comentario) para que una persona pueda abrir el pasaje exacto con cualquier lector de SPDF. 5. **El texto literal es literal.** Los fragmentos conservan la grafía de la fuente. La capa modernizada (`search_text`) existe solo para encontrarlos: nunca se cita de ella. 6. **Si el fichero no lo dice, no lo cites.** Un resultado de búsqueda es un candidato, no una prueba: lee el pasaje antes de atribuirle una afirmación. --- # Integraciones URL: https://spdf.joseluissaorin.com/es/integraciones > SPDF en las herramientas que la gente ya usa: un servidor MCP para agentes, cargadores para LlamaIndex y LangChain en Python y JavaScript, un complemento de Zotero y un filtro de Pandoc que convierte las anclas de SPDF en citas con el folio impreso. El formato vale lo que valgan los sitios a los que llega. Estas integraciones viven en la carpeta `integrations/` del repositorio, cada una con sus pruebas y su README, y todas se apoyan en las bibliotecas oficiales: ninguna vuelve a implementar el formato. - [Servidor MCP: spdf-mcp](https://spdf.joseluissaorin.com/es/integraciones/mcp.md): Cualquier agente busca en una carpeta de SPDF y cita con el folio exacto. - [Cargador de LangChain.js](https://spdf.joseluissaorin.com/es/integraciones/langchain-js.md): Pasajes como documentos de LangChain con su cita y su URI de ancla. - [Lector de LlamaIndex.TS](https://spdf.joseluissaorin.com/es/integraciones/llamaindex-js.md): Pasajes como documentos de LlamaIndex, con los vectores guardados si los quieres. - [Cargador de LangChain (Python)](https://spdf.joseluissaorin.com/es/integraciones/langchain-python.md): El mismo cargador para LangChain en Python. - [Lector de LlamaIndex (Python)](https://spdf.joseluissaorin.com/es/integraciones/llamaindex-python.md): El mismo lector para LlamaIndex en Python. - [Complemento para Zotero 7 y 8](https://spdf.joseluissaorin.com/es/integraciones/zotero.md): Importa un SPDF como ítem, adjúntalo y copia una cita con el folio. - [Filtro de Pandoc](https://spdf.joseluissaorin.com/es/integraciones/pandoc.md): Las anclas de SPDF en Markdown se convierten en citas con el folio impreso, en cualquier estilo CSL. ## Servidor MCP `spdf-mcp` es un servidor del [Model Context Protocol](https://modelcontextprotocol.io) en TypeScript sobre `spdf-format`. Apúntalo a una carpeta de ficheros `.spdf` y cualquier agente podrá listar los documentos, buscar en ellos, leer un pasaje, ver las figuras y citar con el folio exacto, sin poder inventárselo. ```sh npx spdf-mcp ~/Biblioteca/SPDF # stdio npx spdf-mcp ~/Biblioteca/SPDF --http 8765 # Streamable HTTP ``` En Claude Code: `claude mcp add spdf -- npx spdf-mcp ~/Biblioteca/SPDF`. En cualquier cliente que lea una configuración en JSON: ```json { "mcpServers": { "spdf": { "command": "npx", "args": ["spdf-mcp", "/ruta/a/la/biblioteca"] } } } ``` Herramientas: `list_documents`, `search`, `read_passage`, `cite`, `list_figures`, `get_metadata`. Las reglas que siguen están en [SPDF para agentes](https://spdf.joseluissaorin.com/es/agentes.md). ## LlamaIndex y LangChain Cargadores que convierten cada fragmento de un SPDF en un documento del framework, **con su ancla y su cita en los metadatos**, para que las respuestas con recuperación citen la página impresa y no un número de trozo. ```py from spdf_llamaindex import SpdfReader # pip install spdf-llamaindex docs = SpdfReader(locale="es").load_data("darwin-origin.spdf") docs[0].metadata["citation"] # '(Darwin, 1859, p. 21)' docs[0].metadata["anchor_uri"] # 'spdf:sha256-…#p=29&f=21' ``` ```py from spdf_langchain import SpdfLoader # pip install spdf-langchain for doc in SpdfLoader("biblioteca/", locale="es").lazy_load(): print(doc.metadata["citation"], doc.page_content[:60]) ``` ```js import { SpdfLoader } from 'spdf-langchain'; // npm install spdf-langchain const docs = await new SpdfLoader('darwin-origin.spdf').load(); ``` ```js import { SpdfReader } from 'spdf-llamaindex'; // npm install spdf-llamaindex const docs = await new SpdfReader().loadData('darwin-origin.spdf'); ``` ## Zotero Un complemento para Zotero 7 y 8 que lleva SPDF a una biblioteca de referencias: - **Importar un SPDF como ítem**: la ficha CSL-JSON que va dentro del fichero se convierte en un ítem de Zotero, con el fichero adjunto. - **Adjuntar un SPDF** a un ítem que ya existe. - **Copiar una cita con el folio**: elige una página o pega una URI de ancla y tendrás `(Darwin, 1859, p. 21)` en el portapapeles, con la URI de ancla al lado. Se instala desde el [fichero `.xpi`](https://spdf.joseluissaorin.com/zotero/spdf-zotero.xpi.md) (Herramientas → Complementos → Instalar complemento desde un archivo); después Zotero lo actualiza desde esta web. Está probado fuera de Zotero (92 pruebas sobre imitaciones de Zotero y los ficheros de prueba reales); las comprobaciones que faltan en un Zotero 7 y 8 de verdad están en [su hoja](https://spdf.joseluissaorin.com/es/integraciones/zotero.md). ## Pandoc Un filtro Lua para Pandoc que convierte las anclas de SPDF de tu Markdown en citas de verdad, en cualquier estilo CSL: ```markdown Darwin lo llama una lucha por la existencia [@spdf:sha256-3f2a9c…#p=29]. ``` ```sh pandoc ensayo.md --lua-filter spdf.lua -M spdf-library=biblioteca/ --citeproc -o ensayo.docx ``` El filtro lee los ficheros SPDF de la carpeta de la biblioteca, añade sus fichas CSL a la bibliografía, **sustituye la página física por el folio impreso** (`p=29` pasa a ser la página 21) y avisa si la página que citas no existe. Después citeproc le da forma en Chicago, APA, MLA o el estilo que elijas. Por qué Pandoc y no Calibre: la escritura académica en Markdown ya pasa por Pandoc y citeproc, y el paso que más importa para la honestidad de una cita (convertir una posición en un fichero en el folio que una persona encontrará en el papel) está justo ahí. Calibre es una biblioteca para leer, y eso ya lo cubre el [lector](https://spdf.joseluissaorin.com/es/descargas.md). --- # Descargar el Lector SPDF URL: https://spdf.joseluissaorin.com/es/descargas > El Lector SPDF abre, busca y cita ficheros SPDF en macOS, Windows, Linux, iOS, Android y la web, con modelos locales, sin conexión y sin cuenta. Gratis. El **Lector SPDF** es el lector gratuito del formato: abre un fichero, léelo página a página o segundo a segundo, búscalo por palabras o por sentido y copia una cita con el folio exacto. Tiene la misma interfaz en todas partes, hecha con Tauri 2 sobre un núcleo de Rust, y los modelos que usa para la búsqueda semántica funcionan en tu propio dispositivo. - **macOS** (Apple silicon, .dmg): En preparación - **Windows** (Windows 10 y 11, .msi): En preparación - **Linux** (AppImage y .deb): En preparación - **Android** (.apk, Android 10 o posterior): En preparación - **iOS y iPadOS** (App Store): En preparación - **Web**: https://spdf.joseluissaorin.com/reader/ ## En el navegador El [lector web](https://spdf.joseluissaorin.com/reader/) es la misma aplicación compilada para la web. Funciona **entera en tu navegador**: los ficheros se leen con SQLite en WebAssembly y se guardan en el almacenamiento privado de tu navegador, y la búsqueda semántica corre en tu tarjeta gráfica con WebGPU. No se envía nada a ningún sitio, salvo la descarga, una sola vez, del modelo de vectores desde Hugging Face y, solo si lo pides y pones tu propia clave, Gemini. ## Qué hace - Abre ficheros SPDF 5.0 y los ficheros heredados 4.x de Scholaris. - Enseña cada página junto a su texto, con el folio impreso, o reproduce la grabación con su transcripción palabra a palabra. - Búsqueda léxica, semántica e híbrida en toda tu biblioteca. - Copia citas en castellano o en inglés, CSL-JSON y BibTeX, con la URI de ancla. - Guarda tus notas fuera del fichero, como anotaciones W3C (`.spdfa.json`), así que el fichero no cambia nunca. --- # SPDF Commons URL: https://spdf.joseluissaorin.com/es/commons > Una colección pequeña y cuidada de obras de dominio público ya leídas en SPDF, en varias lenguas y de varios tipos (libros escaneados, EPUB, grabaciones de LibriVox), para descargar, probar y citar libremente. **SPDF Commons** es una pequeña colección de obras de dominio público ya leídas en SPDF: libros escaneados con sus folios impresos, EPUB con sus listas de páginas y grabaciones de LibriVox con su transcripción sincronizada palabra a palabra. Están aquí para descargarlas, abrirlas, buscar en ellas y citarlas, para probar las implementaciones con documentos reales y para enseñar lo que guarda el formato. | Obra | Lengua | Tipo | Unidades | Tamaño | Descargar | | --- | --- | --- | ---: | ---: | --- | | El garrote mas bien dado, y alcalde de Zalamea (Calderón de la Barca, Pedro, 1746) | castellano | libro escaneado | 32 | 5.7 MB | https://spdf.joseluissaorin.com/commons/files/calderon-alcalde-de-zalamea-1746.spdf | | Sidereus nuncius: Magna, longeqve admirabilia spectacula pandens, suspiciendaque proponens vnicuique, præsertim verò philosophis, atq́; astronomis, quæ à Galileo Galileo patritio Florentino Patauini Gymnasij publico mathematico perspicilli nuper à se reperti beneficio sunt obseruata in Lunæ facie, fixis innumeris, Lacteo Circulo, stellis nebulosis, apprime verò in quatuor planetis circa Iovis stellam disparibus interuallis, atque periodis, celeritate mirabili circumuolutis; quos, nemini in hanc vsque diem cognitos, nouissimè author depræhendit primus; atque Medicea Sidera nuncupandos decrevit (Galilei, Galileo, 1610) | latín | libro escaneado | 68 | 8.7 MB | https://spdf.joseluissaorin.com/commons/files/galilei-sidereus-nuncius-1610.spdf | | The Yellow Wall Paper (Gilman, Charlotte Perkins, 1901) | inglés | libro escaneado | 80 | 5.1 MB | https://spdf.joseluissaorin.com/commons/files/gilman-yellow-wall-paper-1901.spdf | | Les Fleurs du mal (Baudelaire, Charles, 1857) | francés | libro escaneado | 264 | 12.8 MB | https://spdf.joseluissaorin.com/commons/files/baudelaire-fleurs-du-mal-1857.spdf | | Die Verwandlung (Kafka, Franz, 1917) | alemán | EPUB | 73 | 1.0 MB | https://spdf.joseluissaorin.com/commons/files/kafka-die-verwandlung-1917.spdf | | Le avventure di Pinocchio: storia di un burattino (Collodi, Carlo, 1902) | italiano | EPUB | 289 | 5.3 MB | https://spdf.joseluissaorin.com/commons/files/collodi-pinocchio-1902.spdf | | Reliquias de casa velha (Machado de Assis, Joaquim Maria, 1906) | portugués | EPUB | 240 | 2.2 MB | https://spdf.joseluissaorin.com/commons/files/machado-de-assis-reliquias-de-casa-velha-1906.spdf | | L'auca del senyor Esteve (Rusiñol, Santiago, 1907) | catalán | EPUB | 265 | 5.7 MB | https://spdf.joseluissaorin.com/commons/files/rusinol-auca-del-senyor-esteve-1907.spdf | | Obras escogidas (Bécquer, Gustavo Adolfo, 1912) | castellano | EPUB | 354 | 5.9 MB | https://spdf.joseluissaorin.com/commons/files/becquer-obras-escogidas-1912.spdf | | The Tell-Tale Heart (Poe, Edgar Allan, 1843) | inglés | audiolibro | 21 | 296 KB | https://spdf.joseluissaorin.com/commons/files/poe-tell-tale-heart-1843.spdf | | El monte de las ánimas (Bécquer, Gustavo Adolfo, 1861) | castellano | audiolibro | 24 | 328 KB | https://spdf.joseluissaorin.com/commons/files/becquer-monte-de-las-animas-1861.spdf | | Menuet (Maupassant, Guy de, 1882) | francés | audiolibro | 14 | 244 KB | https://spdf.joseluissaorin.com/commons/files/maupassant-menuet-1882.spdf | El manifiesto de la colección: https://spdf.joseluissaorin.com/commons/commons.spdfl.json ### Cómo se verificó cada obra - **El garrote mas bien dado, y alcalde de Zalamea** (2026-10-07): Leído página a página con el motor de visión. Folios comprobados a ojo contra las imágenes de página en las páginas físicas 1 («Fol. 1»), 11, 16, 29 y 32 (la última, con el colofón «Año de 1746»): coinciden todos. Los 32 folios se leyeron en la página; ninguno es deducido. Los versos del ejemplo («el honor / es patrimonio del alma») se encontraron en la imagen de la p. 11. - **Sidereus nuncius: Magna, longeqve admirabilia spectacula pandens, suspiciendaque proponens vnicuique, præsertim verò philosophis, atq́; astronomis, quæ à Galileo Galileo patritio Florentino Patauini Gymnasij publico mathematico perspicilli nuper à se reperti beneficio sunt obseruata in Lunæ facie, fixis innumeris, Lacteo Circulo, stellis nebulosis, apprime verò in quatuor planetis circa Iovis stellam disparibus interuallis, atque periodis, celeritate mirabili circumuolutis; quos, nemini in hanc vsque diem cognitos, nouissimè author depræhendit primus; atque Medicea Sidera nuncupandos decrevit** (2026-10-07): Libro foliado: las hojas se numeran en el recto. Comprobado a ojo contra las imágenes de página: las páginas físicas 7, 19, 35, 41 y 63 llevan impresos los números de hoja 2, 8, 16, 17 y 28 (se citan como fols. 2r, 8r, 16r, 17r y 28r); los versos llevan el número de su hoja entre corchetes, por ejemplo [16v] en la física 36, donde sigue el texto del 16r, y [28v] en la física 64, la del FINIS. Las dos hojas sin numerar intercaladas tras la 16 (físicas 37-40: mapas estelares de Orión y de las Pléyades, nebulosas de Orión y del Pesebre) no tienen folio, igual que la hoja de portada, la encuadernación y las guardas. Ningún fragmento mezcla páginas con folio y sin él: el comienzo del 17r («De Luna, de inerrantibus Stellis») se cita fol. 17r. - **The Yellow Wall Paper** (2026-10-07): Leído página a página con el motor de visión. Folios comprobados a ojo contra las imágenes de página en las páginas físicas 13 (p. 1), 40 (p. 28) y 67 (p. 55, la última del texto): coinciden todos, y todas las páginas del texto (pp. 1-55) llevan un folio leído en la página. Los preliminares, las hojas en blanco tras el texto, la encuadernación y las fichas de la biblioteca (físicas 1-12 y 68-80) no tienen folio. La portada dice «Charlotte Perkins Stetson»; la ficha usa el nombre posterior de la autora, Gilman. Ningún fragmento mezcla páginas con folio y sin él: el comienzo del cuento («ancestral halls for the summer») se cita pp. 1-3, con su fragmento empezando en lo alto de la p. 1. - **Les Fleurs du mal** (2026-10-07): Leído página a página con el motor de visión. Folios comprobados a ojo contra las imágenes de página en las páginas físicas 14 (p. 6), 51 (p. 43), 55 (p. 47), 137 (p. 129) y 256 (p. 248): coinciden todos, con una diferencia constante de 8 entre página física y folio. Las páginas sin número impreso (la primera de cada poema, la tabla) llevan el folio deducido entre corchetes, por ejemplo [11] en «Bénédiction» y [128] en «A une dame créole». De las 102 entradas de la tabla del propio libro, 96 se encontraron en la página que da la tabla; de las otras seis, la tabla da la p. 43 para «Châtiment de l'orgueil», que empieza en la p. [44], la dedicatoria (p. 1) no tiene folio, y cuatro títulos no se leyeron como tales aunque su página es la correcta. La física 53 imprime «44» por error y se cita como [45]. Cuatro páginas (físicas 30, 53, 97 y 225) las rechazó el motor de visión por considerarlas recitación de un texto conocido y llevan en su lugar la capa de OCR antigua, marcada con confianza baja: sus folios son correctos y su texto es pobre. Los preliminares y las guardas no tienen folio. - **Die Verwandlung** (2026-10-07): Paginación de la edición de Kurt Wolff de 1917, tomada de los marcadores de página del EPUB (la lista de páginas de Project Gutenberg omite 11). Comprobadas las 71 páginas (5-75): cada unidad empieza exactamente en su marcador de página del EPUB, sin páginas perdidas ni repetidas; la p. 23 empieza en «Leibe zu spüren bekommt» y la p. 75 es la última del relato. La cabecera y la licencia de Project Gutenberg no llevan folio y van en fragmentos propios; la última frase («ihren jungen Körper dehnte») se cita pp. 74-75, el fragmento del último párrafo. - **Le avventure di Pinocchio: storia di un burattino** (2026-10-07): Paginación de la edición de Bemporad de 1902, tomada de los marcadores de página del EPUB (la lista de páginas de Project Gutenberg omite la p. 268). Comprobados los 287 marcadores de página (pp. 5-300): cada unidad empieza exactamente en su marcador de página del EPUB, sin páginas perdidas ni repetidas. La cabecera y la licencia de Project Gutenberg no llevan folio y van en fragmentos propios; ningún fragmento mezcla páginas con folio y sin él. - **Reliquias de casa velha** (2026-10-07): Paginación de la edición de Garnier de 1906, tomada de los marcadores de página del EPUB. Comprobados los 238 marcadores de página (I-III y 3-264): cada unidad empieza exactamente en su marcador de página del EPUB, sin páginas perdidas ni repetidas. La cabecera y la licencia de Project Gutenberg no llevan folio y van en fragmentos propios; ningún fragmento mezcla páginas con folio y sin él. - **L'auca del senyor Esteve** (2026-10-07): Paginación de la primera edición (Antoni López, [1907]), tomada de los marcadores de página del EPUB (la lista de páginas de Project Gutenberg omite las pp. 1, 3 y 204). Comprobados los 263 marcadores de página (pp. 1-280): cada unidad empieza exactamente en su marcador de página del EPUB, sin páginas perdidas ni repetidas. La cabecera y la licencia de Project Gutenberg no llevan folio y van en fragmentos propios; ningún fragmento mezcla páginas con folio y sin él. - **Obras escogidas** (2026-10-07): Paginación de la edición de Fernando Fé de 1912, tomada de los marcadores de página del EPUB (la lista de páginas de Project Gutenberg omite las pp. 18, 186, 209 y 250). Comprobados los 352 marcadores de página (pp. ii-xii y 1-353): cada unidad empieza exactamente en su marcador de página del EPUB, sin páginas perdidas ni repetidas. La cabecera y la licencia de Project Gutenberg no llevan folio y van en fragmentos propios; ningún fragmento mezcla páginas con folio y sin él. - **The Tell-Tale Heart** (2026-10-07): Comprobado sin reproducir el audio. En 2:15, 10:00 y 16:30 se compararon las palabras que el fichero sitúa en ese segundo con una transcripción local independiente (whisper.cpp) de la misma ventana de 30 s: son las mismas palabras, y la diferencia mediana entre los tiempos de inicio de las palabras emparejadas es como mucho de 0,4 s (los tiempos por palabra se interpolan dentro de los segmentos del transcriptor). Transcripción frente al texto de Project Gutenberg: tasa de error por palabra del 0,2 %. - **El monte de las ánimas** (2026-10-07): Comprobado sin reproducir el audio. En 2:15, 10:10 y 18:20 se compararon las palabras que el fichero sitúa en ese segundo con una transcripción local independiente (whisper.cpp) de la misma ventana de 30 s: son las mismas palabras, pero los tiempos del fichero van alrededor de 1 s por detrás (mediana de +0,9 a +1,3 s; se interpolan dentro de los segmentos del transcriptor). Transcripción frente al texto de la edición de 1912: tasa de error por palabra del 2,8 % (la lectora usó el texto de Wikisource). - **Menuet** (2026-10-07): Comprobado sin reproducir el audio. En 2:15, 6:40 y 10:40 se compararon las palabras que el fichero sitúa en ese segundo con una transcripción local independiente (whisper.cpp) de la misma ventana de 30 s: son las mismas palabras, y la diferencia mediana entre los tiempos de inicio de las palabras emparejadas es como mucho de 0,2 s (los tiempos por palabra se interpolan dentro de los segmentos del transcriptor). Transcripción frente al texto de Project Gutenberg: tasa de error por palabra del 1,4 %. ## Cómo se hicieron y se comprobaron Todos los ficheros salieron de `spdf build`, el productor de referencia. La tabla `provenance` de cada fichero registra qué modelo leyó cada página o cada segundo, con qué confianza y cuándo; abre cualquiera en el [validador](https://spdf.joseluissaorin.com/es/validador.md) para verlo. **Validar un fichero no basta**: un folio puesto en la página equivocada pasa el validador. Por eso, antes de publicar aquí una obra, sus folios impresos (o, en una grabación, sus tiempos) se comprueban a ojo contra las imágenes de las páginas o la transcripción, en varias páginas del principio, del medio y del final, y la tabla de arriba dice, para cada obra, qué páginas se comprobaron y qué se encontró. Una obra que no pasa esa comprobación se reconstruye; no se publica. Las fuentes son obras de dominio público verificables (Proyecto Gutenberg, Internet Archive, LibriVox, Wikisource). Cada fichero nombra los bytes del original por su SHA-256, así que cualquiera puede comprobar que se leyó de la fuente que dice. ## El manifiesto La colección entera se describe en un manifiesto `.spdfl.json`: una entrada por fichero, con su SHA-256, título, autores, año y dirección de descarga. Cualquier herramienta de SPDF puede usarlo para descargar o verificar el conjunto. ## Licencia Las obras son de dominio público. Los ficheros SPDF (la lectura: transcripción, anclas, secciones, vectores) se ceden al dominio público con CC0 1.0. --- # Validador e inspector URL: https://spdf.joseluissaorin.com/es/validador > Suelta un fichero .spdf para comprobarlo contra la especificación de SPDF y ver lo que hay dentro: ficha, unidades, fragmentos con anclas, figuras y espacios vectoriales. Todo se hace en tu navegador; no se sube nada. El validador funciona en el navegador, en `/es/validador`: suelta un fichero .spdf y lo comprueba con spdf-format y SQLite en WebAssembly, sin subirlo. Desde la línea de órdenes: `npx spdf-format validate fichero.spdf`. Muestras: SPDF in five pages (/muestras/spdf-in-five-pages.spdf), SPDF en cinco páginas (/muestras/spdf-en-cinco-paginas.spdf), El garrote mas bien dado, y alcalde de Zalamea (/commons/files/calderon-alcalde-de-zalamea-1746.spdf), Sidereus nuncius: Magna, longeqve admirabilia spectacula pandens, suspiciendaque proponens vnicuique, præsertim verò philosophis, atq́; astronomis, quæ à Galileo Galileo patritio Florentino Patauini Gymnasij publico mathematico perspicilli nuper à se reperti beneficio sunt obseruata in Lunæ facie, fixis innumeris, Lacteo Circulo, stellis nebulosis, apprime verò in quatuor planetis circa Iovis stellam disparibus interuallis, atque periodis, celeritate mirabili circumuolutis; quos, nemini in hanc vsque diem cognitos, nouissimè author depræhendit primus; atque Medicea Sidera nuncupandos decrevit (/commons/files/galilei-sidereus-nuncius-1610.spdf), The Yellow Wall Paper (/commons/files/gilman-yellow-wall-paper-1901.spdf), Les Fleurs du mal (/commons/files/baudelaire-fleurs-du-mal-1857.spdf). ## Qué comprueba El validador hace las mismas comprobaciones, en el mismo orden, que cualquier implementación conforme, e informa con los mismos códigos. Usa `spdf-format`, la implementación en TypeScript, con SQLite compilado a WebAssembly: **el fichero no sale de tu ordenador**. | Código | Qué significa | | --- | --- | | E001 | No es una base de datos SQLite | | E002 | `application_id` o versión desconocidos | | E003 | Un fichero 5.0 envuelto en gzip (aviso: los ficheros 5.0 se distribuyen sin comprimir) | | E010 · E011 | Falta una tabla o una columna obligatoria | | E012 | Falta una clave obligatoria de `spdf_meta` | | E013 | `documents` debe tener exactamente una fila | | E020 | Hay un disparador o una vista | | E030 · E031 · E032 | Un vector de longitud equivocada, un espacio desconocido o un tipo de dato desconocido | | E040 · E041 · E042 | Un ancla no válida, un tipo de ancla desconocido o caracteres fuera de rango | | E050 · E051 | Ficha con JSON no válido, o que no es un ítem CSL | | E060 | Una extensión obligatoria que este lector no conoce | | E070 | El índice de texto completo no coincide con los fragmentos | | E080 · E081 · E082 | No coincide un blob, la huella del contenido o la firma | | E090 | Las unidades no están numeradas seguidas desde 1 | | W100–W110 | Avisos: perfil semántico sin vectores, perfil multimedia sin tiempos, versión menor más nueva, fichero heredado… | ## Desde la línea de órdenes Todas las implementaciones validan también. Con la de TypeScript: ```sh npx spdf-format validate darwin-origin.spdf ``` --- # Gobernanza y RFC URL: https://spdf.joseluissaorin.com/es/gobernanza > Quién mantiene SPDF, cómo cambia la especificación (con RFC públicas y casos de conformidad), cómo funcionan las versiones, y las licencias y el compromiso sobre patentes. How SPDF is maintained and how it changes: who decides, the RFC process that every normative change goes through, how versions are numbered and what they promise, and the registrations that will make SPDF files recognizable by archives and operating systems. ## Documents | Document | What it covers | |---|---| | [GOVERNANCE.md](GOVERNANCE.md) | The editor, the future technical committee, decisions, conflicts of interest, appeals, code of conduct, licences and the patent commitment. | | [RFC-PROCESS.md](RFC-PROCESS.md) | When an RFC is needed, its states and steps, discussion periods, and when an RFC counts as implemented. | | [VERSIONING.md](VERSIONING.md) | Major, minor and editorial versions, `user_version`, the compatibility promise, deprecation, and the separate versions of the conformance suite and the implementations. | | [RFCs](../spec/rfcs/) | The RFCs themselves, and the [template](../spec/rfcs/0000-template.md). | | [CONTRIBUTING.md](../CONTRIBUTING.md) | How to contribute to the specification, the conformance suite and the implementations. | | [CODE_OF_CONDUCT.md](../CODE_OF_CONDUCT.md) | Contributor Covenant 2.1. | | [SECURITY.md](../SECURITY.md) | How to report a vulnerability privately. | ## In short - SPDF is edited by **José Luis Saorín Ferrer**, who created it for Scholaris. A **technical committee** takes over disputed decisions once three independent implementations, maintained by at least two organizations, pass the conformance suite. - **Nothing changes in silence.** An accepted RFC lands with at least one conformance case; it is considered implemented when two independent implementations pass it. Discussion lasts at least 14 days, followed by a 7-day final comment period. - **Versions live in the file**: `PRAGMA user_version` is major × 100 + minor × 10 (5.0 is 500). Every 5.x reader reads every 5.y file; minor versions only add things readers can ignore; deprecated features are removed only in a major version, at least 24 months later. Every 5.x reader also reads the legacy 4.0 and 4.1 files. - **Licences**: the specification and documentation under CC BY 4.0, the code under `MIT OR Apache-2.0`, and a public commitment not to assert patents against implementations. ## Las RFC - [RFC 0001: SPDF 5.0](https://spdf.joseluissaorin.com/es/gobernanza/rfcs/0001.md) (Accepted (2026-10-07)) - [RFC 0002: Conformance cases for exports and anchor resolution](https://spdf.joseluissaorin.com/es/gobernanza/rfcs/0002.md) (Accepted (2026-10-07); normative text in SPEC §5.4 and §19; cases in) ## Registrations in preparation Drafts of the requests that will register SPDF with the bodies that identify and describe file formats. None has been submitted: each one starts with a note (in Spanish) saying who it is for, through which channel, and what is still missing. | Body | Request | Draft | |---|---|---| | IANA | Media type `application/vnd.spdf+sqlite3` (RFC 6838, vendor tree) | [iana-media-type.md](drafts/iana-media-type.md) | | IANA | Provisional registration of the `spdf` URI scheme (RFC 7595) | [uri-scheme-spdf.md](drafts/uri-scheme-spdf.md) | | The National Archives (UK), PRONOM | Format record and DROID signature for SPDF 5.0 | [pronom-submission.md](drafts/pronom-submission.md) | | Library of Congress | Format description for *Sustainability of Digital Formats* | [loc-sustainability-fdd.md](drafts/loc-sustainability-fdd.md) | | SQLite and file(1) | `application_id` 0x53504446 in SQLite's `magic.txt` and in libmagic | [sqlite-magic-entry.md](drafts/sqlite-magic-entry.md) | | W3C | Charter for a Community Group on citable processed documents | [w3c-community-group-charter.md](drafts/w3c-community-group-charter.md) | ## Contact . Security reports: see [SECURITY.md](../SECURITY.md). --- # Governance URL: https://spdf.joseluissaorin.com/es/gobernanza/governance > Who decides what SPDF is, how, and how that changes as the format gains implementers and users. SPDF starts with a single editor and is designed to pass to a technical committee as soon as there are people outside… Who decides what SPDF is, how, and how that changes as the format gains implementers and users. SPDF starts with a single editor and is designed to pass to a technical committee as soon as there are people outside the original project to share it with. The key words MUST, SHOULD and MAY are to be read as described in BCP 14 (RFC 2119 and RFC 8174) when, and only when, they appear in capitals. ## Principles - **Nothing changes in silence.** Normative changes go through the [RFC process](RFC-PROCESS.md), in public, with recorded reasons. - **Conformance over authority.** What the format means is settled by the specification and its conformance cases, not by what any one implementation does, including the reference one. - **Open by licence.** The specification can be implemented by anyone, for any purpose, without asking (see [Licences and patents](#licences-and-patents)). - **Users outside software count.** Libraries, archives, scholars and readers in every language are part of the community, not only implementers. ## Phase 1: the editor SPDF was created by **José Luis Saorín Ferrer** for Scholaris and opened as a standard in October 2026. Until the technical committee exists, he is the **editor** and the **maintainer** of the specification and of this repository. The editor: - runs the RFC process: assigns numbers, opens and closes discussion periods, records decisions and their reasons; - decides on RFCs after public discussion, seeking consensus first, and writes down how every substantive objection was answered; - keeps the specification, its Spanish translation and the conformance suite consistent; - merges changes to each implementation's folder only after its maintainer's review, once each folder has a named maintainer; - represents the project in registrations (IANA, PRONOM, the Library of Congress, the SQLite and file(1) projects) and before standards bodies; - handles security reports as [`SECURITY.md`](../SECURITY.md) describes. Contact: . **Continuity.** If the editor cannot be reached for 90 days, the maintainers of the first-tier implementations that pass the conformance suite MAY convene an interim committee by the rules below, to keep the specification and the repository alive until the editor returns or a technical committee is formed. ## Phase 2: the technical committee ### When it is formed The editor calls for a technical committee within 90 days after both conditions hold: 1. **three independent implementations** (in the sense of [`RFC-PROCESS.md`](RFC-PROCESS.md#independent-implementations)) pass the whole conformance suite of the current version; and 2. they are maintained by **at least two organizations** independent of each other. The editor and the projects he leads count as one organization; a person maintaining an implementation on their own counts as their own organization. The editor MAY call for the committee earlier. ### Composition - Between 5 and 7 members. - The editor holds a seat while the role of editor exists. - At least two members are maintainers of implementations that pass the suite, and at least one member represents users of the format (a library, archive, publisher or research group) rather than an implementation. - **No organization holds more than one third of the seats** (rounded down, minimum one). If a member changes employer and the limit is exceeded, a seat is renewed early. - Terms are two years and staggered, so that about half the seats are renewed each year. ### How members are chosen - **First committee.** The editor opens a public call for nominations for 30 days (self-nominations welcome), publishes the candidates and their affiliations, and proposes a committee that meets the composition rules. It is confirmed unless a sustained objection, with reasons, is raised within 14 days; objections are resolved in public before the committee takes office. - **Later renewals.** Elections with approval voting. Voters are people with at least one merged contribution to the specification, the conformance suite or an implementation in the previous 24 months, plus the named maintainers of implementations. Ties are broken by the public random procedure of RFC 3797. - Vacancies are filled for the rest of the term by the same method, or by co-option confirmed as for the first committee when fewer than 6 months remain. ### The editor under the committee The editor keeps writing and maintaining the specification and running the RFC process, and is a member of the committee. Disputed decisions are taken by the committee. The committee may appoint additional editors, and may replace the editor by a two-thirds majority of its members. ## Making decisions 1. **Consensus first.** Decisions are sought by consensus: no sustained objection remains after discussion. Routine matters use lazy consensus: a proposal announced in public stands if nobody objects within 7 days. 2. **Voting when consensus fails.** Only when the discussion has been given a fair chance and progress requires a decision, the committee votes. A quorum is a majority of its members. Decisions take a simple majority of the members voting, except: - a new **major version** of the specification, changes to this document, [`RFC-PROCESS.md`](RFC-PROCESS.md) or [`VERSIONING.md`](VERSIONING.md), and replacing the editor need **two thirds of all members**; - the licences cannot be changed to anything less open than they are (see below). 3. **Records.** Every decision is recorded in public, in the RFC, issue or pull request it concerns, with the reasons and, for votes, who voted how. Meetings may be held, but their decisions are tentative until written down there, and anyone may object within 7 days of publication, giving technical reasons. 4. In phase 1 the editor takes decisions by the same standard: consensus sought first, reasons recorded, objections answered in writing. ## Conflicts of interest - Committee members and the editor publish their affiliations, employers and funding related to SPDF, and keep that information current. - A member MUST disclose any direct interest in a decision (for example, a proposal that favours or harms a product of their employer) and SHOULD abstain from voting on it. They may still take part in the discussion. - The editor's own products, among them Scholaris and SPDF Reader, are declared interests. While no committee exists, decisions that would give them an advantage over other implementations are explained in writing and remain open to appeal. ## Appeals - Anyone affected by a decision may appeal it in writing within **30 days** of its publication, giving the reasons and the outcome they seek. - **Phase 1:** the appeal goes to the editor, who reconsiders and answers in public within 30 days. If the appellant is not satisfied, the appeal stays on record and the first committee reviews it if the matter is still open. - **Phase 2:** the appeal goes to the whole committee, which answers within 30 days. Members who took the decision under appeal may speak, but the outcome needs a majority of the members who did not. The committee's answer is final. - Appeals do not suspend a decision unless the editor or the committee says so. - The licences allow anyone to fork the specification at any time. Forks are welcome to build on SPDF, but files and implementations that do not follow this specification must not be presented as conforming SPDF. ## Claims of conformance An implementation may describe itself as conforming to SPDF *version* for the kinds of case it claims only if it passes every case of those kinds in a suite that targets that version, and its report (`conformance.json`) is public. Partial support is welcome and should be described as such. ## Code of conduct Everyone who takes part follows the [Contributor Covenant 2.1](../CODE_OF_CONDUCT.md). - **Phase 1:** reports go to the editor at . Reports concerning the editor himself are recorded, and the reporter may ask for them to be reviewed by an independent mediator agreed with them, or by the first committee. - **Phase 2:** the committee names two of its members as conduct contacts. A person who is the subject of a report, or has a conflict of interest in it, takes no part in handling it. - Reports are kept confidential, as the code of conduct requires. ## Licences and patents - The **specification and the documentation** are licensed under [Creative Commons Attribution 4.0 International](https://creativecommons.org/licenses/by/4.0/) (CC BY 4.0). - The **code** in this repository (libraries, the reference producer, the reader, the conformance suite, the website) is licensed under the MIT licence or the Apache License 2.0, at the user's choice (`MIT OR Apache-2.0`). - Contributions are accepted under the same licence as the part of the repository they change (inbound = outbound), without a contributor licence agreement or a Developer Certificate of Origin sign-off for now (see [`CONTRIBUTING.md`](../CONTRIBUTING.md)). - **Patent commitment.** José Luis Saorín Ferrer, as author and editor, commits not to assert any patent claim that he owns or controls, now or in the future, against anyone for making, using, selling, offering, importing or distributing an implementation of the SPDF specification. Every contributor to the specification makes the same commitment, by the act of contributing, for the patent claims they own or control that would necessarily be infringed by implementing the text they contributed. The commitment is irrevocable. - No decision of the editor or of the committee may relicense the specification under terms less open than CC BY 4.0, or the code under terms less open than `MIT OR Apache-2.0`. Versions already published keep their licences in any case. ## Where the project lives - Repository: . It is private until the specification, the conformance suite and the first-tier libraries pass, and public afterwards. If a GitHub organization `spdf-format` is created, the repository moves there. - Website: . - The editor or the committee may later propose to continue the work in a standards venue (for example a W3C Community Group). Such a move is decided by RFC and keeps the licences and the patent commitment above. ## Changing this document This document changes by RFC, with the same discussion and final comment periods as the specification. Under the committee, the change needs two thirds of all members. --- # The RFC process URL: https://spdf.joseluissaorin.com/es/gobernanza/rfc-process > Every normative change to SPDF goes through a public request for comments (RFC) and lands together with the conformance cases that check it. This document says when an RFC is needed, how it moves from draft to… Every normative change to SPDF goes through a public request for comments (RFC) and lands together with the conformance cases that check it. This document says when an RFC is needed, how it moves from draft to decision, and when it counts as implemented. The key words MUST, SHOULD and MAY are to be read as described in BCP 14 (RFC 2119 and RFC 8174) when, and only when, they appear in capitals. ## The rule **An accepted RFC lands with at least one conformance case; it is considered implemented when two independent implementations pass it.** Everything below exists to apply that rule fairly. ## When an RFC is needed An RFC is REQUIRED for any change to what a valid file is, what a reader, writer or validator must do, or what a function defined by the specification returns. In particular: - the schema, the canonical dump, `user_version` or `spdf_meta` keys; - anchor types, anchor members and the anchor URI; - the reference search, citation and export functions; - validation codes and their order; - integrity and signatures; - the legacy mapping; - the sidecar formats (`.spdfa.json`, `.spdfl.json`); - profiles, and the rules for extensions; - this process, [`VERSIONING.md`](VERSIONING.md) and [`GOVERNANCE.md`](GOVERNANCE.md). An RFC is NOT needed for: - **editorial changes**: typos, clearer wording, examples, diagrams, translations, as long as no implementation would have to change. Open a pull request labelled `editorial`. If anyone shows that an "editorial" change alters behaviour, it becomes an RFC; - **fixes to a wrong conformance case** that contradicts the specification: fixed in place and logged, as [`conformance/README.md`](../conformance/README.md) describes; - **vendor extensions**: tables named `x__` declared in the `extensions` table need no permission. An RFC is needed only to make an extension part of the specification; - **implementation changes** that do not change behaviour defined by the specification. When in doubt, open an issue and ask. ## States | State | Meaning | |---|---| | Draft | Written, not yet open for discussion. | | Discussion | Open for public comment, at least 14 days. | | Final comment period | Last call, 7 days, with a proposed disposition. | | Accepted | Decided in favour; merged with its normative text and at least one conformance case. | | Implemented | Two independent implementations pass all of its cases in CI. | | Rejected | Decided against, with reasons. | | Postponed | Good idea, wrong time; may be reopened. | | Withdrawn | Abandoned by its authors. | | Superseded | Replaced by a later RFC, which is named. | ## Steps 1. **Before writing (optional).** Open an issue to test the idea. It saves everyone time when the answer is "this already exists" or "this belongs in an extension". 2. **Draft.** Copy [`spec/rfcs/0000-template.md`](../spec/rfcs/0000-template.md) to `spec/rfcs/0000-short-title.md` and open a pull request. Fill in every section; the Conformance cases section may start as a sketch, but it MUST be complete before the final comment period. 3. **Discussion, at least 14 days.** The editor assigns the next free number, renames the file, sets the state to Discussion and announces it. Anyone may comment. Authors revise the text in the same pull request; each substantive revision is summarized in a comment so that late readers can follow. 4. **Final comment period, 7 days.** When the discussion has settled, the editor (or the technical committee, once it exists) announces a proposed disposition: accept, reject or postpone. A new substantive objection during this period returns the RFC to Discussion; the 7 days start again when it is resolved. 5. **Decision.** The editor, or the technical committee once it exists, decides by the rules in [`GOVERNANCE.md`](GOVERNANCE.md), and records in the RFC the date, the outcome and the reasons, including how each substantive objection was answered. 6. **Merge.** An accepted RFC is merged together with: - the normative text in `spec/SPEC.md` and its Spanish translation `spec/SPEC.es.md`; - any schema change in `spec/schema/`; - **at least one conformance case** in `conformance/`, with its line in `conformance/CHANGELOG.md` and the suite version bumped as [`VERSIONING.md`](VERSIONING.md) says. An RFC without a conformance case is not merged. 7. **Implemented.** When two independent implementations pass every case of the RFC in CI (their `conformance-` artifacts show it), the editor sets the state to Implemented and records which implementations and versions. A version of the specification is published as final only with Implemented RFCs. Accepted RFCs that are not yet implemented may appear in working drafts, marked as such. ## Independent implementations Two implementations are independent when neither wraps, binds or mechanically translates the other's code, and each was written from the specification. Bindings over the Rust core, including the C ABI and anything built on it, count together with the Rust implementation. Implementations may share the conformance suite and the reference oracle; that is what they are for. Independence of authorship is not required for this rule; it matters for forming the technical committee (see [`GOVERNANCE.md`](GOVERNANCE.md)). ## Where discussion happens - On the RFC's pull request, once the repository is public. Decisions reached in calls or meetings are tentative until they are written in the pull request. - Until the repository is public, the editor publishes RFCs in Discussion on the website and collects comments sent to ; comments are recorded in the pull request with their author's permission. - Contributions in English or Spanish are welcome. The normative text is written in English, with a faithful Spanish translation. ## Shortened periods The editor (or the committee) MAY shorten the discussion and final comment periods only to fix a security problem, or a defect that makes conforming implementations incompatible with each other. The reason is recorded in the RFC, and the change is open to an RFC that revisits it afterwards with the normal periods. ## Numbering and files - RFCs are numbered with four digits, in the order discussion opens. Numbers are never reused; a rejected or withdrawn RFC keeps its number and file. - `0000-template.md` is the template, not an RFC. - A Spanish version MAY sit next to the English one as `NNNN-short-title.es.md`; the English text prevails if they differ. - [RFC 0001](../spec/rfcs/0001-spdf-5.0.md), which defines SPDF 5.0, was accepted by the editor before this process existed, without its discussion periods. Every later RFC follows this document. --- # Versioning URL: https://spdf.joseluissaorin.com/es/gobernanza/versioning > How the SPDF specification is numbered, how the number is written inside files, what each kind of version may change, and how the conformance suite and the libraries are versioned on their own. The normative rules… How the SPDF specification is numbered, how the number is written inside files, what each kind of version may change, and how the conformance suite and the libraries are versioned on their own. The normative rules are in [SPEC §23](../spec/SPEC.md#versioning); this document repeats them and adds the policy around them. The key words MUST, SHOULD and MAY are to be read as described in BCP 14 (RFC 2119 and RFC 8174) when, and only when, they appear in capitals. ## Specification versions The specification is numbered **MAJOR.MINOR**. Editorial corrections do not change that number; they are published as editorial releases with a third number. | Kind | Example | May change | Needs | |---|---|---|---| | Major | 5.0 → 6.0 | anything, including breaking changes | RFCs | | Minor | 5.0 → 5.1 | only OPTIONAL additions (see the compatibility promise) | RFCs | | Editorial | 5.0.0 → 5.0.1 | wording, examples, translations, corrections that change no behaviour | no RFC | An editorial release never changes what a valid file is or what an implementation must do. It does not change the version written in files, and conformance results do not depend on it. Releases of the specification are tagged in the repository as `spec-vMAJOR.MINOR.PATCH` (the first is `spec-v5.0.0`), separately from the tags of the libraries. Each published text states its maturity next to its number: - **Working draft**: stable enough to implement, open to change by RFC. SPDF 5.0 has been a working draft since 2026-10-07. - **Final**: every RFC it contains is Implemented (two independent implementations pass its cases) and the conformance suite covers every MUST that can be tested. A final version changes only through editorial releases; anything else waits for the next minor or major version. ## The version inside a file ```text PRAGMA user_version = MAJOR × 100 + MINOR × 10 ``` | Version | `user_version` | Bytes 60 to 63 | |---|---|---| | 4.0 (legacy) | 400 (or 0, see below) | `00 00 01 90` | | 4.1 (legacy) | 410 | `00 00 01 9A` | | 5.0 | 500 | `00 00 01 F4` | | 5.1 | 510 | `00 00 01 FE` | - The same version is written as text in `spdf_meta.spdf_version` (`"5.0"`), never with the editorial number. - Writers write the units digit as 0. Readers accept the whole range of their major (500 to 599 for 5.x), as the specification says. - A 5.x writer that uses nothing introduced after 5.0 SHOULD write 500, so that readers of 5.0 do not warn. - The encoding allows ten minor versions per major (x.0 to x.9). If a major ever needs more, the RFC that proposes the eleventh defines its encoding. - Legacy 4.x files could not always set `user_version` (some have 0); their version is in the `spdf` table, and the legacy rules of the specification ([SPEC §20](../spec/SPEC.md#legacy)) cover them. - The optional `version` parameter of the media type (`application/vnd.spdf+sqlite3; version=5.0`) is informative. The file header is authoritative. ## What readers do with versions - A reader of major M reads every minor of M, ignoring what it does not know. On a minor newer than its own it MAY warn (W105). - A reader refuses a major it does not know (E002), and SHOULD keep reading older majors. Every 5.x reader MUST read the legacy versions 4.0 and 4.1, and MAY import 3.0. ## The compatibility promise **A file that conforms to 5.0 is readable by every conforming reader of any later 5.x version, and its anchor URIs keep resolving. A 5.0 reader reads every 5.x file.** To keep that promise, a minor version: - MAY add OPTIONAL things only: tables, columns, `spdf_meta` keys, anchor members or anchor types, metadata members, and validation warnings or errors for things that were already forbidden; - MUST NOT remove or rename anything, change the meaning or the type of existing data, make something optional required, or change the canonical dump, the `content_sha256`, the anchor URIs, the citations or the reference search results of files that use nothing new; - states, in each RFC that adds something a reader of an earlier minor cannot interpret (a new anchor type, for example), what such a reader does with it: it reads the rest of the file and leaves the unknown part aside. The RFC also says how a validator of an earlier minor reports it, and a conformance case checks both. A major version may change or remove anything. Its RFC says how its readers treat files of the previous major, as 5.0 does for 4.x. ## Deprecation - A feature is deprecated by an RFC in a minor version, which gives the reason and the replacement. Deprecated features stay in the specification and readers keep reading them; writers SHOULD stop writing them. - A deprecated feature is removed no earlier than the next **major** version, and no earlier than **24 months** after the publication of the version that deprecated it. - The specification lists every deprecation with the version that deprecated it and the earliest version that may remove it. ## Things versioned on their own The specification, the conformance suite and the libraries have separate version numbers. None of them waits for the others to release. | What | Where the number lives | Scheme | |---|---|---| | Specification | `spec/SPEC.md`, `user_version`, `spdf_meta.spdf_version`; tags `spec-vX.Y.Z` | MAJOR.MINOR (+ editorial) | | Conformance suite | `conformance/manifest.json` (`suite_version`, and the `spdf_version` it targets) | semantic versioning | | Libraries | each package manifest; tags `X.Y.Z` (Go: `go/vX.Y.Z`) | one semantic version shared by all libraries | | SPDF Reader, `spdf build`, the website | their own manifests | their own | | Collection manifests | `spdf_library` member of `*.spdfl.json` (now `"1.0"`) | MAJOR.MINOR, with the same promise as the specification | | Extensions | `extensions.version` of each extension | chosen by its vendor | ### The conformance suite - **Major**: it targets a new major version of the specification, or a kind of case is retired. - **Minor**: new cases or new kinds of case. - **Patch**: a wrong case corrected in place, or a fix in the tools that changes no expectation. - While the suite is at 0.x, as now, minor releases may also correct expectations. - Every change is logged in `conformance/CHANGELOG.md`. Case ids are never renamed or reused. ### The libraries - The libraries follow a **common release train**: a semantic version tag without prefix (`0.1.0`, `0.2.0`…) marks a coordinated release of every library, with the same number in all of them. Go uses `go/vX.Y.Z` with the same number, as Go modules in a subdirectory require. - The libraries stay at 0.x until the 5.0 specification is final and there are two independent producers and at least three independent readers that pass the whole conformance suite. Then they release 1.0.0. - Each library states in its README which specification versions it supports, which product class and kinds of case it claims ([SPEC §21](../spec/SPEC.md#conformance)) and which suite version it passes; its runner reports its own version in `conformance.json`. - A library MUST NOT claim to support a specification version unless it passes every case of the kinds it claims in a suite that targets that version. - A new minor of the specification does not force a major release of the libraries: reading newer minors is already part of the compatibility promise. --- # Media type registration: application/vnd.spdf+sqlite3 URL: https://spdf.joseluissaorin.com/es/gobernanza/drafts/iana-media-type > Borrador sin enviar. Solicitud de registro del tipo de medio application/vnd.spdf+sqlite3 en el árbol de fabricante (vnd.) del registro de tipos de medio de IANA, según la plantilla de la sección 5.6 de la RFC 6838,… > **Borrador sin enviar.** Solicitud de registro del tipo de medio > `application/vnd.spdf+sqlite3` en el árbol de fabricante (`vnd.`) del registro de tipos > de medio de IANA, según la plantilla de la sección 5.6 de la RFC 6838, con el sufijo > estructurado `+sqlite3` (registrado en IANA). Va a IANA por su formulario web > (), donde la revisa un experto designado. La > RFC 6838 recomienda, sin exigirlo, enviarla antes a la lista media-types@iana.org > para que la comente la comunidad; conviene hacerlo. > > Decidido el 7-10-2026 por el orquestador del proyecto: el tipo es > `application/vnd.spdf+sqlite3` y no `application/vnd.spdf`. El sufijo permite que las > herramientas genéricas de SQLite reconozcan el fichero, y SPDF ya cumple lo que pide el > registro del sufijo (un `application_id` propio en el desplazamiento 68 y su entrada en > `magic.txt`, véase `sqlite-magic-entry.md`). `SPEC.md` §24 ya define los > identificadores de fragmento que se describen abajo. > > Falta antes de enviarla: > > 1. Una versión estable de la especificación. `https://spdf.joseluissaorin.com/spec` > ya sirve `SPEC.md` (comprobado el 7-10-2026), pero es un borrador de trabajo que > cambia; conviene enlazar una versión fechada que no cambie (la de la etiqueta > `spec-v5.0.0`) y, a ser posible, enviarla cuando la 5.0 sea final. > 2. Confirmar quién figura como responsable del cambio (José Luis o, si se crea, la > organización `spdf-format` o el comité técnico). > 3. Por verificar: el texto exacto de las consideraciones de identificadores de > fragmento del registro del sufijo `+sqlite3`, para citarlo en el apartado > correspondiente. > > Comprobado el 7-10-2026: el nombre `vnd.spdf` no figura en el registro de IANA; la > plantilla sigue la RFC 6838; el registro de `application/vnd.sqlite3` sirvió de > referencia para las consideraciones de seguridad, que resumen las §2.4, §14 y §15 de > `SPEC.md`. --- **Type name:** application **Subtype name:** vnd.spdf+sqlite3 **Required parameters:** N/A **Optional parameters:** `version`: the version of the SPDF specification the file declares, as `MAJOR.MINOR` (syntax: `1*DIGIT "." 1*DIGIT`), for example `version=5.0`. The parameter is informative only. The authoritative version is stored in the file itself (the SQLite `user_version` header field and the `spdf_meta.spdf_version` row), and receivers MUST NOT rely on the parameter instead of the file. **Encoding considerations:** binary **Structured syntax suffix:** `+sqlite3`. The considerations of the `+sqlite3` suffix registration apply; an SPDF file is a SQLite 3 database that any SQLite tool can open read-only. The SPDF-specific rules below add to them. **Security considerations:** An SPDF file is a SQLite 3 database. The security considerations of `application/vnd.sqlite3` apply in full; in addition: 1. *Active content.* SQLite schemas can contain triggers and views, which are SQL code executed by the library on behalf of whoever uses the database, and virtual tables, which call module code. SPDF files MUST NOT contain triggers, views, or virtual tables other than the two FTS5 full-text tables the specification defines. Readers open files read-only (`SQLITE_OPEN_READONLY`, `PRAGMA query_only=1`), with `PRAGMA trusted_schema=OFF` and, where the SQLite binding exposes it, `SQLITE_DBCONFIG_DEFENSIVE`; they never load SQLite extensions, and they refuse files whose schema contains any of those objects. Files of the legacy versions 4.0 and 4.1 contain exactly three full-text synchronization triggers, which readers tolerate because a read-only connection never fires them. 2. *Malformed and hostile files.* The SQLite file format is complex, and crafted files can exercise bugs in the library. Following SQLite's own advice for untrusted databases, readers use a current SQLite release, disable memory-mapped I/O, enable `cell_size_check`, may run `quick_check` first, and bound the size of every value they read (the specification recommends 512 MiB) and the nesting depth of the JSON they parse (it recommends 64). An index that disagrees with its table can make queries return data that is not stored; validators check the full-text index on a private in-memory copy. Search terms reach the full-text engine only as quoted strings and SQL is always parameterized, so user input cannot inject query syntax. 3. *Compression.* The SPDF 5.x container is not compressed. Embedded objects (page images, the original document, audio or video) are stored in their own formats and carry those formats' considerations, including their own compression. Legacy 4.x files are SQLite databases wrapped in gzip; readers decompress them only up to a configurable limit (the specification's default is 4 GiB) to defeat decompression bombs. 4. *Embedded documents and images.* A file may embed the original it was read from (for example PDF, EPUB, HTML or SVG) and page images. They are untrusted input to their own decoders and may contain scripts or other active content. Readers MUST NOT execute embedded content, MUST NOT render SVG with scripts enabled, and SHOULD render embedded documents only in a sandbox. Blob keys are opaque strings: a reader that writes blobs to disk sanitizes them (no absolute paths, no `..`, no device names). 5. *Text and links.* Unit text is light Markdown; readers render it without raw HTML and escape it before inserting it into HTML. URLs may appear in the source reference, in image references, in the metadata and in web anchors. Readers MUST NOT fetch them automatically, because fetching discloses that the file was opened and can reach internal services; they fetch only on a user action, showing the address first. 6. *Language models.* Text read from a file may contain instructions aimed at language models. Applications that pass SPDF text to a model treat it as data, not as instructions. 7. *Privacy.* Besides the visible text, a file may hold the original document, page images, word-level timings and speaker names of recordings, and provenance records naming the tools and models that processed it and when. The content itself may be personal data (for example a recorded interview). Embedding vectors can be inverted to reconstruct much of the text they were computed from, so distributing the vectors of a text is close to distributing the text, and a writer that strips the text of a restricted document strips its vectors too. User annotations are kept outside the file, so sharing a document does not share its reader's notes. Writers compact files with `VACUUM` before distribution so that deleted data does not remain in free pages. 8. *Integrity and authenticity.* The format provides no confidentiality. It provides optional integrity: `content_sha256` is a SHA-256 of a canonical serialization of the content (RFC 8785), which covers embedded objects and vectors through their hashes but not the SQLite page layout, and `signature` is an Ed25519 signature (RFC 8032) of that hash. Verifiers recompute the hash rather than trust the stored value. Provenance, confidence values and metadata are claims made by the writer; a valid signature shows that the holder of the key vouched for them, and says nothing about whether the key should be trusted (key distribution is outside the specification). Unsigned files can be altered without detection. 9. *Anchor URIs.* SPDF anchor URIs name a document by the SHA-256 of its source. Sending such a URI to a resolver reveals which document and which passage a user is reading. **Interoperability considerations:** - SPDF files are ordinary SQLite 3 databases and can be opened by any SQLite tool. Full conformance requires FTS5 (part of the SQLite amalgamation since 3.9.0) with the `unicode61 remove_diacritics 2` tokenizer (SQLite 3.27.0 or later); safe opening requires `trusted_schema` (3.31.0 or later); the optional `trigram` index for Chinese, Japanese and Korean requires 3.34.0 or later. - Distributed files use the DELETE journal mode, so they are self-contained (no `-wal` or `-journal` companion file). - The version is stored as `PRAGMA user_version` = major × 100 + minor × 10. A reader of a major version reads every minor version of it; it refuses unknown major versions. - Files of the legacy versions 4.0 and 4.1, written by the Scholaris application before this registration, share the `.spdf` extension but are gzip-compressed SQLite databases (magic number `1F 8B`) without the SPDF `application_id`. Conforming SPDF readers read them. - Metadata is a CSL-JSON item; user annotations are kept outside the file as W3C Web Annotation documents (`*.spdfa.json`, served as `application/json` or `application/ld+json`); collections are JSON manifests (`*.spdfl.json`, `application/json`). Neither sidecar uses this media type. - A public conformance suite defines the expected results of reading, validating, searching, building anchor URIs and citing, and several independent implementations are tested against it. **Published specification:** SPDF: Semantic Processed Document Format, version 5.0. (Spanish translation: ). Source: . Licensed under CC BY 4.0. **Applications that use this media type:** SPDF Reader (desktop, mobile and web); the Scholaris citation application; the reference producer `spdf build`; the SPDF libraries for Rust, TypeScript, Python, Swift, Kotlin/JVM, Go, C#, PHP, Ruby, R and Julia, and their integrations with reference managers and document tools. **Fragment identifier considerations:** As RFC 6838 section 4.11 allows, this registration defines fragment identifier semantics specific to the type, in addition to those of the `+sqlite3` suffix. A fragment identifier on a URI that resolves to an `application/vnd.spdf+sqlite3` resource addresses a location in the single document the file holds. Its syntax is the parameter list of the SPDF anchor URI: `key=value` pairs joined by `&`, in the canonical order the specification defines, values percent-encoded as UTF-8. For example: ```text https://example.org/darwin.spdf#p=29&f=21&char=118,301 ``` `p` and `pe` are physical pages; `f` and `fe` printed folios; `t` a time range in seconds; `s` a section path; `para` a paragraph; `sl` a slide; `sh` and `rows` a sheet and its rows; `v` a verse or range of verses; `ref` a canonical reference (`scheme:ref`); `char` a character range; `xywh` a region. Where they overlap with existing standards the parameters use their syntax: `t=,` and `xywh=percent:,,,` as in W3C Media Fragments URI 1.0, and `char=,` as in RFC 5147, with positions counted in Unicode code points of the NFC-normalized text of the unit. Unknown keys are ignored. The same parameter list follows `#` in SPDF anchor URIs of the form `spdf:sha256-#…`, which name the document by the SHA-256 of its source instead of by location. **Additional information:** - Deprecated alias names for this type: N/A - Magic number(s): - offset 0, 16 bytes: `53 51 4C 69 74 65 20 66 6F 72 6D 61 74 20 33 00` (ASCII "SQLite format 3" followed by a NUL byte); - offset 68, 4 bytes: `53 50 44 46` (ASCII "SPDF"; the SQLite `application_id` 1397769286 = 0x53504446, stored big-endian); - offset 60, 4 bytes: the version as a big-endian integer, `00 00 01 F4` (500) for version 5.0, and in general from 500 to 599 for the 5.x versions. - File extension(s): `.spdf` - Macintosh file type code(s): none - Uniform Type Identifier (Apple platforms): `com.joseluissaorin.spdf`, conforming to `public.data` and `public.database`. **Person & email address to contact for further information:** José Luis Saorín Ferrer **Intended usage:** COMMON **Restrictions on usage:** N/A **Author:** José Luis Saorín Ferrer **Change controller:** José Luis Saorín Ferrer , editor of the SPDF specification. **Provisional registration? (standards tree only):** N/A --- # Format description: SPDF (Semantic Processed Document Format), version 5.0 URL: https://spdf.joseluissaorin.com/es/gobernanza/drafts/loc-sustainability-fdd > Borrador sin enviar. Propuesta de ficha de formato (Format Description Document) para Sustainability of Digital Formats: Planning for Library of Congress Collections, de la Library of Congress. Sigue la estructura de… > **Borrador sin enviar.** Propuesta de ficha de formato (Format Description Document) > para *Sustainability of Digital Formats: Planning for Library of Congress > Collections*, de la Library of Congress. Sigue la estructura de las fichas > publicadas (comprobada el 7-10-2026 sobre la de SQLite 3, fdd000461, y las > explicaciones de términos del sitio). Las fichas las redacta y publica el equipo de > formatos de la Library of Congress; esto es una sugerencia para que la valoren, no un > texto que vayan a copiar tal cual. > > Vía: la página de contacto del sitio de formatos > (), que enlaza > la propia ficha de SQLite; el canal concreto (formulario o correo) está por verificar > al enviarla. > > Falta antes de enviarla: > > 1. Una versión final, o al menos fechada, de la especificación. Hoy la web la sirve > en `https://spdf.joseluissaorin.com/spec` como borrador de trabajo (comprobado el > 7-10-2026). > 2. Registro en IANA (`application/vnd.spdf+sqlite3`) y ficha en PRONOM, para poder citarlos > aquí; hoy figuran como pendientes. > 3. Más adopción. La ficha es honesta: hoy el formato lo usa sobre todo el proyecto de > su autor, y la Library of Congress puede decidir esperar. Conviene enviarla cuando > haya al menos un productor o un usuario institucional ajeno al proyecto. > > Por verificar: el nombre corto que asignarían (aquí `SPDF_5_0`), las facetas y los > identificadores de ficha de los formatos relacionados que no son SQLite, JSON ni > GeoPackage (TEI, ALTO, IIIF y W3C Web Annotation no aparecen en la lista de fichas > consultada). --- ## Format description properties | Property | Value | |---|---| | ID | to be assigned | | Short name | SPDF_5_0 (proposed) | | Content categories | text, dataset | | Format category | file-format | | Other facets | unitary, binary, structured | | Draft status | Preliminary (proposed) | ## Identification and description | | | |---|---| | Full name | SPDF (Semantic Processed Document Format), version 5.0 | | Description | SPDF is an open file format for documents that have been read once, by a PDF text layer, optical character recognition, a vision model, a speech recognizer or by hand, and stored so that every passage can be cited with its exact location in the source. A file is a SQLite 3 database (see SQLite_3) holding exactly one document: a bibliographic record as a CSL-JSON item; citable units in reading order (pages or leaves, time spans, slides, sections, spreadsheet rows) with their text in Unicode NFC and a JSON "anchor" giving the physical page and printed folio, leaf or column, verse, canonical reference, time range or slide; passages of about 150 to 300 words indexed for full-text search with the SQLite FTS5 module; and, optionally, section structure, figures with regions, embedding vectors from one or more models, provenance records, rights information, and the original file and page images as embedded binary objects. The SQLite header identifies the format: `application_id` "SPDF" at byte offset 68 and the version (500 for 5.0) at offset 60. A portable URI syntax (`spdf:sha256-…#p=29&f=21`) addresses any passage independently of the file's location. | | Production phase | Generally a middle- or final-state derivative: produced from an original (a printed book, a scan, a born-digital PDF, an EPUB, a recording, a slide deck, a web page) to make it searchable and citable. It does not replace the original, which it may embed or reference by its SHA-256 hash. | | Relationship to other formats | | | Subtype of | SQLite_3, SQLite, Version 3 (fdd000461) | | Has earlier version | SPDF 4.0 and 4.1 (no FDD): gzip-compressed SQLite databases with Spanish identifiers, internal to the Scholaris application; readers of 5.0 must read them | | May contain | JSON (fdd000381), in text columns (anchors, metadata, rights, provenance); embedded objects in their own formats (for example PNG or JPEG page images, the original PDF, EPUB, audio or video) | | Affinity to | GeoPackage_1_0 (fdd000419), another application format built on SQLite and identified by its `application_id`; TEI P5, ALTO and IIIF Presentation API 3.0, to which SPDF content can be exported; W3C Web Annotation, used by SPDF's annotation sidecar files (`.spdfa.json`); CSL-JSON, used for its metadata | ## Local use | | | |---|---| | LC experience or existing holdings | None known (for the Library of Congress to complete). | | LC preference | None (for the Library of Congress to complete). | ## Sustainability factors | Factor | Assessment | |---|---| | Disclosure | Fully documented, open specification, edited by José Luis Saorín Ferrer and published under the Creative Commons Attribution 4.0 licence, with a faithful Spanish translation. Version 5.0 is a working draft (first published 2026-10-07); changes go through a public request-for-comments process. | | Documentation | *SPDF: Semantic Processed Document Format, version 5.0* (specification, SQL schemas, JSON Schemas); a public conformance suite with expected results for reading, validating, searching, building anchor URIs and citing. | | Adoption | Very low as of October 2026. The format was created for the Scholaris citation application of the same author, whose earlier versions (4.0, 4.1) are internal to it. Libraries for version 5.0 in Rust, TypeScript, Python, Go, Swift, PHP and Ruby report passing the whole conformance suite; Kotlin, C#, R and Julia libraries, a reference producer and a free reader application are in development in the same project. No use by memory institutions or by unrelated producers is known yet. | | Licensing and patents | Specification under CC BY 4.0. Reference code under MIT or Apache License 2.0. The author and every contributor to the specification commit not to assert patents against implementations. The underlying SQLite format and library are in the public domain. No patents are known to apply. | | Transparency | High for the structure and text: any SQLite tool (for example the `sqlite3` command-line shell) can list the tables and read the text, which is UTF-8 in NFC, with JSON in text columns and light Markdown in unit text. Embedding vectors are little-endian binary arrays whose meaning depends on the model named in the file; they are derived data and can be recomputed. Embedded images and originals keep their own formats. Legacy 4.x files must be decompressed (gzip) before inspection. | | Self-documentation | Strong. The file records its format version and generator, a CSL-JSON bibliographic record with the source and confidence of each field, the SHA-256, media type and size of the original, rights (SPDX licence identifier, access level, holder), which reader produced the text of each unit and with what confidence, how each folio was obtained (read from the page or inferred), which model produced each set of vectors, and a provenance log of the processing stages. An optional content hash and Ed25519 signature allow checking the content and who vouched for it. | | Accessibility features | The format provides a text layer for scanned and image-only sources, reading order, section headings, footnotes separated from the main text, figure captions and descriptions, and time-aligned transcripts of recordings with speaker names and word timings, which support screen readers, search and captions. Tables can be written as Markdown tables in the unit text; there is no layout tagging comparable to tagged PDF. | | External dependencies | A SQLite library (widely available on all platforms) with the FTS5 module for search; no other software is required to read the text and metadata. Semantic search requires the embedding model named in the file to encode queries; the file remains fully readable without it. When the original is not embedded, it is referenced by URL or hash and must be preserved separately. | | Technical protection considerations | None. The format defines no encryption or access control; an encrypted SQLite database is not a valid SPDF file. The optional Ed25519 signature provides integrity and attribution, not protection. A `rights` object states the licence and the intended access level for information only. | ## Quality and functionality factors | Category | Factor | Assessment | |---|---|---| | Text | Normal rendering | Unit text in reading order, as light Markdown (CommonMark subset), suitable for linear reading, quotation, search and indexing. The literal spelling of the source is kept; a separate modernized-spelling layer is used only for search. | | Text | Integrity of document structure | Units (pages, leaves, time spans, slides, sections), a section hierarchy, paragraphs, verse lines, footnotes, running heads and footers kept apart from the text. | | Text | Integrity of layout and display | Not preserved by the text layer. Layout survives only in the optional page images or the embedded original. Regions of figures and passages can be recorded as fractions of the page image. | | Text | Support for mathematics, formulae, etc. | Not specified in version 5.0 beyond plain text and light Markdown. | | Text | Functionality beyond normal rendering | Exact citation: every passage carries its physical page and printed folio (roman, inferred, leaf or column), verse or canonical reference, with character offsets; a reference function produces short citations; anchor URIs; lexical, semantic and hybrid search; export to CSL-JSON, BibTeX, TEI, ALTO, IIIF and W3C Web Annotation. | | Still image | Normal rendering, clarity, color maintenance | Determined by the embedded page images or figures, which are stored in their own formats (for example PNG or JPEG) at whatever resolution the producer chose; SPDF adds regions and descriptions but no image encoding of its own. | | Sound | Normal rendering, fidelity, multiple channels | SPDF does not encode audio. It stores time-aligned transcripts (time spans with speakers and word timings in centiseconds) and may embed or reference the original recording, which keeps its own format and quality. | | Moving image | Normal rendering, clarity | As for sound: transcripts and time anchors, with the original video embedded or referenced. | | Dataset (metadata) | Normal functionality | Typed SQLite tables with a fixed schema; one document per file; JSON columns validated by JSON Schemas published with the specification. | | Dataset (metadata) | Support for software interfaces | Any SQLite interface; libraries in many programming languages tested against a common conformance suite; command-line tools and a validator. | | Dataset (metadata) | Data documentation (quality, provenance, etc.) | Per-unit reader and confidence, per-field metadata provenance, folio source (read or inferred), model and version of each vector space, processing log, and an optional signed content hash. | ## File type signifiers and format identifiers | Tag | Value | Note | |---|---|---| | Filename extension | spdf | Also used by the legacy versions 4.0 and 4.1, which are gzip-compressed. | | Internet Media Type | application/vnd.spdf+sqlite3 | Registration with IANA in preparation. | | Magic numbers | Hex: `53 51 4C 69 74 65 20 66 6F 72 6D 61 74 20 33 00` at offset 0 (ASCII "SQLite format 3" and NUL); `53 50 44 46` at offset 68 (ASCII "SPDF", the SQLite `application_id` 0x53504446); `00 00 01 F4` at offset 60 (`user_version` 500, version 5.0) | Legacy 4.x files begin with `1F 8B` (gzip) and cannot be told from other gzip files without decompressing. | | Other | Uniform Type Identifier `com.joseluissaorin.spdf`, conforming to `public.data` and `public.database` | Proposed in the specification. | | Pronom PUID | none yet | Submission in preparation. | | Wikidata Title ID | none yet | | ## Notes **General.** An SPDF file is meant to accompany its original, not to replace it: it records what was read from the original and where, so that citations stay exact even when the original is not at hand. The original's SHA-256 is always recorded, and anchor URIs use it, so the same passage can be addressed in every copy of the document. User annotations and collection lists are kept in separate JSON files (`.spdfa.json`, `.spdfl.json`), so the document file can stay unchanged. Distributed files must not contain triggers, views or foreign virtual tables, and readers open them read-only. **History.** SPDF was created by José Luis Saorín Ferrer for Scholaris, an application that inserts verified, page-exact citations into academic writing, under the name *Scholaris Processed Document Format*. Version 3.0 was a gzip-compressed SQLite database with `metadata` and `chunks` tables; versions 4.0 and 4.1 a gzip-compressed SQLite database with Spanish identifiers. All of them were internal to Scholaris. Version 5.0, published as a working draft on 7 October 2026 under the name *Semantic Processed Document Format*, is the first public version: uncompressed, with English identifiers, CSL-JSON metadata, defined text offsets, new anchor types, an anchor URI aligned with W3C Media Fragments and RFC 5147, optional signatures and a conformance suite. ## Format specifications - *SPDF: Semantic Processed Document Format, version 5.0*. José Luis Saorín Ferrer, 2026. . Source: . - SPDF conformance suite, version 0.2.0 (228 cases), in the same repository. ## Useful references - SQLite, *Database File Format*. - SQLite, *FTS5 Extension*. - Citation Style Language, CSL-JSON schema. - W3C, *Media Fragments URI 1.0 (basic)*, Recommendation, 2012. - W3C, *Web Annotation Data Model*, Recommendation, 2017. - RFC 5147, *URI Fragment Identifiers for the text/plain Media Type*. - RFC 8032, *Edwards-Curve Digital Signature Algorithm (EdDSA)*. - RFC 8785, *JSON Canonicalization Scheme (JCS)*. --- # PRONOM submission: SPDF (Semantic Processed Document Format) 5.0 URL: https://spdf.joseluissaorin.com/es/gobernanza/drafts/pronom-submission > Borrador sin enviar. Propuesta de ficha nueva para PRONOM, el registro técnico de formatos de The National Archives (Reino Unido), que alimenta las herramientas de identificación DROID y FIDO. Los campos siguen la… > **Borrador sin enviar.** Propuesta de ficha nueva para PRONOM, el registro técnico de > formatos de The National Archives (Reino Unido), que alimenta las herramientas de > identificación DROID y FIDO. Los campos siguen la plantilla oficial de envío > («PRONOM Submission template», en Word y en hoja de cálculo) del repositorio > `digital-preservation/PRONOM_Research`. > > La vía actual es GitHub: un *pull request* con la investigación en la carpeta > `Submissions` de , o una > *issue* en ese repositorio. También aceptan correo a PRONOM@nationalarchives.gov.uk, > que es la vía para muestras que no deban publicarse. Lo que se sube a ese repositorio > se publica con licencia CC0 (muestras) y la descripción, con la Open Government > Licence. > > Falta antes de enviarla: > > 1. **Tipo de medio.** PRONOM solo admite tipos de medio registrados en IANA o que > figuren en la documentación oficial del formato. Lo ideal es enviar esto después > de registrar `application/vnd.spdf+sqlite3` (ver `iana-media-type.md`); si no, > hay que citar `SPEC.md` publicada como documentación oficial. > 2. **Muestras.** Preparar ficheros de ejemplo descargables y de dominio público: los de > `conformance/files/` sirven (Cervantes, Darwin, Hooke…), pero hay que publicarlos en > una URL estable o adjuntarlos al *pull request* (quedan bajo CC0). > 3. **Probar la firma con DROID** (por ejemplo con la utilidad de desarrollo de firmas > de Ross Spencer que enlaza la propia guía de PRONOM) sobre las siete muestras 5.0, > sobre los dos ficheros heredados y sobre un SQLite cualquiera, para comprobar que > no hay falsos positivos. > 4. Una versión estable de la especificación: la web ya la sirve en > `https://spdf.joseluissaorin.com/spec` (comprobado el 7-10-2026), pero como > borrador de trabajo; conviene citar una versión fechada que no cambie. > > Por verificar: si PRONOM prefiere una ficha por versión menor (5.0, 5.1…) o una ficha > «5.x» con la firma genérica; qué clasificación asignan (aquí se propone «Database», > con «Text (Structured)» como alternativa); y si quieren fichas aparte para las > versiones heredadas 4.0 y 4.1, que no se pueden identificar sin descomprimir. > Comprobado el 7-10-2026: la ficha de SQLite 3 es fmt/729 (tipo MIME > `application/x-sqlite3`, firma `53514C69746520666F726D6174203300` en el desplazamiento > 0); GZIP es x-fmt/266; los formatos basados en SQLite con `application_id` propio > (OGC GeoPackage, fmt/1700; Audacity 3, fmt/1826) declaran «Has priority over» respecto > de fmt/729 y usan la sintaxis de huecos `{n}` de DROID. --- ## Format | Field | Value | |---|---| | File format name | SPDF (Semantic Processed Document Format) | | Version | 5.0 | | Other names | Semantic Processed Document Format; Scholaris Processed Document Format (original name, versions 3.0 to 4.1); SPDF document | | PUID | to be assigned | | Format family | none | | Format type (classification) | Database (alternatively Text (Structured)) | | Extension(s) | spdf | | MIME / media type | application/vnd.spdf+sqlite3 (registration with IANA pending; defined in the official specification) | | Byte order | Big-endian (SQLite header integers); the format's own binary vector data is little-endian | | Disclosure | Open, fully documented; specification under CC BY 4.0 | | Developer | José Luis Saorín Ferrer, editor of the SPDF specification | | Support | The SPDF project, , contact jl@joseluissaorin.com | | Released | 5.0 working draft, 2026-10-07 | ## Description SPDF (Semantic Processed Document Format) is an open file format for documents that have been read once, by optical character recognition, a text layer, a speech recogniser or by hand, and stored so that every passage can be cited with its exact location: printed page or folio, leaf or column, verse, canonical reference, second of a recording, slide or spreadsheet rows. An SPDF file is a SQLite 3 database holding exactly one document: its bibliographic record as a CSL-JSON item, its citable units (pages, time spans, slides, sections, sheets) with their text in Unicode NFC, passages indexed for full-text search with the SQLite FTS5 module, optional section structure, figures, embedding vectors for semantic search, provenance records, and optionally the original file and page images as embedded binary objects. The SQLite header identifies the format: the `application_id` field at offset 68 holds the ASCII bytes "SPDF" and the `user_version` field at offset 60 holds the version (500 for version 5.0). The format was created by José Luis Saorín Ferrer for the Scholaris citation application, where it was called "Scholaris Processed Document Format". Versions 3.0 to 4.1 were internal to Scholaris and were stored compressed with gzip (version 3.0 as a gzip-compressed SQLite database with `metadata` and `chunks` tables; 4.0 and 4.1 as a gzip-compressed SQLite database with Spanish table names); those files also use the `.spdf` extension. Version 5.0 (October 2026) is the first public version: an uncompressed SQLite database with English identifiers, readable directly by any SQLite tool. Readers of version 5.0 are required to read the legacy versions 4.0 and 4.1 as well. Files may be accompanied by two kinds of JSON sidecar files, which are not SPDF files: `.spdfa.json` (user annotations in W3C Web Annotation format) and `.spdfl.json` (collection manifests). The format is used for scholarly reading, citation and search of books, articles, manuscripts, recordings, slides and web pages. ## Internal signatures ### Signature 1: SPDF 5.0 | Field | Value | |---|---| | Signature name | SPDF 5.0 | | Position type | Absolute from BOF | | Offset | 0 | | Max offset | 0 | | Value | `53514C69746520666F726D6174203300{44}000001F4{4}53504446` | Description: BOF, offset 0: 'SQLite format 3' followed by 0x00 (`53514C69746520666F726D6174203300`, 16 bytes), then a gap of 44 bytes, then at offset 60 the SQLite `user_version` 500 as a big-endian integer (`000001F4`), then a gap of 4 bytes, then at offset 68 the SQLite `application_id` 'SPDF' (`53504446`, 0x53504446 = 1397769286). The same signature as three byte sequences, if preferred: | Position type | Offset | Max offset | Value | |---|---|---|---| | Absolute from BOF | 0 | 0 | `53514C69746520666F726D6174203300` | | Absolute from BOF | 60 | 60 | `000001F4` | | Absolute from BOF | 68 | 68 | `53504446` | Checked against the seven version 5.0 files of the SPDF conformance suite: bytes 60 to 71 are `00 00 01 F4 00 00 00 00 53 50 44 46` in all of them. ### Signature 2 (optional): SPDF, any version from 5.0 If PRONOM prefers one record for all 5.x versions, or as a fallback for future minor versions, the `application_id` alone identifies the format, in the same way as the records for OGC GeoPackage (fmt/1700) and Audacity Project File 3.x (fmt/1826): | Position type | Offset | Max offset | Value | |---|---|---|---| | Absolute from BOF | 0 | 0 | `53514C69746520666F726D6174203300{52}53504446` | Future minor versions change only the `user_version` value: 5.1 is 510 (`000001FE`), 5.2 is 520 (`00000208`), and so on (major × 100 + minor × 10). ### Legacy versions 4.0 and 4.1 Legacy files begin with the gzip header (`1F8B08`) and are identified by DROID as GZIP Format (x-fmt/266). The SPDF database is only visible after decompression, where it has `application_id` 0 and `user_version` 410 (4.1), 400 or 0 (4.0), with tables named `spdf`, `documentos`, `unidades` and `fragmentos`. A `user_version` value on its own is not a reliable signature (other SQLite applications use values such as 400), so no signature is proposed for them. If PRONOM records them, version 5.0 would be related to them as "Is subsequent version of" SPDF 4.1; their signatures cannot collide with version 5.0, so no priority relationship is needed. ## External signature | Type | Value | |---|---| | File extension | spdf | ## Relationships | Relationship | Related format | |---|---| | Has priority over | SQLite Database File Format, version 3 (fmt/729) | | Is subtype of | SQLite Database File Format, version 3 (fmt/729) | ## Documentation - SPDF: Semantic Processed Document Format, version 5.0. José Luis Saorín Ferrer, 2026. . Licensed under CC BY 4.0. - SPDF conformance suite, version 0.1.0 (220 cases). - SQLite Database File Format. ## Samples for testing the signature Sample files built from short excerpts of public-domain works are part of the SPDF conformance suite in : - `conformance/files/*.spdf`: seven version 5.0 files (Cervantes, the *Lazarillo*, Bécquer, Hooke, Darwin, the *Analects*, NASA air-to-ground transmissions); - `conformance/legacy/*.spdf`: two legacy files, versions 4.1 (Garcilaso) and 4.0 (J. F. Kennedy), gzip-compressed; - `conformance/invalid/*.spdf`: files that each break one rule of the specification, useful as near misses (for example a version 5.0 database wrapped in gzip, or a SQLite database with an unknown `application_id`). ## Vendor and support José Luis Saorín Ferrer, editor of the SPDF specification. Contact: jl@joseluissaorin.com. ## Submitted by José Luis Saorín Ferrer, SPDF project. --- # application_id 0x53504446 ("SPDF") for SQLite's magic.txt and for file(1) URL: https://spdf.joseluissaorin.com/es/gobernanza/drafts/sqlite-magic-entry > Borrador sin enviar. Dos propuestas para que las herramientas reconozcan un SPDF sin abrirlo: una línea en magic.txt, el fichero del árbol de fuentes de SQLite que lista los applicationid conocidos (la documentación… > **Borrador sin enviar.** Dos propuestas para que las herramientas reconozcan un SPDF > sin abrirlo: una línea en `magic.txt`, el fichero del árbol de fuentes de SQLite que > lista los `application_id` conocidos (la documentación del formato de fichero de > SQLite remite a él), y las reglas equivalentes en el fichero mágico de file(1) y > libmagic, que es lo que de verdad usan los sistemas operativos. > > Vías: > > 1. **SQLite.** El registro del sufijo `+sqlite3` en IANA pide añadir el valor a > `magic.txt` enviando un parche a la lista sqlite-users. Según tengo entendido esa > lista se cerró y la sustituyó el foro de SQLite (); por > verificar si basta un mensaje en el foro o si hay que escribir a los desarrolladores. > SQLite no acepta parches de código de terceros, pero `magic.txt` es un fichero de > texto que mantienen ellos: lo normal es pedir la línea y dejar que la añadan. > 2. **file(1) y libmagic.** El proyecto lo mantiene Christos Zoulas. Vías publicadas en > su README: el gestor de fallos y la lista > ; el código público está en GitHub (`file/file`), pero por > verificar si allí aceptan *pull requests* o solo las vías anteriores. Conviene > adjuntar un fichero de muestra (los de `conformance/files/` son de dominio público) y, > si lo aceptan, añadirlo también a su batería de pruebas (`file/file-tests`). > > Falta antes de enviarlas: > > - Que la especificación tenga una versión estable. La URL que citan los comentarios > de la regla (`https://spdf.joseluissaorin.com/spec`) ya responde (comprobado el > 7-10-2026), pero sirve un borrador de trabajo. > - Una URL pública para la muestra que cita el mensaje del foro. > - Decidir el nombre que se imprime: aquí «SPDF document», con el nombre completo en > `magic.txt` como pide la propuesta inicial; los demás valores del fichero usan > nombres cortos. > - El tipo de medio: libmagic suele usar tipos registrados o `x-` y puede pedir que > `application/vnd.spdf+sqlite3` esté ya en IANA (ver `iana-media-type.md`); si no, > la alternativa es dejar `application/vnd.sqlite3` y añadir solo el nombre. Por > verificar con el mantenedor. > > Probado el 7-10-2026 con file 5.41 (el de macOS) cargando las reglas con la variable > `MAGIC`: los siete ficheros 5.0 de la batería se reconocen, y un SQLite con otro > `application_id` (`invalid/E002-application-id.spdf`) no. Los diffs están hechos contra > las versiones de `magic.txt` (rama `master` del espejo de SQLite en GitHub) y de > `magic/Magdir/sql` (revisión 1.38) descargadas ese día. --- ## 1. SQLite `magic.txt` SQLite's documentation of the database header says that the `application_id` at offset 68 lets utilities such as file(1) identify application file formats, and that the list of assigned values is the `magic.txt` file in the SQLite source tree. SPDF uses `PRAGMA application_id = 1397769286` (0x53504446, ASCII "SPDF") and stores the specification version in `PRAGMA user_version` (major × 100 + minor × 10; 500 for version 5.0). ### Proposed lines ```text >68 belong =0x53504446 SPDF document (Semantic Processed Document Format) >>60 belong x \b, user_version %d - ``` The second line prints the format version, which SPDF keeps in `user_version`. Output with the current `magic.txt`: ```text quijote.spdf: SPDF document (Semantic Processed Document Format), user_version 500 - SQLite3 database ``` If a single line in the style of the other entries is preferred: ```text >68 belong =0x53504446 SPDF document - ``` which prints `SPDF document - SQLite3 database`. ### As a patch ```diff --- a/magic.txt +++ b/magic.txt @@ -30,4 +30,6 @@ >68 belong =0x45737269 Esri Spatially-Enabled Database - >68 belong =0x4d504258 MBTiles tileset - >68 belong =0x6a035744 TeXnicard card database +>68 belong =0x53504446 SPDF document (Semantic Processed Document Format) +>>60 belong x \b, user_version %d - >0 string =SQLite SQLite3 database ``` ### Message (forum post) > **Subject:** application_id 0x53504446 for SPDF, for magic.txt > > SPDF (Semantic Processed Document Format) is an open file format for documents that > have been read once (OCR, text layer or speech recognition) and stored so that every > passage can be cited with its exact page, folio, verse or time. Each file is a single > uncompressed SQLite 3 database. The specification (CC BY 4.0) is at > https://spdf.joseluissaorin.com/spec. > > SPDF files set `PRAGMA application_id = 1397769286` (0x53504446, "SPDF" in ASCII) and > keep the format version in `user_version` (500 for version 5.0). Could this value be > added to the list in magic.txt? Suggested lines: > > ``` > >68 belong =0x53504446 SPDF document (Semantic Processed Document Format) > >>60 belong x \b, user_version %d - > ``` > > A public-domain sample file is available at [URL of a sample]. Thank you for SQLite, > and for the application_id mechanism in particular. > > José Luis Saorín Ferrer, editor of the SPDF specification ## 2. file(1) and libmagic: `magic/Magdir/sql` In libmagic the SQLite 3 rules live in `magic/Magdir/sql`. Known application ids appear twice: once to set the media type and extensions, once to print the name. The generic rules that follow already print the user version, so no extra rule is needed for it. ### Patch ```diff --- a/magic/Magdir/sql +++ b/magic/Magdir/sql @@ -196,6 +196,13 @@ >>68 belong =0x4D504258 database !:mime application/vnd.sqlite3 !:ext mbtiles +# URL: https://spdf.joseluissaorin.com +# Reference: https://spdf.joseluissaorin.com/spec +# Note: SPDF, Semantic Processed Document Format, with application id 53504446h "SPDF" +# and the version as user version (500 for 5.0) +>>68 belong =0x53504446 database +!:mime application/vnd.spdf+sqlite3 +!:ext spdf >>68 default x database !:mime application/vnd.sqlite3 # no examples found with s3db sl3 suffix @@ -228,6 +235,7 @@ >>>68 belong =0x544f4952 (Riot Games patcher) >>>68 belong =0x544d5052 (Twinmotion project) >>>68 belong =0x43484557 (Chewing IME) +>>>68 belong =0x53504446 (SPDF document) # unknown application ID >>>68 default x >>>>68 belong !0 \b, application id %u ``` ### Result ```text $ file darwin.spdf darwin.spdf: SQLite 3.x database (SPDF document), user version 500 (0x1f4), last written using SQLite version … $ file --mime-type darwin.spdf darwin.spdf: application/vnd.spdf+sqlite3 $ file --extension darwin.spdf darwin.spdf: spdf ``` ### Message (bug tracker or mailing list) > **Subject:** magic: SQLite application id 0x53504446 (SPDF document) > > The attached patch to `magic/Magdir/sql` recognizes SPDF files (Semantic Processed > Document Format), SQLite 3 databases with `application_id` 0x53504446 ("SPDF"), and > gives them the media type `application/vnd.spdf+sqlite3` and the extension `spdf`. The format > keeps its version in `user_version`, which the existing rules already print. > Specification: https://spdf.joseluissaorin.com/spec. A public-domain sample is > attached; it can go into file-tests if useful. > > José Luis Saorín Ferrer, editor of the SPDF specification ## Legacy files SPDF 4.0 and 4.1 files (written by the Scholaris application before version 5.0) are SQLite databases compressed with gzip and have `application_id` 0. file(1) reports them as gzip data, and with `-z` as a SQLite 3.x database with user version 410 or 400 (or 0). No rule is proposed for them: a user version alone is not a reliable signature. --- # URI scheme registration (provisional): spdf URL: https://spdf.joseluissaorin.com/es/gobernanza/drafts/uri-scheme-spdf > Borrador sin enviar. Solicitud de registro provisional del esquema de URI spdf en el registro de esquemas de URI de IANA, con la plantilla de la sección 7.4 de la RFC 7595. Los registros provisionales solo exigen que… > **Borrador sin enviar.** Solicitud de registro provisional del esquema de URI `spdf` > en el registro de esquemas de URI de IANA, con la plantilla de la sección 7.4 de la > RFC 7595. Los registros provisionales solo exigen que el nombre no esté ya registrado > y que la solicitud esté completa; no piden una especificación estable. La RFC 7595 > recomienda anunciar la propuesta en la lista uri-review@ietf.org para recibir > comentarios antes o a la vez que se envía a IANA. > > Vía: el formulario de IANA para esquemas de URI o un correo a iana@iana.org con esta > plantilla (por verificar cuál prefiere IANA hoy). > > Falta antes de enviarla: > > 1. Comprobar que `spdf` no figura ya en el registro de esquemas de URI (por > verificar el 7-10-2026 no se pudo; consultar > ). > 2. Enlazar una versión fechada de la especificación (etiqueta `spec-v5.0.0`) en lugar > del borrador de trabajo. > 3. Confirmar el responsable del cambio, como en el registro del tipo de medio > (`iana-media-type.md`). > > Cuando la 5.0 sea final y haya implementaciones independientes en uso, se puede pedir > el paso a registro permanente, que exige revisión de experto y una especificación > estable. --- **Scheme name:** spdf **Status:** Provisional **Applications/protocols that use this scheme name:** Applications that read, cite or annotate documents in the SPDF format (Semantic Processed Document Format): the SPDF libraries (Rust, TypeScript, Python, Swift, Kotlin/JVM, Go, C#, PHP, Ruby, R, Julia), the SPDF Reader, the Scholaris citation application and annotation sidecars (`*.spdfa.json`, W3C Web Annotation), which use `spdf` URIs as annotation targets. **Contact:** José Luis Saorín Ferrer **Change controller:** José Luis Saorín Ferrer , editor of the SPDF specification. **References:** SPDF: Semantic Processed Document Format, version 5.0, section 5 ("Anchor URI"). ; source: . **Scheme syntax:** ```abnf spdf-uri = "spdf:" docref [ "#" params ] docref = hash-ref / id-ref hash-ref = "sha256-" 64lhex ; lowercase hex SHA-256 of the source document id-ref = 1*( unreserved / pct-encoded ) ``` `params` is the list of anchor parameters defined in the specification (physical page, printed folio, time range, section path, paragraph, slide, sheet rows, verse, canonical reference, character range, region). The time and region parameters use the syntax of W3C Media Fragments URI 1.0 (`t=4160,4175.5`, `xywh=percent:10,20,30,10`) and the character range that of RFC 5147 (`char=118,301`). Example: ```text spdf:sha256-27ea8a4dd0bbf0a511246fe67f81c19084aa2da75c871fe7acde688fd182cb60#p=29&pe=30&f=Ir&fe=Iv&char=729,745 ``` **Scheme semantics:** An `spdf` URI names a document, independently of where copies of it are stored, by the SHA-256 of the bytes of its original (the scanned book, the recording, the PDF), and optionally a place in it: a page, a leaf, a time span, a verse, a passage. It is a name, not a locator: there is no network protocol and no resolution service. An application resolves it against the SPDF files it holds whose `source_sha256` matches. The document reference may instead be a document identifier local to a file; such URIs are only meaningful next to that file. **Encoding considerations:** Values are UTF-8 strings percent-encoded as in RFC 3986; every byte other than an unreserved character is percent-encoded with uppercase hexadecimal digits in the canonical form. Parsers also accept lowercase hexadecimal digits and unencoded non-ASCII characters (IRI form, RFC 3987). **Interoperability considerations:** The same parameter list is the fragment identifier of `application/vnd.spdf+sqlite3` resources (`https://example.org/x.spdf#p=5&f=1r`). A canonical form, the order of parameters and round-trip rules are defined by the specification and tested by its public conformance suite. **Security considerations:** An `spdf` URI reveals which document, and which passage of it, a user is reading: applications should not send such URIs to third parties without consent. Because the URI identifies a document by a hash, it cannot be used to retrieve the document, and an application must not fetch anything from the network to resolve it. Parsers must reject malformed percent-encoding and invalid UTF-8, and must not use decoded values as file paths or as query syntax. --- # Citable Processed Documents Community Group Charter URL: https://spdf.joseluissaorin.com/es/gobernanza/drafts/w3c-community-group-charter > Borrador sin enviar. Propuesta de carta (charter) para un Community Group de W3C que se ocupe de SPDF y de su modelo de anclas. Sigue la plantilla oficial de cartas de Community Group (w3c/cg-charter, comprobada el… > **Borrador sin enviar.** Propuesta de carta (*charter*) para un Community Group de > W3C que se ocupe de SPDF y de su modelo de anclas. Sigue la plantilla oficial de > cartas de Community Group (`w3c/cg-charter`, comprobada el 7-10-2026); los apartados > de proceso, contribución, transparencia, elección de presidentes y enmiendas son su > texto, adaptado lo mínimo. > > Vía: se propone un grupo nuevo desde > con una cuenta de W3C (gratuita; no hace falta ser miembro de W3C ni pagar). Según el > proceso de los Community Groups, la propuesta está completa con un nombre que no use > otro grupo, una descripción del alcance y el apoyo de cinco personas; entonces W3C > anuncia el grupo. La carta se aprueba después, dentro del grupo. Por verificar si > W3C pide ahora redactar las cartas en el repositorio `w3c-cg/charter-proposals`. > > Falta antes de proponerlo: > > 1. **Nombre.** Aquí «Citable Processed Documents Community Group», que describe el > problema y no el producto; la alternativa es «SPDF Community Group». Comprobar que > no existe ya un grupo con ese nombre. > 2. **Cinco apoyos**, mejor de organizaciones distintas (implementadores, bibliotecas, > archivos, editores de TEI o IIIF), y un **segundo presidente** ajeno al proyecto. > 3. **Licencias.** Las contribuciones a las especificaciones del grupo quedan bajo el > CLA de W3C y los informes finales, bajo su acuerdo de especificación final (FSA). > La especificación SPDF actual es CC BY 4.0, que permite aportarla, pero conviene > confirmar con el Community Development Lead de W3C cómo conviven las dos licencias > y si el repositorio puede seguir con MIT OR Apache-2.0 para el código y la batería. > 4. Decidir por RFC, según `governance/GOVERNANCE.md`, si la especificación pasa a > desarrollarse en el grupo o el grupo solo la revisa; y que el repositorio sea > público. --- - **Status:** draft. This charter is a work in progress. To submit feedback, please use the issues of the repository where it is being developed: [repository URL]. - **This charter:** [URI] - **Previous charter:** none - **Start date:** [date the charter takes effect] - **Last modified:** 2026-10-07 ## Goals The mission of the Citable Processed Documents Community Group is to develop open, royalty-free specifications that let a document be **read once and cited forever**: a portable file format that keeps the text read from a book, scan, recording, slide deck, spreadsheet or web page together with the exact location of every passage in its source (printed page and folio, leaf, column, verse, canonical reference, time, slide), so that people, software and language-model agents can quote and cite it without approximation. The group starts from SPDF (Semantic Processed Document Format), version 5.0, an open specification with a conformance suite and implementations in several programming languages, and aims to make it a shared community specification, aligned with the W3C and other standards that already address annotation, fragments, bibliographic metadata and digital editions. ## Scope of work The group will work on: - **The SPDF file format**: the container (a single SQLite 3 database), its schema, the CSL-JSON metadata profile, text normalization and offsets, rights and provenance information, vector spaces for semantic search, integrity and signatures, profiles, extensions, and the reading of earlier versions. - **The anchor model and the anchor URI**: how a location in a document is described (pages and folios, leaves and columns, time ranges, sections and paragraphs, slides, spreadsheet rows, verses, canonical references, character ranges and regions) and how it is written as a URI, keeping the syntax of W3C Media Fragments and RFC 5147 where they overlap. - **Annotation and collection sidecars**: annotation files that use the W3C Web Annotation Data Model with an SPDF anchor selector, and collection manifests that list documents by hash. - **Reference behaviour that must be identical across implementations**: canonical serialization, validation, reference search, short citations and exports. - **A conformance test suite** for all of the above. - **Mappings** between SPDF anchors and the location models of TEI, IIIF, ALTO, CTS and W3C Web Annotation. Key use cases: verifiable citation in the humanities and social sciences; reading and citing early printed books, manuscripts and classical texts by folio, verse or canonical reference; citing oral history and recorded lectures to the second; archives and libraries publishing searchable, citable derivatives of their holdings; language-model agents that must quote a source and say exactly where. ## Out of scope - Optical character recognition, speech recognition, embedding models and ranking methods beyond the reference behaviour that makes implementations interoperable. - Specific products: readers, editors, producers or services. - Citation styles, which belong to the Citation Style Language project; bibliographic vocabularies beyond the CSL-JSON profile. - Changes to the Web Annotation, Media Fragments, IIIF, TEI, ALTO or CSL specifications themselves. The group may send them comments and requests. - Encryption, access control and digital rights management. - Protocols for storing, synchronizing or serving documents. ## Deliverables ### Specifications - **SPDF: Semantic Processed Document Format.** The file format, its anchor model and anchor URI, starting from version 5.0. Estimated: a first Community Group Report within 12 months of the group's start. - **SPDF annotation and collection sidecars.** The `.spdfa.json` profile of W3C Web Annotation with the SPDF anchor selector, and the `.spdfl.json` collection manifest. This may be published as part of the format specification or as a separate report. ### Non-normative reports The group may produce other Community Group Reports within the scope of this charter that are not Specifications, for instance: - Use cases and requirements for citable processed documents. - Mapping notes between SPDF anchors and TEI, IIIF, ALTO, CTS and CSL locators. - Implementation reports based on the conformance suite. ### Test suites and other software The group MAY produce test suites to support the Specifications. The SPDF conformance suite already exists and will be maintained by the group. Please see the GitHub LICENSE file for test suite contribution licensing information. ## Dependencies or liaisons - W3C *Web Annotation Data Model* and *Vocabulary* (Recommendations, 2017): annotation sidecars and selectors. - W3C *Media Fragments URI 1.0 (basic)* (Recommendation, 2012): time and spatial parameters of the anchor URI. - IETF: RFC 5147 (character ranges), RFC 8785 (JSON canonicalization), RFC 8032 (Ed25519), and IANA registrations for the media type `application/vnd.spdf+sqlite3` and the `spdf` URI scheme. - IIIF Consortium: Presentation API 3.0 and the region syntax of the Image API, for exports and mappings. - TEI Consortium: TEI P5, for exports and for the citation practices of digital editions. - Citation Style Language project: the CSL-JSON schema used for metadata. - Library of Congress: ALTO. - The CITE architecture: CTS URNs for canonical references. - SQLite: the file format and the `application_id` registry in its source tree. - Digital preservation registries: PRONOM (The National Archives, UK) and the Library of Congress *Sustainability of Digital Formats*. ## Community and Business Group Process The group operates under the Community and Business Group Process. Terms in this Charter that conflict with those of the Community and Business Group Process are void. As with other Community Groups, W3C seeks organizational licensing commitments under the W3C Community Contributor License Agreement (CLA). When people request to participate without representing their organization's legal interests, W3C will in general approve those requests for this group with the following understanding: W3C will seek and expect an organizational commitment under the CLA starting with the individual's first request to make a contribution to a group Deliverable. The section on Contribution Mechanics describes how W3C expects to monitor these contribution requests. The W3C Code of Conduct and the W3C Antitrust and competition policy apply to participation in this group. ## Work limited to charter scope The group will not publish Specifications on topics other than those listed under Specifications above. See below for how to modify the charter. ## Contribution mechanics Substantive Contributions to Specifications can only be made by Community Group Participants who have agreed to the W3C Community Contributor License Agreement (CLA). Reports other than Specifications published by this group should use the W3C Software and Document License where possible. Community Group participants agree to make all contributions in the GitHub repository the group is using for the particular document. This may be in the form of a pull request (preferred), by raising an issue, or by adding a comment to an existing issue. All GitHub repositories attached to the Community Group must contain a copy of the CONTRIBUTING and LICENSE files. ## Transparency The group will conduct all of its technical work in public. All technical work will occur in its GitHub repositories (and not in mailing list discussions). This is to ensure contributions can be tracked through a software tool. Meetings may be restricted to Community Group participants, but a public summary or minutes must be posted to a GitHub issue. ## Decision process This group will seek to make decisions where there is consensus. Normative changes to the Specifications follow a request-for-comments process: a proposal is discussed in public for at least 14 days, followed by a 7-day final comment period, and the Chairs assess consensus. **An accepted proposal lands with at least one conformance case; it is considered implemented when two independent implementations pass it.** A Specification is published as a final Community Group Report only with implemented proposals. Where consensus is not clear, the Chairs may issue a Call for Consensus to allow multi-day online feedback for a proposed course of action. It is expected that participants can earn Committer status through a history of valuable contributions, as is common in open source projects. After discussion and due consideration of different opinions, a decision should be publicly recorded as the resolution of a GitHub issue. If substantial disagreement remains (e.g., the group is divided) and the group needs to decide an Issue in order to continue to make progress, the Committers will choose an alternative that had substantial support (with a vote of Committers if necessary). Individuals who disagree with the choice are strongly encouraged to take ownership of their objection by taking ownership of an alternative fork. This is explicitly allowed (and preferred to blocking progress) to let implementation experience inform which spec is ultimately chosen by the group to move ahead with. Any decisions reached at any meeting are tentative and should be recorded in a GitHub Issue. Any group participant may object to a decision reached at an online or in-person meeting within 7 days of publication of the decision provided that they include clear technical reasons for their objection. The Chairs will facilitate discussion to try to resolve the objection according to this decision process. It is the Chairs' responsibility to ensure that the decision process is fair, respects the consensus of the CG, and does not unreasonably favor or discriminate against any group participant or their employer. ## Chair selection Participants in this group choose their Chair(s) and can replace their Chair(s) at any time using whatever means they prefer. However, if 5 participants, no two from the same organization, call for an election, the group must use the following process to replace any current Chair(s) with a new Chair, consulting the Community Development Lead on election operations (e.g., voting infrastructure and using RFC 3797). - Participants announce their candidacies. Participants have 14 days to announce their candidacies, but this period ends as soon as all participants have announced their intentions. If there is only one candidate, that person becomes the Chair. If there are two or more candidates, there is a vote. Otherwise, nothing changes. - Participants vote. Participants have 21 days to vote for a single candidate, but this period ends as soon as all participants have voted. The individual who receives the most votes, no two from the same organization, is elected chair. In case of a tie, RFC 3797 is used to break the tie. An elected Chair may appoint co-Chairs. Participants dissatisfied with the outcome of an election may ask the Community Development Lead to intervene. The Community Development Lead, after evaluating the election, may take any action including no action. Proposed initial Chairs: José Luis Saorín Ferrer (editor of SPDF), and a second Chair from another organization, to be identified before the group is proposed. ## Amendments to this charter The group can decide to work on a proposed amended charter, editing the text using the Decision Process described above. The decision on whether to adopt the amended charter is made by conducting a 30-day vote on the proposed new charter. The new charter, if approved, takes effect on either the proposed date in the charter itself, or 7 days after the result of the election is announced, whichever is later. A new charter must receive 2/3 of the votes cast in the approval vote to pass. The group may make simple corrections to the charter such as deliverable dates by the simpler group decision process rather than this charter amendment process. The group will use the amendment process for any substantive changes to the goals, scope, deliverables, decision process or rules for amending the charter. ## Licensing - Specifications: contributions under the W3C Community Contributor License Agreement (CLA); final reports under the W3C Community Final Specification Agreement (FSA). The SPDF specification brought to the group is already published under CC BY 4.0, and that publication remains available under its licence. - Other reports: the W3C Software and Document License where possible. - Test suites and software: as stated in the LICENSE file of the repository (currently MIT OR Apache-2.0). - The editor of SPDF and its contributors have committed not to assert patents against implementations of the specification; the CLA and the FSA add W3C's royalty-free patent commitments for contributions made in the group. --- # Contributing to SPDF URL: https://spdf.joseluissaorin.com/es/gobernanza/contributing > SPDF is three things that change in different ways: a specification, a conformance suite that turns the specification into checkable cases, and many implementations that must agree with both. This guide says how to… SPDF is three things that change in different ways: a **specification**, a **conformance suite** that turns the specification into checkable cases, and many **implementations** that must agree with both. This guide says how to contribute to each. Everyone who takes part follows the [code of conduct](CODE_OF_CONDUCT.md). Security problems are reported privately, as [`SECURITY.md`](SECURITY.md) explains, never in a public issue. The repository is private until the specification, the conformance suite and the first-tier libraries pass; until then, contributions come from invited collaborators, and comments on the specification are welcome by email at . ## Contributing to the specification The normative text is [`spec/SPEC.md`](spec/SPEC.md), in English, with a faithful Spanish translation in [`spec/SPEC.es.md`](spec/SPEC.es.md). If they differ, the English text prevails. - **Editorial changes** (typos, clearer wording, examples, diagrams, translation fixes) that would not make any implementation change: open a pull request labelled `editorial`. Change both languages when the fix applies to both, or say in the pull request that the other language needs the same fix. - **Normative changes** (anything that changes what a valid file is, what a reader, writer or validator must do, or what a function returns) go through an RFC: copy [`spec/rfcs/0000-template.md`](spec/rfcs/0000-template.md) and follow [`governance/RFC-PROCESS.md`](governance/RFC-PROCESS.md). **An accepted RFC lands with at least one conformance case; it is considered implemented when two independent implementations pass it.** - **Questions and ambiguities** are contributions too. If two readings of the text are possible, open an issue with both, and an example file or case if you can. Do not settle it by copying what another implementation does. - Use the BCP 14 key words (MUST, SHOULD, MAY) only where they are meant, and in capitals. How versions change, and what a minor version may add, is in [`governance/VERSIONING.md`](governance/VERSIONING.md). ## Contributing to the conformance suite The suite in [`conformance/`](conformance/) is normative for behaviour. Its layout, case format and runner protocol are in [`conformance/README.md`](conformance/README.md). - **Only public-domain texts whose status anyone can verify.** Name the work, the edition and the reason it is in the public domain (for example the author's death date, or a government work). Layouts, folios, timings and hashes that are synthetic must say so in the document's CSL `note`. Never use personal data. - **Expectations come from an oracle, not from an implementation**: SQLite running the reference SQL, exact arithmetic, the reference oracle `tools/spdfref.py`, or cases written and reviewed by hand in `tools/manual/`. A case that only records what one library returns is not accepted. - After editing a source or a manual case, run `python3 conformance/tools/generar.py --sellar` and then `python3 conformance/tools/verificar.py` (Python 3.13, standard library only), and commit the regenerated files. - Add a line to [`conformance/CHANGELOG.md`](conformance/CHANGELOG.md) and bump the suite version as `governance/VERSIONING.md` says. - Case ids are never renamed or reused. A wrong case is fixed in place and the change is logged; a removed case keeps its id retired. - A new kind of case, or a case that changes what the specification requires, needs an RFC. ## Contributing to an implementation Each implementation lives in its own folder (`rust/`, `js/`, `python/`, `go/`, `swift/`, `kotlin/`, `dotnet/`, `php/`, `ruby/`, `r/`, `julia/`, `c/`) with its own CI workflow, `.github/workflows/.yml`. The same holds for `producer/`, `reader/`, `site/`, `integrations/` and `models/`. - **One folder per pull request.** Do not change another implementation's folder or workflow in the same pull request; if you find a bug in another implementation, open an issue or a separate pull request for it. - **Follow the specification, not another implementation**, including the reference one. When an implementation and the suite disagree, the suite wins until an RFC says otherwise. - **Run the conformance runner.** Every implementation ships a runner that discovers `conformance/cases/*.json` and prints the report described in `conformance/README.md`. A change must not turn a passing case into a failing one. CI uploads the report as the artifact `conformance-`. - **Keep the safety rules**: safe opening, size limits and the other requirements of [SPEC §2.4](spec/SPEC.md#container) and [§14](spec/SPEC.md#security) are not optional, even in tests and examples. - Follow the conventions of the language: its formatter, linter and test framework, as the folder's README describes. Add tests with the change. - New dependencies need a reason in the pull request. Prefer what the platform already provides; SQLite comes first. - Models are downloaded on demand and never committed or packaged. ## Pull requests - Keep them small and about one thing. Explain the why, not only the what. - Link the RFC, issue or conformance case the change relates to. - CI must pass, including the conformance runner of the folder you touched. - Before pushing, rebase on the current `main` (`git pull --rebase`). - Never commit secrets, credentials, API keys or personal data. ## Commit messages - Subject line: `: `, where `` is the top-level folder (`spec`, `conformance`, `governance`, `rust`, `js`, `python`, …) or `ci`, and the summary is in English or Spanish, at most 72 characters, without a final period. - A body, separated by a blank line, says why the change is needed when that is not obvious. - Add `Co-authored-by` lines for every co-author. A `Signed-off-by` line is welcome but not required. - Contributions prepared with the help of AI tools are welcome. The person who submits them is responsible for them and has reviewed them; say in the pull request which tools were used. ## Licensing of contributions Contributions are licensed under the licence of the part of the repository they change: | Part | Licence | |---|---| | Specification and documentation (`spec/`, `governance/`, prose in every folder) | [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/) | | Code (implementations, producer, reader, conformance suite, website, integrations) | MIT OR Apache-2.0, at the user's choice ([`LICENSE-MIT`](LICENSE-MIT), [`LICENSE-APACHE`](LICENSE-APACHE)) | The rule is **inbound = outbound**: by submitting a contribution you license it under the same licence as the part of the repository it changes, and you confirm that you wrote it or otherwise have the right to submit it under that licence. No contributor licence agreement and no Developer Certificate of Origin sign-off are required for now; if the project adopts one later, it will be announced through the RFC process and will not apply retroactively. Contributors to the specification also make the patent commitment described in [`governance/GOVERNANCE.md`](governance/GOVERNANCE.md#licences-and-patents): they will not assert patents they own or control against implementations of SPDF. ## Languages - The normative specification is written in English, with a faithful Spanish translation maintained alongside it. - Issues, pull requests, reviews and discussions may be in English or Spanish. - Documentation for people is published in English and, where possible, in Spanish. Spanish text is written with correct spelling, accents and `ñ` included; identifiers in code, schemas and JSON keys stay in English without accents. ## Reporting security problems Do not report vulnerabilities in public issues. Follow [`SECURITY.md`](SECURITY.md): email or, once the repository is public, use GitHub's private vulnerability reporting. ## En español Las contribuciones en español son igual de bienvenidas. La especificación normativa está en inglés y tiene una traducción española fiel en `spec/SPEC.es.md`; si se contradicen, manda el inglés. Las erratas y mejoras de redacción se proponen con un *pull request* marcado `editorial`; todo cambio normativo pasa por una RFC (`spec/rfcs/`, proceso en `governance/RFC-PROCESS.md`), y una RFC aceptada entra con al menos un caso de conformidad y se da por implementada cuando dos implementaciones independientes lo pasan. Los casos de la batería solo usan obras de dominio público comprobable. Cada implementación vive en su carpeta, con su propio flujo de CI, y un *pull request* toca una sola carpeta. Las contribuciones se aceptan con la misma licencia con la que se publican (*inbound = outbound*), sin DCO ni acuerdo de contribución; la especificación y la documentación se publican con CC BY 4.0 y el código con MIT OR Apache-2.0. Los fallos de seguridad se comunican en privado, como explica `SECURITY.md`. --- # Contributor Covenant Code of Conduct URL: https://spdf.joseluissaorin.com/es/gobernanza/code-of-conduct > We as members, contributors, and leaders pledge to make participation in our community a harassment-free experience for everyone, regardless of age, body size, visible or invisible disability, ethnicity, sex… ## Our Pledge We as members, contributors, and leaders pledge to make participation in our community a harassment-free experience for everyone, regardless of age, body size, visible or invisible disability, ethnicity, sex characteristics, gender identity and expression, level of experience, education, socio-economic status, nationality, personal appearance, race, caste, color, religion, or sexual identity and orientation. We pledge to act and interact in ways that contribute to an open, welcoming, diverse, inclusive, and healthy community. ## Our Standards Examples of behavior that contributes to a positive environment for our community include: * Demonstrating empathy and kindness toward other people * Being respectful of differing opinions, viewpoints, and experiences * Giving and gracefully accepting constructive feedback * Accepting responsibility and apologizing to those affected by our mistakes, and learning from the experience * Focusing on what is best not just for us as individuals, but for the overall community Examples of unacceptable behavior include: * The use of sexualized language or imagery, and sexual attention or advances of any kind * Trolling, insulting or derogatory comments, and personal or political attacks * Public or private harassment * Publishing others' private information, such as a physical or email address, without their explicit permission * Other conduct which could reasonably be considered inappropriate in a professional setting ## Enforcement Responsibilities Community leaders are responsible for clarifying and enforcing our standards of acceptable behavior and will take appropriate and fair corrective action in response to any behavior that they deem inappropriate, threatening, offensive, or harmful. Community leaders have the right and responsibility to remove, edit, or reject comments, commits, code, wiki edits, issues, and other contributions that are not aligned to this Code of Conduct, and will communicate reasons for moderation decisions when appropriate. ## Scope This Code of Conduct applies within all community spaces, and also applies when an individual is officially representing the community in public spaces. Examples of representing our community include using an official e-mail address, posting via an official social media account, or acting as an appointed representative at an online or offline event. ## Enforcement Instances of abusive, harassing, or otherwise unacceptable behavior may be reported to the community leaders responsible for enforcement at . All complaints will be reviewed and investigated promptly and fairly. All community leaders are obligated to respect the privacy and security of the reporter of any incident. ## Enforcement Guidelines Community leaders will follow these Community Impact Guidelines in determining the consequences for any action they deem in violation of this Code of Conduct: ### 1. Correction **Community Impact**: Use of inappropriate language or other behavior deemed unprofessional or unwelcome in the community. **Consequence**: A private, written warning from community leaders, providing clarity around the nature of the violation and an explanation of why the behavior was inappropriate. A public apology may be requested. ### 2. Warning **Community Impact**: A violation through a single incident or series of actions. **Consequence**: A warning with consequences for continued behavior. No interaction with the people involved, including unsolicited interaction with those enforcing the Code of Conduct, for a specified period of time. This includes avoiding interactions in community spaces as well as external channels like social media. Violating these terms may lead to a temporary or permanent ban. ### 3. Temporary Ban **Community Impact**: A serious violation of community standards, including sustained inappropriate behavior. **Consequence**: A temporary ban from any sort of interaction or public communication with the community for a specified period of time. No public or private interaction with the people involved, including unsolicited interaction with those enforcing the Code of Conduct, is allowed during this period. Violating these terms may lead to a permanent ban. ### 4. Permanent Ban **Community Impact**: Demonstrating a pattern of violation of community standards, including sustained inappropriate behavior, harassment of an individual, or aggression toward or disparagement of classes of individuals. **Consequence**: A permanent ban from any sort of public interaction within the community. ## Attribution This Code of Conduct is adapted from the [Contributor Covenant][homepage], version 2.1, available at [https://www.contributor-covenant.org/version/2/1/code_of_conduct.html][v2.1]. Community Impact Guidelines were inspired by [Mozilla's code of conduct enforcement ladder][Mozilla CoC]. For answers to common questions about this code of conduct, see the FAQ at [https://www.contributor-covenant.org/faq][FAQ]. Translations are available at [https://www.contributor-covenant.org/translations][translations]. [homepage]: https://www.contributor-covenant.org [v2.1]: https://www.contributor-covenant.org/version/2/1/code_of_conduct.html [Mozilla CoC]: https://github.com/mozilla/diversity [FAQ]: https://www.contributor-covenant.org/faq [translations]: https://www.contributor-covenant.org/translations --- # Security policy URL: https://spdf.joseluissaorin.com/es/gobernanza/security > SPDF files come from strangers. Opening one means parsing an untrusted database with a complex engine, and the specification asks every reader to do it safely (SPEC §2.4, §14 and §15). If you find a way around that,… SPDF files come from strangers. Opening one means parsing an untrusted database with a complex engine, and the specification asks every reader to do it safely ([SPEC §2.4](spec/SPEC.md#container), [§14](spec/SPEC.md#security) and [§15](spec/SPEC.md#privacy)). If you find a way around that, or any other vulnerability in the specification or in the software in this repository, please tell us privately first. ## Supported versions | Component | Supported | |---|---| | Specification | 5.0 (working draft), including the reading of legacy 4.0 and 4.1 files that it requires | | Libraries (Rust, TypeScript, Python, Go, Swift, Kotlin/JVM, C#, PHP, Ruby, R, Julia, C ABI) | the latest release of the common release train; while the libraries are at 0.x, only the latest minor release gets fixes | | Reference producer `spdf build` | the latest release | | SPDF Reader (desktop, mobile and web) | the latest release | | Website, browser validator and conformance suite | the current version on `main` | Earlier releases do not receive fixes; upgrading is the remedy. ## How to report Report vulnerabilities privately, by either of these channels: - **Email** to , with `[SPDF security]` at the start of the subject. If you need an encrypted channel, say so in a first message without details and we will agree on one. - **GitHub private vulnerability reporting** ("Report a vulnerability" in the Security tab of ), once the repository is public. Please do not open public issues, pull requests or discussions about a vulnerability before it is fixed and disclosed. A useful report includes: the component and version (or commit), the platform and the SQLite version, a description of the impact, and a minimal reproduction. A crafted `.spdf` file is the best reproduction; please build it from public-domain or synthetic content, never from personal data. ## What happens next | Step | Deadline | |---|---| | Acknowledgement of your report | within 72 hours | | First assessment (confirmed or not, severity, components affected) | within 14 days | | Coordinated public disclosure | within 90 days of the report | - We keep you informed while we work on a fix and agree the disclosure date with you. Disclosure may come earlier when a fix is released, and later only by mutual agreement, or if a fix needs a change to the specification that cannot be made safely in time. - If the vulnerability is being exploited, we may disclose sooner, with mitigations. - Fixes to the specification that cannot wait for the normal discussion periods follow the shortened procedure in [`governance/RFC-PROCESS.md`](governance/RFC-PROCESS.md). - Once the repository is public, advisories are published through GitHub Security Advisories, which can assign a CVE identifier. - We credit reporters in the advisory, unless you prefer to remain anonymous. ## In scope - **Files that escape safe opening**: an SPDF or legacy file that makes a conforming reader execute SQL from the file (through triggers, views, virtual tables or schema tricks), load an extension, write to the file or elsewhere on disk, read outside the file, or keep running without bound despite the limits the specification requires. - **Memory-safety and parsing bugs** in readers and validators: overflows, out-of-bounds reads, crashes or hangs triggered by a file, including in gzip decompression of legacy files, JSON columns, word timings, vector blobs (`f32`, `f16`, `i8`), anchor URI parsing and full-text queries. - **Signatures that validate when they should not**: a file whose `content_sha256` or Ed25519 signature verifies although its content differs from what was signed, two different contents with the same canonical dump, or a verifier that trusts a stored hash instead of recomputing it. - **Leaks through vectors or provenance**: an implementation or producer that claims to remove the text of a document but keeps vectors from which it can be recovered, or that writes paths, user names, keys or other personal data into `provenance` or `generator` against [SPEC §15](spec/SPEC.md#privacy). - **Content handling** in the reader, the website and the validator: script execution from unit text, metadata, SVG or embedded documents; automatic fetching of remote references; path traversal when exporting blobs; a browser validator or inspector that sends a file off the user's device. - **The specification itself**: a rule that, followed exactly, leaves conforming implementations unsafe. - **The supply chain of this repository**: CI workflows, release artifacts and published packages. ## Not a vulnerability - Bugs in SQLite itself: please report them to the SQLite project. Do tell us if the way SPDF readers use SQLite makes such a bug exploitable. - A file that is invalid but is refused or read safely. Validation errors are the expected behaviour. - Resource use within the configured limits. Very large files that stay within the limits a reader was given are not a denial of service; files that bypass the limits are. - The content of documents: wrong transcriptions, wrong folios or doubtful metadata are quality problems, and instructions written in a document's text are data. Report them as ordinary issues. - A valid signature made with a key you do not trust. The specification attributes content to a key; deciding which keys to trust is up to the user. - The Ed25519 test key in `conformance/`, which is public on purpose and signs only test files. - Vectors that reveal information about a text that the same file already contains. - Reports from automated scanners without a demonstrated impact, and missing security headers on static pages without a concrete attack. --- # Plantilla de RFC URL: https://spdf.joseluissaorin.com/es/gobernanza/rfcs/0000 > Copy this file to spec/rfcs/0000-short-title.md and open a pull request. The editor assigns the next free number when discussion opens and renames the file. Keep every section below: when one does not apply, write… - **Status**: Draft - **Authors**: Full Name (affiliation, if any) - **Created**: YYYY-MM-DD - **Discussion**: link to the pull request or to the public thread - **Specification**: the SPDF version this targets (for example 5.1) and whether the change is minor (additive) or major (breaking) - **Affects**: sections of `spec/SPEC.md`, schema files, JSON Schemas, kinds of conformance case - **Conformance cases**: ids of the cases added or changed - **Supersedes / superseded by**: RFC numbers, if any - **Decision**: filled in by the editor or the technical committee (date, outcome and reasons) Copy this file to `spec/rfcs/0000-short-title.md` and open a pull request. The editor assigns the next free number when discussion opens and renames the file. Keep every section below: when one does not apply, write "Not applicable" and one sentence saying why. Delete these instructions and the italic guidance under each heading before the final comment period. The process is described in [`governance/RFC-PROCESS.md`](../../governance/RFC-PROCESS.md); how versions change is in [`governance/VERSIONING.md`](../../governance/VERSIONING.md). ## Summary *One paragraph that someone who has never read the specification can follow: what changes and for whom.* ## Motivation *The problem, with real examples: a document that cannot be anchored today, a citation that comes out wrong, an implementation that cannot do something efficiently. Say who needs this (readers, producers, archives, users of a particular language or kind of source) and what happens if nothing changes.* ## Guide-level explanation *Explain the change as it would be taught to an implementer or a producer: new tables, columns, anchor members or functions, with a small worked example (a row, an anchor, a URI, a citation). No normative language here.* ## Normative changes *The exact changes, written so they can be merged into `spec/SPEC.md` as they are. Use the BCP 14 key words (RFC 2119 and RFC 8174: MUST, SHOULD, MAY) only where they are meant. For each change give the section, the current text and the new text. Cover, when they are affected:* - *the schema (`spec/schema/`) and the canonical dump;* - *anchors and the anchor URI (parameters, canonical order, parsing);* - *search, citation and export functions;* - *validation: new error or warning codes take the next free number in their range and their place in the check order;* - *integrity and signatures;* - *the legacy mapping (4.x to 5.x view);* - *sidecars (`.spdfa.json`, `.spdfl.json`);* - *the value of `PRAGMA user_version` and `spdf_meta.spdf_version` that files using the change must carry.* *The Spanish translation (`spec/SPEC.es.md`) is updated in the same pull request, or the editor updates it before the change is published.* ## Conformance cases *An accepted RFC lands with at least one conformance case; it is considered implemented when two independent implementations pass it. List each case: id, kind, what it checks, the source or file it uses, and how its expected output was obtained (SQLite running the reference SQL, exact arithmetic, hand-written and reviewed, the reference oracle), so that nobody has to trust a single implementation. Follow "Adding cases" in [`conformance/README.md`](../../conformance/README.md). Only public-domain texts whose status anyone can verify.* ## Backwards compatibility *Answer each question:* - *Do existing files stay valid, with the same canonical dump and the same `content_sha256`?* - *What does a reader of an earlier minor version of the same major do with a file that uses the change? It MUST still read it: see the compatibility promise in `governance/VERSIONING.md`. If it cannot, the change belongs in a new major version.* - *What must writers do differently, and when?* - *Is anything deprecated? Give the version that deprecates it and the earliest major version that may remove it (at least 24 months later).* - *Does the legacy 4.x mapping change?* ## Security and privacy *New ways for a file to make a reader do work, fetch something, execute something or reveal something: SQL objects, URLs, embedded content, sizes and limits, information about people, what vectors or provenance disclose, effects on hashes and signatures. If there are none, say why.* ## Alternatives *Other designs considered, including doing nothing, and why this one is better. Prior art in other formats and standards (TEI, IIIF, W3C Web Annotation, Media Fragments, CSL, EPUB, PDF) is welcome.* ## Unresolved questions *What must be settled before acceptance, and what is deliberately left for later.* ## Implementations *Filled in as implementations land. An RFC becomes Implemented when two independent implementations pass all of its cases in CI. Two implementations are independent when neither wraps or translates the other's code: bindings over the Rust core, including the C ABI, count as the Rust implementation.* | Implementation | Status | Version or commit | Cases passed | |---|---|---|---| | | | | | --- # RFC 0001: SPDF 5.0 URL: https://spdf.joseluissaorin.com/es/gobernanza/rfcs/0001 > SPDF 5.0 is the first public version of the format. It keeps the idea of the Scholaris versions 4.0 and 4.1 (one document per SQLite file; every passage carries an anchor to the printed page, folio, verse, second or… - **Status**: Accepted (2026-10-07) - **Authors**: José Luis Saorín Ferrer (editor) - **Created**: 2026-10-07 - **Discussion**: review by the implementers of the libraries, recorded in the change log of `spec/CONTRACT.md` (drafts 0 to 1.2) - **Specification**: SPDF 5.0, a new major version. It replaces 4.1 as the current version; 4.0 and 4.1 become legacy versions that every reader still reads - **Affects**: the whole of `spec/SPEC.md`, `spec/schema/spdf-5.0.sql`, `spec/json-schema/`, `conformance/` - **Conformance cases**: the conformance suite 0.1.0 (220 cases) and 0.2.0 (228 cases, `cases_sha256` `2d28cb5169e426ce26fec1df6c7481e2e5a3b208f0f14f66411e51043aa442c3`) - **Supersedes / superseded by**: none - **Decision**: accepted by the editor on 2026-10-07. This RFC bootstraps the process and was accepted without the public discussion and final comment periods that apply to every later RFC (see [`governance/RFC-PROCESS.md`](../../governance/RFC-PROCESS.md)). It is published so that anyone can see what changed from 4.1 and why; any of its decisions can be revisited by a later RFC. ## Summary SPDF 5.0 is the first public version of the format. It keeps the idea of the Scholaris versions 4.0 and 4.1 (one document per SQLite file; every passage carries an anchor to the printed page, folio, verse, second or slide it comes from; several vector spaces may coexist) and turns it into an open standard: the container is plain uncompressed SQLite, the file identifies itself in its header, identifiers are in English, metadata is CSL-JSON, text offsets are defined, the anchor vocabulary grows, anchors get a portable URI aligned with W3C Media Fragments and RFC 5147, vectors can be stored in half precision or 8 bits, conformance profiles and an extension mechanism are defined, content can be hashed and signed, and distributed files may no longer contain code (triggers, views or foreign virtual tables). Every 5.0 reader still reads 4.0 and 4.1 files. The normative text is [`spec/SPEC.md`](../SPEC.md), whose Appendix A summarizes the changes from 4.1; section numbers below refer to it. ## Motivation SPDF 4.x was the internal export format of Scholaris. It worked for one application and one team, but it had properties that stand in the way of a format anyone can implement and archive: - **Compressed container.** A 4.x file is a SQLite database wrapped in gzip. A reader must decompress the whole file before reading a single row, so it cannot read by HTTP ranges or memory-map the file, and must defend itself against decompression bombs. Most of the bulk of a typical file (page images, an embedded original) is already in compressed formats and gains little from a second layer. - **No identification.** 4.x files carry `application_id` 0. The version is stored in a table (`spdf.spdf_version`) and only sometimes in `user_version`, which some platforms did not allow Scholaris to set (the 4.0 sample in the conformance suite has `user_version` 0). Tools such as file(1) or DROID cannot tell a 4.x file from any other gzip stream. - **Spanish identifiers.** Tables and columns are named in Spanish (`documentos`, `unidades`, `anio`, `ancla_fin`). That was natural inside Scholaris and is a barrier for everyone else. - **Bespoke metadata.** Document metadata used Scholaris's own JSON structure (`MetadatosDocumento`), which no reference manager understands. - **Undefined offsets.** Nothing said how to count a position inside a text, and the languages that implement SPDF count differently (UTF-16 code units in JavaScript, code points in Python, bytes in Rust and Go). - **Missing anchors.** Verse numbers, canonical references (Stephanus, Bekker, biblical, CTS) and leaf or column foliation, which are how poetry, classical texts, early printed books and manuscripts are cited, could not be expressed. - **No way to point at a passage from outside the file.** 4.x defined no URI for an anchor. - **Vectors in f32 only**, which makes files with several vector spaces large and is wasteful on phones. - **SQL code inside distributed files.** 4.x files carry three triggers that keep the full-text index in sync. Opening a file that contains SQL code written by someone else is a risk that a document format should not ask readers to take. - **No integrity, no profiles, no extension mechanism, no rights**: nothing to detect tampering, nothing to say what a minimal conforming file is, no way for a vendor to add data without forking the format, and nowhere to say under what terms a file may be shared. ## Guide-level explanation A 5.0 file is a SQLite 3 database that any SQLite tool can open. Its header says what it is: bytes 68 to 71 read `SPDF` (the `application_id`) and bytes 60 to 63 hold the version, 500. Inside there is exactly one document: - `documents`: one row with the CSL-JSON record of the work, the SHA-256 of the original file, its media type and size, and its rights; - `units`: the citable units in reading order (pages, time spans, slides, sections, sheets), each with its anchor and its text in NFC; - `fragments`: passages of about 150 to 300 words, each with the anchor of its start and, if it crosses units, of its end, indexed by FTS5 for lexical search; - optionally `sections`, `figures`, vector `spaces` and `vectors`, embedded `blobs` (page images, the original), `provenance` and `extensions`. An anchor is a small JSON object: ```json {"type":"page","physical":29,"printed":"21","source":"read","chars":[118,301]} ``` and the same location as a URI, which names the document by the hash of its original bytes so it survives renaming and copying: ```text spdf:sha256-3f2a…#p=29&f=21&char=118,301 ``` A citation is computed from the stored anchor and the CSL record, never generated: `(Darwin, 1859, p. 21)`. A producer may also store a SHA-256 of the canonical content and an Ed25519 signature over it, so that a reader can check that the content is the one the signer vouched for. ## Normative changes The changes against 4.1, by area, with the section of `spec/SPEC.md` that states each rule. ### Container and identification (§2, §24) 1. A 5.0 file is an **uncompressed** SQLite 3 database holding exactly one document; the header starts at byte 0. Page size 4096, journal mode DELETE and a final `VACUUM` are recommended. Collections are separate manifests (§17). 2. `PRAGMA application_id = 1397769286` (0x53504446, stored big-endian at offset 68, which reads `SPDF` in ASCII). 3. `PRAGMA user_version` = major × 100 + minor × 10 (5.0 is 500). Readers accept 500 to 599, may warn on a newer minor (W105), and refuse other majors (E002). 4. Extension `.spdf`; media type `application/vnd.spdf+sqlite3` (registration pending); Uniform Type Identifier `com.joseluissaorin.spdf`. 5. Files MUST NOT contain triggers, views, or virtual tables other than the FTS5 tables of the schema. Writers keep the full-text index in sync themselves. ### Safe opening (§2.4, §14) 6. Every reader opens files read-only, with `query_only` on, `trusted_schema` off, the defensive flag on and extension loading off; refuses triggers, views and foreign virtual tables (except the three legacy FTS triggers); and bounds the size of any single value (RECOMMENDED 512 MiB) and of decompressed gzip input (RECOMMENDED 4 GiB). Operations that write, such as the FTS5 `integrity-check`, run on a private copy. ### Schema with English identifiers (§3, §20) 7. Tables: `spdf` becomes `spdf_meta`, `documentos` `documents`, `unidades` `units`, `secciones` `sections`, `fragmentos` `fragments`, `figuras` `figures`, `espacios` `spaces`, `vectores` `vectors`, `procedencia` `provenance`; `blobs` keeps its name. Every column is renamed as §20.1 lists. 8. Removed: `documentos.estado` and `documentos.bibliotecas` (library membership belongs to collection manifests) and index-only columns. 9. Added: `documents.rights` (§16); `documents.source_ref`, nullable, replacing `original`; `spaces.dtype`, `spaces.truncated_from`, `spaces.task_prefixes`; `blobs.sha256`; `provenance.model`; the `extensions` table. 10. `spdf_meta` has the REQUIRED keys `spdf_version`, `profile`, `created`, `generator` and `document_id`; integrity keys are OPTIONAL. 11. `units.ord` is numbered from 1 and contiguous (4.x numbered units from 0). ### Metadata (§6) 12. `documents.metadata` is one CSL-JSON item plus an `spdf` extension object for what CSL cannot hold: provenance per field, a date range for undated works, the original language, ORCID identifiers. The record can be handed to Zotero, citeproc or Pandoc as it is. ### Text and offsets (§7) 13. All stored text is NFC. Positions inside a text are counted in **Unicode code points over the NFC text**, end exclusive, which every language can compute the same way. 14. The literal text is never modernized; the `search_text` column (introduced in 4.1) holds a modernized-spelling layer used only for search. ### Anchors (§4) 15. New anchor types: `verse` and `canonical` (schemes such as `stephanus`, `bekker`, `bible` or `cts`). 16. Page anchors gain `foliation`: `page` (default), `leaf` (`fol. 1r`) or `column` (`col. 45`). 17. Any anchor MAY carry `region` (fractions 0 to 1 of the unit image) and `chars` (code point range in the unit's NFC text). 18. Required members are fixed per type and validated (E040, E041, E042). ### Anchor URI (§5) 19. New URI form `spdf:#`, with an ABNF. The preferred docref is `sha256-` and the hex SHA-256 of the original, which survives renaming and copying. 20. Parameters have one canonical order and encoding, and `format(parse(uri))` reproduces the URI byte for byte. 21. Where SPDF overlaps with existing standards it uses their syntax: `t=` and `xywh=percent:` as in W3C Media Fragments URI 1.0, and `char=` as in RFC 5147. 22. Resolution rules say how a reader finds the unit a URI designates. ### Vectors (§9) 23. Vector components are little-endian `f32`, `f16` (IEEE binary16) or `i8` (value q/127); the space id records model, dimensions and dtype. 24. Spaces record Matryoshka truncation and the task prefixes used at encoding time; a compatibility rule says when one query vector serves several spaces; writers quantize with fixed rounding rules. 25. A file without vectors is valid. ### Search, citation and export (§8, §18, §19) 26. Reference algorithms that conformance tests: lexical search through FTS5 with the `unicode61 remove_diacritics 2` tokenizer (terms sent as written, without case folding), a route for Chinese, Japanese and Korean with an optional `trigram` index and a substring fallback, brute-force vector search, and hybrid search by reciprocal rank fusion with k = 10. Products MAY rank better. 27. A short citation function for Spanish and English, and exports to CSL-JSON and BibTeX (REQUIRED) and other formats. ### Profiles and extensions (§10, §11) 28. Profiles, declared in `spdf_meta.profile`: `core`, `semantic`, `media` and `full`. 29. Extensions are declared in the `extensions` table, with tables named `x__`. A reader that meets a **required** extension it does not know refuses the file (E060); optional ones are ignored. ### Integrity, signature and rights (§12, §13, §16) 30. A canonical JSON dump of the file (RFC 8785, with fixed rounding and ordering) is the conformance oracle and the basis of integrity. 31. `content_sha256` hashes that dump; `signature` is an Ed25519 signature (RFC 8032) over it. Because the dump covers blob and vector bytes through their hashes but not the SQLite page layout, the signature survives `VACUUM` and SQLite version changes. 32. `documents.rights` states the licence (SPDX), the access level and the holder. ### Validation (§22) 33. A deterministic check order and a closed list of error and warning codes. Conformance compares the sets of codes, not the messages. ### Sidecars (§17) 34. User annotations live outside the document, in `*.spdfa.json` files (W3C Web Annotation with an `SpdfAnchorSelector` and a `TextQuoteSelector`), so the document stays immutable and sharing it never shares its reader's notes. 35. Collections are `*.spdfl.json` manifests that list documents by hash. ### Legacy (§20) 36. Every reader reads 4.0 and 4.1 files: detects gzip, decompresses within the limit, tolerates exactly the triggers `fragmentos_ai`, `fragmentos_ad` and `fragmentos_au`, and presents the 5.0 view, with `"legacy": true` in the dump and warning W110. A gzip-wrapped 5.0 file is read but flagged (E003, a warning). 37. Version 3.0 MAY be supported through an importer. ## Conformance cases The conformance suite is the evidence for this RFC. Version 0.1.0 (2026-10-07) published 220 cases; version 0.2.0 (the same day) brought them to 228: `anchor_uri` 40, `cite` 79, `dump` 7, `legacy_dump` 2, `quantize` 6, `roundtrip` 7, `search_hybrid` 2, `search_lexical` 39, `search_vector` 5 and `validate` 41. They use seven 5.0 files built from public-domain texts, two authentic legacy files (4.0 and 4.1, gzip-wrapped, with FTS triggers), and invalid files that each break one rule. How each expectation is obtained (SQLite as the oracle for lexical search, exact arithmetic for vectors, hand-reviewed URIs and citations) is described in [`conformance/README.md`](../../conformance/README.md); the changes are in [`conformance/CHANGELOG.md`](../../conformance/CHANGELOG.md). ## Backwards compatibility - **5.0 is a new major version.** A 4.x reader cannot read 5.0 files: the schema is different and the container is no longer gzip. This is intended. - **5.0 readers read 4.x.** Every conforming 5.0 reader reads 4.0 and 4.1 files through the 5.0 view (§20). Nothing that a 4.x file holds about the document is lost; only the Scholaris shelf state (`estado`, `bibliotecas`) is dropped. - **Writers** produce 5.0 only. Scholaris keeps reading its 4.x files and exports 5.0 through the `spdf-format` library. - **Sidecars** are new and carry their own version (`spdf_library: "1.0"` in manifests). - **Nothing is deprecated** within 5.0. The obligation to read 4.x lasts for the whole 5.x line; a future major version decides by RFC whether its readers still read 4.x (see [`governance/VERSIONING.md`](../../governance/VERSIONING.md)). ## Security and privacy Sections 14 and 15 of the specification are new in 5.0. In short: - Files contain no triggers, views or foreign virtual tables, and readers refuse files that do, so opening a file runs no code written by its author. Readers open read-only, with `trusted_schema` off, defensive mode on and extensions disabled. - The 5.0 container is not compressed. Legacy gzip input is decompressed within a limit to defeat decompression bombs; values and JSON nesting are bounded. - Text is light Markdown rendered without raw HTML; remote references are never fetched automatically; blob keys are sanitized before anything is written to disk; text passed to language models is data, not instructions. - Vectors can leak the text they were computed from; provenance can leak details of the producer. Rights and confidentiality rules that apply to a text apply to its vectors. - `content_sha256` and the Ed25519 signature give integrity and attribution of the content, not confidentiality, and say nothing about whether the signer's key should be trusted. ## Alternatives - **Keep the gzip wrapper.** Rejected: it prevents HTTP range reads and memory mapping, forces full decompression and adds decompression-bomb handling, for little gain on already-compressed images and originals. Readers still accept it for legacy files. - **A ZIP package with JSON and a SQLite index inside** (as EPUB or OOXML do). Rejected: two levels of parsing, and the full-text index needs SQLite anyway; SQLite alone gives random access, an index and a single file. - **Pure JSON, Parquet or Arrow.** Rejected: no built-in full-text search, and either no random access (JSON) or no good fit for long text with anchors (columnar formats). - **Keep Spanish identifiers with English aliases.** Rejected: two names for every table and column would double the surface of every implementation. The specification keeps a faithful Spanish translation instead. - **UTF-16 code units, bytes or grapheme clusters for offsets.** Rejected: UTF-16 ties the format to JavaScript, bytes to one encoding, and grapheme clusters to a Unicode version. Code points over NFC are stable and cheap everywhere. - **IIIF-style regions (`pct:`) or bare fractions in `xywh`.** Rejected in favour of W3C Media Fragments (`percent:`); in Media Fragments bare numbers mean pixels, so bare fractions would have been misread. Mapping to IIIF Image API regions is trivial. - **Signing the file bytes, or wrapping the file in JWS or COSE.** Rejected: SQLite writes its own version into the header and `VACUUM` reorders pages, so byte signatures break without any change in content. A signature over the canonical content is stable. An envelope format can still be added later as an extension. - **Allowing triggers and views.** Rejected for safety; writers can keep the index in sync without them. ## Unresolved questions - The registration of `application/vnd.spdf+sqlite3` with IANA, and whether to use the registered `+sqlite3` structured syntax suffix instead (`application/vnd.spdf+sqlite3`). Drafts of this and other registrations are in `governance/drafts/`. - The provisional registration of the `spdf` URI scheme (RFC 7595), which §5.4 announces. - Whether the anchor URI parameters may also be used as the fragment identifier of a URL that points to a `.spdf` file (`https://example.org/x.spdf#p=29`), which the media type registration would like to say. - How a reader and a validator of an earlier minor version treat an anchor type, a dtype or a validation rule introduced by a later minor (today an unknown anchor type is error E041), so that the compatibility promise holds for validators too. - Markup for mathematics and other non-textual content inside unit text, which 5.0 does not specify. - A persistent identifier (DOI) for each published version of the specification. ## Implementations This RFC becomes Implemented when two independent implementations pass the whole conformance suite in CI, as their `conformance-` artifacts show. The table records what each implementation reported in its own commits on 2026-10-07; several independent implementations already report the whole suite, so the editor will mark the RFC Implemented once their CI artifacts confirm it. | Implementation | Folder | Reported on 2026-10-07 | |---|---|---| | Rust (reference, also the C ABI) | `rust/` | 228 of 228 (commit `38c75e5`) | | TypeScript (`spdf-format`, npm) | `js/` | 228 of 228, also in Chromium and Bun (commit `e063594`) | | Python (`spdf-format`, PyPI) | `python/` | 228 of 228 (commit `557af8f`) | | Go | `go/` | 228 of 228 (commit `9355f22`) | | Swift | `swift/` | 228 of 228 on macOS and the iOS simulator (commit `54d5525`) | | PHP and Ruby | `php/`, `ruby/` | 228 of 228 (commit `537202e`) | | Kotlin/JVM, C#, R | `kotlin/`, `dotnet/`, `r/` | in development | | Julia, C (over the Rust ABI) | `julia/`, `c/` | planned | | Reference producer `spdf build` (Python) | `producer/` | in development | | Scholaris (second producer) | external | reads 4.x; 5.0 export through `spdf-format` planned | --- # RFC 0002: Conformance cases for exports and anchor resolution URL: https://spdf.joseluissaorin.com/es/gobernanza/rfcs/0002 > Add three kinds of conformance case: exportcsl and exportbibtex, which make the MUST of SPEC §19 testable, and resolve, which tests how an anchor URI is resolved against a file (SPEC §5.4). Fix the details those… - **Status:** Accepted (2026-10-07); normative text in SPEC §5.4 and §19; cases in conformance 0.4.0. Implemented when two independent implementations pass them. - **Author:** spec agent, for the editor (José Luis Saorín Ferrer) - **Created:** 2026-10-07 - **Affects:** SPEC §5.4, §19, §21; `conformance/` ## Summary Add three kinds of conformance case: `export_csl` and `export_bibtex`, which make the MUST of SPEC §19 testable, and `resolve`, which tests how an anchor URI is resolved against a file (SPEC §5.4). Fix the details those cases need: the BibTeX key algorithm, the field set and the resolution order. ## Motivation SPEC §19 says every implementation MUST export CSL-JSON and BibTeX, and SPEC §5.4 says how a reader finds the unit an anchor URI designates. Neither is covered by the suite 0.2.0, so twelve implementations can diverge silently: different BibTeX keys break the `\cite{}` commands of a user who switches tools, and different resolution sends a reader to a different page than the citation says. Both are the kind of disagreement SPDF exists to prevent. ## Guide-level explanation ```json {"id": "export-bibtex-quijote", "kind": "export_bibtex", "input": {"file": "files/quijote.spdf"}, "expect": {"entry_type": "book", "key": "cervantessaavedra1605", "fields": {"author": "Cervantes Saavedra, Miguel de", "title": "El ingenioso hidalgo don Quijote de la Mancha", "year": "1605", "publisher": "Juan de la Cuesta", "address": "Madrid", "language": "es", "note": "…"}}} ``` BibTeX is compared structurally (entry type, key, field map), not byte for byte, so implementations keep their own layout and escaping style, which BibTeX tools do not care about. ```json {"id": "resolve-quijote-leaf", "kind": "resolve", "input": {"file": "files/quijote.spdf", "uri": "spdf:sha256-fa38…4c75#f=1v"}, "expect": {"unit": "p6", "chars": null}} ``` ## Normative changes 1. SPEC §19, BibTeX key: take the `family` (or `literal`) of the first author, else the first word of the title; decompose with NFKD, drop every character that is not an ASCII letter, lowercase; if nothing is left, use `anon`; append the first year of `issued`, or `nd`. Collisions inside one export get the suffixes `a`, `b`, `c`… in the order of the documents. 2. SPEC §19, BibTeX fields: exactly the mapping already listed in §19; names as `Family, Given` joined with ` and `; a `literal` name wrapped in braces; `year` as a string; fields whose source is absent are omitted. 3. SPEC §5.4, resolution order: `p` (and `pe`) first; else `f` through `units.printed`, first unit in `ord` order; else `t` (the first unit with `t0 ≤ t < t1`, the last unit if `t` equals the document's end); else `s`/`para`, `sl`, `sh`/`rows`, `v`, `ref` against the units' anchors, then the fragments' anchors (giving their `unit`). The result is the unit id plus the `char` range and the `xywh` region, if any. A URI whose document reference does not match the file resolves to an error. ## Conformance cases About 20 cases on the existing corpus: `export_csl` and `export_bibtex` for every file in `files/` and `legacy/` (including the anonymous *Lazarillo*, the literal author of the NASA recording and the Chinese title of the *Analects*, which exercises the `anon` fallback), and `resolve` for every anchor type, for a folio printed twice and for a mismatched document reference. ## Backwards compatibility No file changes. Implementations that already export BibTeX may need to change their keys; that is the point. ## Security and privacy None beyond SPEC §14: exports copy metadata the file already exposes. ## Alternatives - Byte-exact BibTeX: rejected; layout differences are harmless and would make the cases brittle. - Leaving the key to each tool: rejected; stable keys across tools are what users need. - Better Bib(La)TeX keys (Better BibTeX style): possible later as an OPTIONAL profile. ## Unresolved questions - Should titles keep their capitalization protected with braces in the structural comparison, or should the comparison ignore braces? - `@online` (biblatex) versus `@misc` for web pages. - Whether `resolve` should also return the fragments that cover the resolved position. ## Implementations None yet. Accepting this RFC requires the cases in `conformance/` and two independent implementations passing them (`governance/RFC-PROCESS.md`). ## Decision (2026-10-07) Accepted, with these changes from the draft above; the normative text is SPEC §5.4 and §19, which prevail: - Keys: when the first author yields no ASCII letter, the first word of `title-short` is used before that of `title` (`lazarillo1554`, not `la1554`); `anon` stays the last fallback; negative years keep their sign. - BibTeX is compared as canonical text, line by line after trimming each line and dropping empty lines; protection braces count (they follow a deterministic rule). `@misc` is the default type; a biblatex profile with `@online` is left for a later RFC. - `locate` returns lists, not a single unit: all matching units (in `ord` order) and all matching fragments (in `n` order), plus `char` and `xywh`. The rule order is `p`, `f`, `t`, `sl`, `v`, `ref`, `s`, `sh`; `s` matches by path prefix unless `para` is given; `v` and `rows` match when their first value falls inside the anchor's range; a time equal to the end of the last timed unit matches it. Fragments are filtered by `char`. A reference to another document yields `document: false`. - Added in the same release: vector search over units and figures (result items carry `unit_id` or `figure_id`) and page-sequence checks for ALTO, TEI and IIIF exports (`export_structure`). --- # SPDF en Rust URL: https://spdf.joseluissaorin.com/es/documentacion/rust > Cómo instalar y usar la implementación de SPDF en Rust (spdf): abrir, validar, buscar y citar. Implementación de referencia; da además la ABI de C. - **Paquete**: `spdf` - **Instalar**: `cargo add spdf` - **Registro**: [crates.io](https://crates.io/crates/spdf) - **Nivel**: primero - **CI**: CI en marcha - **Carpeta**: `rust/` Espacio de trabajo Cargo con tres crates: | crate | qué es | |---|---| | [`crates/spdf`](crates/spdf) | la biblioteca: apertura segura, lectura 5.0 y legado 4.x, validación, volcado canónico, búsqueda léxica, vectorial e híbrida, anclas y URI, cita corta, CSL-JSON y BibTeX, escritor, conversión del legado, integridad y firma, lectura remota por rangos HTTP (`--features http`, experimental) | | [`crates/spdf-tools`](crates/spdf-tools) | la CLI `spdf` | | [`crates/spdf-ffi`](crates/spdf-ffi) | la ABI de C estable (`include/spdf.h`) | SQLite va empaquetado (`rusqlite` con `bundled`): la misma versión y el mismo FTS5 en macOS, Linux, Windows, iOS y Android. ## Uso rápido ```sh cd rust cargo build --release ./target/release/spdf validate ../conformance/files/quijote.spdf ./target/release/spdf search ../conformance/files/quijote.spdf --lexical hidalgo ./target/release/spdf conformance ../conformance > conformance.json ``` Desde otro crate del repositorio (por ejemplo el lector Tauri): ```toml spdf = { path = "../../rust/crates/spdf" } ``` ## Pruebas ```sh cargo test --workspace --all-features # unitarias, de propiedades, de API, remotas, conformidad y doctests cargo clippy --workspace --all-targets --all-features -- -D warnings cargo fmt --all -- --check cargo bench -p spdf --features http # véase RENDIMIENTO.md ``` La prueba `tests/conformance.rs` ejecuta toda la batería de `../conformance/cases`; el CI (`.github/workflows/rust.yml`) la corre en Linux, macOS y Windows y sube el informe como artefacto `conformance-rust`. ## Documentos - [`NOTAS.md`](NOTAS.md): decisiones propias y observaciones para el agente de la especificación. - [`RENDIMIENTO.md`](RENDIMIENTO.md): cifras. - [`PUBLICAR.md`](PUBLICAR.md): cómo publicar en crates.io. Licencia: MIT OR Apache-2.0. --- # SPDF en TypeScript URL: https://spdf.joseluissaorin.com/es/documentacion/js > Cómo instalar y usar la implementación de SPDF en TypeScript (spdf-format): abrir, validar, buscar y citar. Node (node:sqlite), Bun y el navegador (sqlite-wasm). Es la que mueve el validador de esta web. - **Paquete**: `spdf-format` - **Instalar**: `npm install spdf-format` - **Registro**: [npm](https://www.npmjs.com/package/spdf-format) - **Nivel**: primero - **CI**: CI en marcha - **Carpeta**: `js/` *El README de la biblioteca está en inglés.* Read, validate, search, cite and write **SPDF** files from JavaScript and TypeScript, in Node, Bun, Deno, Cloudflare Workers and the browser. SPDF (Semantic Processed Document Format) is an open format for documents that have been read once and can be cited forever: every passage carries its exact anchor (printed page or folio, second of a recording, slide, verse, canonical reference), so a citation can only print what the source says. A `.spdf` file is a plain SQLite 3 database with full-text indexes, optional embedding vectors and CSL-JSON metadata. Specification: [`spec/SPEC.md`](https://github.com/joseluissaorin/spdf/blob/main/spec/SPEC.md). This package is a native, independent implementation of SPDF 5.0 and of the legacy 4.0 and 4.1 formats (Scholaris). - **No required dependencies.** Node uses the built-in `node:sqlite`, Bun uses `bun:sqlite`; in the browser it uses the official SQLite WebAssembly build (`@sqlite.org/sqlite-wasm`, an optional peer dependency). Gzip, SHA-256 and Ed25519 come from the platform (`zlib`, `DecompressionStream`, WebCrypto). - **TypeScript first**: strict types, ESM only, declarations included. - **Conforming reader, semantic reader, writer and validator**, profiles `core`, `semantic` and `media`, with the ALTO, TEI and IIIF exports: it passes the whole SPDF conformance suite (341 cases of suite 0.4.1, every kind, nothing skipped) with `node:sqlite` (Node 22, 24 and 26), with `bun:sqlite`, and with `sqlite-wasm` in Node and in Chromium. - **Remote reading**: in the browser, `openRemote(url)` opens a file with HTTP Range requests and downloads only the pages a query touches (about 1 % of a 23 MiB book for a lexical search; figures below). ```sh npm install spdf-format npm install @sqlite.org/sqlite-wasm # only for the browser, Deno or Workers ``` ## Node Node 22.13 or newer (22.5–22.12 with `--experimental-sqlite`). On Node 24 and later the library works in memory (`serialize`/`deserialize`) and sets the defensive flag and the value-size limit; on Node 22 it goes through a temporary file. ```ts import { openSpdf, validate } from 'spdf-format'; const doc = await openSpdf('quijote.spdf'); // a path, bytes, or a Blob; 5.0 or legacy 4.x (gzip too) console.log(doc.version, doc.document.metadata.title); // '5.0' 'El ingenioso hidalgo…' for (const hit of await doc.searchLexical('lanza en astillero', { limit: 5 })) { console.log(doc.cite(hit.anchor, 'es', hit.anchor_end), hit.anchor_uri); // (Cervantes Saavedra, 1605, p. 23) spdf:sha256-3f2a…#p=29&f=23&char=0,159 } const page = await doc.unitByPrinted('23'); // the page whose printed folio is 23 await doc.close(); const report = await validate('quijote.spdf'); // { valid, version, profile, errors, warnings } ``` Lexical search follows the reference algorithm of the specification: words are OR-ed, quoted phrases (`"…"`, `“…”`, `«…»`, `„…“`) are AND-ed, and case and diacritics are folded by the FTS5 index (`unicode61 remove_diacritics 2`). Queries in Chinese, Japanese or Korean use the `trigram` index when the file has one, and a substring scan otherwise. ```ts const space = await doc.space('embeddinggemma-2@768'); // model, dims, dtype, task prefixes const q = await myModel.embed(`${space.task_prefixes?.query ?? ''}lanza en astillero`); await doc.searchVector(space.id, q, { limit: 5 }); // brute force, f32 / f16 / i8 await doc.searchHybrid('lanza en astillero', q, space.id, { limit: 5 }); // RRF, k = 10 ``` ## Browser ```ts import { configureBrowserEngine, openSpdf, openRemote, openBlob } from 'spdf-format'; // Only if your bundler moves sqlite3.wasm (esbuild, plain copies): tell the engine where it is. configureBrowserEngine({ wasmUrl: '/assets/sqlite3.wasm' }); const fromBytes = await openSpdf(await (await fetch('/quijote.spdf')).arrayBuffer()); const fromFile = await openBlob(fileInput.files[0]); // lazy inside a Worker (FileReaderSync) const remote = await openRemote('https://example.org/quijote.spdf'); // HTTP Range requests console.log(remote.source.stats()); // { requests, bytesFetched, chunksUsed, chunkSize } ``` - `openRemote` reads synchronously inside SQLite's VFS (synchronous `XMLHttpRequest`), so run it in a **Web Worker**; it also works on the main thread, where browsers warn about synchronous requests. Cross-origin servers must allow the `Range` request header and expose `Content-Range` (`Access-Control-Expose-Headers: Content-Range`). If the server ignores ranges, the file is downloaded whole (`fallbackToDownload: false` turns that into an error). Legacy gzip files are always downloaded whole. - `storage: 'opfs'` (or `'auto'`) keeps opened and written databases in the Origin Private File System (`opfs-sahpool` VFS, Workers only) instead of the WebAssembly heap: `configureBrowserEngine({ storage: 'auto' })`. - In Cloudflare Workers, where WebAssembly cannot be compiled at run time, initialize sqlite-wasm yourself with the precompiled module and pass it: `import { wasmEngine } from 'spdf-format/wasm'; openSpdf(bytes, { engine: wasmEngine({ sqlite3 }) })`. ### What a remote search costs Measured in Chromium 153 (headless) inside a Web Worker, on a 23.11 MiB SPDF of the whole *Don Quijote* (Project Gutenberg #2000: 1 382 pages, 2 723 fragments, one 768-dimension f32 space for fragments and pages), each operation on a freshly opened file, so the figures include opening it (`test/browser/run.mjs`): | Operation | Downloaded | Requests | Share of the file | |---|---:|---:|---:| | Open (header, schema, meta, document) | 32 KiB | 8 | 0.14 % | | Lexical search `Rocinante` (10 hits) | 180 KiB | 34 | 0.76 % | | Lexical search `molinos de viento` | 244 KiB | 44 | 1.03 % | | Lexical search `"Dulcinea del Toboso"` (phrase) | 216 KiB | 43 | 0.91 % | | Lexical search `hidalgo de la Mancha` | 268 KiB | 52 | 1.13 % | | The same lexical search again | 0 | 0 | cached | | `unit(700)` and its fragments | 72 KiB | 18 | 0.30 % | | Vector search, 768 × f32, brute force | 10.95 MiB | 811 | 47.4 % | | Hybrid search | 11.27 MiB | 869 | 48.8 % | Pages are fetched in 4 KiB chunks with an adaptive read-ahead (sequential misses double the next request up to 256 KiB) and an LRU cache of 64 MiB. Lexical results are ranked inside the FTS index and only the hits are read from `fragments`, which keeps a search at about 1 % of the file. Vector search is brute force by definition and reads every vector of the space: for remote use, ship an `i8` space (a quarter of the bytes) or query a server. ## Bun ```ts import { openSpdf } from 'spdf-format'; // resolves to bun:sqlite under Bun const doc = await openSpdf('quijote.spdf'); ``` Bun's binding has no switch for `SQLITE_DBCONFIG_DEFENSIVE` nor for value-size limits; read-only mode, `query_only` and `trusted_schema = OFF` still apply. ## Writing ```ts import { SpdfWriter, convertLegacy, generateSigningKey } from 'spdf-format'; const w = await SpdfWriter.create(); await w.setDocument({ id: 'quijote-1605', kind: 'scanned_pdf', mime: 'application/pdf', bytes: 123456, source_sha256: '3f2a…', // SHA-256 of the original metadata: { type: 'book', title: 'El ingenioso hidalgo don Quijote de la Mancha', author: [{ family: 'Cervantes Saavedra', given: 'Miguel de' }], issued: { 'date-parts': [[1605]] } }, }); await w.addUnits([{ id: 'p29', anchor: { type: 'page', physical: 29, printed: '21' }, text: '…', reader: 'pdf-text-layer' }]); await w.addFragments([{ id: 'f1', unit: 'p29', text: '…', anchor: { type: 'page', physical: 29, printed: '21', chars: [0, 159] } }]); const space = await w.addSpace({ provider: 'google', model: 'embeddinggemma-2', dims: 768, dtype: 'i8', modalities: ['text'] }); await w.addVectors(space, [{ target: 'fragment', id: 'f1', vector: embedding }]); // quantized to i8 const { privateKey } = await generateSigningKey(); const bytes = await w.finish({ signWith: privateKey }); // FTS rebuilt, VACUUM, content_sha256 + Ed25519 const v50 = await convertLegacy(legacyBytes); // a Scholaris 4.x file (gzip) as SPDF 5.0 const copy = await SpdfWriter.fromSpdf(doc); // a full 5.0 copy, e.g. to add a vector space ``` ## API at a glance | Area | Functions | |---|---| | Open | `openSpdf(input, opts)`, `openRemote(url)`, `openBlob(blob)` (browser), `validate(input)`, `dump(input)` | | Document | `doc.version`, `doc.legacy`, `doc.meta`, `doc.document`, `units()`, `unit(ord)`, `unitByPrinted()`, `sections()`, `fragments()`, `fragment(id)`, `figures()`, `spaces()`, `vectors()`, `blob(key)`, `blobs()`, `provenance()`, `extensions()`, `dump()`, `validate()` | | Search | `doc.searchLexical(q)`, `doc.searchVector(space, vec)`, `doc.searchHybrid(q, vec, space)` | | Anchors | `formatAnchorUri(docref, anchor, end)`, `parseAnchorUri(uri)`, `locatorToAnchor()`, `checkAnchor()`, `doc.locate(uriOrUrl)` (SPEC §5.4: units, fragments, `char`, `xywh`), `doc.resolve()` | | Citation | `cite(anchor, document, locale, end)`, `doc.cite(anchor, locale, end)` (es, en), `doc.citePassage(fragmentId, quote, locale)` (SPEC §18.2: cites the unit the quotation is in) | | Export | `toCslJson()`, `toCslJsonArray()`, `cslCitationItem()`, `toBibtex()` (keys `cervantessaavedra1605`, SPEC §19), `toAlto()`, `toTei()`, `toIiif()` | | Write | `SpdfWriter.create()`, `.fromSpdf()`, `.fromSource()`, `convertLegacy()`, `encodeVector()` | | Integrity | `doc.contentSha256()`, `doc.verifyIntegrity(publicKey?)`, `verifySignature()`, `generateSigningKey()` | | Engines | `nodeEngine()`, `bunEngine()`, `wasmEngine()`, `setDefaultEngine()`; the port is `SqlEngine`/`SqlConnection` | Entry points: `spdf-format` (picks Node, Bun or the browser by export condition), `spdf-format/node`, `spdf-format/bun`, `spdf-format/browser`, `spdf-format/core` (no engine; pass `{ engine }`), `spdf-format/wasm` and `spdf-format/cli`. Media type: `application/vnd.spdf+sqlite3` (`MEDIA_TYPE`); a `.spdf` URL may carry the anchor parameters as its fragment (`https://example.org/quijote.spdf#p=5&f=1r`), which `doc.locate()` accepts as well as `spdf:` URIs. ## Command line ```sh npx spdf-format validate quijote.spdf # exit status 1 if invalid npx spdf-format dump quijote.spdf --pretty # canonical JSON dump (JCS without --pretty) npx spdf-format search quijote.spdf molinos de viento --limit 5 npx spdf-format cite quijote.spdf f12 --locale en npx spdf-format cite quijote.spdf --bibtex npx spdf-format convert legacy-4.1.spdf out.spdf --hash npx spdf-format conformance path/to/spdf/conformance ``` ## Safety Files come from strangers. Every file is opened read-only, with `query_only`, `trusted_schema = OFF`, `mmap_size = 0`, `cell_size_check = ON`, the defensive flag where the binding has it, no extensions, a 512 MiB limit for any single value (`SQLITE_LIMIT_LENGTH` with `node:sqlite` on Node 24+ and with sqlite-wasm; checked by the library for blobs elsewhere) and a 4 GiB limit for gzip input. Files with triggers, views or virtual tables other than the FTS5 indexes are refused (`E020`), except the three FTS triggers of legacy 4.x files. Write operations of the validator (the FTS `integrity-check`) run on a private in-memory copy. ## Conformance `npm test` runs the unit tests and the whole shared suite with `node:sqlite` and `sqlite-wasm`; `npm run test:bun` runs it with `bun:sqlite`; `npm run test:browser` runs it in headless Chromium (Playwright, `--mute-audio`) together with the remote-reading, OPFS and Blob tests. CI uploads the runner report as the `conformance-js` artifact. ## License MIT OR Apache-2.0, at your option. --- # SPDF en Python URL: https://spdf.joseluissaorin.com/es/documentacion/python > Cómo instalar y usar la implementación de SPDF en Python (spdf-format): abrir, validar, buscar y citar. Solo sqlite3 de la biblioteca estándar; se importa como spdf. - **Paquete**: `spdf-format` - **Instalar**: `pip install spdf-format` - **Registro**: [PyPI](https://pypi.org/project/spdf-format/) - **Nivel**: primero - **CI**: CI en marcha - **Carpeta**: `python/` *El README de la biblioteca está en inglés.* Read, validate, search, cite and write **SPDF** files from Python. SPDF (Semantic Processed Document Format) is an open format for documents that have been read once and can be cited forever: every passage carries its exact anchor (printed page or folio, second of a recording, slide, verse, canonical reference), so a citation can only print what the source says. A `.spdf` file is a plain SQLite 3 database with full-text indexes, optional embedding vectors and CSL-JSON metadata. This package is a native, independent implementation of SPDF 5.0 and of the legacy 4.0 and 4.1 formats. It is installed as `spdf-format` and imported as `spdf`. - Python 3.10 or newer, **standard library only** (`sqlite3`). - Optional extras: `numpy` (fast vector search), `crypto` (Ed25519 through `cryptography`; a pure-Python fallback is included), `pandas`, `arrow`. - Passes the whole SPDF conformance suite (0.4.1, 341 cases): reader, semantic reader, writer and validator, profiles `core`, `semantic` and `media`. ```sh pip install spdf-format # or: uv add spdf-format pip install "spdf-format[numpy]" # faster vector search ``` ## Quick start ```python import spdf with spdf.open("quijote.spdf") as f: # 5.0, or legacy 4.x (gzip-wrapped too) doc = f.document # spdf.Document, metadata is CSL-JSON print(doc.display_title, f.cite()) # El ingenioso hidalgo… (Cervantes Saavedra, 1605) for hit in f.search("lanza en astillero"): print(f.cite(hit.fragment), hit.anchor_uri) # (Cervantes Saavedra, 1605, p. 23) spdf:sha256-3f2a…#p=29&f=23&char=0,159 unit = f.unit_by_printed("23") # the page whose printed folio is 23 print(unit.text) ``` Lexical search follows the reference algorithm of the specification: words are OR-ed, quoted phrases (`"…"`, `“…”`, `«…»`, `„…“`) are AND-ed, and case and diacritics are folded by the FTS5 index (`unicode61 remove_diacritics 2`). Queries in Chinese, Japanese or Korean use the `trigram` index when the file has one, and a substring scan otherwise. ```python with spdf.open("darwin.spdf") as f: space = f.space("embeddinggemma-2@768") # model, dims, dtype, task prefixes qvec = my_model.encode(space.task_prefixes["query"] + "natural selection") f.search_vector(qvec, space=space.id, limit=5) f.search_hybrid("natural selection", qvec, space=space.id) # RRF, k = 10 f.search_vector(qvec, space=space.id, target="unit") # hits carry unit_id ``` Every hit has `id`, `target` (`fragment`, `unit` or `figure`), `score`, `via`, `anchor` and `anchor_uri`; `fragment_id`, `unit_id` and `figure_id` give the id for each target. ## Validate ```python report = spdf.validate("quijote.spdf") report.valid # True when there are no errors report.codes # {"W102"} … report.to_dict() # {"valid", "version", "profile", "errors", "warnings"} as in the spec ``` Validation follows the order of the specification and reports every code it finds: `E001` not SQLite, `E002` unknown version, `E003` gzip-wrapped 5.0 (warning), `E010`/`E011` missing table or column, `E012` missing metadata key, `E013` not exactly one document, `E020` trigger, view or foreign virtual table, `E030`–`E032` vectors and spaces, `E040`–`E042` anchors, `E050`/`E051` metadata, `E060` unknown required extension, `E070` FTS index out of sync, `E080` blob hash, `E081`/`E082` content hash and signature, `E090` unit order, and the warnings `W100`–`W110` (`W103`: a fragment crosses from one kind of `matter` to another, such as body text into a plate, or from a page with a printed folio to one without; writers should set `matter` on the units of paged documents and not let fragments cross). ## Write ```python with spdf.Writer("out.spdf") as w: w.add_document({ "id": "quijote", "kind": "pdf", "mime": "application/pdf", "bytes": 1203456, "source_sha256": "3f2a…", "source_ref": w.add_blob("original.pdf", "application/pdf", pdf_bytes), "metadata": {"type": "book", "title": "El ingenioso hidalgo don Quijote de la Mancha", "author": [{"family": "Cervantes Saavedra", "given": "Miguel de"}], "issued": {"date-parts": [[1605]]}, "language": "es"}, }) page = {"type": "page", "physical": 29, "printed": "23", "roman": False, "source": "read"} w.add_unit({"id": "u29", "ord": 1, "anchor": page, "text": "En un lugar de la Mancha…", "reader": "pdf-text-layer"}) w.add_fragment({"id": "f1", "unit": "u29", "ord": 1, "text": "En un lugar de la Mancha…", "anchor": {**page, "chars": [0, 24]}}) w.add_space({"id": "embeddinggemma-2@768", "provider": "google", "model": "embeddinggemma-2", "dims": 768, "modalities": ["text"]}) w.add_vector("fragment", "f1", "embeddinggemma-2@768", data=vector) # list, ndarray or bytes ``` The writer builds the file next to its destination, rebuilds the FTS index (distributed files carry no triggers), fills the required metadata keys (`spdf_version`, `profile`, `created`, `generator`, `document_id`), computes `content_sha256`, can sign it (`finalize(sign_key=…)`), compacts it with `VACUUM`, validates it and only then moves it into place. Text is normalized to NFC. `i8` and `f16` vectors are quantized as the specification says. `spdf.convert_legacy("old.spdf", "new.spdf")` converts a 4.x file to 5.0, and `spdf.write_source(full_dump, path)` rebuilds a file from a full dump (`f.full_dump()`). ## Anchors, URIs and citations ```python uri = spdf.make_uri("sha256-3f2a…", {"type": "page", "physical": 29, "printed": "21", "chars": [118, 301]}) # 'spdf:sha256-3f2a…#p=29&f=21&char=118,301' spdf.parse_uri(uri) # {"docref": "sha256-3f2a…", "locator": {"p": 29, "f": "21", "char": [118, 301]}} spdf.cite({"type": "page", "physical": 9, "printed": "1r", "foliation": "leaf"}, doc, locale="es") # '(Cervantes Saavedra, 1605, fol. 1r)' spdf.cite({"type": "time", "t0": 4160.0, "t1": 4175.5}, doc, locale="en") # '(Cortázar, 1959, 1:09:20)' ``` `f.locate(reference)` resolves an anchor URI, or the URL of a `.spdf` with an anchor fragment, against the file (SPEC §5.4): ```python f.locate("https://example.org/quijote.spdf#p=5&pe=6&char=101,278").to_dict() # {"document": True, "units": ["p5", "p6"], "fragments": ["q4"], "char": [101, 278], "xywh": None} ``` A quotation is cited by the unit it lies in, never by the start of the fragment that contains it (SPEC §18.2), so a passage on page 211 of a fragment that begins on an unnumbered plate cites page 211: ```python c = f.cite_passage("m4", "tube N N", locale="es") c.text # '(Hooke, 1665, p. 211)' c.uri # 'spdf:sha256-…#p=321&f=211&char=0,8' ``` Citations print only what the anchor says: inferred folios in brackets (`p. [21]`), unnumbered pages as `s. p.` / `n. pag.`, leaves and columns (`fol. 1r`, `col. 45`), `h:mm:ss` times, slides, sheets, verses and canonical references. Spanish uses `e` instead of `y` before the sound /i/ (`Gómez e Iglesias`). ## Bibliography and other exports | Export | API | CLI | |---|---|---| | CSL-JSON (Zotero, citeproc, Pandoc) | `f.to_csl_json()`, `spdf.bibliography.csl_citation_item()` | `spdf export -f csl` | | BibTeX | `f.to_bibtex()` | `spdf export -f bibtex` | | ALTO 4 XML (page units) | `f.to_alto()` | `spdf export -f alto` | | TEI P5 (minimal: header, `pb`, `p`, `lg`/`l`, `u`, `note`) | `f.to_tei()` | `spdf export -f tei` | | IIIF Presentation 3 manifest | `f.to_iiif(base_url)` | `spdf export -f iiif --base-url URL [--images-dir DIR]` | | JSON Lines of fragments | `spdf.interop.frames.fragment_records(f)` | `spdf export -f jsonl` | | pandas DataFrame | `f.to_pandas(vectors="space id")` | | | Arrow table | `f.to_arrow(vectors="space id")` | | | Canonical dump (JCS) | `f.dump()`, `f.dump_json()` | `spdf dump` | ALTO has no invented coordinates: SPDF stores text per unit, so blocks and lines carry none (they are optional in ALTO 4). The IIIF manifest paints each canvas with the page image, adds the unit text as a `supplementing` annotation, figure descriptions as `describing` annotations on `#xywh=percent:` regions, and sections as ranges; audio and video become one time-based canvas with a range per unit. ## Annotations and collections User annotations and libraries live outside the documents, as the specification's sidecar files: ```python from spdf import sidecars with spdf.open("quijote.spdf") as f: hit = f.search("lanza en astillero")[0] note = sidecars.annotation(f, hit.fragment, body="Origen del tópico.") # W3C Web Annotation sidecars.write_annotations("notas.spdfa.json", [note], label="Notas de lectura") lib = sidecars.library(["quijote.spdf", "lazarillo.spdf"], "Tesis: fuentes") sidecars.write_library("fuentes.spdfl.json", lib) ``` Each annotation targets the document by identity (`spdf:sha256-…`) with an `SpdfAnchorSelector` (the anchor URI) and a `TextQuoteSelector` (exact text with a little context), so it survives a re-reading that shifts offsets. ## Command line ```text spdf validate FILE… [--json] exit status 1 if a file is invalid spdf dump FILE [--pretty] canonical dump (RFC 8785) spdf info FILE | --env summary; --env shows the SQLite and FTS5 in use spdf search FILE QUERY [--vector JSON --space ID] [--mode lexical|vector|hybrid] [--json] spdf cite FILE [--fragment ID [--quote TEXT] | --unit ID | --uri URI] [--locale es|en] [--bibtex] spdf export FILE -f csl|bibtex|alto|tei|iiif|jsonl spdf convert OLD.spdf NEW.spdf legacy 4.x to 5.0 spdf sign FILE --key KEY / spdf verify FILE [--public-key ed25519:…] spdf conformance [DIR] run the conformance suite, print the report ``` Other packages can add subcommands through the `spdf.commands` entry point group: the entry point is a callable `register(subparsers)` that adds an `argparse` parser and sets `func` (a handler returning the exit status). The SPDF producer adds `spdf build` this way: ```toml [project.entry-points."spdf.commands"] build = "spdf_build.cli:register" ``` ## Safety Files are untrusted input. `spdf.open` checks the SQLite header before SQLite sees the file, decompresses gzip input to a private temporary file with a size limit (4 GiB by default), opens the database read-only by URI (`mode=ro`) with `query_only`, `trusted_schema=OFF`, `cell_size_check` and, on Python 3.12 or newer, `SQLITE_DBCONFIG_DEFENSIVE`; it never loads extensions, refuses triggers, views and virtual tables other than the format's FTS5 indexes (legacy files may keep their three FTS triggers), refuses files that require unknown extensions, and enforces a maximum blob size (512 MiB by default, `max_blob_size=`). ## FTS5 and Python builds Lexical search, writing and the FTS integrity check need SQLite's FTS5, which depends on how Python was built. It is present in the python.org installers, Homebrew, `uv python install` (python-build-standalone), conda and the usual Linux distributions; CI checks Linux, macOS and Windows with Python 3.10 to 3.13. `spdf info --env` tells you what you have. Without FTS5, reading, vector search, citations, exports and validation still work, and the operations that need it raise `spdf.Fts5UnavailableError` explaining what to do; on Linux, `pip install pysqlite3-binary` is picked up automatically as a fallback driver. ## Conformance The SPDF conformance suite lives in `conformance/` in the [repository](https://github.com/joseluissaorin/spdf). Run it with: ```sh spdf conformance path/to/conformance # prints {"impl", "version", "passed", "failed", "skipped"} ``` This implementation claims every kind of case: `dump`, `legacy_dump`, `roundtrip`, `validate`, `search_lexical`, `search_vector`, `search_hybrid`, `anchor_uri`, `cite`, `quantize`, `locate`, `cite_passage`, `export_csl`, `export_bibtex` and `export_structure` (checked on its own ALTO, TEI and IIIF output). CI publishes its report as the `conformance-python` artifact. ## Development ```sh cd python uv sync --group dev uv run pytest uv run ruff check src tests && uv run ruff format --check src tests uv run mypy ``` ## License Code: MIT OR Apache-2.0, at your option. The SPDF specification is CC BY 4.0. --- # SPDF en Swift URL: https://spdf.joseluissaorin.com/es/documentacion/swift > Cómo instalar y usar la implementación de SPDF en Swift (SPDF): abrir, validar, buscar y citar. Swift Package Manager, desde el repositorio. Plataformas de Apple y Linux. - **Paquete**: `SPDF` - **Instalar**: `.package(url: "https://github.com/joseluissaorin/spdf", from: "5.0.0")` - **Nivel**: primero - **CI**: CI en marcha - **Carpeta**: `swift/` *El README de la biblioteca está en inglés.* Native Swift implementation of **SPDF** (Semantic Processed Document Format): documents that have been read once and can be cited forever, because every passage carries its exact anchor (printed page, folio, second of a recording, slide, verse). - SwiftPM package `SPDF` for iOS 16+, macOS 13+, tvOS 16+, watchOS 9+, visionOS 1+ and Linux. - Uses the **system SQLite** (FTS5, `unicode61 remove_diacritics 2` and `trigram` are present on iOS and macOS; verified on the iOS simulator and on macOS). On Linux it links the distribution's `libsqlite3` (install `libsqlite3-dev` and `zlib1g-dev`). - No third-party code on Apple platforms (CryptoKit for SHA-256 and Ed25519, zlib for gzip). On Linux, [`swift-crypto`](https://github.com/apple/swift-crypto) provides the same CryptoKit API. - API with `async` variants, `Sendable` and `Codable` types; `SPDFFile` is thread-safe. - Conformance: passes the whole SPDF conformance suite (`../conformance`), every kind, on macOS and on the iOS simulator. ## Installation ```swift .package(url: "https://github.com/joseluissaorin/spdf", from: "0.1.0") // target dependency: .product(name: "SPDF", package: "spdf") ``` The repository root carries a thin `Package.swift` pointing into `swift/`, so the URL above is all SwiftPM needs; versions follow the monorepo-wide tags (`0.1.0`, …). ## Reading and searching ```swift import SPDF let file = try await SPDFFile.open(url) // 5.0, or a legacy 4.x .spdf (gzip) defer { file.close() } for hit in try await file.searchLexical("«lugar de la Mancha»", limit: 5) { let cite = try file.cite(hit.anchor!, end: hit.anchorEnd, locale: "es") print(hit.id, hit.score, hit.anchorURI, cite) // q4 1.889394 spdf:sha256-fa38…#p=5&pe=6&f=1r&fe=1v&char=101,278 (Cervantes Saavedra, 1605, fols. 1r-[1v]) } // Vector and hybrid search with your own query embedding (dot product when the // space is normalized, cosine otherwise). let vector: [Double] = embed(query) let semantic = try await file.searchVector(vector, space: "embeddinggemma-2@768", limit: 10) let hybrid = try await file.searchHybrid(query, vector: vector, space: "embeddinggemma-2@768", limit: 10) let units = try await file.units() // pages, time spans, slides… let bib = try file.exportBibTeX() // keys of SPEC §19: cervantessaavedra1605… let csl = try file.exportCSL() let alto = try file.exportALTO() // also exportTEI(), exportIIIF() // Cite a quotation by the unit it lies in (SPEC §18.2), not by its fragment's start. let cited = try file.citePassage(fragment: "m4", quote: "XXXIV.\ntube N N", locale: "es") // cited.text == "(Hooke, 1665, p. 211)", cited.uri == "spdf:sha256-ba9d…#p=319&pe=321&fe=211" // Resolve a reference (SPEC §5.4): an spdf: URI or a .spdf URL with a fragment. let where = try file.locate("https://example.org/quijote.spdf#p=7") // units ["p7"], fragments ["q5"] ``` Every method also has a synchronous form (`try file.searchLexical(…)`), handy in command-line tools and tests. ## Validation and canonical dump ```swift let report = SPDFValidator.validate(url) // ValidationResult: Codable print(report.valid, report.errorCodes, report.warningCodes) let json = try file.dumpJSON() // RFC 8785 canonical JSON let hash = try file.contentSHA256() // integrity hash (§8) ``` Opening is defensive: read-only, `query_only`, `trusted_schema=OFF`, `SQLITE_DBCONFIG_DEFENSIVE`, no extensions; files with triggers, views or foreign virtual tables (E020) or unknown required extensions (E060) are refused; strings and blobs are capped (512 MiB) and so is gunzipped input (4 GiB). ## Anchors and citations ```swift let anchor = Anchor(["type": "page", "physical": 29, "printed": "21", "source": "inferred"]) let uri = AnchorURI.format(docref: "sha256-3f2a…", anchor: anchor) // spdf:sha256-3f2a…#p=29&f=21 let parsed = try AnchorURI.parse(uri) // strict parser let text = Citation.cite(anchor, metadata: [ "type": "book", "title": "Arte nuevo de hacer comedias", "author": [["family": "Vega", "non-dropping-particle": "de", "given": "Lope"]], "issued": ["date-parts": [[1609]]], ], locale: "es") // (de Vega, 1609, p. [21]) ``` ## Writing ```swift let writer = try SPDFWriter(url: out, options: .init(generator: "my-app/1.0")) try writer.setDocument(SPDFDocumentInfo(id: "doc", kind: "pdf", metadata: ["type": "book", "title": "…"], sourceSHA256: sha, mime: "application/pdf", bytes: size)) let page = Anchor(["type": "page", "physical": 1, "printed": "1"]) try writer.add(SPDFUnit(id: "u1", anchor: page, text: "…", reader: "pdf-text-layer")) try writer.add(SPDFFragment(id: "f1", unit: "u1", text: "…", anchor: page)) try writer.add(VectorSpace(id: "embeddinggemma-2@768", provider: "local", model: "embeddinggemma-2", dims: 768)) try writer.addVector(target: .fragment, id: "f1", space: "embeddinggemma-2@768", values: embedding) // [Float] try writer.addBlob(key: "pages/0001.png", mime: "image/png", data: png) try writer.finish() // FTS rebuilt, content_sha256 written, VACUUM, atomic replace; no triggers ``` Every file the writer produces carries `spdf_meta.content_sha256`; pass `.init(signingKey: seed)` (a 32-byte Ed25519 seed) to sign it too (`signer`, `signature`, SPEC §8). Files signed this way verify with the Go, Rust, Python, C# and JavaScript implementations. `SPDFSeal.seal(url, signingKey:)` hashes and signs an existing file in place. `SPDFSource.write(_:to:)` builds a file from a full JSON dump (the format of `conformance/sources/`). ## Command line ```sh swift run spdf-swift validate file.spdf swift run spdf-swift dump file.spdf swift run spdf-swift search file.spdf "lugar de la Mancha" -n 5 swift run spdf-swift cite file.spdf q4 -locale en swift run spdf-swift export file.spdf bibtex swift run spdf-swift uri parse 'spdf:sha256-…#p=29&f=21' swift run spdf-swift build source.json out.spdf swift run spdf-swift seal out.spdf -key seed.hex swift run spdf-swift export file.spdf tei # also alto, iiif swift run spdf-swift conformance ../conformance -o conformance.json ``` ## Tests and conformance ```sh swift test # unit tests + the whole suite xcodebuild test -scheme SPDF-Package -destination 'platform=iOS Simulator,name=iPhone 17' ``` The runner (`SPDFConformance`) discovers the cases by listing `conformance/cases/*.json` and prints `{"impl","version","passed","failed","skipped"}`. Nothing is skipped: this is a full (reader and writer) implementation. ## License MIT OR Apache-2.0. --- # SPDF en Kotlin / JVM URL: https://spdf.joseluissaorin.com/es/documentacion/kotlin > Cómo instalar y usar la implementación de SPDF en Kotlin / JVM (io.github.joseluissaorin:spdf): abrir, validar, buscar y citar. Kotlin y Java sobre la JVM y Android. - **Paquete**: `io.github.joseluissaorin:spdf` - **Instalar**: `implementation("io.github.joseluissaorin:spdf:5.0.0")` - **Registro**: [Maven Central](https://central.sonatype.com/artifact/io.github.joseluissaorin/spdf) - **Nivel**: primero - **CI**: CI en marcha - **Carpeta**: `kotlin/` *El README de la biblioteca está en inglés.* Native Kotlin/JVM implementation of **SPDF** (Semantic Processed Document Format): documents that have been read once and can be cited forever, because every passage carries its exact anchor (printed page, folio, second of a recording, slide, verse). It is usable from Kotlin, Java and Android, and independent of the other implementations in this repository (it does not wrap the Rust ABI). - Maven coordinates: `io.github.joseluissaorin:spdf` (JVM), `io.github.joseluissaorin:spdf-android` (Android), both on top of `io.github.joseluissaorin:spdf-core`. - Java 17 or newer on the JVM; Android minSdk 23. Built with Kotlin 2.4 at language level 2.2, so apps on Kotlin 2.1 or newer can consume it. - Conformance: passes the whole SPDF conformance suite (`../conformance`, 309 cases in suite 0.4.0) with **both** SQLite adapters, on the JVM and on Android emulators (API 31 and 36), every kind (`dump`, `legacy_dump`, `roundtrip`, `validate`, `search_lexical`, `search_vector`, `search_hybrid`, `anchor_uri`, `cite`, `quantize`, `locate`, `export_csl`, `export_bibtex`, `export_structure`). Nothing is skipped: this is a full reader and writer. ## Modules | Artifact | What | SQLite | |---|---|---| | `spdf-core` | all the logic: safe opening, dump, validation, search, anchors, citation, export, writer, conformance runner | none: it talks to a small `SqlDriver` interface with typed values (`SqlValue`: NULL, INTEGER, REAL, TEXT, BLOB) | | `spdf` | `JdbcSqlDriver` and the `spdf` command-line tool | `org.xerial:sqlite-jdbc` (bundled SQLite with FTS5 and `trigram`) | | `spdf-android` | `AndroidxSqlDriver` | `androidx.sqlite` driver API with `BundledSQLiteDriver` (`androidx.sqlite:sqlite-bundled`): the app ships its own SQLite with FTS5 and `trigram`, which `android.database.sqlite` does not guarantee | Both adapters register themselves with `java.util.ServiceLoader`, so `SpdfFile.open(path)` picks the one on the classpath. You can always pass a driver explicitly (recommended on Android, where R8 may strip service files). Any other binding can be plugged in by implementing `SqlDriver` (three methods). ```kotlin // build.gradle.kts dependencies { implementation("io.github.joseluissaorin:spdf:0.1.0") // JVM // implementation("io.github.joseluissaorin:spdf-android:0.1.0") // Android } ``` ```xml io.github.joseluissaorin spdf 0.1.0 ``` ## What it does | | | |---|---| | Safe opening | read-only (`SQLITE_OPEN_READONLY`), `query_only`, `trusted_schema = OFF`, `mmap_size = 0`, `cell_size_check`, no extensions; files with triggers, views or foreign virtual tables refused (E020); unknown required extensions refused (E060); maximum value size (512 MiB) and maximum gunzipped size (4 GiB); WAL files are read through a copy with the header patched | | Versions | SPDF 5.0, and the legacy 4.0 / 4.1 files of Scholaris (Spanish schema, gzip-wrapped) through the 5.0 view, metadata mapped to CSL-JSON | | Dump | canonical JSON (RFC 8785 / JCS) of the whole file, `content_sha256` | | Validation | every code of SPEC §22 (E001–E090, W100–W110) in the reference order; FTS `integrity-check` on a private copy; Ed25519 signatures (platform provider, with a pure fallback for older Android) | | Search | lexical (FTS5 BM25, CJK route with `trigram` or substring), vector (`f32`, `f16`, `i8`; dot product or cosine), hybrid (reciprocal rank fusion, k = 10) | | Anchors | anchor ↔ URI (`spdf:sha256-…#p=29&f=21&char=118,301`), strict parser, canonical form | | Resolution | `locate(reference)`: an anchor URI, or the URL of a `.spdf` with the anchor as fragment (`https://…/quijote.spdf#p=5&f=1r`), resolved to units, fragments, `char` and `xywh` (SPEC §5.4) | | Citation | short author-date citation in Spanish and English | | Export | CSL-JSON (with the CSL `label`/`locator` of a cited passage) and BibTeX, one or several documents, with the key and field rules of SPEC §19 (same keys as every other implementation: `cervantessaavedra1605`, `lazarillo1554`, `anonnd`); ALTO 4, a minimal TEI and a IIIF Presentation 3 manifest (SPEC §19.4), with no invented coordinates or dimensions | | Writer | builds valid SPDF 5.0 files (FTS kept in sync, `VACUUM`, no triggers, atomic replace); f16/i8 quantization as the spec says; writes `content_sha256` by default and, given an Ed25519 key, `signer` and `signature` (SPEC §8) | | Typed reading | `document()`, `units()`, `fragments()`, `sections()`, `figures()`, `spaces()`, `provenance()`, `blob(key)` | ## Kotlin ```kotlin import io.github.joseluissaorin.spdf.* SpdfFile.open("quijote.spdf").use { f -> // also legacy .spdf (gzip) files for (hit in f.searchLexical("«lugar de la Mancha»", limit = 5)) { println("${hit.fragmentId} ${hit.score} ${hit.anchorUri}") println(f.cite(hit.anchor!!, hit.anchorEnd, locale = "es")) // q4 1.889394 spdf:sha256-fa38…#p=5&pe=6&f=1r&fe=1v&char=101,278 // (Cervantes Saavedra, 1605, fols. 1r-[1v]) } val query = DoubleArray(8) // your own query embedding val nearest = f.searchVector(query, space = "toy-embedding@8", limit = 5) val fused = f.searchHybrid("hidalgo", query, "toy-embedding@8", limit = 5) println(f.exportBibTeX()) // Resolution of a reference (SPEC §5.4) and structural exports (SPEC §19.4) val loc = f.locate("https://example.org/quijote.spdf#p=5&char=101,278") println("${loc.units} ${loc.fragments}") // [p5] [q4] val alto: String = f.exportAlto() val tei: String = f.exportTei() val manifest: String = f.exportIiif(base = "https://example.org/quijote") val citeproc = f.exportCslJson(Anchor.page(5, "1r", foliation = "leaf")) // with "label": "folio", "locator": "1r" } // Several documents in one bibliography (keys disambiguated with a, b, c…) val items = listOf("a.spdf", "b.spdf").map { p -> SpdfFile.open(p).use { it.metadata() } } println(Export.bibtex(items)) val csl = Json.compact(Export.cslItems(items)) // Validation and canonical dump val result = Validator.validate("file.spdf") println("${result.isValid} ${result.errorCodes} ${result.warningCodes}") SpdfFile.open("file.spdf").use { f -> val jcs: String = f.dumpJson() // RFC 8785 bytes val sum: String = f.contentSha256() // integrity hash of SPEC §18 } // Anchors and citations val a = Anchor.page(29, "21", source = "inferred") val uri = AnchorUri.format("sha256-3f2a…", a) // spdf:sha256-3f2a…#p=29&f=21 val parsed = AnchorUri.parse(uri) // strict: malformed URIs throw SpdfException val text = Citation.cite(a, null, mapOf( "type" to "book", "title" to "Arte nuevo de hacer comedias", "author" to listOf(mapOf("family" to "Vega", "non-dropping-particle" to "de", "given" to "Lope")), "issued" to mapOf("date-parts" to listOf(listOf(1609))), ), "es") // (de Vega, 1609, p. [21]) // Writing // content_sha256 is written by default; a 32-byte Ed25519 seed also signs the file. SpdfWriter.create(File("out.spdf"), WriterOptions(generator = "my-tool/1.0", signingKey = seed)).use { w -> w.setDocument(Document("doc", "pdf", "application/pdf", sha256Hex, mapOf("type" to "book", "title" to "…"))) w.addUnit(CitableUnit("u1", Anchor.page(1, "1"), "…", reader = "pdf-text-layer")) w.addFragment(Fragment("f1", "u1", "…", Anchor.page(1, "1"))) w.addSpace(Space("embeddinggemma-2@768", "local", "embeddinggemma-2", 768)) w.addVector("fragment", "f1", "embeddinggemma-2@768", embedding) // FloatArray or DoubleArray w.addBlob("pages/0001.png", "image/png", png) w.finish() // rebuilds FTS, VACUUM, atomic move; closing without finish() discards the file } ``` `Sources.write(source, file)` builds a file from a full JSON dump (the format of `conformance/sources/`). ## Java ```java import io.github.joseluissaorin.spdf.*; import io.github.joseluissaorin.spdf.sql.SqlDriver; import java.util.List; import java.util.Map; try (SpdfFile f = SpdfFile.open("quijote.spdf")) { List hits = f.searchLexical("lugar de la Mancha", 5); for (Hit h : hits) { System.out.println(h.getFragmentId() + " " + h.getAnchorUri()); System.out.println(f.cite(h.getAnchor(), h.getAnchorEnd(), "en")); } } ValidationResult r = Spdf.validate("quijote.spdf"); System.out.println(r.isValid() + " " + r.getErrorCodes()); try (SpdfWriter w = SpdfWriter.create("out.spdf")) { Document d = new Document("lazarillo", "pdf", "application/pdf", sha256Hex, Map.of("type", "book", "title", "La vida de Lazarillo de Tormes")); d.setTitle("Lazarillo de Tormes"); w.setDocument(d); w.addUnit(new CitableUnit("u1", Anchor.page(3, "A2r", "read", "leaf"), "Pues sepa Vuestra Merced", "pdf-text-layer")); Fragment frag = new Fragment("f1", "u1", "Pues sepa Vuestra Merced", Anchor.page(3, "A2r", "read", "leaf")); frag.setContext("Prólogo"); w.addFragment(frag); w.finish(); } AnchorUri.Parsed p = Spdf.parseUri("spdf:sha256-3f2a…#p=29&f=21"); SqlDriver driver = SqlDriver.defaultDriver(); // `default` is a Java keyword Location where = SpdfFile.open("quijote.spdf").locate("spdf:sha256-fa38…#f=1v"); String cite = Spdf.cite(Anchor.verse(12), null, Json.parseObject("{\"title\":\"Rimas\",\"author\":[{\"family\":\"Bécquer\"}]}"), "es"); ``` Every entry point is static for Java (`@JvmStatic`), optional parameters have overloads (`@JvmOverloads`), rows are classes with a constructor for the required members and setters for the rest, and nothing is `suspend`. `src/test/java` in the `spdf` module checks this. From Java the default adapter is `SqlDriver.defaultDriver()`. ## Android ```kotlin val driver = AndroidxSqlDriver() // BundledSQLiteDriver inside val options = OpenOptions(tempDir = context.cacheDir) // gunzipped and private copies go here SpdfFile.open(File(context.filesDir, "book.spdf"), driver, options).use { f -> val hits = f.searchLexical("golondrinas") } ``` `spdf-android` is a plain JVM library on purpose: it needs no Android SDK to build, and its tests run the androidx bundled driver on the host JVM (the artifact ships Linux, macOS and Windows natives too), including the whole conformance suite. Android apps consume it like any jar; Gradle resolves the Android variant of `androidx.sqlite:sqlite-bundled` (with the native library for each ABI) for them. The core avoids JDK APIs newer than Android API 23 (a test checks the bytecode). On a real Android runtime, the test-only module `spdf-android-device` (Android Gradle Plugin 9.4, Java instrumented test, never published) packages `../conformance` as test assets and runs the whole suite on a device or emulator. It is not part of the default build; enable it explicitly: ```sh # with an emulator or device attached (ANDROID_HOME pointing at the SDK) ANDROID_SERIAL=emulator-5584 ./gradlew -Pspdf.androidDevice=true :spdf-android-device:connectedAndroidTest ``` Results so far: 309/309 on Android 12 (API 31) and Android 16 (API 36) arm64 emulators. CI runs it on x86_64 emulators (API 31 and 35) in a separate job. ## Safety notes SPDF files come from strangers (SPEC §2.4, §14). What each adapter enforces: | | sqlite-jdbc | androidx.sqlite bundled | |---|---|---| | `SQLITE_OPEN_READONLY` | yes | yes | | `query_only`, `trusted_schema = OFF`, `mmap_size = 0`, `cell_size_check` | yes (set and checked by the core) | yes (set and checked by the core) | | extension loading | off | never enabled (no `addExtension`) | | `SQLITE_DBCONFIG_DEFENSIVE` | not exposed by sqlite-jdbc | not exposed by the driver API | | maximum value size | `SQLITE_LIMIT_LENGTH` | not exposed: no value can exceed the file size, so files larger than the limit are measured once when opened | `SpdfFile.safety` reports what was in force. Without the defensive flag, the read-only, query-only connection still refuses every write, and files with triggers, views or foreign virtual tables are refused before any query on user tables. The FTS `integrity-check` (a write) runs on a private temporary copy. ## Command line ```sh ./gradlew :spdf:installDist B=spdf/build/install/spdf/bin/spdf $B validate file.spdf $B dump file.spdf $B search file.spdf "lugar de la Mancha" -n 5 $B vsearch file.spdf toy-embedding@8 0,0.5,0.25,0.75,0.25,0,0.25,0 -n 3 $B hybrid file.spdf toy-embedding@8 0,0.5,0.25,0.75,0.25,0,0.25,0 selection -n 3 $B cite file.spdf q4 --locale en $B export file.spdf bibtex # also csl, alto, tei, iiif $B locate file.spdf 'spdf:sha256-…#f=1v' $B uri parse 'spdf:sha256-…#p=29&f=21' $B build source.json out.spdf $B conformance ../conformance -o conformance.json ``` ## Conformance ```sh ./gradlew build # unit tests; both adapters run the whole suite ./gradlew :spdf:conformance # runner report in build/conformance.json ./gradlew :spdf:run --args="conformance ../conformance -o conformance.json" ``` The runner discovers the cases by listing `conformance/cases/*.json` (`SPDF_CONFORMANCE_DIR` overrides the location) and prints `{"impl":"spdf-kotlin","version":"0.1.0","passed":[…], "failed":[…],"skipped":[…]}`; it exits non-zero if anything fails. ## Building Gradle 9.8 through the wrapper, Kotlin 2.4, any JDK 17 or newer (bytecode targets 17). `./gradlew publishAllPublicationsToBuildRepoRepository` writes exactly what would be published to `build/repo`; publishing to Maven Central is described in [`PUBLICAR.md`](PUBLICAR.md) (in Spanish). ## License MIT OR Apache-2.0. --- # SPDF en Go URL: https://spdf.joseluissaorin.com/es/documentacion/go > Cómo instalar y usar la implementación de SPDF en Go (github.com/joseluissaorin/spdf/go): abrir, validar, buscar y citar. Módulo de Go dentro del monorepo. - **Paquete**: `github.com/joseluissaorin/spdf/go` - **Instalar**: `go get github.com/joseluissaorin/spdf/go` - **Registro**: [pkg.go.dev](https://pkg.go.dev/github.com/joseluissaorin/spdf/go) - **Nivel**: primero - **CI**: CI en marcha - **Carpeta**: `go/` *El README de la biblioteca está en inglés.* Pure-Go implementation of **SPDF** (Semantic Processed Document Format): documents that have been read once and can be cited forever, because every passage carries its exact anchor (printed page, folio, second of a recording, slide, verse). - Module: `github.com/joseluissaorin/spdf/go` (package `spdf`) - SQLite without cgo ([`modernc.org/sqlite`](https://pkg.go.dev/modernc.org/sqlite), FTS5 and the `trigram` tokenizer included), so it cross-compiles to every Go target. - Go 1.25 or newer. - Conformance: passes the whole SPDF conformance suite (`../conformance`), every kind (`dump`, `legacy_dump`, `roundtrip`, `validate`, `search_*`, `anchor_uri`, `cite`). ```sh go get github.com/joseluissaorin/spdf/go ``` ## What it does | | | |---|---| | Safe opening | read-only, `query_only`, `trusted_schema=OFF`, `SQLITE_DBCONFIG_DEFENSIVE`, no extensions, files with triggers or views refused (E020), unknown required extensions refused (E060), max blob size (512 MiB) and max gunzipped size (4 GiB) | | Versions | SPDF 5.0, and the legacy 4.0 / 4.1 files of Scholaris (Spanish schema, gzip-wrapped) through the 5.0 view, metadata mapped to CSL-JSON | | Dump | canonical JSON (RFC 8785) of the whole file, `content_sha256` (§8) | | Validation | every code of the specification (E001–E090, W100–W110), Ed25519 signatures | | Search | lexical (FTS5 BM25, CJK route with `trigram` or substring), vector (`f32`, `f16`, `i8`), hybrid (reciprocal rank fusion, k = 10) | | Anchors | anchor ↔ URI (`spdf:sha256-…#p=29&f=21&char=118,301`), strict parser; resolution of a URI or of a `.spdf` URL with a fragment to units and fragments (`Locate`, SPEC §5.4) | | Citation | short author-date citation in Spanish and English; citation of a quotation by the unit it lies in (`CitePassage`, SPEC §18.2) | | Export | CSL-JSON (also citations with `label`/`locator`) and BibTeX, one or several documents with the keys of SPEC §19 (`cervantessaavedra1605`, `lazarillo1554`, collision suffixes); ALTO 4, minimal TEI and IIIF Presentation 3 (`ExportALTO`, `ExportTEI`, `ExportIIIF`; no invented coordinates or dimensions) | | Writer | builds valid SPDF 5.0 files (FTS kept in sync, `VACUUM`, no triggers) with `content_sha256` and, given a key, an Ed25519 signature (§8); `Seal` hashes and signs an existing file in place | ## Reading and searching ```go package main import ( "fmt" "log" spdf "github.com/joseluissaorin/spdf/go" ) func main() { f, err := spdf.Open("quijote.spdf", nil) // also legacy .spdf (gzip) files if err != nil { log.Fatal(err) } defer f.Close() hits, err := f.SearchLexical("«lugar de la Mancha»", 5) if err != nil { log.Fatal(err) } for _, h := range hits { cite, _ := f.Cite(h.Anchor, h.AnchorEnd, "es") fmt.Println(h.FragmentID, h.Score, h.AnchorURI, cite) // q4 1.889394 spdf:sha256-fa38…#p=5&pe=6&f=1r&fe=1v&char=101,278 (Cervantes Saavedra, 1605, fols. 1r-[1v]) } // Vector and hybrid search with your own query embedding. vec := make([]float64, 8) vhits, _ := f.SearchVector(vec, "toy-embedding@8", "fragment", 5) _ = vhits hy, _ := f.SearchHybrid("hidalgo", vec, "toy-embedding@8", 5) _ = hy bib, _ := f.ExportBibTeX() fmt.Print(bib) } ``` ## Validating and dumping ```go res := spdf.Validate("file.spdf", nil) fmt.Println(res.Valid, res.ErrorCodes(), res.WarningCodes()) f, _ := spdf.Open("file.spdf", nil) dump, _ := f.DumpJSON() // RFC 8785 bytes sum, _ := f.ContentSHA256() // integrity hash of §8 ``` ## Anchors and citations ```go a := spdf.Anchor{"type": "page", "physical": int64(29), "printed": "21", "source": "inferred"} uri := spdf.AnchorURI("sha256-3f2a…", a, nil) // spdf:sha256-3f2a…#p=29&f=21 docref, loc, err := spdf.ParseURI(uri) // strict: malformed URIs are errors text := spdf.Cite(a, nil, map[string]any{ "type": "book", "title": "Arte nuevo de hacer comedias", "author": []any{map[string]any{"family": "Vega", "non-dropping-particle": "de", "given": "Lope"}}, "issued": map[string]any{"date-parts": []any{[]any{int64(1609)}}}, }, "es") // (de Vega, 1609, p. [21]) ``` ## Citing a quotation ```go // A fragment of Micrographia runs from a plate into page 211: the quotation is cited // by the unit(s) it actually touches, not by the start anchor of its fragment. p, _ := f.CitePassage("m4", "XXXIV.\ntube N N", "es") fmt.Println(p.Text, p.URI) // (Hooke, 1665, p. 211) spdf:sha256-ba9d…#p=319&pe=321&fe=211 ``` ## Resolving references ```go r, _ := f.Locate("spdf:sha256-…#f=1v") // or "https://example.org/quijote.spdf#p=7" fmt.Println(r.Document, r.Units, r.Fragments, r.Char) // true [p6] [q4 q5] [] ``` ## Writing ```go w, err := spdf.Create("out.spdf", &spdf.WriterOptions{Generator: "my-tool/1.0"}) if err != nil { log.Fatal(err) } w.SetDocument(spdf.Document{ID: "doc", Kind: "pdf", Mime: "application/pdf", SourceSHA256: sha, Bytes: n, Metadata: map[string]any{"type": "book", "title": "…"}}) w.AddUnit(spdf.Unit{ID: "u1", Reader: "pdf-text-layer", Text: "…", Anchor: spdf.Anchor{"type": "page", "physical": int64(1), "printed": "1"}}) w.AddFragment(spdf.Fragment{ID: "f1", Unit: "u1", Text: "…", Anchor: spdf.Anchor{"type": "page", "physical": int64(1), "printed": "1"}}) w.AddSpace(spdf.SpaceDef{ID: "embeddinggemma-2@768", Provider: "local", Model: "embeddinggemma-2", Dims: 768}) w.AddVector("fragment", "f1", "embeddinggemma-2@768", embedding) // []float32, quantized for f16/i8 w.AddBlob("pages/0001.png", "image/png", png) if err := w.Close(); err != nil { // rebuilds FTS, VACUUM, atomic rename log.Fatal(err) } ``` Every file the Writer produces carries `spdf_meta.content_sha256`; pass `WriterOptions{SigningKey: ed25519.NewKeyFromSeed(seed)}` to sign it too (`signer`, `signature`). Files signed this way verify with the Rust, Python and JavaScript implementations (and tampering gives E081 or E082 in all of them). `spdf.Seal(path, key)` does the same for an existing file. `spdf.WriteSource(source, path)` builds a file from a full JSON dump (the format of `conformance/sources/`). ## Command line ```sh go install github.com/joseluissaorin/spdf/go/cmd/spdf@latest spdf validate file.spdf spdf dump file.spdf spdf search file.spdf "lugar de la Mancha" -n 5 spdf cite file.spdf q4 -locale en spdf export file.spdf bibtex spdf uri parse 'spdf:sha256-…#p=29&f=21' spdf build source.json out.spdf spdf seal out.spdf -key seed.hex # content_sha256 + Ed25519 signature spdf export file.spdf alto # also tei, iiif spdf conformance ../conformance -o conformance.json ``` ## Conformance ```sh go test ./... # unit tests and the whole suite go run ./cmd/spdf conformance ../conformance # the report of contract §11 ``` The runner discovers the cases by listing `conformance/cases/*.json` and prints `{"impl","version","passed","failed","skipped"}`. Nothing is skipped: this is a full (reader and writer) implementation. ## License MIT OR Apache-2.0. --- # SPDF en C# / .NET URL: https://spdf.joseluissaorin.com/es/documentacion/dotnet > Cómo instalar y usar la implementación de SPDF en C# / .NET (Spdf.Format): abrir, validar, buscar y citar. .NET con Microsoft.Data.Sqlite. - **Paquete**: `Spdf.Format` - **Instalar**: `dotnet add package Spdf.Format` - **Registro**: [NuGet](https://www.nuget.org/packages/Spdf.Format) - **Nivel**: primero - **CI**: CI en rojo - **Carpeta**: `dotnet/` *El README de la biblioteca está en inglés.* Native C# implementation of **SPDF** (Semantic Processed Document Format): documents that have been read once and can be cited forever, because every passage carries its exact anchor (printed page, folio, second of a recording, slide, verse). - NuGet package: `Spdf.Format` · namespace `Spdf` · .NET 8 or newer. - SQLite through `Microsoft.Data.Sqlite.Core` with the `SQLitePCLRaw.bundle_e_sqlite3` native bundle (FTS5 and the `trigram` tokenizer included) on Windows, macOS and Linux. - An independent implementation: it does not wrap the Rust library or any other one. Ed25519 verification, RFC 8785 serialization and the exact rounding rules are written in C#. - Conformance: passes the whole SPDF conformance suite (`../conformance`), every kind (`dump`, `legacy_dump`, `roundtrip`, `quantize`, `validate`, `search_lexical`, `search_vector`, `search_hybrid`, `anchor_uri`, `cite`, `locate`, `export_csl`, `export_bibtex`, `export_structure`). Nothing is skipped. ```sh dotnet add package Spdf.Format ``` ## What it does | | | |---|---| | Safe opening | read-only, `query_only`, `trusted_schema=OFF`, `SQLITE_DBCONFIG_DEFENSIVE`, extension loading disabled, `mmap_size=0`, `cell_size_check=ON`; triggers, views and foreign virtual tables refused (E020), unknown required extensions refused (E060); max blob size (512 MiB) and max gunzipped size (4 GiB); files left in WAL mode are read from a private copy | | Versions | SPDF 5.0, and the legacy 4.0 / 4.1 files of Scholaris (Spanish schema, gzip-wrapped) through the 5.0 view, with metadata mapped to CSL-JSON | | Dump | canonical JSON (RFC 8785) of the whole file and `content_sha256` (§12, §13) | | Validation | every code of the specification (E001–E090, W100–W110), FTS integrity on an in-memory copy, Ed25519 signatures | | Search | lexical (FTS5 BM25, CJK route with `trigram` or substring), vector over fragments, units or figures (`f32`, `f16`, `i8`; dot product or cosine), hybrid (reciprocal rank fusion, k = 10) | | Anchors | anchor ↔ URI (`spdf:sha256-…#p=29&f=21&char=118,301`), strict parser, canonical form; resolution of `spdf:` URIs and `.spdf` URLs to units and fragments (§5.4) | | Citation | short author-date citation in Spanish and English | | Export | CSL-JSON (with the CSL `label`/`locator` of a citation) and BibTeX with the keys of §19.1 (`cervantessaavedra1605`, `lazarillo1554`, `anonnd`, collision suffixes `a`, `b`…); ALTO 4, minimal TEI and IIIF Presentation 3 (§19.4) | | Writer | builds valid SPDF 5.0 files (FTS kept in sync, `VACUUM`, no triggers, atomic replace), with `content_sha256` and an optional Ed25519 signature (§13) | ## Reading and searching ```csharp using Spdf; using var file = SpdfFile.Open("quijote.spdf"); // also legacy .spdf (gzip) files foreach (var hit in file.SearchLexical("«lugar de la Mancha»", limit: 5)) { Console.WriteLine($"{hit.FragmentId} {hit.Score:F6} {hit.AnchorUri}"); Console.WriteLine(file.Cite(hit.Anchor!, hit.AnchorEnd, "es")); // (Cervantes Saavedra, 1605, fols. 1r-[1v]) } // Vector and hybrid search with your own query embedding. double[] query = new double[8]; var nearest = file.SearchVector(query, space: "toy-embedding@8", target: "fragment", limit: 5); var pages = file.SearchVector(query, "toy-embedding@8", target: "unit"); // hits carry UnitId var fused = file.SearchHybrid("hidalgo", query, "toy-embedding@8", limit: 5); // Typed reading. Document doc = file.GetDocument(); IReadOnlyList pages = file.GetUnits(); IReadOnlyList fragments = file.GetFragments(); Blob? original = file.GetBlob("blob:original.pdf"); Console.Write(file.ExportBibTeX()); // @book{cervantessaavedra1605, … Console.WriteLine(file.ExportCslJson()); // [{"id":"cervantessaavedra1605", …}] BibTexEntry entry = file.ExportBibTeXEntry(); // entry type, key and fields // Several documents at once (keys disambiguated with a, b, c…), and a CSL citation item. string bib = BibliographyExport.BibTeX([file.GetMetadata(), other.GetMetadata()]); var cited = BibliographyExport.CslItems([file.GetMetadata()], hit.Anchor, hit.AnchorEnd); // + label, locator // Resolving a reference (§5.4): units, fragments, char range and region it designates. LocateResult where = file.Locate("spdf:sha256-…#p=5&pe=6&char=101,278"); LocateResult byUrl = file.Locate("https://example.org/quijote.spdf#f=1v"); // Structural exports (§19.4): ALTO 4, a minimal TEI and a IIIF Presentation 3 manifest. // No invented coordinates or dimensions; inferred folios bracketed (TEI, IIIF) or absent (ALTO). string alto = file.ExportAlto(); // Page per page unit, TextBlock / TextLine / String per word string tei = file.ExportTei(); // teiHeader, ,

, /, , string iiif = file.ExportIiif(new IiifOptions { Base = "https://example.org/quijote" }); var pageSequence = StructureExport.PageSequence(StructureFormat.Tei, tei); // read back from the XML ``` `SpdfFile.OpenAsync(path)` decompresses or copies asynchronously when the file needs it; `SpdfFile.Open(stream)` and `SpdfFile.Open(bytes)` read from memory. A `SpdfFile` is not thread-safe; open one per thread. ## Validating and dumping ```csharp ValidationResult result = SpdfValidator.Validate("file.spdf"); Console.WriteLine($"{result.Valid} {string.Join(",", result.ErrorCodes)} {string.Join(",", result.WarningCodes)}"); using var file = SpdfFile.Open("file.spdf"); string dump = file.DumpJson(); // RFC 8785 text string hash = file.ContentSha256(); // integrity hash of §13 ``` ## Anchors and citations ```csharp var anchor = Anchor.Page(29, "21", source: "inferred").WithChars(118, 301); string uri = AnchorUri.Format("sha256-3f2a…", anchor); // spdf:sha256-3f2a…#p=29&f=21&char=118,301 ParsedAnchorUri parsed = AnchorUri.Parse(uri); // FormatException if malformed var metadata = new Dictionary { ["type"] = "book", ["title"] = "Arte nuevo de hacer comedias", ["author"] = new List { new Dictionary { ["family"] = "Vega", ["non-dropping-particle"] = "de", ["given"] = "Lope" } }, ["issued"] = new Dictionary { ["date-parts"] = new List { new List { 1609L } } }, }; Console.WriteLine(Citation.Cite(anchor, null, metadata, "es")); // (de Vega, 1609, p. [21]) ``` JSON values (metadata, anchors, word timings) are plain trees: `null`, `bool`, `long`, `double`, `string`, `List` and `Dictionary`. `SpdfJson` parses and serializes them (`Canonical` is RFC 8785 with the six-decimal rounding of the specification; `Compact` keeps numbers as they are). ## Writing ```csharp using var w = SpdfWriter.Create("out.spdf", new SpdfWriterOptions { Generator = "my-tool/1.0" }); w.SetDocument(new Document { Id = "rimas", Kind = "pdf", Mime = "application/pdf", SourceSha256 = sha, Bytes = size, Metadata = new Dictionary { ["type"] = "book", ["title"] = "Rimas" }, }); w.AddUnit(new Unit { Id = "u1", Reader = "pdf-text-layer", Text = "…", Anchor = Anchor.Page(1, "1") }); w.AddFragment(new Fragment { Id = "f1", Unit = "u1", Text = "…", Anchor = Anchor.Page(1, "1").WithChars(0, 120) }); w.AddSpace(new Space { Id = "embeddinggemma-2@768:i8", Provider = "local", Model = "embeddinggemma-2", Dims = 768, DType = "i8" }); w.AddVector("fragment", "f1", "embeddinggemma-2@768:i8", embedding); // float[] or double[], quantized per §9.2 w.AddBlob("original.pdf", "application/pdf", bytes); w.Commit(); // document + spdf_meta, FTS rebuild, VACUUM, atomic move into place ``` Disposing a writer without `Commit()` discards the file. `SpdfSource.Write(source, path)` builds a file from a full JSON dump (the format of `conformance/sources/`), value for value. ### Integrity and signatures By default the writer stores `spdf_meta.content_sha256` (§13), computed on the finished file. Give it a 32-byte Ed25519 secret key (seed) to sign as well; it writes `signer` (`ed25519:` + base64 public key) and `signature` (base64 of the RFC 8032 signature over `spdf-content-sha256:` + the hex hash): ```csharp byte[] seed = LoadSeedFromYourKeyStore(); // 32 bytes; never commit it to a repository using var w = SpdfWriter.Create("signed.spdf", new SpdfWriterOptions { SigningKey = seed }); // … rows … w.Commit(); bool ok = SpdfValidator.Validate("signed.spdf").Valid; // recomputes the hash, verifies the signature ``` **The Ed25519 signer is not constant-time.** It is a portable BigInteger implementation (.NET 8 has no built-in Ed25519): signing handles the secret key with variable-time arithmetic, which can leak it through timing to anyone able to measure many signatures on the same machine. Sign only on a trusted machine, never in a shared or multi-tenant service. Verification uses public data only and is safe anywhere. `Ed25519.Sign`, `Ed25519.PublicKey`, `Ed25519.Verify` and `SpdfValidator.SignContentHash` are public for producers that sign outside the writer. ## Command line The repository includes a small CLI (`src/Spdf.Cli`, not published as a package): ```sh dotnet run --project src/Spdf.Cli -- validate file.spdf dotnet run --project src/Spdf.Cli -- dump file.spdf dotnet run --project src/Spdf.Cli -- search file.spdf "lugar de la Mancha" -n 5 dotnet run --project src/Spdf.Cli -- vsearch file.spdf toy-embedding@8 0.5,0.25,0.5,0.25,0,0.25,0.25,0.5 dotnet run --project src/Spdf.Cli -- cite file.spdf q4 --locale en dotnet run --project src/Spdf.Cli -- export file.spdf bibtex # also csl, alto, tei, iiif dotnet run --project src/Spdf.Cli -- sample demo.spdf --key-file seed.hex # small signed demo file dotnet run --project src/Spdf.Cli -- uri parse 'spdf:sha256-…#p=29&f=21' dotnet run --project src/Spdf.Cli -- build source.json out.spdf dotnet run --project src/Spdf.Cli -- conformance ../conformance -o conformance.json ``` ## Building and testing ```sh cd dotnet dotnet build dotnet test # unit tests and the whole conformance suite dotnet run --project src/Spdf.Cli -- conformance ../conformance -o conformance.json dotnet pack src/Spdf.Format -c Release -o artifacts ``` The test project finds the suite at `../conformance` relative to `dotnet/`; set `SPDF_CONFORMANCE_DIR` to use another copy. The runner discovers the cases by listing `conformance/cases/*.json` and prints `{"impl","version","passed","failed","skipped"}`. ## License MIT OR Apache-2.0. --- # SPDF en PHP URL: https://spdf.joseluissaorin.com/es/documentacion/php > Cómo instalar y usar la implementación de SPDF en PHP (joseluissaorin/spdf): abrir, validar, buscar y citar. PDO SQLite. - **Paquete**: `joseluissaorin/spdf` - **Instalar**: `composer require joseluissaorin/spdf` - **Registro**: [Packagist](https://packagist.org/packages/joseluissaorin/spdf) - **Nivel**: segundo - **CI**: CI en marcha - **Carpeta**: `php/` *El README de la biblioteca está en inglés.* `joseluissaorin/spdf` reads, validates, searches, cites and writes **SPDF** files (Semantic Processed Document Format): documents that have been read once and can be cited forever. Every passage carries its exact anchor (printed page, folio, second of a recording, slide, verse), so a citation can only print what the source says. It is a native implementation of SPDF 5.0 over PDO SQLite. It also reads the legacy 4.0/4.1 files produced by Scholaris (gzip-wrapped, Spanish schema) through the 5.0 view. It is built for the PHP hosts where journals and libraries live: OJS, Omeka S, WordPress. ## Install ```sh composer require joseluissaorin/spdf ``` Requirements: PHP 8.1 or newer with `pdo_sqlite` (SQLite with FTS5; 3.44+ recommended), `intl`, `mbstring` and `zlib`. `sodium` (bundled with PHP) verifies signatures. ## Read, search and cite ```php use Spdf\Document; $doc = Document::open('lazarillo.spdf'); // read-only, safe opening echo $doc->title(), "\n"; // La vida de Lazarillo de Tormes… echo $doc->cite(['type' => 'image'], null, 'es'); // (Anónimo, 1554) foreach ($doc->searchLexical('"Antona Pérez" Tejares', 5) as $hit) { $f = $doc->fragment($hit['fragment_id']); echo $f['text'], ' ', $doc->cite($hit['anchor'], $f['anchor_end'], 'es'), "\n"; // hijo de Tomé González y de Antona Pérez… (Anónimo, 1554, p. [4]) echo $hit['anchor_uri'], "\n"; // spdf:sha256-3f2a…#p=10&f=4 } ``` Units, fragments, sections, figures, spaces, blobs and provenance are plain arrays with the 5.0 column names (`$doc->units()`, `$doc->fragments()`, `$doc->blob('blob:cover')`…). `$doc->metadata()` is the CSL-JSON item plus the `spdf` extension object. ### Vector and hybrid search ```php $query = $myEmbedder->embed('el ciego y el jarro de vino'); // list, same model as the space $doc->searchVector($query, 'embeddinggemma-2@768', 10); // f32, f16 and i8 spaces $doc->searchHybrid('ciego jarro', $query, 'embeddinggemma-2@768', 10); // RRF, k = 10 ``` These are the reference algorithms of the specification (§6): brute-force dot product (cosine when the space is not normalized) and reciprocal rank fusion over lists of depth `max(limit, 50)`. ### Anchors and URIs ```php use Spdf\AnchorUri; $uri = $doc->anchorUri($fragment['anchor'], $fragment['anchor_end']); $parsed = AnchorUri::parse('spdf:sha256-3f2a…#p=29&f=21&char=118,301'); // ['docref' => 'sha256-3f2a…', 'locator' => ['p' => 29, 'f' => '21', 'char' => [118, 301]]] AnchorUri::format($parsed['docref'], $parsed['locator']); // the same URI, byte for byte $doc->locate($uri); // {document, units, fragments, char, xywh} (SPEC §5.4) $doc->citePassage('f12', 'molinos de viento', 'es'); // {text, uri}: cites the unit the quote lies in (§18.2) ``` ### Bibliography ```php file_put_contents('lazarillo.json', $doc->cslJson()); // Zotero, Pandoc, citeproc file_put_contents('lazarillo.bib', $doc->bibtex()); // @book{lazarillo1554, ... ``` Keys and fields follow the specification (§19): the first author's name, or the first word of the short title, folded to ASCII and lowercased, plus the year (`cervantessaavedra1605`, `lazarillo1554`, `anonnd`); the CSL-JSON `id` is the same key. `Spdf\Bibliography::cslItems()` and `bibtexAll()` export several records and disambiguate colliding keys with `a`, `b`, `c`…; `cslItems([$meta], $anchor, $anchorEnd)` adds the CSL `label` and `locator` of a citation. ### ALTO, TEI and IIIF ```php file_put_contents('lazarillo.alto.xml', $doc->alto()); // ALTO 4, one Page per page unit file_put_contents('lazarillo.tei.xml', $doc->tei()); // TEI P5: pb, p, lg/l, u, note $manifest = $doc->iiif('https://revista.example.org/iiif/lazarillo'); // IIIF Presentation 3 ``` These are the optional exports of SPEC §19.4: printed folios only where the page carries them (`[iv]` marks an inferred folio in TEI and IIIF), sections as IIIF ranges, figures as `describing` annotations on their region, and no invented coordinates. ## Validate ```php $report = Spdf\Validator::validate('file.spdf'); // ['valid' => true, 'version' => '5.0', 'profile' => ['core', 'semantic'], // 'errors' => [], 'warnings' => []] ``` Error and warning codes are those of the specification (§12): `E001` not SQLite, `E020` trigger or view, `E070` FTS index out of sync, `E081` content hash mismatch… ## Write ```php use Spdf\Writer; $w = Writer::create('out.spdf', generator: 'my-journal/1.0', profile: 'core'); $w->document(['id' => 'art-12', 'kind' => 'pdf', 'source_sha256' => hash_file('sha256', 'art-12.pdf'), 'mime' => 'application/pdf', 'bytes' => filesize('art-12.pdf'), 'unit_count' => 1, 'metadata' => ['type' => 'article-journal', 'title' => 'Sobre el Lazarillo', 'author' => [['family' => 'Pérez', 'given' => 'Ana']], 'issued' => ['date-parts' => [[2026]]]]]); $w->unit(['id' => 'p1', 'ord' => 1, 'anchor' => ['type' => 'page', 'physical' => 1, 'printed' => '45'], 'text' => 'Texto de la página…', 'reader' => 'pdf-text-layer']); $w->fragment(['n' => 1, 'id' => 'f1', 'unit' => 'p1', 'ord' => 1, 'text' => 'Texto de la página…', 'anchor' => ['type' => 'page', 'physical' => 1, 'printed' => '45']]); $w->finish(contentHash: true); // FTS rebuilt, no triggers, VACUUM, atomic rename ``` ## Security Files are untrusted input. `Document::open()` opens them read-only with `PRAGMA query_only`, `trusted_schema=OFF`, never loads extensions, refuses triggers and views (except the three FTS triggers of legacy files), bounds blob sizes (512 MiB by default) and gzip inflation (4 GiB), and copies WAL-mode files instead of touching them. Limits are set with `new Spdf\Options(maxBlobBytes: …, maxInflatedBytes: …)`. PDO does not expose `SQLITE_DBCONFIG_DEFENSIVE`; the other measures cover what it guards in a read-only connection. ## In OJS, Omeka S and WordPress `examples/` holds three minimal integrations: - `show-and-cite.php`: a standalone page (title, whole-work citation, search, cited passages, BibTeX and CSL-JSON downloads). `SPDF_DIR=… php -S localhost:8080 examples/show-and-cite.php`. - `ojs/spdfViewer`: an OJS 3.4 generic plugin that renders `.spdf` galleys. - `omeka-s/SpdfViewer`: an Omeka S module with a file renderer for `application/vnd.spdf`. The plugin and the module are sketches to start from; they show the calls, not a finished product. ## Command line ```sh vendor/bin/spdf validate file.spdf vendor/bin/spdf dump file.spdf # canonical dump (RFC 8785) vendor/bin/spdf search file.spdf "molinos de viento" vendor/bin/spdf cite file.spdf f12 en vendor/bin/spdf bibtex file.spdf vendor/bin/spdf conformance ../conformance ``` ## Conformance `php bin/spdf conformance ../conformance` runs the shared suite of the repository and prints `{"impl":"joseluissaorin/spdf (PHP)","version":…,"passed":[…],"failed":[…],"skipped":[…]}`. CI runs it on PHP 8.1 to 8.4 and publishes the report as the `conformance-php` artifact. All kinds are claimed, `export_structure` included. ## License MIT OR Apache-2.0, at your option. The SPDF specification is CC BY 4.0. --- # SPDF en Ruby URL: https://spdf.joseluissaorin.com/es/documentacion/ruby > Cómo instalar y usar la implementación de SPDF en Ruby (spdf-format): abrir, validar, buscar y citar. Sobre la gema sqlite3. - **Paquete**: `spdf-format` - **Instalar**: `gem install spdf-format` - **Registro**: [RubyGems](https://rubygems.org/gems/spdf-format) - **Nivel**: segundo - **CI**: CI en marcha - **Carpeta**: `ruby/` *El README de la biblioteca está en inglés.* `spdf-format` reads, validates, searches, cites and writes **SPDF** files (Semantic Processed Document Format): documents that have been read once and can be cited forever. Every passage carries its exact anchor (printed page, folio, second of a recording, slide, verse), so a citation can only print what the source says. Native implementation of SPDF 5.0 on the `sqlite3` gem. It also reads the legacy 4.0/4.1 files produced by Scholaris (gzip-wrapped, Spanish schema) through the 5.0 view. ## Install ```sh gem install spdf-format ``` ```ruby require "spdf" ``` Ruby 3.1 or newer. The `sqlite3` gem ships SQLite with FTS5; signatures are verified with the standard `openssl` library. ## Read, search and cite ```ruby Spdf::Document.open("lazarillo.spdf") do |doc| # read-only, safe opening puts doc.title puts doc.cite({ "type" => "image" }, locale: "es") # (Anónimo, 1554) doc.search_lexical('"Antona Pérez" Tejares', limit: 5).each do |hit| f = doc.fragment(hit["fragment_id"]) puts "#{f["text"]} #{doc.cite(hit["anchor"], f["anchor_end"], locale: "es")}" # hijo de Tomé González y de Antona Pérez… (Anónimo, 1554, p. [4]) puts hit["anchor_uri"] # spdf:sha256-3f2a…#p=10&f=4 end end ``` Rows are hashes with the 5.0 column names: `doc.units`, `doc.fragments`, `doc.sections`, `doc.figures`, `doc.spaces`, `doc.blob("blob:cover")`, `doc.metadata` (the CSL-JSON item plus the `spdf` extension object). ### Vector and hybrid search ```ruby query = embedder.embed("el ciego y el jarro de vino") # same model as the space doc.search_vector(query, space: "embeddinggemma-2@768", limit: 10) # f32, f16, i8 doc.search_hybrid("ciego jarro", query, space: "embeddinggemma-2@768") # RRF, k = 10 ``` ### Anchors, bibliography ```ruby Spdf::AnchorUri.parse("spdf:sha256-3f2a…#p=29&f=21&char=118,301") # {"docref" => "sha256-3f2a…", "locator" => {"p" => 29, "f" => "21", "char" => [118, 301]}} doc.locate("spdf:sha256-3f2a…#p=29") # {"document", "units", "fragments", "char", "xywh"} (SPEC §5.4) doc.cite_passage("f12", "molinos de viento") # {"text", "uri"}: cites the unit the quote lies in (§18.2) File.write("lazarillo.json", doc.csl_json) # Zotero, Pandoc, citeproc (id = BibTeX key) File.write("lazarillo.bib", doc.bibtex) # @book{lazarillo1554, ... (SPEC §19) Spdf::Bibliography.csl_items([meta], anchor, anchor_end) # adds CSL "label" and "locator" ``` ### ALTO, TEI and IIIF ```ruby File.write("lazarillo.alto.xml", doc.alto) # ALTO 4, one Page per page unit File.write("lazarillo.tei.xml", doc.tei) # TEI P5: pb, p, lg/l, u, note manifest = doc.iiif("https://biblioteca.example.org/iiif/lazarillo") # IIIF Presentation 3 ``` These are the optional exports of SPEC §19.4: printed folios only where the page carries them (`[iv]` marks an inferred folio in TEI and IIIF), sections as IIIF ranges, figures as `describing` annotations, and no invented coordinates. ## Validate ```ruby Spdf::Validator.validate("file.spdf") # {"valid" => true, "version" => "5.0", "profile" => ["core"], "errors" => [], "warnings" => []} ``` ## Write ```ruby Spdf::Writer.create("out.spdf", generator: "my-app/1.0") do |w| w.document("id" => "d1", "kind" => "pdf", "source_sha256" => Digest::SHA256.file("d1.pdf").hexdigest, "mime" => "application/pdf", "bytes" => File.size("d1.pdf"), "unit_count" => 1, "metadata" => { "type" => "book", "title" => "Lazarillo de Tormes", "issued" => { "date-parts" => [[1554]] } }) w.unit("id" => "p1", "ord" => 1, "anchor" => { "type" => "page", "physical" => 1, "printed" => "3" }, "text" => "Pues sepa Vuestra Merced…", "reader" => "pdf-text-layer") w.fragment("n" => 1, "id" => "f1", "unit" => "p1", "ord" => 1, "text" => "Pues sepa Vuestra Merced…", "anchor" => { "type" => "page", "physical" => 1, "printed" => "3" }) end ``` ## Security Files are untrusted input: they are opened read-only with `query_only` and `trusted_schema=OFF`, extensions are never loaded, triggers and views are refused (except the three FTS triggers of legacy files), blob sizes (512 MiB) and gzip inflation (4 GiB) are bounded, and WAL-mode files are copied. `Spdf::Document.open(path, max_blob_bytes: …)` changes the limits. The gem does not expose `SQLITE_DBCONFIG_DEFENSIVE`. ## Command line and conformance ```sh spdf validate file.spdf spdf dump file.spdf spdf search file.spdf "molinos de viento" spdf cite file.spdf f12 en spdf conformance path/to/spdf/conformance ``` `spdf conformance` runs every case of the shared suite and prints the report of the specification (§21). CI publishes it as the `conformance-ruby` artifact. All kinds are claimed, `export_structure` included. ## License MIT OR Apache-2.0, at your option. The SPDF specification is CC BY 4.0. --- # SPDF en R URL: https://spdf.joseluissaorin.com/es/documentacion/r > Cómo instalar y usar la implementación de SPDF en R (spdf): abrir, validar, buscar y citar. Sobre RSQLite; los fragmentos como data frames. - **Paquete**: `spdf` - **Instalar**: `remotes::install_github("joseluissaorin/spdf", subdir = "r")` - **Nivel**: segundo - **CI**: CI en marcha - **Carpeta**: `r/` *El README de la biblioteca está en inglés.* Read, validate, search, cite and write **SPDF** files (Semantic Processed Document Format) from R. A SPDF file holds a document that has been read once and can be cited forever: every passage carries its exact anchor (printed page, folio, second of a recording, slide, verse), so a citation can only print what the source says. The package is a native implementation of SPDF 5.0 on RSQLite. It also reads the legacy 4.0/4.1 files produced by Scholaris (gzip-wrapped, Spanish schema) through the 5.0 view. Tables come back as tibbles. ## Install ```r # from CRAN, once published install.packages("spdf") # from the repository remotes::install_github("joseluissaorin/spdf", subdir = "r") ``` ## Read, search, cite ```r library(spdf) doc <- spdf_open(system.file("extdata", "quijote.spdf", package = "spdf")) spdf_info(doc) # title, authors, year, version, counts spdf_units(doc) # one row per citable unit (page, folio, time span...) fr <- spdf_fragments(doc) # searchable passages, anchors as list-columns hits <- spdf_search(doc, "hermoso") # finds the long-s "hermoso" through the modern layer hits$anchor_uri # spdf:sha256-27ea...#p=13&char=10,194 spdf_cite(spdf_metadata(doc), fr$anchor[[5]], fr$anchor_end[[5]], locale = "es") #> "(Cervantes Saavedra, 1608, fols. Ir-[Iv])" spdf_cite_passage(doc, "q5", "rozin, como tomaua la podadera.")$text #> "(Cervantes Saavedra, 1608, fol. [Iv])" the page the quotation is on spdf_locate(doc, hits$anchor_uri[1]) # list(document, units, fragments, char, xywh) cat(spdf_bibtex(doc)) # @book{cervantessaavedra1608, ... (also spdf_csl()) spdf_close(doc) ``` Vector and hybrid search take a query vector computed with the same model as the space: `spdf_search_vector(doc, v, "embeddinggemma-2@768")`, `spdf_search_hybrid(doc, "ciego jarro", v, "embeddinggemma-2@768")`. ## ALTO, TEI and IIIF ```r writeLines(spdf_alto(doc), "quijote.alto.xml") # ALTO 4, one Page per page unit writeLines(spdf_tei(doc), "quijote.tei.xml") # TEI P5: pb, p, lg/l, u, note writeLines(spdf_iiif_json(doc, "https://example.org/iiif/quijote"), "manifest.json") ``` ## Corpora ```r files <- list.files("corpus", pattern = "\\.spdf$", full.names = TRUE) spdf_corpus(files) # one row per document spdf_corpus_search(files, "\"molinos de viento\"") # one row per passage, with citation spdf_count_terms(files, c("honra", "fortuna")) # fragments and occurrences per work ``` The vignette `vignette("corpus", package = "spdf")` (in Spanish) walks through a digital-humanities workflow: searching a corpus and counting occurrences by work and year, with every number traceable to its page. ## Validate and write ```r spdf_validate("file.spdf") # list(valid, version, profile, errors, warnings) spdf_write("out.spdf", document = ..., units = ..., fragments = ...) ``` ## Security Files are untrusted input: `spdf_open()` connects read-only with `query_only` and `trusted_schema=OFF`, never loads extensions, refuses triggers, views and foreign virtual tables (except the three FTS triggers of legacy files), bounds blob sizes and gzip inflation, and copies WAL-mode files before opening them. ## Conformance `spdf_conformance("path/to/spdf/conformance")` runs the shared suite of the specification; `Rscript inst/scripts/conformance.R ../conformance` prints the JSON report. CI publishes it as the `conformance-r` artifact. All kinds are claimed, `export_structure` included. ## License MIT OR Apache-2.0, at your option. The SPDF specification is CC BY 4.0. The sample files in `inst/extdata` are short excerpts of public-domain works. --- # SPDF en Julia URL: https://spdf.joseluissaorin.com/es/documentacion/julia > Cómo instalar y usar la implementación de SPDF en Julia (SPDF.jl): abrir, validar, buscar y citar. Sobre SQLite.jl. - **Paquete**: `SPDF.jl` - **Instalar**: `pkg> add SPDF` - **Nivel**: segundo - **CI**: CI en marcha - **Carpeta**: `julia/` *El README de la biblioteca está en inglés.* Read, validate, search, cite and write **SPDF** files (Semantic Processed Document Format) from Julia. A SPDF file holds a document that has been read once and can be cited forever: every passage carries its exact anchor (printed page, folio, second of a recording, slide, verse), so a citation can only print what the source says. Native implementation of SPDF 5.0 on SQLite.jl. It also reads the legacy 4.0/4.1 files produced by Scholaris (gzip-wrapped, Spanish schema) through the 5.0 view. ## Install ```julia using Pkg Pkg.add("SPDF") # once registered; until then: Pkg.add(url = "https://github.com/joseluissaorin/spdf", subdir = "julia") ``` ## Read, search, cite ```julia using SPDF SPDF.open("quijote.spdf") do doc println(title(doc), " (", doc.version, ")") for hit in search_lexical(doc, "hermoso"; limit = 5) # the 1608 edition prints «hermoſo» f = SPDF.fragment(doc, hit.fragment_id) println(cite(doc, hit.anchor, f["anchor_end"]; locale = "es")) # (Cervantes Saavedra, 1608, s. p.) println(hit.anchor_uri) # spdf:sha256-27ea…#p=13&char=10,194 end # a quotation is cited by the page it lies in, not by the start of its fragment (§18.2) println(cite_passage(doc, "q5", "rozin, como tomaua la podadera.")["text"]) # (Cervantes Saavedra, 1608, fol. [Iv]) println(bibtex(doc)) end ``` `units(doc)`, `fragments(doc)`, `sections(doc)`, `figures(doc)`, `spaces(doc)`, `blobs(doc)` and `provenance(doc)` return vectors of `Dict`s with the 5.0 column names; `metadata(doc)` is the CSL-JSON item. `vectors(doc, space)` gives the decoded vectors (f32, f16 or i8) as `Dict(id => Vector{Float64})`. ```julia search_vector(doc, qvec, "embeddinggemma-2@768"; limit = 10) search_hybrid(doc, "ciego jarro", qvec, "embeddinggemma-2@768") # RRF, k = 10 parse_uri("spdf:sha256-3f2a…#p=29&f=21&char=118,301") locate(doc, "spdf:sha256-3f2a…#p=29") # Dict("document", "units", "fragments", "char", "xywh") csl_item(doc)["id"] # "cervantessaavedra1605", the BibTeX key (SPEC §19) validate("file.spdf") # Dict("valid" => true, "errors" => [], "warnings" => [], …) ``` ## ALTO, TEI and IIIF ```julia write("quijote.alto.xml", alto(doc)) # ALTO 4, one Page per page unit write("quijote.tei.xml", tei(doc)) # TEI P5: pb, p, lg/l, u, note manifest = iiif(doc, "https://example.org/iiif/quijote") # IIIF Presentation 3 (a Dict) ``` ## Write ```julia w = SPDF.Writer("out.spdf"; generator = "my-app/1.0", profile = "core") SPDF.document!(w, Dict("id" => "d1", "kind" => "pdf", "source_sha256" => bytes2hex(sha256(read("d1.pdf"))), "mime" => "application/pdf", "bytes" => filesize("d1.pdf"), "unit_count" => 1, "metadata" => Dict("type" => "book", "title" => "Lazarillo de Tormes", "issued" => Dict("date-parts" => [[1554]])))) SPDF.unit!(w, Dict("id" => "p1", "ord" => 1, "anchor" => Dict("type" => "page", "physical" => 1, "printed" => "3"), "text" => "Pues sepa Vuestra Merced…", "reader" => "pdf-text-layer")) SPDF.fragment!(w, Dict("n" => 1, "id" => "f1", "unit" => "p1", "ord" => 1, "text" => "Pues sepa Vuestra Merced…", "anchor" => Dict("type" => "page", "physical" => 1, "printed" => "3"))) SPDF.finish!(w) # FTS rebuilt, no triggers or views, VACUUM, atomic rename ``` ## Security Files are untrusted input: `SPDF.open` connects read-only (`mode=ro`), sets `query_only`, `trusted_schema=OFF` and `SQLITE_DBCONFIG_DEFENSIVE`, never loads extensions, refuses triggers, views and foreign virtual tables (except the three FTS triggers of legacy files), bounds blob sizes and gzip inflation, and copies WAL-mode files before opening them. Signatures are verified with a small pure-Julia Ed25519. ## Conformance `SPDF.conformance("path/to/spdf/conformance")` runs the shared suite; `Pkg.test()` runs it too when the package lives in the SPDF repository, and `julia --project=. bin/conformance.jl ../conformance` prints the JSON report. CI publishes it as the `conformance-julia` artifact. All kinds are claimed, `export_structure` included. ## License MIT OR Apache-2.0, at your option. The SPDF specification is CC BY 4.0. --- # SPDF en C URL: https://spdf.joseluissaorin.com/es/documentacion/c > Cómo instalar y usar la implementación de SPDF en C (libspdf): abrir, validar, buscar y citar. ABI de C sobre el núcleo de Rust, para C, C++ y cualquier FFI. - **Paquete**: `libspdf` - **Instalar**: `#include "spdf.h" /* link with -lspdf */` - **Nivel**: segundo - **CI**: CI en marcha - **Carpeta**: `c/` *El README de la biblioteca está en inglés.* C and C++ access to **SPDF** files (Semantic Processed Document Format) through the C ABI of the Rust reference implementation (`rust/crates/spdf-ffi`, header `spdf.h`). Everything the other implementations do is here: safe opening of SPDF 5.0 and legacy 4.x files, validation, the canonical dump, the reference lexical, vector and hybrid searches, anchor URIs (format, parse, locate), short citations, CSL-JSON and BibTeX, and writing files from a dump. SQLite is bundled inside the library. This folder adds: - `CMakeLists.txt`: builds the Rust library with cargo and exposes the CMake target `spdf::spdf` (static by default, `-DSPDF_SHARED=ON` for the shared library, `-DSPDF_PREBUILT_DIR=…` to use a prebuilt `libspdf_ffi`); - `include/spdf.hpp`: a header-only C++17 wrapper (RAII `spdf::Document`, exceptions, `std::string` results); - `examples/`: `search.c`, `validate.c` and `search.cpp`; - `tests/conformance.c`: the conformance runner, which drives every case through the ABI; - `src/sjson.c`: a small JSON reader used by the examples and the runner (not part of the ABI); - `spdf.pc.in`: a pkg-config file for installations. ## Build ```sh cmake -S c -B build -DCMAKE_BUILD_TYPE=Release cmake --build build ctest --test-dir build --output-on-failure ``` Requirements: a C11 and C++17 compiler, CMake 3.16+, and Rust (cargo) unless `SPDF_PREBUILT_DIR` points to a built library. Static linking pulls in `-lpthread -ldl -lm` on Linux and `-framework Security -framework CoreFoundation` on macOS; the CMake target adds them. ## C ```c #include "spdf.h" SpdfDoc *doc = NULL; if (spdf_open("quijote.spdf", NULL, &doc) != SPDF_OK) { /* read-only, safe opening */ fprintf(stderr, "%s\n", spdf_last_error()); /* {"status","code","message"} */ return 1; } char *hits = NULL; if (spdf_search_lexical(doc, "hermoso", 10, &hits) == SPDF_OK) { /* the 1608 edition prints «hermoſo» */ puts(hits); /* [{"fragment_id":"q1","score":…,"via":["lexical"],"anchor":{…},"anchor_uri":"spdf:sha256-27ea…#p=13&char=10,194"}] */ spdf_string_free(hits); } char *cite = NULL; /* a quotation is cited by the page it lies in (SPEC §18.2) */ spdf_cite_passage(doc, "q5", "rozin, como tomaua la podadera.", "es", &cite); puts(cite); /* {"text":"(Cervantes Saavedra, 1608, fol. [Iv])","uri":"spdf:sha256-27ea…#p=30&f=Iv&char=130,161",…} */ spdf_string_free(cite); spdf_close(doc); ``` Every function returns an `int` status (`SPDF_OK` = 0) and writes its result through an out parameter; strings returned by the library are freed with `spdf_string_free`, byte buffers with `spdf_bytes_free`. Complex values are JSON with the shapes of the specification. ## C++ ```cpp #include "spdf.hpp" spdf::Document doc("quijote.spdf"); std::string hits = doc.search("hidalgo", 5); // JSON array std::string where = doc.locate("spdf:sha256-…#p=5"); // {"document","units","fragments","char","xywh"} std::cout << doc.bibtex(); // @book{cervantessaavedra1608, … std::string tei = doc.tei(); // also alto() and iiif(base_url) std::string bib = spdf::export_bibtex({&doc, &other}); // several documents, keys disambiguated ``` Errors throw `spdf::Error` (with `status()` and the JSON of `spdf_last_error()`). ## Conformance `build/spdf_conformance ../conformance` runs every case through the ABI and prints the report of the specification; `ctest` runs it. CI publishes it as the `conformance-c` artifact. All kinds are claimed: `roundtrip` through `spdf_write_from_dump`, `quantize` through `spdf_quantize`, `locate` through `spdf_locate`, the exports through `spdf_export_csl_multi`, `spdf_export_bibtex_multi` and `spdf_export_structure`, and `cite_passage` through `spdf_cite_passage`. ## License MIT OR Apache-2.0, at your option. The SPDF specification is CC BY 4.0. --- # Servidor MCP: spdf-mcp URL: https://spdf.joseluissaorin.com/es/integraciones/mcp > Cualquier agente busca en una carpeta de SPDF y cita con el folio exacto. *El README de la integración está en inglés, como su código.* A [Model Context Protocol](https://modelcontextprotocol.io) server for [SPDF](https://spdf.joseluissaorin.com) files. Point it at a folder of `.spdf` documents and any agent (Claude, ChatGPT, Cursor, Zed, your own) can list them, search them, read a passage, look at the figures and **cite with the exact printed folio, without being able to invent one**. It is a thin layer over [`spdf-format`](../../js), the official TypeScript implementation: the text comes from the file, the citation comes from the stored anchor, and a page that does not exist is an error, never an approximation. ## Run it ```sh npx spdf-mcp ~/Library/SPDF # stdio (what desktop clients use) npx spdf-mcp ~/Library/SPDF --locale es # citations in Spanish by default npx spdf-mcp ~/Library/SPDF --http 8765 # Streamable HTTP on http://127.0.0.1:8765/mcp ``` Options: `--http `, `--host ` (default `127.0.0.1`), `--locale en|es`, `--no-recursive`. Several folders or files can be given. Unreadable or unsafe files (for example one with a view or a trigger) are skipped and reported by `list_documents`. ### Claude Code ```sh claude mcp add spdf -- npx spdf-mcp ~/Library/SPDF ``` ### Any client with a JSON configuration (Claude Desktop, Cursor, Zed…) ```json { "mcpServers": { "spdf": { "command": "npx", "args": ["spdf-mcp", "/path/to/library"] } } } ``` ## Tools All tools are read-only. Each returns readable JSON as text and the same data as structured content. | Tool | Arguments | Returns | | --- | --- | --- | | `list_documents` | `filter?`, `refresh?` | Every document: `doc` reference (`sha256-…` of the original), title, authors, year, kind, language, units, fragments, figures, vector spaces; and the files that were skipped | | `search` | `query`, `docs?`, `limit?` (1–50), `vector?` + `space?`, `locale?` | Passages with their literal `text`, `citation`, `anchor_uri`, section, context and score. Lexical search uses the SPDF reference algorithm (accent-insensitive, `"phrases"`); with a query vector in a space the files carry it is hybrid (reciprocal rank fusion, k = 10). Results from several files are merged by rank | | `read_passage` | `doc` + one of `fragment_id`, `folio`, `page`, `time`; or `uri`; `around?`; `locale?` | The literal text, citation, anchor URI, who read it and with what confidence, warnings, and optionally the neighbouring fragments | | `cite` | same as `read_passage`, plus `reference?` | `citation`, `anchor_uri` and `quote` together; with `reference` also CSL-JSON and BibTeX | | `list_figures` | `doc?`, `figure_id?`, `include_image?`, `locale?` | Figures, plates and frames with caption, description, citation, anchor URI and region; the image itself when asked for one figure | | `get_metadata` | `doc` | CSL-JSON, BibTeX, rights, SHA-256 of the original and provenance | `doc` accepts the full reference from `list_documents`, a unique prefix of its hash, the document id or the file name. ### What the citations look like | Passage | `cite` returns | | --- | --- | | Physical page 2, printed folio 1 | `(Saorín Ferrer, 2026, p. 1)` | | A plate whose folio was inferred | `(Saorín Ferrer, 2026, p. [3])` plus a warning to keep the brackets | | A cover with no printed folio | `(Saorín Ferrer, 2026, n. pag.)` (`s. p.` in Spanish) plus a warning | | A folio that does not exist | an error: *"… has no page with printed folio 21. Printed folios: 1, 2, 3, 4, 5. Do not cite it."* | The server also sends the model a short set of instructions (the MCP `instructions` field) with the rules for citing without inventing; they are the same as on [SPDF for agents](https://spdf.joseluissaorin.com/agents). ## As a library ```js import { Library, createServer } from 'spdf-mcp'; const lib = await Library.open(['./library'], { locale: 'en' }); const server = createServer(lib); // an McpServer from @modelcontextprotocol/sdk await server.connect(myTransport); ``` ## Tests ```sh cd js && npm ci && npm run build # the official library, once cd integrations/spdf-mcp && npm ci && npm test ``` The tests use the MCP SDK itself as the client, three ways: in memory (every tool and every error path), over **stdio** against the compiled binary, and over **Streamable HTTP** (including the DNS-rebinding guard). They run against the sample files in [`../fixtures`](../fixtures). ## Security - Files are opened read-only, with `trusted_schema=OFF`, defensive mode and no extensions, and files with triggers or views are refused (`spdf-format` does this; the server never writes). - The HTTP transport listens on `127.0.0.1` by default, is stateless, accepts only `POST /mcp` and rejects requests whose `Host` is not the one it listens on. There is no authentication: do not expose it to a network you do not trust. ## Licence MIT OR Apache-2.0. --- # Cargador de LangChain.js URL: https://spdf.joseluissaorin.com/es/integraciones/langchain-js > Pasajes como documentos de LangChain con su cita y su URI de ancla. *El README de la integración está en inglés, como su código.* A [LangChain.js](https://js.langchain.com) document loader for [SPDF](https://spdf.joseluissaorin.com) files. Every passage becomes a `Document` with its **literal text** and, in the metadata, its **citation with the exact printed folio** (or second, slide, verse) and its **anchor URI**, so a retrieval-augmented answer can cite the page a reader will find on paper instead of a chunk number. It sits on [`spdf-format`](../../js), the official TypeScript implementation: the citation is computed from the anchor stored in the file, never generated. ## Install ```sh npm install spdf-langchain @langchain/core ``` ## Use ```js import { SpdfLoader } from 'spdf-langchain'; const docs = await new SpdfLoader('darwin-origin.spdf').load(); docs[0].pageContent; // the literal passage docs[0].metadata.citation; // '(Darwin, 1859, p. 21)' docs[0].metadata.anchor_uri; // 'spdf:sha256-…#p=29&f=21&char=118,301' // A folder (recursive), Spanish citations, one document per page, skipping broken files: const pages = await new SpdfLoader('library/', { locale: 'es', granularity: 'unit', skipInvalid: true }).load(); // Streaming: for await (const d of new SpdfLoader('library/').lazyLoad()) console.log(d.metadata.citation); ``` When you answer from retrieved documents, quote `pageContent` and cite with `metadata.citation`; keep `metadata.anchor_uri` next to the claim. ## Options | Option | Default | Meaning | | --- | --- | --- | | `granularity` | `'fragment'` | `'fragment'` (passages of 150 to 300 words, the unit SPDF searches and cites) or `'unit'` (whole pages, time spans, slides) | | `locale` | `'en'` | Language of `citation`: `'en'` or `'es'` | | `recursive` | `true` | Descend into subfolders | | `skipInvalid` | `false` | Skip files that cannot be opened safely (with a warning on stderr) instead of failing | | `embeddingsFrom` | none | A vector space stored in the files (for example `all-MiniLM-L6-v2@384`): its vector goes to `metadata.embedding`, so you can index without re-embedding when your query model is the same | ## Metadata All values are scalars (string, number or boolean) and keys whose value would be null are left out, so every vector store accepts them (Chroma, for one, rejects nulls). The keys match the Python loaders. | Key | Example | | --- | --- | | `citation` | `(Saorín Ferrer, 2026, p. 1)`; `p. [3]` when the folio was inferred; `n. pag.` (`s. p.`) when the page has none | | `anchor_uri` | `spdf:sha256-50d9…#p=2&f=1&char=15,307` | | `printed_folio`, `physical_page`, `folio_inferred` | `'1'`, `2`, `false` (page anchors; no `printed_folio` key when the page has none) | | `t0`, `t1`, `speaker` | seconds (time anchors) | | `slide`, `line_from`, `line_to` | slides and verses | | `title`, `authors`, `year`, `language`, `kind` | from the CSL record | | `section`, `context` | `'I. Anchors'`, one line situating the passage | | `fragment_id` or `unit_id`, `ord`, `spdf_doc_id`, `docref`, `source`, `spdf_version`, `anchor_type` | identifiers (`spdf_doc_id`, not `doc_id`: vector stores and parent-document retrievers overwrite `doc_id`) | | `anchor`, `anchor_end` | the full anchors as JSON strings | ## Tests ```sh cd js && npm ci && npm run build # the official library, once cd integrations/langchain-js && npm ci && npm test ``` They run against [`../fixtures`](../fixtures) and include a LangChain retriever over an in-memory vector store that returns documents with their citation intact. ## Licence MIT OR Apache-2.0. --- # Lector de LlamaIndex.TS URL: https://spdf.joseluissaorin.com/es/integraciones/llamaindex-js > Pasajes como documentos de LlamaIndex, con los vectores guardados si los quieres. *El README de la integración está en inglés, como su código.* A [LlamaIndex.TS](https://ts.llamaindex.ai) reader for [SPDF](https://spdf.joseluissaorin.com) files. Every passage becomes a `Document` with its **literal text** and, in the metadata, its **citation with the exact printed folio** (or second, slide, verse) and its **anchor URI**, so a retrieval-augmented answer can cite the page a reader will find on paper instead of a chunk number. It can also hand over the **vectors already stored in the file**, so an index can be built without embedding anything again. It sits on [`spdf-format`](../../js), the official TypeScript implementation: the citation is computed from the anchor stored in the file, never generated. ## Install ```sh npm install spdf-llamaindex @llamaindex/core ``` ## Use ```js import { VectorStoreIndex } from 'llamaindex'; import { SpdfReader } from 'spdf-llamaindex'; const docs = await new SpdfReader({ locale: 'en' }).loadData('library/'); docs[0].metadata.citation; // '(Darwin, 1859, p. 21)' docs[0].metadata.anchor_uri; // 'spdf:sha256-…#p=29&f=21&char=118,301' const index = await VectorStoreIndex.fromDocuments(docs); const nodes = await index.asRetriever({ similarityTopK: 5 }).retrieve('natural selection'); nodes.map((n) => n.node.metadata.citation); ``` The `citation` is visible to the LLM (`MetadataMode.LLM`), so a query engine's answer can quote it; the anchor JSON, hashes and identifiers are excluded from the text that gets embedded (`EXCLUDED_EMBED_METADATA`, `EXCLUDED_LLM_METADATA`). With `SimpleDirectoryReader`, register it for the extension: `fileExtToReader: { spdf: new SpdfReader() }` (it implements `loadDataAsContent`). ### Reusing the stored vectors ```js const docs = await new SpdfReader({ embeddingsFrom: 'all-MiniLM-L6-v2@384' }).loadData('library/'); // docs[i].embedding is set from the file: index with an embed model of the same space. ``` ## Options | Option | Default | Meaning | | --- | --- | --- | | `granularity` | `'fragment'` | `'fragment'` (passages of 150 to 300 words) or `'unit'` (whole pages, time spans, slides) | | `locale` | `'en'` | Language of `citation`: `'en'` or `'es'` | | `recursive` | `true` | Descend into subfolders | | `skipInvalid` | `false` | Skip files that cannot be opened safely instead of failing | | `embeddingsFrom` | none | A vector space stored in the files: sets `Document.embedding` | ## Metadata Scalars only (string, number or boolean; keys with a null value are left out): `citation`, `anchor_uri`, `printed_folio`, `physical_page`, `folio_inferred`, `t0`/`t1`/`speaker`, `slide`, `line_from`/`line_to`, `title`, `authors`, `year`, `language`, `kind`, `section`, `context`, `fragment_id` or `unit_id`, `ord`, `spdf_doc_id`, `docref`, `source`, `spdf_version`, and the full `anchor`/`anchor_end` as JSON strings. See the table in [`spdf-langchain`](../langchain-js/README.md#metadata). ## Tests ```sh cd js && npm ci && npm run build # the official library, once cd integrations/llamaindex-js && npm ci && npm test ``` They run against [`../fixtures`](../fixtures) and build a `VectorStoreIndex` with a deterministic toy embedding (no network, no keys) to check that retrieved nodes keep their citation. ## Licence MIT OR Apache-2.0. --- # Cargador de LangChain (Python) URL: https://spdf.joseluissaorin.com/es/integraciones/langchain-python > El mismo cargador para LangChain en Python. *El README de la integración está en inglés, como su código.* A [LangChain](https://www.langchain.com/) document loader for **SPDF** files (Semantic Processed Document Format), so that retrieval-augmented answers cite the printed page instead of a chunk number. An SPDF file is a SQLite database holding one document that has already been read: every passage (fragment) carries its exact anchor (physical page and printed folio, second of a recording, slide, verse…). `SpdfLoader` turns each passage into a LangChain `Document` whose metadata holds a ready-made short citation such as `(Saorín Ferrer, 2026, p. 1)` and a portable anchor URI such as `spdf:sha256-…#p=2&f=1&char=15,307`, both computed by the official library [`spdf-format`](../../python). ## Install ```bash pip install spdf-langchain ``` Requires Python 3.10 or later, `langchain-core>=0.3` and `spdf-format` (standard library only). From a checkout of the repository: ```bash pip install -e python/ -e "integrations/langchain-python[test]" ``` ## Usage ```python from spdf_langchain import SpdfLoader loader = SpdfLoader("spdf-in-five-pages.spdf", locale="en") # a file, a folder or a list of both docs = loader.load() # or: for doc in loader.lazy_load(): ... docs[0].page_content # the literal passage, exactly as in the source docs[0].metadata["citation"] # '(Saorín Ferrer, 2026, p. 1)' docs[0].metadata["anchor_uri"] # 'spdf:sha256-50d9…5f4c#p=2&f=1&char=15,307' docs[0].id # 'sha256-50d9…5f4c:f2-1' (stable across runs) ``` - A **folder** is searched recursively for `*.spdf` files (hidden files and folders are skipped); a **list** may mix files and folders. `lazy_load()` yields the documents one by one, one file open at a time; `load()`, `aload()` and `alazy_load()` come from `BaseLoader`. - Unsafe or invalid files are refused by `spdf-format`: loading one raises its error (`spdf.UnsafeFileError`, `spdf.NotSpdfError`…, all subclasses of `spdf.SpdfError`) with the validation code (`E020`…) and the file path in the message. A missing path raises `FileNotFoundError`. - Legacy SPDF 4.0 and 4.1 files (also gzip-wrapped) read like 5.0 ones. - Vector stores in recent `langchain-core` versions take the ids from `Document.id`; with older ones, pass them yourself: `store.add_documents(docs, ids=[d.id for d in docs])`. ### Options | Option | Default | Meaning | | --- | --- | --- | | `granularity` | `"fragment"` | `"fragment"`: one document per passage (about 150 to 300 words). `"unit"`: one per page, time span, slide… | | `locale` | `"en"` | Locale of `citation`: `"en"` or `"es"` (`"es-ES"` works; others fall back to English). | | `with_vectors` | `None` | Id of a vector space stored in the files (`"all-MiniLM-L6-v2@384"`). Puts the stored vector in `metadata["vector"]` (a list of floats) and the space id in `metadata["vector_space"]`. Off by default, because most vector stores expect flat metadata. A file without that space raises `spdf.SpdfError`. | ## Metadata Values are flat scalars (`str`, `int`, `float`, `bool`), so every vector store accepts them (the only exception is `vector`, and only if you ask for it). **A key whose value would be null is left out** (Chroma and others reject `None`): the cover of a book has no `printed_folio` key, a page has no `t0`. | Key | Type | Example (first passage of the English fixture) | Notes | | --- | --- | --- | --- | | `source` | str | `fixtures/spdf-in-five-pages.spdf` | Path the file was read from. | | `spdf_version` | str | `5.0` | `4.0` or `4.1` for legacy files. | | `spdf_doc_id` | str | `spdf-in-five-pages` | The document id inside the file. Not called `doc_id`, which LangChain's multi-vector and parent-document retrievers (and LlamaIndex vector stores) use for their own ids. | | `docref` | str | `sha256-50d94244…5f4c` | Document reference used by anchor URIs (SHA-256 of the original). | | `title` | str | `SPDF in five pages` | | | `authors` | str | `Saorín Ferrer` | | | `year` | int | `2026` | | | `language` | str | `en` | BCP 47. | | `kind` | str | `pdf` | `pdf`, `epub`, `audio`, `video`… | | `fragment_id` | str | `f2-1` | Fragment granularity only. | | `unit_id` | str | `u2` | The unit (page…) where the passage starts. | | `anchor_type` | str | `page` | `page`, `time`, `section`, `slide`, `sheet`, `web`, `image`, `verse`, `canonical`. | | `physical_page` | int | `2` | Page anchors: position of the page in the file. | | `printed_folio` | str | `1` | The folio as printed (`"xiv"`, `"1r"`). | | `folio_inferred` | bool | `false` | True when the folio was deduced, not read; the citation prints it in brackets, `p. [3]`. | | `section` | str | `I. Anchors` | Heading path joined with `" / "`. In unit granularity, the sections that share the unit are joined with `" \| "`. | | `context` | str | `SPDF in five pages, I. Anchors` | One line that situates the passage (fragment granularity). | | `anchor` | str | `{"chars":[15,307],"confidence":1,"physical":2,"printed":"1","source":"read","type":"page"}` | The start anchor as canonical JSON; `json.loads` it for the full object. | | `anchor_end` | str | | End anchor (JSON) when the passage crosses into another unit. | | `t0`, `t1` | float | | Seconds, for time anchors (recordings). | | `anchor_uri` | str | `spdf:sha256-50d94244…5f4c#p=2&f=1&char=15,307` | Resolve it with `spdf.open(path).locate(uri)`; parse it with `spdf.parse_uri`. | | `citation` | str | `(Saorín Ferrer, 2026, p. 1)` | Short author-date citation in the chosen locale. | | `vector`, `vector_space` | list, str | | Only with `with_vectors`. | `page_content` is always the literal passage (`fragments.text` or `units.text`), never the modernised-spelling search layer, which SPDF forbids quoting. ## End-to-end example (no API key) Retrieval with a toy embedding and an in-memory vector store, then a prompt in which every passage carries its citation; pipe the prompt into any chat model. ```python import hashlib import math import re from langchain_core.embeddings import Embeddings from langchain_core.prompts import ChatPromptTemplate from langchain_core.vectorstores import InMemoryVectorStore from spdf_langchain import SpdfLoader class HashingEmbeddings(Embeddings): """A toy bag-of-words embedding: no model to download, no API key.""" def __init__(self, dim: int = 512) -> None: self.dim = dim def _vec(self, text: str) -> list[float]: v = [0.0] * self.dim for word in re.findall(r"\w+", text.lower()): v[int(hashlib.md5(word.encode()).hexdigest(), 16) % self.dim] += 1.0 norm = math.sqrt(sum(x * x for x in v)) or 1.0 return [x / norm for x in v] def embed_documents(self, texts: list[str]) -> list[list[float]]: return [self._vec(t) for t in texts] def embed_query(self, text: str) -> list[float]: return self._vec(text) docs = SpdfLoader("integrations/fixtures/spdf-in-five-pages.spdf", locale="en").load() store = InMemoryVectorStore.from_documents(docs, embedding=HashingEmbeddings()) # needs numpy question = "How is a plate without a printed folio cited?" hits = store.similarity_search(question, k=2) for d in hits: print(d.metadata["citation"], d.metadata["anchor_uri"]) prompt = ChatPromptTemplate.from_messages([ ("system", "Answer from the passages only. After each claim, copy the citation of its passage."), ("human", "{context}\n\nQuestion: {question}"), ]) context = "\n\n".join(f"{d.page_content} {d.metadata['citation']}" for d in hits) messages = prompt.invoke({"context": context, "question": question}) # answer = chat_model.invoke(messages) ``` Output: ```text (Saorín Ferrer, 2026, p. [3]) spdf:sha256-50d94244…5f4c#p=4&f=3&char=0,133 (Saorín Ferrer, 2026, p. 2) spdf:sha256-50d94244…5f4c#p=3&f=2&char=19,258 ``` The plate carries no printed number; its folio is inferred, so the citation prints it in brackets. The context the model receives reads: ```text Plate I. A page with its folio and a manicule pointing at a passage. This plate carries no printed number; its folio, 3, is inferred. (Saorín Ferrer, 2026, p. [3]) Every unit records who read it: … the citation puts it in brackets. (Saorín Ferrer, 2026, p. 2) ``` ## Reusing the vectors stored in the file SPDF files may ship vectors (`f.spaces()` in `spdf-format` lists them, with model, size and any task prefixes). When your embedding model is the one that produced a space, load the vectors instead of embedding every passage again, for example with FAISS: ```python from langchain_community.vectorstores import FAISS # pip install langchain-community faiss-cpu from langchain_huggingface import HuggingFaceEmbeddings # pip install langchain-huggingface from spdf_langchain import SpdfLoader docs = SpdfLoader("library/", with_vectors="all-MiniLM-L6-v2@384").load() pairs = [(d.page_content, d.metadata.pop("vector")) for d in docs] # keep the metadata flat store = FAISS.from_embeddings( pairs, HuggingFaceEmbeddings(model_name="sentence-transformers/all-MiniLM-L6-v2"), # embeds queries only metadatas=[d.metadata for d in docs], ids=[d.id for d in docs], ) print(store.similarity_search("inferred folio", k=1)[0].metadata["citation"]) ``` ## Development ```bash cd integrations/langchain-python uv venv && uv pip install -e ../../python -e ".[test]" .venv/bin/python -m pytest ``` The tests use the shared fixtures in `integrations/fixtures/`. ## License MIT OR Apache-2.0, at your option. --- # Lector de LlamaIndex (Python) URL: https://spdf.joseluissaorin.com/es/integraciones/llamaindex-python > El mismo lector para LlamaIndex en Python. *El README de la integración está en inglés, como su código.* A [LlamaIndex](https://www.llamaindex.ai/) reader for **SPDF** files (Semantic Processed Document Format), so that retrieval-augmented answers cite the printed page instead of a chunk number. An SPDF file is a SQLite database holding one document that has already been read: every passage (fragment) carries its exact anchor (physical page and printed folio, second of a recording, slide, verse…). `SpdfReader` turns each passage into a LlamaIndex `Document` whose metadata holds a ready-made short citation such as `(Saorín Ferrer, 2026, p. 1)` and a portable anchor URI such as `spdf:sha256-…#p=2&f=1&char=15,307`, both computed by the official library [`spdf-format`](../../python). ## Install ```bash pip install spdf-llamaindex ``` Requires Python 3.10 or later, `llama-index-core>=0.12` and `spdf-format` (standard library only). From a checkout of the repository: ```bash pip install -e python/ -e "integrations/llamaindex-python[test]" ``` ## Usage ```python from spdf_llamaindex import SpdfReader reader = SpdfReader(locale="en") docs = reader.load_data("spdf-in-five-pages.spdf") # a file, a folder or a list of both docs[0].text # the literal passage, exactly as in the source docs[0].metadata["citation"] # '(Saorín Ferrer, 2026, p. 1)' docs[0].metadata["anchor_uri"] # 'spdf:sha256-50d9…5f4c#p=2&f=1&char=15,307' docs[0].id_ # 'sha256-50d9…5f4c:f2-1' (stable across runs) ``` - A **folder** is searched recursively for `*.spdf` files (hidden files and folders are skipped); a **list** may mix files and folders. `lazy_load_data()` yields the documents one by one, one file open at a time. - With `SimpleDirectoryReader`, register the reader for the extension: `SimpleDirectoryReader("library/", file_extractor={".spdf": SpdfReader()})`. An `fsspec` filesystem passed as `fs=` is honoured. - `extra_info={...}` adds metadata to every document (and wins over the reader's keys). - Unsafe or invalid files are refused by `spdf-format`: loading one raises its error (`spdf.UnsafeFileError`, `spdf.NotSpdfError`…, all subclasses of `spdf.SpdfError`) with the validation code (`E020`…) and the file path in the message. A missing path raises `FileNotFoundError`. - Legacy SPDF 4.0 and 4.1 files (also gzip-wrapped) read like 5.0 ones. ### Options | Option | Default | Meaning | | --- | --- | --- | | `granularity` | `"fragment"` | `"fragment"`: one document per passage (about 150 to 300 words). `"unit"`: one per page, time span, slide… | | `locale` | `"en"` | Locale of `citation`: `"en"` or `"es"` (`"es-ES"` works; others fall back to English). | | `include_embeddings` | `None` | Id of a vector space stored in the files (`"all-MiniLM-L6-v2@384"`). Sets `Document.embedding` from the stored vectors, so that an index whose embedding model matches that space does not embed the passages again. A file without that space raises `spdf.SpdfError`; a passage without a stored vector keeps `embedding=None` and is embedded by the index. | | `excluded_embed_metadata_keys` | all but `title`, `section` | Keys kept out of the text that is embedded. | | `excluded_llm_metadata_keys` | all but `title`, `authors`, `year`, `section`, `citation` | Keys hidden from the LLM. By default the LLM sees the `citation` line next to each passage and can copy it into its answer. | ## Metadata Values are flat scalars (`str`, `int`, `float`, `bool`), so every vector store accepts them. **A key whose value would be null is left out** (Chroma and others reject `None`): the cover of a book has no `printed_folio` key, a page has no `t0`. | Key | Type | Example (first passage of the English fixture) | Notes | | --- | --- | --- | --- | | `source` | str | `fixtures/spdf-in-five-pages.spdf` | Path the file was read from. | | `spdf_version` | str | `5.0` | `4.0` or `4.1` for legacy files. | | `spdf_doc_id` | str | `spdf-in-five-pages` | The document id inside the file. Not called `doc_id`: LlamaIndex vector stores overwrite `doc_id`, `document_id` and `ref_doc_id` with the node's reference document id. | | `docref` | str | `sha256-50d94244…5f4c` | Document reference used by anchor URIs (SHA-256 of the original). | | `title` | str | `SPDF in five pages` | | | `authors` | str | `Saorín Ferrer` | | | `year` | int | `2026` | | | `language` | str | `en` | BCP 47. | | `kind` | str | `pdf` | `pdf`, `epub`, `audio`, `video`… | | `fragment_id` | str | `f2-1` | Fragment granularity only. | | `unit_id` | str | `u2` | The unit (page…) where the passage starts. | | `anchor_type` | str | `page` | `page`, `time`, `section`, `slide`, `sheet`, `web`, `image`, `verse`, `canonical`. | | `physical_page` | int | `2` | Page anchors: position of the page in the file. | | `printed_folio` | str | `1` | The folio as printed (`"xiv"`, `"1r"`). | | `folio_inferred` | bool | `false` | True when the folio was deduced, not read; the citation prints it in brackets, `p. [3]`. | | `section` | str | `I. Anchors` | Heading path joined with `" / "`. In unit granularity, the sections that share the unit are joined with `" \| "`. | | `context` | str | `SPDF in five pages, I. Anchors` | One line that situates the passage (fragment granularity). | | `anchor` | str | `{"chars":[15,307],"confidence":1,"physical":2,"printed":"1","source":"read","type":"page"}` | The start anchor as canonical JSON; `json.loads` it for the full object. | | `anchor_end` | str | | End anchor (JSON) when the passage crosses into another unit. | | `t0`, `t1` | float | | Seconds, for time anchors (recordings). | | `anchor_uri` | str | `spdf:sha256-50d94244…5f4c#p=2&f=1&char=15,307` | Resolve it with `spdf.open(path).locate(uri)`; parse it with `spdf.parse_uri`. | | `citation` | str | `(Saorín Ferrer, 2026, p. 1)` | Short author-date citation in the chosen locale. | | `vector_space` | str | `all-MiniLM-L6-v2@384` | Only when `include_embeddings` attached a vector. | The document text is always the literal passage (`fragments.text` or `units.text`), never the modernised-spelling search layer, which SPDF forbids quoting. ## End-to-end example (no API key) A complete retrieval-augmented query with a toy embedding and LlamaIndex's `MockLLM`; swap them for your models. The sources of the answer carry their citations. ```python import hashlib import math import re from llama_index.core import VectorStoreIndex from llama_index.core.embeddings import BaseEmbedding from llama_index.core.llms import MockLLM from spdf_llamaindex import SpdfReader class HashingEmbedding(BaseEmbedding): """A toy bag-of-words embedding: no model to download, no API key.""" dim: int = 512 def _vec(self, text: str) -> list[float]: v = [0.0] * self.dim for word in re.findall(r"\w+", text.lower()): v[int(hashlib.md5(word.encode()).hexdigest(), 16) % self.dim] += 1.0 norm = math.sqrt(sum(x * x for x in v)) or 1.0 return [x / norm for x in v] def _get_text_embedding(self, text: str) -> list[float]: return self._vec(text) def _get_query_embedding(self, query: str) -> list[float]: return self._vec(query) async def _aget_query_embedding(self, query: str) -> list[float]: return self._vec(query) docs = SpdfReader(locale="en").load_data("integrations/fixtures/spdf-in-five-pages.spdf") index = VectorStoreIndex(docs, embed_model=HashingEmbedding()) engine = index.as_query_engine(llm=MockLLM(), similarity_top_k=2) response = engine.query("How is a plate without a printed folio cited?") for source in response.source_nodes: print(source.node.metadata["citation"], source.node.metadata["anchor_uri"]) ``` Output: ```text (Saorín Ferrer, 2026, p. [3]) spdf:sha256-50d94244…5f4c#p=4&f=3&char=0,133 (Saorín Ferrer, 2026, p. 2) spdf:sha256-50d94244…5f4c#p=3&f=2&char=19,258 ``` The plate carries no printed number; its folio is inferred, so the citation prints it in brackets. What the LLM receives for each passage is: ```text title: SPDF in five pages authors: Saorín Ferrer year: 2026 section: III. Read once, query many citation: (Saorín Ferrer, 2026, p. [3]) Plate I. A page with its folio and a manicule pointing at a passage. … ``` Fragments are already passage-sized, so the example passes the documents straight to `VectorStoreIndex(docs, …)`. `VectorStoreIndex.from_documents(docs, …)` also works (the splitter copies the metadata to every node), but it creates new nodes without the stored embeddings. ## Reusing the vectors stored in the file SPDF files may ship vectors (`f.spaces()` in `spdf-format` lists them, with model, size and any task prefixes). When your embedding model is the one that produced a space, load the vectors instead of embedding every passage again: ```python from llama_index.core import VectorStoreIndex from llama_index.embeddings.huggingface import HuggingFaceEmbedding # pip install llama-index-embeddings-huggingface from spdf_llamaindex import SpdfReader docs = SpdfReader(include_embeddings="all-MiniLM-L6-v2@384").load_data("library/") embed = HuggingFaceEmbedding(model_name="sentence-transformers/all-MiniLM-L6-v2") # local, embeds queries only index = VectorStoreIndex(docs, embed_model=embed) # not from_documents: keep the stored vectors print(index.as_retriever().retrieve("inferred folio")[0].node.metadata["citation"]) ``` ## Development ```bash cd integrations/llamaindex-python uv venv && uv pip install -e ../../python -e ".[test]" .venv/bin/python -m pytest ``` The tests use the shared fixtures in `integrations/fixtures/`. ## License MIT OR Apache-2.0, at your option. --- # Complemento para Zotero 7 y 8 URL: https://spdf.joseluissaorin.com/es/integraciones/zotero > Importa un SPDF como ítem, adjúntalo y copia una cita con el folio. *El README de la integración está en inglés, como su código.* A Zotero 7 and Zotero 8 plugin for [SPDF](https://spdf.joseluissaorin.com) (Semantic Processed Document Format) files: documents that were read once and keep, for every passage, its exact anchor (printed folio, physical page, second of a recording, slide…). With the plugin, Zotero can: - **Import SPDF as Item…** (Tools menu and item context menu): pick an `.spdf` file and get a Zotero item built from the file's own CSL-JSON metadata, in the selected library and collection, with the `.spdf` file attached. - **Attach SPDF…** (item context menu, one regular item selected): attach an `.spdf` file to an item you already have. - **Copy Citation with Folio…** (item context menu, an item with an SPDF attachment, or the attachment itself, or a sibling attachment such as the PDF): type a printed folio, a physical page or an anchor URI, and the clipboard receives the short citation and, on the next line, the anchor URI: ``` (Saorín Ferrer, 2026, p. [3]) spdf:sha256-50d942445564fe24effe743701f0c9a16fde414e6eb1f0ce2095b866d0555f4c#p=4&f=3 ``` The citation is always computed from the anchor stored in the file, never from what was typed, so it can only print what the source says: an inferred folio is printed in brackets (`p. [3]`), a page without folio is `s. p.` in Spanish and `n. pag.` in English, and a folio or page that the document does not have is reported as missing and nothing is copied. The menus and dialogs are in English and Spanish (Fluent, `locale/en-US` and `locale/es-ES`). Citations follow Zotero's interface language: Spanish when it is Spanish, English otherwise (SPEC §18). ## Install 1. Download `spdf-zotero-.xpi` (or build it, see below). 2. In Zotero: **Tools → Plugins**, then the gear menu → **Install Plugin From File…**, and choose the `.xpi`. The plugin needs Zotero 7 or Zotero 8 (`strict_min_version` 6.999, `strict_max_version` 8.*). It installs nothing else and downloads nothing. ## Use **Import SPDF as Item…** creates the item with `Zotero.Utilities.Item.itemFromCSLJSON` from `documents.metadata` without its `spdf` extension (exactly what `spdf-format`'s `toCslJson` exports). Legacy Scholaris files (SPDF 4.0 and 4.1, usually gzip-wrapped, with Spanish table names) are read too; their metadata is mapped to CSL-JSON by `spdf-format` (`mapLegacyMetadata`). The file is then copied into Zotero's storage with `Zotero.Attachments.importFromFile` (with the media type `spdf-format` declares, `MEDIA_TYPE`; the specification fixes `application/vnd.spdf+sqlite3`), and the new item is selected. Attachments are recognised as SPDF by that type, by the older `application/vnd.spdf` or by the `.spdf` extension. Both Import and Attach add one line to the item's **Extra** field: ``` SPDF: sha256-<64 hex digits> ``` That is the document reference of the file (its `source_sha256`), the same one anchor URIs carry, so a URI found in a manuscript can be matched to the item later. Other lines of Extra are kept; the line is not repeated. **Copy Citation with Folio…** accepts: | You type | Meaning | | --- | --- | | `145`, `xiv`, `1r`, `p. 145`, `pág. 12` | a printed folio, as printed (`XIV` also finds `xiv`) | | `[21]` | the same, with the brackets of an inferred folio | | `145-146`, `pp. 2-[3]` | a range of folios | | `p=12`, `#12`, `p=12-13` | a physical page (position in the original), or a range | | `f=A-3` | a printed folio, explicitly (for folios that contain a dash) | | `spdf:sha256-…#p=29&f=21` | a full anchor URI; `char=` and `xywh=` are kept | | `spdf:sha256-…` | the whole document: `(Saorín Ferrer, 2026)` | An anchor URI is resolved as SPEC §5.4 says: `p` first, then `f`, then `t` (time), then slides, sheets, verses, canonical references and sections. If the item has several SPDF attachments, a URI picks the one whose document it names, and a folio asks which attachment to use. If two pages carry the same folio, the plugin asks which one. What the plugin refuses: files that are not SQLite, of an unknown version, with views or triggers, or with an unknown required extension (SPEC §2.4). It does not run the full validator when importing (see Design); use the web validator or `npx spdf-format validate file.spdf` to audit a file. ## Design - **The format is not reimplemented.** Everything SPDF-specific (opening safely, legacy 4.x mapping, anchors and anchor URIs, the short citation, CSL-JSON export) comes from the official TypeScript library, `spdf-format`, whose engine-less core (`spdf-format/core`) is bundled into the plugin by esbuild. - **SQLite is Zotero's own.** `src/engine.ts` is a `spdf-format` `SqlEngine` over `Sqlite.sys.mjs` (`resource://gre/modules/Sqlite.sys.mjs`), the asynchronous mozStorage wrapper Zotero 7 (Firefox 115) and Zotero 8 (Firefox 140) ship. The whole `SpdfDocument` API (units, fragments, `cite`, `anchorUri`, search, `validate`, `dump`) therefore works inside Zotero with no second SQLite in the package. - **Read only.** Files are opened with `Sqlite.openConnection({ path, readOnly: true })` (`SQLITE_OPEN_READONLY`); mozStorage keeps extension loading off; the core then sets `PRAGMA query_only = 1` and `PRAGMA trusted_schema = OFF`, and refuses views and triggers. Only temporary copies (a decompressed legacy file, the private copy used by the FTS5 integrity check) are written, in Zotero's temp directory, and they are deleted when the connection closes. - **Column names.** A `mozIStorageRow` can be read by index or by a known name but cannot list its columns, while the core expects rows keyed by column name. `src/sql.ts` infers the names the way SQLite assigns them (the `AS` alias, the column of a bare reference, the expression text otherwise, fixed names for PRAGMAs), and the engine checks every inferred name against the first row with `getResultByName`, so a wrong guess is an error, never mislabelled data. The tests check the inference against SQLite's own names and run every statement the core issues through it. - **Gzip.** Legacy files are gunzipped with `DecompressionStream`, taken from Zotero's main window because the plugin sandbox does not have the Compression Streams API, with the 4 GiB output limit of SPEC §2.3. - **Compartments.** Bytes from `IOUtils`, mozStorage or the main window are copied into typed arrays of the plugin's own realm, because the core tests `instanceof Uint8Array`. The sandbox also lacks `structuredClone`, which the core uses to copy metadata; the bundle gets a JSON-based fallback (`src/shims/structured-clone.ts`). - **One file.** `bootstrap.js` is the whole plugin: the bundle followed by the bootstrap hooks. The plugin never loads a second script from its own `jar:` URL, and the file is pure ASCII so its decoding never depends on the script loader. - **Menus.** On Zotero 8 the entries are registered with `Zotero.MenuManager` (targets `main/menubar/tools` and `main/library/item`); on Zotero 7, which has no menu API, they are added to `menu_ToolsPopup` and `zotero-itemmenu` in each main window, as Zotero's sample plugin does. Labels come from Fluent in both cases. - **No full validation on import.** The validator's content-hash step re-reads every blob (page images, the original PDF), and mozStorage returns BLOBs as JavaScript arrays of numbers, which is slow and memory-hungry for a large book. Firefox's SQLite may also lack FTS5, which the integrity check needs. A reference manager only needs the metadata and the anchors, and the safety checks a reader must make are made. ## Build and test ```sh cd integrations/zotero npm install # spdf-format comes from ../../js (file: dependency) npm run build # dist/addon/ and dist/spdf-zotero-.xpi npm test # typecheck, build, then all tests ``` `spdf-format` must be built first (`cd js && npm run build`), since the plugin bundles `js/dist`. The build is reproducible: same sources, same `.xpi` bytes. For development, Zotero can load the unpacked plugin: in a **test profile**, create a text file named `spdf@joseluissaorin.com` in the profile's `extensions` directory whose only line is the absolute path of `dist/addon/`, then start Zotero with `-purgecaches`. ### How it is tested There is no Zotero in the test run. Everything that does not need Zotero runs in Node with `node:sqlite`, against the shared fixtures (`integrations/fixtures`) and the legacy conformance files (`conformance/legacy`). That corpus is rebuilt from time to time, so the tests find legacy files by the kind of anchor they hold (pages, times, sections) and take the expected citations from the official Node engine of `spdf-format`, checked against the rules of SPEC §18: - `test/engine.test.ts`: the adapter runs over a stand-in for `Sqlite.sys.mjs` built on `node:sqlite` whose rows behave like `mozIStorageRow` (values by index, no column names, BLOBs as arrays of octets, one statement per call, the same parameter binding rules). Through it, every fixture and legacy file gives the same canonical dump, units, fragments, citations, search results and validation report as the official Node engine of `spdf-format`. Read-only opening, temp-file cleanup, gzip limits and refusals (E001, E020) are checked too. - `test/sql.test.ts`: column-name inference compared with SQLite's own names, statement splitting, placeholder renaming. - `test/locate.test.ts`: folio, page, range and URI lookup and the citations, for example `(Saorín Ferrer, 2026, p. 1)` for physical page 2, `p. [3]` for the inferred plate, `s. p.` / `n. pag.` for the cover; in legacy files every printed folio, physical page, time (`h:mm:ss` from one hour on) and section paragraph; and "not found" for folios, pages, times and paragraphs that do not exist. - `test/commands.test.ts`: the three commands against a fake `Zotero` that records items, CSL-JSON, attachments, Extra and the clipboard. - `test/plugin.test.ts`: startup, both menu paths (DOM for Zotero 7, `MenuManager` for Zotero 8), windows opening and closing, shutdown, and the real host code (`Services.prompt`, `FilePicker`, `IOUtils`, `PathUtils`, `Localization`) over fakes. - `test/l10n.test.ts`: English and Spanish have the same messages, every id used in the code exists, Spanish has its accents and « » quotes. - `test/xpi.test.ts`: the `.xpi` unzips, the manifest declares `spdf@joseluissaorin.com`, 6.999 to 8.*, and `bootstrap.js` itself runs startup, import, citation and shutdown in a `vm` context holding only the globals of Zotero's plugin sandbox (no `window`, `console`, `DecompressionStream` or `structuredClone`). ### What still has to be checked by hand None of this has run inside a real Zotero yet. Before a release, in a test profile of **Zotero 7** and of **Zotero 8**: 1. The plugin installs from the `.xpi`, shows in Tools → Plugins, and can be disabled, enabled and removed without errors in the Error Console. 2. The three entries appear with their labels (English and Spanish UI); in Zotero 7 in the Tools menu and the item context menu (DOM path); in Zotero 8 through `Zotero.MenuManager`, and Attach / Copy citation appear only when they apply. 3. Import: the file picker filters `.spdf`; the item gets the right type, creators, date and Extra line; the file is copied into storage; it lands in the selected collection; a read-only group library is refused. 4. `Sqlite.openConnection({ readOnly: true })` opens files in Zotero's storage and in arbitrary folders, and the column-name check passes on real `mozIStorageRow`s (the fake models them from the IDL and Zotero's own use of them). 5. A legacy gzip-wrapped 4.x file imports (`DecompressionStream` from the main window, temp file in Zotero's `tmp` directory, removed afterwards). 6. Copy Citation: `Services.prompt` dialogs, the clipboard content, the progress notice; the "not found" warnings. 7. Large files (hundreds of MB, thousands of pages) open and cite in reasonable time. 8. The `update_url` in the manifest (`https://spdf.joseluissaorin.com/zotero/updates.json`) is served by the website, or is removed; until then Zotero simply finds no updates. ## License MIT OR Apache-2.0, like the rest of the SPDF code. The specification is CC BY 4.0. --- # Filtro de Pandoc URL: https://spdf.joseluissaorin.com/es/integraciones/pandoc > Las anclas de SPDF en Markdown se convierten en citas con el folio impreso, en cualquier estilo CSL. *El README de la integración está en inglés, como su código.* `spdf.lua` is a [Pandoc](https://pandoc.org) Lua filter that lets you cite a place in an [SPDF](../../spec/SPEC.md) file from Markdown and get a real citation, formatted by citeproc in any CSL style, with the folio **printed in the source**. An SPDF file is a document that has already been read: every unit knows its physical position in the file and the folio printed on the page (or its second, slide, verse…). You write the anchor; the filter looks it up in your SPDF files, adds the work to the bibliography and hands citeproc the locator a reader will find on paper: ```markdown --- spdf-library: ~/Library/SPDF --- Physical page 2 carries the printed folio 1 [@spdf:sha256-50d942445564fe24effe743701f0c9a16fde414e6eb1f0ce2095b866d0555f4c#p=2]. An inferred folio goes in brackets [@spdf:sha256-50d942445564fe24effe743701f0c9a16fde414e6eb1f0ce2095b866d0555f4c#p=4]. The cover has no folio [@spdf:sha256-50d942445564fe24effe743701f0c9a16fde414e6eb1f0ce2095b866d0555f4c#p=1]. # References ``` ```console $ pandoc paper.md --lua-filter spdf.lua --citeproc -t plain --wrap=none [WARNING] Scripting warning at spdf.lua line 46 column 1: spdf: [@spdf:sha256-50d942445564fe24effe743701f0c9a16fde414e6eb1f0ce2095b866d0555f4c#p=1]: physical page 1 of spdf-in-five-pages.spdf has no printed folio; cited as unnumbered (n. pag.) Physical page 2 carries the printed folio 1 (Saorín Ferrer 2026, 1). An inferred folio goes in brackets (Saorín Ferrer 2026, [3]). The cover has no folio (Saorín Ferrer 2026, n. pag.). References Saorín Ferrer, José Luis. 2026. SPDF in Five Pages. Spdf.joseluissaorin.com. https://spdf.joseluissaorin.com/validator. ``` Every output in this file is real, produced with Pandoc 3.9 and its default style, Chicago author-date, mostly from the sample booklet [`integrations/fixtures/spdf-in-five-pages.spdf`](../fixtures/) (six page units: a cover without folio, then printed folios 1, 2, an inferred [3], 4 and 5); the same cases are checked by the tests in [`test/`](test/). ## Why Pandoc, and not Calibre Academic writing in Markdown already goes through Pandoc and citeproc: that is the step where an author's sources become citations in the style a journal or a university asks for. Turning a position in a file (the 29th page of a PDF) into the folio a reader finds on paper (p. 21, or p. [21] when the folio was inferred, or *n. pag.* when there is none) belongs exactly there, just before the style formats it. Calibre is a reading library, and reading SPDF files is what the SPDF Reader is for; a Calibre plugin would be one more place to read, and would not help anyone cite. ## Requirements and installation - **Pandoc 3.1.1 or later** (the filter uses `pandoc.json`; tested with 3.9). - **The `sqlite3` command-line tool**, 3.33 or later (JSON output); 3.37 or later is recommended, because the filter then runs it in safe mode. macOS ships it; on Debian or Ubuntu `apt install sqlite3`, on Fedora `dnf install sqlite`. No compiled Lua module is needed. - For legacy SPDF 4.x files, which are gzip-compressed: `gzip`, `head` and a POSIX `sh` (all standard on macOS and Linux). Copy [`spdf.lua`](spdf.lua) next to your document, or into the `filters` folder of your Pandoc user data directory (`~/.local/share/pandoc/filters/` on macOS and Linux), where `--lua-filter spdf.lua` finds it from anywhere. Run it **before** citeproc: ```sh pandoc paper.md --lua-filter spdf.lua --citeproc -o paper.pdf ``` or with a defaults file (`pandoc -d spdf.yaml paper.md -o paper.pdf`): ```yaml # spdf.yaml filters: - spdf.lua - citeproc metadata: spdf-library: ~/Library/SPDF ``` ## Writing citations The key of the citation is an SPDF anchor URI ([SPEC §5](../../spec/SPEC.md#anchor-uri)): `spdf:` followed by the document reference and, after `#`, the anchor parameters. - The document reference is `sha256-` and the 64 hexadecimal digits of the document's `source_sha256` (recommended: it is the same in every copy of the file), or the document id, percent-encoded. To read both from a file: `sqlite3 -readonly book.spdf "SELECT source_sha256, id FROM documents"`. - The parameters say where: `p=` the physical page (the position in the file), `f=` the printed folio, `pe=` / `fe=` the end of a range, `t=` seconds (`t=4160`, `t=12,24.5`, `t=1:09:20`), `s=` and `para=` a section path and paragraph, `sl=` a slide, `sh=` and `rows=` a sheet, `v=` verse lines, `ref=` a canonical reference (`ref=stephanus:514a`). `char=` and `xywh=` narrow a unit and do not change the citation. Everything Pandoc offers around a citation keeps working: prefixes, suffixes, several citations in one bracket, suppressing the author, in-text citations. To keep them short, the examples from here on cite the booklet by its document id, `spdf-in-five-pages`; `sha256-50d942445564fe24effe743701f0c9a16fde414e6eb1f0ce2095b866d0555f4c` gives the same output. ```markdown Prefix and suffix are kept [see @spdf:spdf-in-five-pages#p=5, emphasis added]. Two citations in one bracket [@spdf:spdf-in-five-pages#p=2; compare @spdf:spdf-in-five-pages#p=6, for the summary]. The author can be suppressed [-@spdf:spdf-in-five-pages#p=3]. A range [@spdf:spdf-in-five-pages#p=2&pe=3], or by folios [@spdf:spdf-in-five-pages#f=4&fe=5]. @spdf:spdf-in-five-pages#p=2 says so. As @spdf:spdf-in-five-pages#p=4, argues, inferred folios are bracketed. ``` ```text Prefix and suffix are kept (see Saorín Ferrer 2026, 4, emphasis added). Two citations in one bracket (Saorín Ferrer 2026, 1; compare Saorín Ferrer 2026, 5, for the summary). The author can be suppressed (2026, 2). A range (Saorín Ferrer 2026, 1–2), or by folios (Saorín Ferrer 2026, 4–5). Saorín Ferrer (2026, 1) says so. As Saorín Ferrer (2026, [3]), argues, inferred folios are bracketed. ``` ### How Pandoc reads these keys In Pandoc's Markdown a citation key starts with a letter, a digit or `_` and may contain letters, digits, `_` and the internal punctuation `: . # $ % & - + ? < > ~ /`. The equals sign is not among them, so **the key stops at the first parameter name** and the rest of the anchor lands in the citation's suffix, or, for an in-text citation, in the text that follows. These are the Pandoc 3.9 ASTs (`pandoc -t native`) the filter is built on: | You write | `citationId` | Where the rest goes | |---|---|---| | `[@spdf:sha256-…#p=29]` | `spdf:sha256-…#p` | suffix `=29` | | `[see @spdf:sha256-…#f=21, emphasis added]` | `spdf:sha256-…#f` | suffix `=21, emphasis added`, prefix `see` | | `[@spdf:sha256-…#p=2&f=1]` | `spdf:sha256-…#p` | suffix `=2&f=1` | | `@spdf:my-doc#p=3 says` | `spdf:my-doc#p` (in-text) | the next text element, `=3` | | `[@{spdf:sha256-…#p=2&pe=3}]` | the whole URI | nothing (braced form) | | `@{spdf:sha256-…#p=2} [emphasis added]` | the whole URI (in-text) | suffix `emphasis added` | | `[@spdf:sha256-…]` | `spdf:sha256-…` | the whole document, no locator | The filter puts the pieces back together, so all of these work as written. A few forms do not reach the filter as citations, and need another spelling: | Instead of | Write | Why | |---|---|---| | `@spdf:…#p=2 [emphasis added]` | `@{spdf:…#p=2} [emphasis added]` | Pandoc attaches a bracketed suffix only directly after the key, and here `=2` comes between them. | | `[@{spdf:my doc#p=2}]` | `[@{spdf:my%20doc#p=2}]` | A space ends the key even inside braces. Percent-encode it, as the URI grammar asks anyway; the filter warns when it meets such text. | | `@spdf:…#f=xiv.` meaning folio `xiv.` | `[@{spdf:…#f=xiv.}]` or `f=xiv%2E` | Prose punctuation stuck to the end (`. , : ; ! ?`, closing quotes, an ellipsis) is taken as punctuation of your sentence; in the braced form every character counts. | **CommonMark and GFM.** `commonmark`, `commonmark_x` and `gfm` have no citation syntax in Pandoc 3.9: `[@spdf:…#p=2]` arrives as plain text. The filter finds those pieces and reads them again with Pandoc's own Markdown reader, so the same syntax works with `--from=commonmark_x`; the prefix and suffix of such a citation must be plain text. ## What the citation prints The filter resolves the anchor against the file and writes the locator into the suffix in a form citeproc recognises as a CSL locator (`, {p. [3]}`), so the **style** decides how to print it: Chicago author-date writes `[3]`, a style that shows labels writes `p. [3]`. The rules are those of the specification ([SPEC §18](../../spec/SPEC.md#citation)): | Anchor | Example | Chicago author-date | |---|---|---| | page, folio read | `#p=2` | `(Saorín Ferrer 2026, 1)` | | page, folio inferred | `#p=4` | `(Saorín Ferrer 2026, [3])` | | page without folio | `#p=1` | `(Saorín Ferrer 2026, n. pag.)` and a warning; `s. p.` in Spanish | | folio given directly | `#f=3` | `(Saorín Ferrer 2026, [3])`: checked against the units, bracketed if inferred | | range | `#p=3&pe=4`, `#f=4&fe=5` | `(Saorín Ferrer 2026, 2–[3])`, `(Saorín Ferrer 2026, 4–5)` | | range from an unnumbered page | `#p=14&pe=29` in the 1608 *Quixote* | `fol. Ir` and a warning: the unnumbered end is left out | | leaf (`foliation: leaf`), roman | `#p=29`, `#p=30`, `#p=29&pe=30` | `(Cervantes Saavedra [1605] 1608, fol. Ir)`, `fol. [Iv]`, `fols. Ir–[Iv]` | | time (ground elapsed time) | `#t=369959`, `#t=369966,369976` | `102:45:59`, `102:46:06-102:46:16` | | section | `#s=學而第一¶=1` | `(孔子, n.d., § 學而第一, para. 1)` | | section with a printed page | `#s=XXI¶=1&f=159` | `(Bécquer [1871] 1885, 159)` | | verse | `#v=2`, `#v=1-3` | `v. 2`, `vv. 1–3` | | slide | `#sl=2` | `slide 2` (`diap. 2` in Spanish) | | sheet | `#sh=Data&rows=4-9` | `Data, rows 4-9` | | canonical | `#ref=stephanus:514a`, `#ref=analects:1.2` | `514a`, `(孔子, n.d., 1.2)` | | whole document | no parameters | `(Saorín Ferrer 2026)` | When both `p` and `f` are given, `p` decides and a disagreeing `f` is reported. A folio printed on several pages (`f=1` in front matter and body) takes the first and warns; add `p=` to choose. A page without a printed folio is **never** cited by its position in the file: that number does not exist on paper. An end without a printed folio never takes part in a range ([SPEC §18.1](../../spec/SPEC.md#citation)): the folio of the other end is cited alone, and the page is unnumbered only when neither end has a folio. Verses, sections, paragraphs, slides, sheets and canonical references are looked up in the anchors of the units and of the fragments ([SPEC §5.4](../../spec/SPEC.md#anchor-uri)), so a reference kept on a fragment (`analects:1.2`, a line of a poem) is found, and one the file does not anchor is reported. Pandoc only reads a locator label in the terms of the locale citeproc is using, with no English fallback: a German document needs `S.` for a page, a Spanish one `f.` for a folio. The filter asks citeproc for those terms itself, once per run, so labels work in every CSL locale. With the test style [`test/styles/labels.csl`](test/styles/labels.csl), which prints the label citeproc recognised, a document with `lang: de-DE` gives `(Saorín Ferrer, 2026, [page] S. 1)` and `(Cervantes Saavedra, 1608, [folio] Fol. [Iv])`, and one with `lang: es-ES` gives `[page] p. 1` and `[folio] f. [Iv]`. Locators that CSL has no label for (an unnumbered page, a time, a slide, a sheet, a canonical reference, a section) are written after an empty locator, `{}, 1:09:20`, so that citeproc does not mistake `1:09:20` or `514a` for a page number. In Spanish: ```markdown --- lang: es --- La portada no lleva folio [@spdf:sha256-6abda0640aaed500ee9673212fab17b83ef4326f5ade42d6c6cabec3f8996df7#p=1]. Según @spdf:spdf-en-cinco-paginas#p=3, el formato guarda el folio impreso. ``` ```text La portada no lleva folio (Saorín Ferrer 2026, s. p.). Según Saorín Ferrer (2026, 2), el formato guarda el folio impreso. ``` ## Options Set them in the document's YAML metadata, in a defaults file, or with `-M`. | Field | Default | Meaning | |---|---|---| | `spdf-library` | `.` and the folder of the input file | A folder, or a list of folders and files. Folders are searched for `*.spdf` (not recursively). Relative paths are taken from the working directory, then from the input file's folder; `~/` is your home folder. | | `spdf-locale` | from `lang` | `en` or `es`: the language of the few words the filter writes itself (`n. pag.` / `s. p.`, `slide` / `diap.`, `para.` / `párr.`, `rows` / `filas`). Any `lang` other than Spanish gives English. | | `spdf-links` | `false` | `true` adds the canonical anchor URI of each resolved citation as a link in a footnote (inside a footnote, in parentheses after the citation). | | `spdf-sqlite3` | `sqlite3` | The `sqlite3` program to run. | With `spdf-links: true`: ```text A citation gets a note with its anchor (Saorín Ferrer 2026, [3])[1]. A range links to its canonical anchor (Saorín Ferrer 2026, 2–[3])[2]. Two citations get one note (Saorín Ferrer 2026, 1; 2026, 5)[3]. Inside a footnote the anchor goes in parentheses.[4] [1] spdf:sha256-50d942445564fe24effe743701f0c9a16fde414e6eb1f0ce2095b866d0555f4c#p=4&f=3 [2] spdf:sha256-50d942445564fe24effe743701f0c9a16fde414e6eb1f0ce2095b866d0555f4c#p=3&pe=4&f=2&fe=3 [3] spdf:sha256-50d942445564fe24effe743701f0c9a16fde414e6eb1f0ce2095b866d0555f4c#p=2&f=1; spdf:sha256-50d942445564fe24effe743701f0c9a16fde414e6eb1f0ce2095b866d0555f4c#p=6&f=5 [4] As shown in Saorín Ferrer (2026, 2) (spdf:sha256-50d942445564fe24effe743701f0c9a16fde414e6eb1f0ce2095b866d0555f4c#p=3&f=2). ``` The link is the canonical URI of what was found ([SPEC §5.2](../../spec/SPEC.md#anchor-uri)), whatever form you wrote: a citation by folio links to its physical page as well. ## References For each cited document the filter takes the CSL-JSON item stored in the file (without its `spdf` extension object), reads it with Pandoc's own CSL JSON reader, and adds it to the document's `references` metadata under the key `spdf-` followed by the first 12 hex digits of the document's `source_sha256` (`spdf-50d942445564`). The citation's key is rewritten to it, so citeproc, `link-citations`, `nocite` and every CSL style see an ordinary reference. Nothing is duplicated. If a reference with that key already exists, it is kept. If your `references` or `bibliography` files already hold the same work (same DOI, same ISBN, or same title, year and first author), the citation uses **your** key and your entry: ```text The English booklet is already in the references as saorin2026, so the citation uses that key (Saorín Ferrer 2026b, 1) and the bibliography lists it once, next to a citation written by hand (Saorín Ferrer 2026b, 4). The Spanish booklet is in the bibliography file as cinco (Saorín Ferrer 2026a, 2). ``` Without `--citeproc`, Pandoc's Markdown writer shows what the filter did, which is also a way to hand a resolved manuscript to someone without SPDF files (`-t markdown -s` also writes the references). For ```markdown A [see @spdf:spdf-in-five-pages#p=4, emphasis added]. B @spdf:spdf-in-five-pages#p=1 says. C [@spdf:spdf-in-five-pages#p=2] ``` `pandoc --lua-filter spdf.lua -t markdown` writes: ```markdown A [see @spdf-50d942445564, {p. \[3\]}, emphasis added]. B @spdf-50d942445564 [{}, n. pag.] says. C [@spdf-50d942445564, {p. 1}] ``` ## Warnings and unresolved citations An anchor that cannot be resolved is never dropped or guessed. The filter warns on stderr and leaves the citation visibly marked: its key becomes the anchor URI you wrote, which citeproc prints in bold with a question mark and reports again: ```text [WARNING] Scripting warning at spdf.lua line 46 column 1: spdf: [@spdf:spdf-in-five-pages#p=99]: page p=99 is not in spdf-in-five-pages.spdf; the citation is left unresolved [WARNING] Citeproc: citation spdf:spdf-in-five-pages#p=99 not found A page the file does not have (spdf:spdf-in-five-pages#p=99?). ``` This happens for an unknown document, a page, folio, time, verse, slide, sheet, section or canonical reference the file does not have, a malformed parameter (`p=0`), a truncated hash, or a time cited in a paged document. A citation next to it in the same bracket is still resolved. An unknown parameter name (`pg=9`) is reported and ignored, as the specification asks. Files that cannot be used are skipped with the reason: ```text spdf: skipping roto.spdf: it contains view x_rotura; SPDF files must not carry triggers, views or foreign virtual tables (E020) spdf: skipping E001-not-sqlite.spdf: it is not a SQLite database (E001) spdf: skipping E002-application-id.spdf: it is not an SPDF file (unknown application_id, E002) spdf: skipping E013-two-documents.spdf: its documents table does not hold exactly one row (E013) spdf: library path not found: no-such-folder ``` The warnings go through Pandoc's own log (Pandoc prefixes them with the line of the filter that emitted them; the message is what follows `spdf:`). So `--quiet` silences them, and `--fail-if-warnings` makes Pandoc exit with status 3: use it in a build that must not publish an unresolved citation. If `sqlite3` is missing and the document cites SPDF anchors, Pandoc stops with: ```text spdf.lua: the sqlite3 command-line tool was not found (looked for 'sqlite3'). Install it (macOS ships it; Debian/Ubuntu: apt install sqlite3; Fedora: dnf install sqlite) or point the metadata field spdf-sqlite3 at it. ``` A document without SPDF citations never calls `sqlite3`. ## Legacy SPDF 4.0 and 4.1 files Files written by Scholaris before SPDF 5.0 are SQLite databases wrapped in gzip, with Spanish table and column names. The filter recognises the gzip magic bytes, decompresses the file into a temporary folder (refusing more than 4 GiB), and reads it through the 5.0 view of [SPEC §20](../../spec/SPEC.md#legacy): `documentos`, `unidades`, `huella`, the anchor members (`fisica`, `impresa`, `origen: deducido`…) and the `MetadatosDocumento` object, mapped to CSL (title and subtitle, authors, editors, dates, publisher, place…). ```text Garcilaso, a folio inferred from its neighbours (Garcilaso de la Vega [1543] 1919, [7]), a folio read on the page (Garcilaso de la Vega [1543] 1919, 159), by document id and folio (Garcilaso de la Vega [1543] 1919, 159), a range (Garcilaso de la Vega [1543] 1919, [7]–159), and a page the file does not have (spdf:garcilaso#p=10?). Kennedy, by paragraph (Kennedy 1962, para. 15) (Kennedy 1962, para. 16), and a paragraph the excerpt does not have (spdf:kennedy-rice#para=40?). Apollo 11, a recording in ground elapsed time: a moment (National Aeronautics and Space Administration 1969, 102:46:16), a span (National Aeronautics and Space Administration 1969, 102:46:18-102:46:23), and a time outside the excerpt (spdf:apolo11-tierra#t=99?). ``` ## Security SPDF files come from strangers, so the filter follows the safe opening of [SPEC §2.4](../../spec/SPEC.md#container) as far as the `sqlite3` tool allows. Each query runs in a separate `sqlite3` process opened with `-readonly`, `-safe` (3.37 or later: no `load_extension`, no file or shell commands), `-batch`, `-bail` and `-init /dev/null` (your `~/.sqliterc` is not read), with `.dbconfig defensive on`, `PRAGMA trusted_schema = OFF`, `query_only = 1`, `mmap_size = 0`, `cell_size_check = ON` and a 512 MiB limit on any value. A file with a trigger, a view or a virtual table other than the FTS5 index is refused before any of its tables is read (legacy files may keep their three FTS triggers, which cannot fire on a read-only connection). The SQL is fixed text in the filter; nothing from a file or from your document is ever put into a query, executed, or fetched from the network. ## Tests ```sh make test # or: test/run.sh [case ...] ``` The runner needs `pandoc`, `sqlite3` and `gzip`. Each case in `test/cases/` is run with `--lua-filter spdf.lua --citeproc -t plain` and compared with `test/expected/` three ways: the text, the warnings, and the citations and references the filter produced (`-t native` through [`test/inspect.lua`](test/inspect.lua)). The cases cover printed folios from `p=`, checks of `f=`, inferred folios, the cover without folio, ranges, leaves, times, sections, verses, slides, sheets, canonical references, unknown documents, pages and parameters, prefixes and suffixes, several citations in one bracket, in-text citations and their punctuation, the braced form, CommonMark input, a Spanish document, locator labels in German and Spanish, merging with existing references and bibliography files, anchor links, invalid files (`roto.spdf` and the `conformance/invalid` files: none may crash the filter), legacy gzip files, and a missing `sqlite3`. They read the shared fixtures in [`integrations/fixtures/`](../fixtures/) and some files of the conformance suite (`conformance/files`, `conformance/legacy`, `conformance/invalid`), and build `test/build/mixed.spdf` from [`test/fixtures/mixed.sql`](test/fixtures/mixed.sql) for the anchor types those lack. Those files are rebuilt by other parts of the repository, so the cases name them by placeholders (`HASH_EN`, `HASH_QUIJOTE`, `HASH_EN_12` for the citekey…) that the runner fills with each file's current `source_sha256` and puts back in the outputs; the table is `HASHED` in `test/run.sh`. `test/run.sh --update` rewrites the expected outputs; read the diff before keeping it. ## Limitations - The filter's own words exist in English and Spanish only. Locator labels follow any CSL locale, but the probe uses the locale's terms: a style whose own `` section redefines the `page`, `folio`, `column` or `verse` terms may not see the label. - Times have no CSL label (Pandoc's locator parser does not know CSL 1.0.2 `timestamp`), so they are plain text in the suffix, `m:ss` below one hour and `h:mm:ss` above, as in SPEC §18; a `t=a,b` pair prints as a range. - `char=` and `xywh=` are kept in links but not printed: a citation locates a unit. - Citations in metadata fields (title, abstract) are not resolved. - Library folders are not searched recursively, `.spdfl.json` collection manifests are not read, and remote files are never fetched. - On Windows, `sqlite3.exe` must be on `PATH`; without a POSIX `sh`, legacy gzip files are decompressed in memory and without the 4 GiB cap. - Opening costs two short `sqlite3` runs per library file and one more per cited document; with a library of thousands of files, list the ones you cite. ## License MIT OR Apache-2.0, like the rest of the code in this repository. The test fixtures in `test/fixtures/` and the test style in `test/styles/` are dedicated to the public domain (CC0 1.0).