Saltar al contenido
SPDF 5.0

SPDF en R

Cómo instalar y usar la implementación de SPDF en R (spdf): abrir, validar, buscar y citar. Sobre RSQLite; los fragmentos como data frames.

Revisado Markdown

  • Paquete: spdf
  • Instalar: remotes::install_github("joseluissaorin/spdf", subdir = "r")
  • Nivel: segundo
  • CI: CI en marcha
  • Carpeta: r/

El README de la biblioteca está en inglés.

Read, validate, search, cite and write SPDF files (Semantic Processed Document Format) from R. A SPDF file holds a document that has been read once and can be cited forever: every passage carries its exact anchor (printed page, folio, second of a recording, slide, verse), so a citation can only print what the source says.

The package is a native implementation of SPDF 5.0 on RSQLite. It also reads the legacy 4.0/4.1 files produced by Scholaris (gzip-wrapped, Spanish schema) through the 5.0 view. Tables come back as tibbles.

Install

# from CRAN, once published
install.packages("spdf")
# from the repository
remotes::install_github("joseluissaorin/spdf", subdir = "r")

Read, search, cite

library(spdf)
doc <- spdf_open(system.file("extdata", "quijote.spdf", package = "spdf"))

spdf_info(doc)                     # title, authors, year, version, counts
spdf_units(doc)                    # one row per citable unit (page, folio, time span...)
fr <- spdf_fragments(doc)          # searchable passages, anchors as list-columns

hits <- spdf_search(doc, "hermoso")    # finds the long-s "hermoso" through the modern layer
hits$anchor_uri                    # spdf:sha256-27ea...#p=13&char=10,194
spdf_cite(spdf_metadata(doc), fr$anchor[[5]], fr$anchor_end[[5]], locale = "es")
#> "(Cervantes Saavedra, 1608, fols. Ir-[Iv])"
spdf_cite_passage(doc, "q5", "rozin, como tomaua la podadera.")$text
#> "(Cervantes Saavedra, 1608, fol. [Iv])"   the page the quotation is on

spdf_locate(doc, hits$anchor_uri[1])   # list(document, units, fragments, char, xywh)
cat(spdf_bibtex(doc))              # @book{cervantessaavedra1608, ... (also spdf_csl())
spdf_close(doc)

Vector and hybrid search take a query vector computed with the same model as the space: spdf_search_vector(doc, v, "embeddinggemma-2@768"), spdf_search_hybrid(doc, "ciego jarro", v, "embeddinggemma-2@768").

ALTO, TEI and IIIF

writeLines(spdf_alto(doc), "quijote.alto.xml")    # ALTO 4, one Page per page unit
writeLines(spdf_tei(doc), "quijote.tei.xml")      # TEI P5: pb, p, lg/l, u, note
writeLines(spdf_iiif_json(doc, "https://example.org/iiif/quijote"), "manifest.json")

Corpora

files <- list.files("corpus", pattern = "\\.spdf$", full.names = TRUE)
spdf_corpus(files)                          # one row per document
spdf_corpus_search(files, "\"molinos de viento\"")   # one row per passage, with citation
spdf_count_terms(files, c("honra", "fortuna"))       # fragments and occurrences per work

The vignette vignette("corpus", package = "spdf") (in Spanish) walks through a digital-humanities workflow: searching a corpus and counting occurrences by work and year, with every number traceable to its page.

Validate and write

spdf_validate("file.spdf")         # list(valid, version, profile, errors, warnings)
spdf_write("out.spdf", document = ..., units = ..., fragments = ...)

Security

Files are untrusted input: spdf_open() connects read-only with query_only and trusted_schema=OFF, never loads extensions, refuses triggers, views and foreign virtual tables (except the three FTS triggers of legacy files), bounds blob sizes and gzip inflation, and copies WAL-mode files before opening them.

Conformance

spdf_conformance("path/to/spdf/conformance") runs the shared suite of the specification; Rscript inst/scripts/conformance.R ../conformance prints the JSON report. CI publishes it as the conformance-r artifact. All kinds are claimed, export_structure included.

License

MIT OR Apache-2.0, at your option. The SPDF specification is CC BY 4.0. The sample files in inst/extdata are short excerpts of public-domain works.