SPDF en R
Cómo instalar y usar la implementación de SPDF en R (spdf): abrir, validar, buscar y citar. Sobre RSQLite; los fragmentos como data frames.
Revisado Markdown
- Paquete:
spdf - Instalar:
remotes::install_github("joseluissaorin/spdf", subdir = "r") - Nivel: segundo
- CI:
- Carpeta:
r/
El README de la biblioteca está en inglés.
Read, validate, search, cite and write SPDF files (Semantic Processed Document Format) from R. A SPDF file holds a document that has been read once and can be cited forever: every passage carries its exact anchor (printed page, folio, second of a recording, slide, verse), so a citation can only print what the source says.
The package is a native implementation of SPDF 5.0 on RSQLite. It also reads the legacy 4.0/4.1 files produced by Scholaris (gzip-wrapped, Spanish schema) through the 5.0 view. Tables come back as tibbles.
Install
# from CRAN, once published
install.packages("spdf")
# from the repository
remotes::install_github("joseluissaorin/spdf", subdir = "r")Read, search, cite
library(spdf)
doc <- spdf_open(system.file("extdata", "quijote.spdf", package = "spdf"))
spdf_info(doc) # title, authors, year, version, counts
spdf_units(doc) # one row per citable unit (page, folio, time span...)
fr <- spdf_fragments(doc) # searchable passages, anchors as list-columns
hits <- spdf_search(doc, "hermoso") # finds the long-s "hermoso" through the modern layer
hits$anchor_uri # spdf:sha256-27ea...#p=13&char=10,194
spdf_cite(spdf_metadata(doc), fr$anchor[[5]], fr$anchor_end[[5]], locale = "es")
#> "(Cervantes Saavedra, 1608, fols. Ir-[Iv])"
spdf_cite_passage(doc, "q5", "rozin, como tomaua la podadera.")$text
#> "(Cervantes Saavedra, 1608, fol. [Iv])" the page the quotation is on
spdf_locate(doc, hits$anchor_uri[1]) # list(document, units, fragments, char, xywh)
cat(spdf_bibtex(doc)) # @book{cervantessaavedra1608, ... (also spdf_csl())
spdf_close(doc)Vector and hybrid search take a query vector computed with the same model as the space:
spdf_search_vector(doc, v, "embeddinggemma-2@768"),
spdf_search_hybrid(doc, "ciego jarro", v, "embeddinggemma-2@768").
ALTO, TEI and IIIF
writeLines(spdf_alto(doc), "quijote.alto.xml") # ALTO 4, one Page per page unit
writeLines(spdf_tei(doc), "quijote.tei.xml") # TEI P5: pb, p, lg/l, u, note
writeLines(spdf_iiif_json(doc, "https://example.org/iiif/quijote"), "manifest.json")Corpora
files <- list.files("corpus", pattern = "\\.spdf$", full.names = TRUE)
spdf_corpus(files) # one row per document
spdf_corpus_search(files, "\"molinos de viento\"") # one row per passage, with citation
spdf_count_terms(files, c("honra", "fortuna")) # fragments and occurrences per workThe vignette vignette("corpus", package = "spdf") (in Spanish) walks through a
digital-humanities workflow: searching a corpus and counting occurrences by work and
year, with every number traceable to its page.
Validate and write
spdf_validate("file.spdf") # list(valid, version, profile, errors, warnings)
spdf_write("out.spdf", document = ..., units = ..., fragments = ...)Security
Files are untrusted input: spdf_open() connects read-only with query_only and
trusted_schema=OFF, never loads extensions, refuses triggers, views and foreign
virtual tables (except the three FTS triggers of legacy files), bounds blob sizes and
gzip inflation, and copies WAL-mode files before opening them.
Conformance
spdf_conformance("path/to/spdf/conformance") runs the shared suite of the
specification; Rscript inst/scripts/conformance.R ../conformance prints the JSON
report. CI publishes it as the conformance-r artifact. All kinds are claimed,
export_structure included.
License
MIT OR Apache-2.0, at your option. The SPDF specification is CC BY 4.0. The sample
files in inst/extdata are short excerpts of public-domain works.