Manages PDFs, images, and text files in a local SQLite catalog with stable node IDs, SHA-256 content addressing, ranked full-text search, tagging, and incremental backup with verification.
A self-hosted internet archiving tool that takes URLs and saves them in multiple formats – HTML, PDF, screenshot, WARC, media files – for long-term preservation.
Discovers citations across CrossRef, OpenCitations, DataCite, and OpenAlex; syncs with Zotero; acquires PDFs with git-annex provenance tracking; and stores everything in a DataLad dataset using a LinkML schema aligned with CiTO and FaBiO ontologies.