Imported from Blackcat-Informatics/purrdf (
AGENTS.md). Install upstream withnpx skills add Blackcat-Informatics/purrdf. Copyright stays with the author.
AI Developer Agent Guide (AGENTS.md)
Welcome, AI Agent! This file is your behavioral contract and instruction manual for contributing to the PurRDF repository.
Deficiency emergency ledger (non-negotiable)
.deficiencies is the log of last resort for critically undone work. Every entry
below its marker is 100% unauthorized, is 100% a bug, and means its
originating issue or pull request failed. It is used for work misrepresented to
pass PR gates, work misrepresented by an agent, or a discovery that an agent was
fundamentally defective. An entry is literally a cry for help from a failing
agent; it is never an accepted risk, authorized descope, backlog, or success with
caveats.
The only normal contents are the tracked notice and marker, with no entries below them. An entry blocks completion, PR creation, and merge of the work that produced it. Immediately verify the defect against current code and give it a durable, visible remediation owner before removing the emergency entry. Removing an entry does not resolve the bug or retroactively make the failed work successful. Never add ledger text to make incomplete work appear complete.
1. What this repository is
PurRDF is the RDF 1.2 toolkit that several downstream projects (notably
gmeow-ontology) use as
their data-carrier backbone. It must stay fast, deterministic, and boring:
one engine, one behavior, carried verbatim into Rust, Python, WebAssembly, and C.
Crate map (all under crates/, published names in Cargo.toml):
| Crate | Role |
|---|---|
purrdf |
Umbrella facade (RDF surface at root; slice/shapes as modules) |
purrdf-rdf (crates/rdf) |
Native text/XML/JSON-LD codecs, GTS adapters, describe, canonicalization |
purrdf-core (crates/rdf-core) |
Interned IR kernel, diagnostics, store traits, provenance, RDFC-1.0 |
purrdf-columnar (crates/columnar) |
Bidirectional five-table Parquet codec for RDF 1.2 + blobs |
purrdf-gts (crates/gts) |
GTS container engine (CBOR log, BLAKE3, COSE) |
purrdf-sparql-{algebra,eval,results} |
SPARQL 1.1/1.2 parser, evaluator, results |
purrdf-shapes (crates/shapes) |
SHACL validation (full Core + SHACL-SPARQL + SHACL-AF) and SHACL Rules (sh:rule inference) |
purrdf-shex (crates/shex) |
ShEx 2.1 schemas + validation |
purrdf-slice (crates/slice) |
Slice catalog, artifacts, ownership analysis |
purrdf-datalog (crates/datalog) |
Deterministic semi-naive Datalog substrate beneath every rule-driven engine |
purrdf-entail (crates/entail) |
Entailment regimes: RDF/RDFS/OWL 2 RL/D materialization, OWL-Direct, RIF |
purrdf-geo (crates/geo) |
GeoSPARQL 1.1: exact float-free WKT/GeoJSON geometry and the geof: family over both extension seams |
purrdf-text (crates/text) |
Deterministic full-text search over literals: exact fixed-point BM25, ranked rows through the property-function seam |
purrdf-retrieval (crates/retrieval) |
Composition layer over the ranked producers: plan → compile → execute → fuse, with a canonical BLAKE3 plan identity and an exact, content-addressed fusion law; producers, strata and weights are caller-supplied |
purrdf-validate (crates/validate) |
Shared string boundary every language binding routes through |
purrdf-json (crates/json) |
Ordered JSON byte-cover codec with queryable occurrences, strict reconstruction and caller-selected profile; sole runtime dependency is purrdf-core |
purrdf-markdown (crates/markdown) |
Structural Markdown-to-RDF 1.2 slicer under a shipped specification: a typed stand-off model over verbatim byte spans, projected to claims; sole runtime dependency is purrdf-core |
purrdf-iri, purrdf-xsd, purrdf-events |
Zero-dependency foundations |
purrdf-cdt (crates/cdt) |
SPARQL composite datatypes (SEP-0009 cdt:List/cdt:Map): closed leaf over purrdf-iri + purrdf-xsd only |
purrdf-wasm, purrdf-capi, bindings/python |
WASM, C-ABI, and PyO3 bindings |
purrdf-cli (crates/cli) |
The purrdf command-line surface (publish = false) |
purrdf-envelope-probe (crates/envelope-probe) |
The micro-hardware envelope capture tool (publish = false) |
purrdf-alloc-probe (crates/alloc-probe) |
The shared counting allocator + per-thread/whole-process measurement windows every allocation test and bench measures with (publish = false, [dev-dependencies] only, path-only with no version) |
purrdf-bench (crates/bench) |
Benchmark tooling: the scale-corpus generator (publish = false) |
2. Hard constraints (violating these fails CI or review)
- NO semantic Cargo features, ever. The sole exception is the empty,
non-semantic
purrdf-capi:capi = []marker thatcargo-crequires. It gates no code and must never appear incfg(feature = ...); CI checks both facts withscripts/check-no-features.py. PurRDF is a carrier; optionality changes semantics per consumer, which is forbidden. Do not add any other feature, optional dependency, or feature-gated behavior. - Kernel ring-fence.
purrdf-coremust never depend on oxigraph or PyO3.purrdf-iri,purrdf-xsd, andpurrdf-eventsmust keep zero runtime dependencies. - Terminal ring-fence: a scanner's character classes are exact, in both
directions. They decide token boundaries, not merely membership, so
substituting a Unicode property for a production's enumerated set does not
just widen the accepted language — it silently re-tokenizes documents both
the liberal and the conforming parser accept.
?s<NBSP>?plexed as one variable, turning a join into a cross product with exit zero and no diagnostic. Every W3C terminal is spelled once, inpurrdf_iri::terminals, with its production cited and its ranges asserted at compile time; scanners call it rather than retyping a table.scripts/check-terminal-predicates.py(inmake check, ormake terminal-hygiene) refuses a Unicode-property test inside a file that holds a character cursor, and names inSCANNERSthe files whose cursor lives in another module. This is not "Unicode properties are bad": a production that names one must be implemented with it, and the gate'sALLOWLISTis the reasoned ledger for exactly those cases. Judge the clause, not the specification — one spec answers this differently in different places. CommonMark defines a "Unicode whitespace character" and uses it for §6.2 emphasis flanking, while its blank line (§2.1), ATX heading, thematic break and GFM table cell all name space-or-tab; citing "CommonMark" alone settles nothing, and doing so once put a false exemption into this file. - Everything is wasm-able. Every release crate (all 25 publishable crates,
purrdf-wasmincluded) must build forwasm32-unknown-unknown— CI hard-fails otherwise (make wasmlocally). Never add a dependency that drags in threads, the filesystem, C toolchains, or wall-clock/RNG syscalls on the wasm path; crypto stays pure-Rust for exactly this reason. - Byte determinism. Serializers and the GTS writer are byte-deterministic.
If your change alters emitted bytes, you must update the affected goldens and
say why in the PR. Never introduce iteration-order, time, or RNG dependence
into output paths (hashers are fixed-key
ahashfor this reason). - Conformance corpora are the contract: W3C SPARQL 1.1
(
crates/sparql-conformance), the W3C SHACL suite (vectors/shacl/), the shexTest v2.1.0 suite (vectors/shexTest/), the first-party SHACL corpus (crates/shapes/corpus/), RDFC-1.0 fixtures (crates/rdf/tests/fixtures/rdfc/), and the frozen GTS vectors invectors/(shared byte-exact with the other GTS engines — never regenerate or "fix" them here; the GTS wire format is governed ingmeow-gts). Harnesses assert exact counts and enforce XPASS discipline on their xfail ledgers — seedocs/CONFORMANCE.mdfor the scoreboard. - PurRDF is NOT an ontology. Structural Markdown and ordered JSON codecs
offer explicitly named standard profiles and vocabularies under their shipped
specifications; callers must select them deliberately or supply a vocabulary.
There is no implicit namespace fallback. Every other
vocabulary the library reads or writes (slice manifests, statement-metadata
downcast, box roles, language retagging, SPARQL extension-function
namespaces, standpoint predicates, json_schema namespaces) is
caller-supplied configuration with no fabricated default: a feature
exercised without its vocabulary hard-errors or stays inactive. Never
hardcode a
blackcatinformatics.canamespace in library code (the GMEOW ontology is a consumer; the dependency arrow never points from purrdf to it). Test fixtures useexample.org. - Generated artifacts under
generated/are projections — never hand-edit; regenerate viamake metadata(scripts/check-generated.shgates drift). - Dependency versions live in one place:
[workspace.dependencies]in the rootCargo.toml. Member crates usedep.workspace = true. Do not pin a version inside a member manifest. - Lints are workspace-inherited (
[workspace.lints], clippy pedantic + nursery).cargo clippy --workspace --all-targetsmust be warning-free. Prefer fixing code over#[allow]; a genuinely-right allow must be tightly scoped and carry a reason comment. - SPDX headers on every source file:
MIT OR Apache-2.0 OR MulanPSL-2.0(docs may beCC-BY-4.0).
3. Commands
make check # the full local gate: fmt, clippy, build, tests, hygiene
make test # cargo test --workspace
make metadata # regenerate + verify generated artifacts
make bench # criterion benchmarks (report-only; not a gate)
make scale-corpus # generate the deterministic scale corpus (streams; stores nothing by default)
make lubm # the LUBM comparison workload, per entailment regime (report-only; network + JRE)
make watdiv # the WatDiv comparison workload over a frozen dataset (report-only; network)
make build-profile-hygiene # prove the gate is compiled the way it claims
scale-corpus, lubm and watdiv are the three comparison lanes. None is a
gate and none runs in make check: lubm needs a JRE and fetches a
GPL-2.0-or-later generator, watdiv fetches a 58 MB frozen dataset that expands
past a gigabyte, and neither vendors a byte. They share one implementation of
the laws that make their numbers evidence — scripts/lane-common.sh — so a
repair to one is a repair to all three. docs/BENCHMARKS.md owns the
parameters, the knobs and the comparison rules.
Toolchain: rust-toolchain.toml names a floating nightly for development and
for every CI gate. That is an analysis decision, not a licence: nightly clippy
and rustdoc carry lints stable lacks, and its default borrow checker is the
stronger one, so a finding is a real finding rather than a channel artifact.
Floating is the point — a dated channel freezes that surface as of one day, and
every check sharpened afterwards stops being a finding and becomes invisible debt
while the gates still report green. Byte-determinism is no argument for freezing:
it is a property of the code — sorted, deduplicated, explicitly ordered output,
identities content-addressed over this workspace's own declared law — and the
goldens and vectors prove it on whatever compiler runs them. A golden that moved
under a compiler bump would be a serializer defect to fix, not a reason to stop
bumping.
The source stays nightly-free. There are zero #![feature(...)] attributes
in crates/ and bindings/, and adding one is forbidden. What consumers need
is rust-version in Cargo.toml (the MSRV, currently 1.96) — a lower floor on
the stable channel — enforced by the dedicated msrv CI job. Never "align" the
MSRV to the dev pin; they answer different questions.
Two traps, both load-bearing:
dtolnay/rust-toolchainselects withrustup default, which ranks belowrust-toolchain.toml. A workflow that installs one toolchain while the repo pins another does not fail — it silently runs the pin. OnlyRUSTUP_TOOLCHAINoutranks the file, which is why themsrvjob and the release lanes set it explicitly, and why themsrvjob also assertsrustc --versionreally is 1.96.x.scripts/check-toolchain-pin.py(inmake checkand CI) fails on any workflow whose install step disagrees with the pin without that explicit escape, and on a floating channel.
Release lanes (release-cargo, release-npm, release-pypi) build on
stable on purpose: nightly's sharper lints buy nothing for an artifact a
consumer installs, and shipping one from an unreleased compiler is risk without
upside.
4. Performance discipline
This library is a hot-path backbone. The IR is an immutable, value-interned
dataset (TermId = niche-optimized NonZeroU32, string arena, store-once
interner). When touching parse/serialize/eval paths:
- Measure first — layout and algorithm choices are justified by the
criterion benches (
crates/rdf-core/benches/ir_layout.rset al.), not by assertion. Add or extend a bench when you claim a win. - Avoid per-token/per-term
Stringallocation; move values out of buffers instead of cloning; pre-size collections in parse loops. - Hot maps use fixed-key
ahash(seecrates/rdf-core/src/ir/builder.rsfor the canonical store-once interner pattern) — never default SipHash in a hot path, and never a randomly-seeded hasher in an output path.
The gate is compute, so the gate is compiled like it. make check runs the
whole test surface and make conformance runs every W3C suite through it, so
both are bounded by the codegen under them: [profile.dev] builds everything —
our crates, dependencies, and (named separately, because they do not inherit the
base profile) build scripts and proc-macros — at opt-level 3, with
debug-assertions and overflow-checks ON. Those two are orthogonal to
opt-level and are not negotiable: first-party code carries
#[cfg(debug_assertions)] bodies that vanish silently with the flag, and
overflow checks are what stop an arithmetic bug in a byte-deterministic codec
becoming a wrong-but-green run. Optimizing the gate does not weaken it: the same
assertions run over the same corpora, on better codegen.
This is not self-enforcing, and it did fail once: four per-crate opt-level = 2
tables made the manifest read as tuned while twenty-one members compiled at
opt-level 0, because package."*" matches dependencies only. So the property is
asserted against the graph Cargo actually resolves, not against the manifest:
make build-profile-hygiene (in make check) reads the --unit-graph of
cargo test and cargo build and checks every unit's effective profile.
Reading effective values is deliberate — it catches a [profile.*] table in
$CARGO_HOME/config.toml or anywhere on the walk up from the workspace (your
home directory is on that walk), a CARGO_PROFILE_* variable, or a --config
override, none of which the manifest can see. Do not add a [profile.test]
block (test inherits dev), and do not set lto or codegen-units = 1 there:
both serialize codegen and inflate link memory, which is what a
constantly-rebuilt, cold-in-CI gate wants least.
5. Brand & naming
The project name is written PurRDF in prose (never PurrDF/PURRDF); all
package/crate/binary identifiers are lowercase purrdf. See
docs/BRAND.md. Logo/social assets follow the shared
black-cat family system — #cat-head-core is shared verbatim; only the
#service-triple group is purrdf-specific.
6. Releases
Tag-driven trusted publishing: rust-v* → crates.io (25 crates, ordered),
py-v* → PyPI (purrdf). See docs/RELEASE.md. Version
is single-sourced in [workspace.package]. Seven members never reach
crates.io: purrdf-capi, purrdf-sparql-conformance, purrdf-cli,
purrdf-envelope-probe, purrdf-bench, purrdf-alloc-probe, and
purrdf-python (PyPI via maturin instead). purrdf-alloc-probe is a
dev-dependency of published crates, so its root [workspace.dependencies] entry
is path-only with no version — cargo then strips it from the packaged
manifest, which is the only way cargo publish's dev-dependency-resolving
verification step can succeed.
7. Provenance
This repo was extracted from gmeow-ontology and gmeow-gts — see
PROVENANCE.md for source commits. PurRDF's replacement
surfaces are being completed here, but the downstream gmeow-ontology cutover
is not yet complete. Its legacy models are migration evidence only and should
be deleted as each PurRDF replacement is integrated.