You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Symmetry tests never run either language's schema validator over the other's documents #155
Joint issue — the Python half is VH-Lab/DID-python#28. The gap is symmetric and the fix has to land on both sides to be meaningful.
The problem
The cross-language symmetry suite verifies that each language can read and round-trip the other's database. It never verifies that either language considers the other's documents schema-valid.
Validation lives in database.add_docs in both languages ('Validate', true here, validate=True in Python — the default on both). Each language calls add_docs only on documents it created itself, during makeArtifacts. The readArtifacts step on both sides opens the other language's SQLite file, re-summarizes it, and compares summaries — it never calls add_docs, so validate_docs / validate_doc_vs_schema never runs.
db=SQLiteDB(db_path)
live_summary=database_summary(db)
report=compare_database_summary(saved_summary, live_summary)
...
doc=db.get_docs(expected_id, OnMissing="ignore") # get_docs does not validate
So a document that one language writes can be invalid under the other's schema validator while the suite stays green.
What the suite does and does not catch
Covered by compareDatabaseSummary: branch names, branch hierarchy, per-branch document count, document IDs, class_name, the .value field of demoA/demoB/demoC, and depends_on name/value pairs.
Not covered — anything validate_doc_vs_schema would check:
field types and ranges outside those three .value fields
required (mustbenotempty) depends_on entries being present and non-empty
inherited superclass property lists being present and well-formed
document_class.validation pointing at a resolvable schema
files / file_info conformance via checkfiles
Two concrete instances
1. Python document IDs were UUID4 (now fixed, but the suite never noticed).base.schema.json types base.id as did_uid — 16 hex digits, an underscore, 16 more, exactly what did.ido.unique_id produces and did.ido.isvalid accepts. Python's IDO.unique_id() returned a UUID4 instead, so every Python-generated document was schema-invalid and this repo's add_docs would have rejected all of them. The symmetry suite was green throughout, because MATLAB only ever read them, never validated them. Fixed in VH-Lab/DID-python#27; the detection gap that let it persist is not.
2. base.datestamp format — believed still live. Python writes it with a space separator:
Python: '2026-08-28 11:14:42.595475' str() of a naive datetime
MATLAB: '2026-08-28T10:11:16.934Z' char(datetime(...,'TimeZone','UTCLeapSeconds'))
base.schema.json types datestamp as timestamp, and did.database.validate_field_type_and_value checks it with:
value = regexprep(value,'Z$','');
java.time.LocalDateTime.parse(java.lang.String(value));
LocalDateTime.parse uses ISO_LOCAL_DATE_TIME, which requires the T separator. If that reading is right, this validator would raise DID:Database:ValidationFieldTimeStamp on every Python-produced document — and the symmetry suite would still be green. Python's own validator accepts both forms (it has strptime fallbacks), so the Python side is quiet too.
This one is reasoning from source, not an observed failure. It was written without a MATLAB runtime, so it needs confirming here before anyone acts on it — the quickest check is to run validate_doc_vs_schema over a document from a Python-produced artifact and see whether base.datestamp trips. Note it is pre-existing and not introduced by the recent Python work.
Proposed fix
Add a validation pass to readArtifacts on both sides, so each language runs its own validator over documents the other language wrote:
MATLAB — in tests_symmetry/+did/+symmetry/+readArtifacts/+database/buildDatabase.m, after opening the database, read the documents and call db.validate_docs(...), failing the test on a validation error.
That turns the suite into a real cross-language contract check rather than a round-trip check, and would have caught instance 1 the day it was introduced.
Per AGENTS.md, anything implemented here needs a MATLAB environment to run and validate.
The problem
The cross-language symmetry suite verifies that each language can read and round-trip the other's database. It never verifies that either language considers the other's documents schema-valid.
Validation lives in
database.add_docsin both languages ('Validate', truehere,validate=Truein Python — the default on both). Each language callsadd_docsonly on documents it created itself, duringmakeArtifacts. ThereadArtifactsstep on both sides opens the other language's SQLite file, re-summarizes it, and compares summaries — it never callsadd_docs, sovalidate_docs/validate_doc_vs_schemanever runs.MATLAB (
tests_symmetry/+did/+symmetry/+readArtifacts/+database/buildDatabase.m):Python (
tests/symmetry/read_artifacts/database/test_build_database.py):So a document that one language writes can be invalid under the other's schema validator while the suite stays green.
What the suite does and does not catch
Covered by
compareDatabaseSummary: branch names, branch hierarchy, per-branch document count, document IDs,class_name, the.valuefield ofdemoA/demoB/demoC, anddepends_onname/value pairs.Not covered — anything
validate_doc_vs_schemawould check:.valuefieldsdid_uid/timestamp/char-length /matrix-shape conformancemustbenotempty)depends_onentries being present and non-emptydocument_class.validationpointing at a resolvable schemafiles/file_infoconformance viacheckfilesTwo concrete instances
1. Python document IDs were UUID4 (now fixed, but the suite never noticed).
base.schema.jsontypesbase.idasdid_uid— 16 hex digits, an underscore, 16 more, exactly whatdid.ido.unique_idproduces anddid.ido.isvalidaccepts. Python'sIDO.unique_id()returned a UUID4 instead, so every Python-generated document was schema-invalid and this repo'sadd_docswould have rejected all of them. The symmetry suite was green throughout, because MATLAB only ever read them, never validated them. Fixed in VH-Lab/DID-python#27; the detection gap that let it persist is not.2.
base.datestampformat — believed still live. Python writes it with a space separator:base.schema.jsontypesdatestampastimestamp, anddid.database.validate_field_type_and_valuechecks it with:LocalDateTime.parseusesISO_LOCAL_DATE_TIME, which requires theTseparator. If that reading is right, this validator would raiseDID:Database:ValidationFieldTimeStampon every Python-produced document — and the symmetry suite would still be green. Python's own validator accepts both forms (it hasstrptimefallbacks), so the Python side is quiet too.This one is reasoning from source, not an observed failure. It was written without a MATLAB runtime, so it needs confirming here before anyone acts on it — the quickest check is to run
validate_doc_vs_schemaover a document from a Python-produced artifact and see whetherbase.datestamptrips. Note it is pre-existing and not introduced by the recent Python work.Proposed fix
Add a validation pass to
readArtifactson both sides, so each language runs its own validator over documents the other language wrote:tests_symmetry/+did/+symmetry/+readArtifacts/+database/buildDatabase.m, after opening the database, read the documents and calldb.validate_docs(...), failing the test on a validation error.test_build_database.py, callingdb.validate_docs(docs)(public since Port document-vs-schema validation from MATLAB to Python DID-python#27).That turns the suite into a real cross-language contract check rather than a round-trip check, and would have caught instance 1 the day it was introduced.
Per
AGENTS.md, anything implemented here needs a MATLAB environment to run and validate.