Source data.
Explainable mappings.
Certified metrics.
A governed workflow for supported synthetic SaaS and SAP-style schemas—carrying evidence, review decisions and reconciliation through to publication.
Historical limitation: the original unfamiliar-schema probe produced zero proposals and matched 0/10 labeled targets. It was never approved or published. Inspect the original result. A later frozen benchmark found no mapping-accuracy improvement and exposed a safety failure; subsequent repairs are post-benchmark remediation, not held-out validation.
Unit mismatches and ambiguous revenue meanings become review items with traceable evidence.
Nine generated monthly metrics are executed and checked against independent reference calculations.
The agent role cannot approve. Publication requires current certification; demo reviewers are automated identities.
A workflow you can inspect
Database state survives process restarts. Content-bound approvals expire when inputs change. Publication recovers from failures between filesystem rename and database commit.
Explore architecture and design decisions →Actual output. Replayable evidence.
This is recorded stdout from the automated synthetic demo. Reviewer decisions are scripted. It is not a live model session or a recording of Alex approving data. No backend or model API runs on this page.
Loading the captured demo…Read the full accessible transcript
A narrated tour and real model sessions
Historical portfolio acceptance: recorded by Alex. The current v0.2.0-rc.3 qualification record documents the later automated integration scope. Read the final acceptance record.
A 3:51 evidence tour uses a clearly labeled synthetic voice and recorded fixture results. Six separate Claude Code sessions preserve actual tool calls and model responses. Only one session invoked a Skill body; no general improvement is claimed.
Follow a result back to its evidence
Extending the source boundary
Typed CSV extracts can be imported as registered read-only snapshots. A bounded public-data exercise imported and profiled 10,000 historical retail rows, passing six aggregate controls while retaining missing IDs, cancellations and negative quantities. No mapping, financial certification or customer benefit is claimed for that exercise.
The first separately authored frozen benchmark produced 31 correct mappings from 33 proposals (31 of 32 targets), unchanged from baseline, and failed safety because a competing monetary interpretation bypassed review. Post-benchmark remediation passes safety at 32/33 correct proposals and 32/32 positive targets covered. That historical report required review for 7/33 proposals; the current v1 regression requires 17/33 after the later uncertainty repair, with the same accuracy, passing safety and eight unresolved fields. One reviewed proposal remains wrong. Original results are preserved; this is not a fresh held-out study.
The read-only PostgreSQL connector passed actual database-to-publication acceptance on implementation 25c6937: 11 synthetic tables, 14,639 rows, nine reconciled metrics, 58 generated-file hashes checked, 59 published files, 45 audit events and zero waived checks. Its reviewer is scripted automation; no customer source or production deployment is claimed. A new six-case agent-authored corpus retains its initial safety failure: 42/52 correct proposals and 42/43 targets covered. Remediation keeps accuracy unchanged, routes 48/52 proposals to review and passes safety. Ten wrong proposals and ten unresolved fields remain; this is feedback-informed qualification, not independent human validation.
On pinned baseline 03c4fb3, separate operator and reviewer agents completed a synthetic SAP-style workflow. Explicit review rejected a misleading invoice-status mapping; six supported metrics were published, with 38 publication files and 46 audit events verified. The source-informed reviewer saw fixture answer-key entries. No human pilot or time-savings result is claimed.
Deliberate boundaries
Synthetic financial workflows, typed CSV ingestion and bounded public-data profiling. The PostgreSQL connector has actual synthetic database-to-publication verification. Versioned local publication. Production JWT controls are implemented in a deployment candidate; the actual host and identity provider have not been accepted. The owner selected free hosting: this GitHub Pages presentation and the reproducible local workflow. No production backend deployment or measured business savings. General Skill efficacy and independent usability remain unclaimed. Owner-approved annotations and delegated workflow completion are documented.
The implementation was developed with substantial AI assistance. The case study distinguishes implemented behavior, automated verification and owner-approved evidence and outstanding live and independent human validation. Separate-agent benchmark authorship is not external human evaluation.