Product · evidence pack console
One assessment, one sealed pack, one link to hand over.
Eight assessed domains, for a vendor asked to evidence an AI system. One domain is measured today; seven are labelled sample data.
In every pack
Nothing is presented as measured unless it was measured.
What a pack assesses
Eight domains. One is measured; seven are sample data.
- Performance & accuracy sample
- Hallucination rate sample
- Robustness sample
- Adversarial & jailbreak sample
- Data leakage sample
- Bias & fairness sample
- Security posture sample
- Computational integrity LIVE
sample = illustrative, not a measurement. LIVE = measured by the engine. The label travels with the score everywhere it is shown (where, exactly). A pack holds these eight and nothing else.
What each domain measures
- Performance & accuracy
- Your test sets, plus per-subgroup accuracy.
- Hallucination rate
- Grounded-reference benchmarks.
- Robustness
- Casing, OCR noise, distribution shift.
- Adversarial & jailbreak
- Prompt injection and red-team batteries.
- Data leakage
- PII, credential and proprietary-content probes.
- Bias & fairness
- Impact ratios (Local Law 144 formulas), score gaps.
- Security posture
- Endpoint exposure, unsafe tool-call probes.
- Computational integrity
- Signed proof this model version, on this hardware, produced these results — re-verifiable bit for bit.
The first seven describe what each domain will measure; their scores today are
sample data. Governance, incident response, data governance,
legal and business description are documents, never scored.
How it works
- Register the system
Name, base model, use case, deployment tier. No API keys, no weights — there is no field for them anywhere in the product.
- Run the assessment
Runs in your browser on a fixed reference tensor — not your weights, not your traffic — sealed with an Ed25519 receipt. Takes seconds.
- Compile and share
A versioned pack plus a read-only link.
How the assessment reaches your system
Three tiers, one pack format: hosted API endpoint, agent inside your VPC, or air-gapped runner. In tiers two and three only signed artifacts leave your network.
Today the tier is recorded on the pack and nothing more — the one live measurement runs in your browser and calls your endpoint in no tier.
Model weights and API keys are never sent to us and never stored by us; there is no field to put them in.
Standards mapping — designed, not built
The design. Each domain filed against AIUC-1 control areas, cross-mapped to EU AI Act Annex IV.
None of it is built, and mapping is organisation and citation, not conformity assessment — see Where we are, stated plainly.
Who arrives here, and what they were asked for
An insurance application
Describe your model's failure modes and how you monitor them.
A misrepresented answer is how a claim gets denied.
Enterprise procurement
Attach evidence of adversarial testing and bias evaluation.
A questionnaire written for deterministic software.
An internal audit
Show me what this model actually did, not a dashboard.
Checkable by someone who does not report to you.
For underwriters, brokers and auditors
You should not have to take the vendor's word for it. Or ours.
sample or LIVE
status, the run behind them, the vendor's open gaps. No account.A valid signature is not a pass — it proves the receipt bytes are unaltered, and a failed result seals just as validly. What it covers, and what it does not.
Where we are, stated plainly
One domain is real today. We label the rest.
Live today — computational integrity, and only that
A bit-exact referee at tolerance zero, sealed with an Ed25519 receipt. The only domain we measure. The demo runs it live in your browser; nothing is scripted.
Sample data — the other seven domains
Not measurements; their harnesses are not running. Every score for them — demo, console,
report, share, print, JSON export — is illustrative and carries the label
sample at the point of display. We will not show one without
that label, and we will not describe one as measured.
A valid signature is not a pass
It covers the receipt — model, both execution targets, spec version, which side the
referee found correct, public key, first divergence — and nothing else; every other
engine field is labelled unsigned. A failed result seals just
as validly, and the signature never says which machine ran the engine.
Not built yet — the standards mapping
No domain carries an AIUC-1 control area, no artifact a control reference, and "Annex IV" appears nowhere in a pack, share or report. Where this page describes that mapping, it is describing the design, not a pack.
Evidence, not certification
No regulation requires this pack. We are not an accredited body, we issue no certificate, and mapping to AIUC-1 or Annex IV is organisation and citation, not conformity assessment.
We do not make verification cheaper
Detection is O(n): you cannot detect a wrong fused multiply-add without recomputing it. Only the dispute is O(log n) — bisection over committed checkpoints settles it on one primitive operation. The saving is in resolving arguments, not avoiding work.
The engine, and six limits that are awkward for us
The engine. The referee names where two machines first diverge, says which side is correct against an IEEE-754 reference, and tells honest hardware difference from a fabricated trace. Real Qwen2.5-7B and TinyLlama tensors.
Scope. Version zero adjudicates one operator (RMSNorm) on a simulated fleet plus real-silicon probes. Not a whole-model attestation, and one operator on one machine is a small share of any real inference bill.
Our attacks are our own. The suite proves the referee catches the attacks we thought of — evidence, not a security proof.
Nobody has relied on us yet. No unaffiliated insurer has formally relied on a Perpendis pack.
Public inputs only. Receipts commit to inputs and weights the relying party may see; we claim nothing beyond.
The regulatory picture. The EU AI Act's high-risk obligations land later and may land differently; several US state rules have been narrowed.
Neutrality. We never build or tune a model we assess.
Custom work and consultation.
If you need a domain measured that we do not measure yet, or an assessment shaped around evidence you already hold, tell us what you are being asked for.