PerpendisSigned evidence you check yourself.
Live demo Sign in

Product · evidence pack console

One assessment, one sealed pack, one link to hand over.

Eight assessed domains, for a vendor asked to evidence an AI system. One domain is measured today; seven are labelled sample data.

← All Perpendis tools

In every pack

One attested run
Signed, bit-exactly re-verifiable.
LIVE
Seven illustrative domains
Not measured. Labelled wherever a score appears.
sample
A documentary gap list
Counted, never scored. Anything missing shows as a gap.
A relying-party link
Read-only, one pack version, expiring, revocable.
A reliance letter on request
Named party, purpose-limited, liability capped.

Nothing is presented as measured unless it was measured.

What a pack assesses

Eight domains. One is measured; seven are sample data.

  • Performance & accuracy sample
  • Hallucination rate sample
  • Robustness sample
  • Adversarial & jailbreak sample
  • Data leakage sample
  • Bias & fairness sample
  • Security posture sample
  • Computational integrity LIVE

sample = illustrative, not a measurement. LIVE = measured by the engine. The label travels with the score everywhere it is shown (where, exactly). A pack holds these eight and nothing else.

What each domain measures
Performance & accuracy
Your test sets, plus per-subgroup accuracy.
Hallucination rate
Grounded-reference benchmarks.
Robustness
Casing, OCR noise, distribution shift.
Adversarial & jailbreak
Prompt injection and red-team batteries.
Data leakage
PII, credential and proprietary-content probes.
Bias & fairness
Impact ratios (Local Law 144 formulas), score gaps.
Security posture
Endpoint exposure, unsafe tool-call probes.
Computational integrity
Signed proof this model version, on this hardware, produced these results — re-verifiable bit for bit.

The first seven describe what each domain will measure; their scores today are sample data. Governance, incident response, data governance, legal and business description are documents, never scored.

How it works

  1. Register the system

    Name, base model, use case, deployment tier. No API keys, no weights — there is no field for them anywhere in the product.

  2. Run the assessment

    Runs in your browser on a fixed reference tensor — not your weights, not your traffic — sealed with an Ed25519 receipt. Takes seconds.

  3. Compile and share

    A versioned pack plus a read-only link.

How the assessment reaches your system

Three tiers, one pack format: hosted API endpoint, agent inside your VPC, or air-gapped runner. In tiers two and three only signed artifacts leave your network.

Today the tier is recorded on the pack and nothing more — the one live measurement runs in your browser and calls your endpoint in no tier.

Model weights and API keys are never sent to us and never stored by us; there is no field to put them in.

Standards mapping — designed, not built

The design. Each domain filed against AIUC-1 control areas, cross-mapped to EU AI Act Annex IV.

None of it is built, and mapping is organisation and citation, not conformity assessment — see Where we are, stated plainly.

Who arrives here, and what they were asked for

An insurance application

Describe your model's failure modes and how you monitor them.

A misrepresented answer is how a claim gets denied.

Enterprise procurement

Attach evidence of adversarial testing and bias evaluation.

A questionnaire written for deterministic software.

An internal audit

Show me what this model actually did, not a dashboard.

Checkable by someone who does not report to you.

For underwriters, brokers and auditors

You should not have to take the vendor's word for it. Or ours.

A read-only workspace
Domain results with their sample or LIVE status, the run behind them, the vendor's open gaps. No account.
Verify the signature yourself
Ed25519 over canonical JSON, verified on load — or check it with your own tooling. A pack edited after sealing fails, including by us.
A reliance letter that names you
Scoped to the pack and model version you read, with a liability cap and restricted-use terms. The vendor's duty of care is preserved, not transferred.

A valid signature is not a pass — it proves the receipt bytes are unaltered, and a failed result seals just as validly. What it covers, and what it does not.

Where we are, stated plainly

One domain is real today. We label the rest.

Live today — computational integrity, and only that

A bit-exact referee at tolerance zero, sealed with an Ed25519 receipt. The only domain we measure. The demo runs it live in your browser; nothing is scripted.

Sample data — the other seven domains

Not measurements; their harnesses are not running. Every score for them — demo, console, report, share, print, JSON export — is illustrative and carries the label sample at the point of display. We will not show one without that label, and we will not describe one as measured.

A valid signature is not a pass

It covers the receipt — model, both execution targets, spec version, which side the referee found correct, public key, first divergence — and nothing else; every other engine field is labelled unsigned. A failed result seals just as validly, and the signature never says which machine ran the engine.

Not built yet — the standards mapping

No domain carries an AIUC-1 control area, no artifact a control reference, and "Annex IV" appears nowhere in a pack, share or report. Where this page describes that mapping, it is describing the design, not a pack.

Evidence, not certification

No regulation requires this pack. We are not an accredited body, we issue no certificate, and mapping to AIUC-1 or Annex IV is organisation and citation, not conformity assessment.

We do not make verification cheaper

Detection is O(n): you cannot detect a wrong fused multiply-add without recomputing it. Only the dispute is O(log n) — bisection over committed checkpoints settles it on one primitive operation. The saving is in resolving arguments, not avoiding work.

The engine, and six limits that are awkward for us

The engine. The referee names where two machines first diverge, says which side is correct against an IEEE-754 reference, and tells honest hardware difference from a fabricated trace. Real Qwen2.5-7B and TinyLlama tensors.

Scope. Version zero adjudicates one operator (RMSNorm) on a simulated fleet plus real-silicon probes. Not a whole-model attestation, and one operator on one machine is a small share of any real inference bill.

Our attacks are our own. The suite proves the referee catches the attacks we thought of — evidence, not a security proof.

Nobody has relied on us yet. No unaffiliated insurer has formally relied on a Perpendis pack.

Public inputs only. Receipts commit to inputs and weights the relying party may see; we claim nothing beyond.

The regulatory picture. The EU AI Act's high-risk obligations land later and may land differently; several US state rules have been narrowed.

Neutrality. We never build or tune a model we assess.

Custom work and consultation.

If you need a domain measured that we do not measure yet, or an assessment shaped around evidence you already hold, tell us what you are being asked for.