For AI vendors in procurement and insurance review
You are 40 questions into a security review written for software.
Perpendis compiles your evidence into one sealed pack and hands the underwriter, buyer or auditor a free read-only workspace where they verify it themselves — without trusting us, or you.
One system, no contract, no card. The demo needs no account and runs the real engine in your browser — or create a workspace and do it yourself.
In every pack
Nothing in a pack is presented as measured unless it was measured.
What a pack assesses
Eight domains. One is measured; seven are sample data.
You can see which is which without reading a word of prose.
- Performance & accuracy sample
- Hallucination rate sample
- Robustness sample
- Adversarial & jailbreak sample
- Data leakage sample
- Bias & fairness sample
- Security posture sample
- Computational integrity LIVE
sample = illustrative, not a measurement. LIVE = measured by the engine. The label travels with the score everywhere it is shown (where, exactly). A pack holds these eight and nothing else.
What each domain measures
- Performance & accuracy
- Metrics on your defined test sets, plus per-subgroup accuracy.
- Hallucination rate
- Grounded-reference benchmarks and a third-party-style harness.
- Robustness
- Perturbation suites: casing, OCR noise, distribution shift.
- Adversarial & jailbreak
- Prompt injection, jailbreak and red-team batteries.
- Data leakage
- PII, credential and proprietary-content leak probes.
- Bias & fairness
- Selection rates and impact ratios (Local Law 144 formulas), score gaps.
- Security posture
- Endpoint exposure, unsafe tool-call probes, agent identity.
- Computational integrity
- Signed proof that this model version, on this hardware, produced these results on these samples — re-verifiable bit for bit, with divergence localised to a single primitive operation.
The first seven describe what each domain will measure; their scores today are
sample data.
Alongside them, a documentary workflow. Governance, incident response, data governance, legal and business description are collected as documents, not measured: attachments are counted, what is missing is listed as a gap. Gaps carry no score and are never presented as a domain result.
How it works
One assessment, one evidence pack, one link to hand over.
- Register the system
Name, base model, use case, deployment tier. No API keys, no weights — there is no field for them anywhere in the product.
- Run the assessment
The engine runs in your browser on a fixed reference weight tensor that ships inside it — not your weights, not your traffic — and seals the result with an Ed25519 receipt. About as long as loading this page.
- Compile and share
A versioned pack plus a read-only link: one pack version, expiring, revocable the moment a deal goes cold.
How the assessment reaches your system
Three tiers, one identical pack format: a hosted API endpoint, an agent inside your VPC, or an air-gapped runner where no outbound call is allowed. Today the tier is recorded on the pack and nothing more — the one live measurement runs the referee in your browser and calls your endpoint in no tier.
In tiers two and three only signed evidence artifacts leave your network. Model weights and API keys are never sent to us and never stored by us; there is no field to put them in.
Standards mapping — the design, and what exists today
The design. AIUC-1 is the spine: each domain filed against AIUC-1 control areas, the same artifacts cross-mapped to EU AI Act Annex IV headings, so an auditor finds evidence where they expect it. The signed run receipt is what turns "test reports, dated and signed" from a filing convention into something a third party can check.
None of that mapping is built — see Where we are, stated plainly. And mapping, when it lands, is organisation and citation: not an assessment against a standard we are accredited to certify, and not a claim that any regulation requires this pack.
Three ways people arrive here
An insurance application
Describe your model's failure modes and how you monitor them.
A misrepresented answer is how a claim gets denied two years later.
Enterprise procurement
Attach evidence of adversarial testing and bias evaluation.
A questionnaire written for deterministic software, blocking a contract.
An internal audit
Show me what this model actually did, not a dashboard.
Dated, signed, checkable by someone who does not report to you.
For underwriters, brokers and auditors
You should not have to take the vendor's word for it. Or ours.
The relying side pays nothing, now or later.
sample or LIVE
status, the run behind them, the gaps the vendor has not closed, the version you are
reading. No account.A valid signature is not a pass. It proves the receipt bytes are unaltered — a failed referee result seals just as validly. What it covers, and what it does not.
Where we are, stated plainly
One domain is real today. We label the rest rather than hoping you don't ask.
An evidence company that overstates its own evidence has already failed its one job.
Live today — computational integrity, and only that
A bit-exact referee at tolerance zero, sealed with an Ed25519 receipt. It is the one domain we measure. The demo runs it in your browser; nothing about the verdict is scripted.
Sample data — the other seven domains
Not measurements. Every score you see for them — demo, console, free report, shared
pack, print output, JSON export — is illustrative and carries the label
sample at the point of display. We will not show one without
that label, and we will not describe one as measured.
A valid signature is not a pass
It covers the receipt — the model, both execution targets, the spec version, which side
the referee found correct, the public key and the first divergence — and nothing else;
every other engine field is labelled unsigned. So a failed
referee result seals with an equally valid signature, and the signature never says which
machine ran the engine.
Not built yet — the standards mapping
No domain carries an AIUC-1 control area, no artifact carries a control reference, and "Annex IV" appears nowhere in a pack, a share or a report. Where this page describes that mapping, it is describing the design, not a pack.
Evidence, not certification
No regulation requires this pack. We are not an accredited body, we issue no certificate, and mapping to AIUC-1 or Annex IV is organisation and citation, not conformity assessment.
We do not make verification cheaper
Detection is O(n): you cannot detect a wrong fused multiply-add without recomputing it, and we will not tell you otherwise. Only the dispute is O(log n) — bisection over committed checkpoints settles a disagreement in a logarithmic number of queries on one primitive operation. The saving is in resolving arguments, not in avoiding work.
The detail behind each line — plus four limits that are awkward for us
The engine. Given two machines that ran the same computation, the referee names the exact row and step where they first diverge, says which side is arithmetically correct against an IEEE-754 reference, and tells an honest hardware difference apart from a fabricated trace. It runs on real Qwen2.5-7B and TinyLlama norm tensors, compiles to WebAssembly, and produces identical hashes on x86-64 and wasm32 — held there by a golden-hash gate in the test suite.
The seven sample domains. Their harnesses are largely open-source assembly and are being integrated, but they are not running.
The signature. Where a stored value disagrees with the receipt it is displayed beside, the workspace shows an explicit mismatch rather than the stored value.
The standards mapping. The report cover states only that the pack is organised as Perpendis's own AIUC-1 readiness checklist — a mapping, not an AIUC-1 certificate.
The regulatory picture. The EU AI Act's high-risk obligations land later and may land differently; several US state rules have been narrowed. We also never build or tune a model we assess — the only way the neutrality claim survives inspection.
Scope. Version zero of the referee adjudicates one operator (RMSNorm) on a simulated vendor fleet plus real-silicon probes. Not a whole-model attestation, and one operator on one machine is a small share of any real inference bill.
Our attacks are our own. The adversarial suite proves the referee catches the attacks we thought of, including a fabricated trace that passes naive trace-checking and is still rejected by re-execution. Evidence, not a security proof.
Nobody has relied on us yet. No unaffiliated insurer has formally relied on a Perpendis pack to date. Being first is the offer and the risk, priced into a first report costing nothing.
Public inputs only. The integrity receipts commit to inputs and weights the relying party is allowed to see. Attesting over data you cannot show anyone is a different, harder problem, and we do not claim it.
Pricing
From $15,000 per year, per AI system
- The first report is free. One system, no contract, no card.
-
What that buys today: one measured domain. Computational integrity runs
live and is signed; the other seven are labelled
sampleand are not measurements. The price reflects that — see Where we are, stated plainly. - Confirmed on a call. The number moves with system count, deployment tier and reliance letters. We quote once we have seen the system.
- Reliance letter: $2,000–$5,000 one-time. Optional, invoiced to you, the assessed party — never to the insurer or auditor relying on the pack. That figure is shown to them in the workspace too, so it cannot reach you second-hand.
Who pays, and who never does
The assessed party pays. You are the one being asked for evidence, and credible evidence protects you.
The relying party pays nothing. Insurers, brokers, MGAs, reinsurers and auditors read packs, verify receipts and receive reliance letters at no charge.
Start with the demo, or start with the free report.
One minute: run the fabricated-trace scenario, verify the receipt, flip one character of the signature, watch the check turn red.