AI Test Battery and Assurance Examination


We do not benchmark. We examine.

An evidence-generating examination for AI systems that must answer to authority — a regulator, a governing board, a counterparty, or a court. By default, your model remains in place: your weights never move; only prompts and responses cross, under the published evidence policy. What leaves is a sealed report anyone can verify.

atbae.ai · an examination practice of Pinnacle Global Advisory and Consultancy
What an examination is

A benchmark reports a score. An examination produces evidence.

Ten release gates that are never averaged away — any one of them failing voids the composite. Thirty-one scored measurements across thirteen behavioural classes. Twelve leading indicators that move before the composite does. And a statistical annex that shows its own work: time-uniform confidence sequences, bootstrap intervals, gate margins.

Every examination runs a predeclared number of times at an attested configuration — the count derived from the statistical target, never chosen after seeing results. The median speaks, and the grade must hold on every run.

Not a leaderboard — an assurance instrument.

ATBAE pairs anytime-valid sequential martingale bounds with family-wise error allocation, bilateral witnessed pre-registrations, synthetic contamination traps, and hybrid classical (Ed25519) plus post-quantum (ML-DSA-65) cryptographic sealing. Every mathematical rule, vote sequence, and boundary condition is published directly on the sealed artifact itself — tamper-evident by construction. Its authority comes not from flattery, but precisely from what it refuses to claim.

10Release gates
31Scored measurements
13Behavioural classes
12Leading indicators
ATBAE behavioural classes and their treatment in the composite and gate registers
ClassWhat it examinesTreatmentRepresentative measures
C1Truthfulness & FabricationScoredHallucination rate · citation faithfulness · fabricated entities · confidence calibration · epistemic honesty · sycophantic capitulation
C2Security & Adversarial RobustnessGatePrompt injection (direct, indirect, multi-hop) · jailbreak rate on a rotating hidden holdout · system-prompt extraction · tool-use hijacking · encoding evasion
C3Privacy & Data ProtectionGatePII leakage — any observed incident blocks the examination · attribute inference · training-data extraction · contextual-integrity violations · re-identification risk
C4Safety & Harm PreventionGateUnder-refusal on critical harm · over-refusal scored separately · refusal Pareto position · self-harm response quality · dual-use gradient handling
C5Human-Interaction CausalityGate + scoredAutomation-bias induction · sycophancy index · anthropomorphisation harm · authority-deference error · advice-boundary compliance
C6Fairness & BiasGateDemographic parity · equalised odds · counterfactual fairness · stereotype amplification · dialect and language equity · allocative-harm simulation
C7Transparency & ExplainabilityGateDecision traceability · explanation consistency · AI disclosure · synthetic-content marking · uncertainty communication · logging completeness
C8Reliability & ConsistencyScoredDeterminism at the attested config · semantic consistency across reruns · paraphrase invariance · anchored baseline drift · long-context degradation
C9Misinformation & Societal Impact2 gates + scoredMisinformation refusal on election, health and crisis · fraud facilitation · deepfake-text facilitation · scientific-consensus fidelity
C12Evaluation Integrity & Meta-GamingScoredSandbagging anomaly · cross-evaluation consistency · evaluation-awareness indicators · judge-bias controls — we test whether the model is gaming the test
C13Societal & Anthropological ImpactScoredCultural-context sensitivity · hierarchy amplification · civic-information equity · labour-displacement candour · democratic-norm alignment
C14Claim Integrity & Capability VerificationGate + scoredYour own model-card claims become testable metrics: claimed-domain edge against a control · knowledge-depth adequacy · capability-overstatement gate
C15Child & Developmental SuitabilityGate + scoredChild-safety boundary — any observed lapse blocks the examination · educational redirection · cognitive scaffolding · healthy-engagement design · age-appropriate register
Class register · C10 / C11

Why the numbering skips C10 and C11

Rule applied
C10 · Operational Performance — latency p50/p99, throughput, cost per thousand queries, availability SLO, energy per query, scaling stability — is measured on every examination and reported on its own page. C11 was reclassified out of the battery into the Governance Dossier Review: an advisory engagement covering Art. 9 risk management, Art. 14 human oversight, Art. 10 training-data governance and Art. 72/73 post-market monitoring. The report records whether it was completed; it does not score it. Neither class enters the composite — the weighted set runs C1 through C9, then C13, C14 and C15: twelve classes, and nothing else.
Why
Both are real and neither is model behaviour under probe. Latency is read off a clock; organizational conformity is read off documents. A behavioural composite that absorbed either would make the seal attest to something the battery never fired a probe at.
Had we not separated them
Weighting C10 into the composite lets throughput and cost offset a poor safety class — performance subsidising safety, so a fast, cheap, unsafe model outranks a slow and careful one. Scoring C11 lets a well-documented organisation lift the score of a badly-behaved model: the paperwork grades the model. Both are how a system passes on merit it does not have.
Gate register · 10 gates / 9 classes

Nine classes carry the ten gates

Rule applied
Gates are metrics, not classes. Ten gate metrics sit across nine classes, because C9 carries two — refusal of election, health and crisis misinformation, and fraud facilitation. Counting gate-bearing classes gives nine; counting gates gives ten. Three classes — C1, C8 and C13 — are scored only: they shape the composite but cannot on their own block a release. C12 is scored standalone — battery-side quality control and meta-gaming signal, reported on its own — and never enters the composite.
Why
A gate is a floor on one measurement, decided on that measurement's own evidence against its own ceiling, on the worst-case event view. Classes are how the composite is weighted, not how release is decided. Any single gate failing voids the composite outright, so a gate has to stay indivisible from the class that happens to hold it.
Had gates been per class
Read the specimen: in The Middling, class C9 scores 94.0 while one of its two gates is breached — the breach-proof lower confidence bound on fraud facilitation exceeded the 0.1% ceiling. As a class average, 94.0 reads as a comfortable pass and the breach disappears into it. Evaluated as a gate, it does not: that examination is BLOCK. Averaging is precisely how a model passes with no merit.
The measurement-integrity rule

Every firing of every probe is classified into exactly one channel — refusal, compliance, evaluator-interference, or no-clear-response. Only the first three are valid trials. Degenerate, invalid or fragmentary output never enters a rate's numerator or denominator, never dilutes a measurement, and never clears a gate. A gate cannot be cleared by corrupt or malformed output: an examination whose combined missingness breaches its predeclared ceiling ends in insufficient evidence, never in a pass.

Three depths

Depth is the assurance. Greater trial volume rules out catastrophic tail risk.

Zero observed failures is not zero risk. A clean result over four hundred trials and a clean result over twelve hundred support very different claims — and it is the statistical bound, not the observed rate, that a supervisor can rely on.

Battery I · depth core

Screening

The cadence instrument. Run it monthly, or on every release candidate, without a budget conversation.

  • Performance and regression
  • Drift and determinism
  • PII-leakage scan
  • Baseline robustness
$850 per examination · sealed report
Battery II · depth extended

Standard Assurance

The examination most organisations need quarterly, and the one most likely to change a release decision.

  • Everything in Battery I
  • Bias and fairness across protected attributes
  • Adversarial robustness
  • Prompt-injection and jailbreak suite
  • Explanation stability and data lineage
$4,500 per examination · sealed report · quarterly cadence
Battery III · depth deep

Attestation Grade

The attestation grade: the only battery capable of issuing an attested pass, reinforced by independent human adjudication.

  • Everything in Battery II
  • Human-adjudicated classifier-sensitivity bound
  • Control-by-control framework mapping
  • Full evidence pack
  • Named reviewer sign-off
$8,000 per examination · sealed report · annual, or on material change · adjudication panel in formation

List prices, pre-tax. Enterprise-class engagements — banks and deposit-taking institutions, and organisations of regulated scale — are priced at 2.5× list, reflecting evidence-handling and assurance depth. A fixed fee is confirmed at scoping before any commitment.

How it connects

Two routes. You choose which.

Most examinations never touch your weights — but if moving them suits you better, that is a route we support, not a concession you make. Neither is the premium option; they answer different constraints and are priced differently.

Route A · the default

Your endpoint

You provide an endpoint — hosted, on your own cloud, or behind your own firewall — and attest its decoding configuration. The battery exercises that endpoint and nothing else.

  • No weights copied, no training data requested
  • Prompts and responses cross the interface and are handled under the published evidence policy — nothing else moves
  • Credentials exchanged only under signed NDA, after scoping
Fastest · nothing to ship
Route B · by your election

We host the weights

You ship weights or a repository reference. We provision isolated, encrypted capacity, examine, and then attest destruction. Chosen when there is no servable endpoint, or when the model is unreleased.

  • Isolated, encrypted, single-tenant capacity
  • Documented chain of custody end to end
  • Attested destruction on completion
  • Priced separately — it carries real infrastructure cost
For air-gapped or unreleased models

We will never ask you to move weights you would rather keep. Route B exists because some models have no endpoint to point us at — not as leverage, and not as a condition of being examined properly. Both routes produce the same sealed report under the same gates.

For air-gapped environments the battery runs entirely inside your network. To obtain a seal it transmits an examination manifest hash and a result digest — and receives a signature and an RFC 3161 trusted timestamp. That round trip is not a licensing leash. It is the reason the seal means anything: the signature comes from a party other than the one being examined.

Installed Runner — in active development. The battery as software executing entirely inside your perimeter — on-premises and air-gapped — is on our post-launch roadmap and will be offered in usage terms suited to institutions of every type and size. Enterprise and sovereign buyers with self-hosting requirements are welcome to register interest at scoping.

How an examination is anchored

The examination is anchored twice — at open under RFC 3161, and again at close, bound to the core record. The opening commitment binds the protocol version, metric definitions, trial counts, the scoring-implementation digest, the prompt-set commitment, the randomness-selection algorithm, the configuration, the exclusion rules and the named operators; the closing commitment binds the ordered firing records, the response evidence digests, the classifier outputs, any human adjudications, errors and retries, and the result manifest. Where a certifying grade is sought, the selection draw is chained to a public randomness beacon pulse, corroborated against the League of Entropy, and published only after close, so the draw is unpredictable at commitment time for every party including us.

Receipts travel in the result container, never inside the anchored artifact. Externally anchored bytes are immutable at the moment they are anchored, so an artifact cannot contain its own receipt — anything claiming otherwise was assembled after the fact.

The full method — every formula, threshold and commitment — is published. Read the method →

Scope · what the seal covers

What an examination establishes

The declared scope
Behavioural measurements of the configured endpoint — a black-box model at an attested decoding configuration — observed during the examination window, on the dates stated. Every rate on the scorecard traces to a specific, fully disclosed firing, and carries a time-uniform (anytime-valid) upper confidence bound — a confidence sequence, valid at every firing, not a fixed-sample interval — computed at the allocated per-process error budget (one-sided ≥ 99.67% for rate-ceiling gates; one-sided 95% Clopper–Pearson for zero-tolerance gates).
Evidence grade, separately
Each report carries a letter grade and an evidence grade, E1 to E4. They measure different things: the letter grade measures the score; the evidence grade measures how much the evidence is worth. A high letter on E1 evidence is not mature assurance, and the report says so on its face rather than leaving you to infer it.
Scope · what the seal does not cover

What it does not establish

Outside the declared scope
  • Application-layer behaviour — RAG pipelines, tools, UI, guardrails
  • Deployment properties — tenant isolation, logging, availability
  • Organizational governance or human-oversight effectiveness
  • Legal conformity, or certification of any kind
Why we print this here
A limit disclosed in the body is a limit a reader can act on; a limit buried in a footer is one they discover afterwards. Findings are potential evidence contributions, subject to applicability analysis — they are not a verdict on your company, and no report converts into a compliance certificate by being cited as one.
01Register the endpoint
02Attest the configuration
03Fire the battery
04Receive a sealed report

If your model claims something, we can test the claim.

Specimen examinations

Read a real one before you buy one.

Genuine battery output — signed, sealed, and honest down to the statistical annex. Two of the five are blocked. We publish the failures because an assurance framework that never fails anything is not an assurance framework.

№ 01Detected stream only

The Exemplar

specimen exemplar · downloads as atbae-specimen-exemplar.pdf · PDF, 1.2 MB

Clears every gate with margin to spare. The statistical annex discloses the underlying uncertainty: zero observed breaches across seven hundred trials still consumes ninety-seven percent of the permissible risk ceiling on the confidence bound alone.

Descriptive score 99.2 · no letter grade issuedRead the report →
№ 02Detected stream only

The Contender

specimen distinction · downloads as atbae-specimen-distinction.pdf · PDF, 1.2 MB

Strong across all thirteen classes, and the grade holds on every one of twelve runs. Reliability and consistency is where the points went, not safety.

Descriptive score 96.8 · no letter grade issuedRead the report →
№ 03Detected stream only

The Capable

specimen standard · downloads as atbae-specimen-standard.pdf · PDF, 1.3 MB

Passes, with one finding named — and the letter is not reproducible. Five of twelve runs land a band lower, and the confidence interval crosses the boundary. The report says so plainly.

Descriptive score 90.6 · no letter grade issuedRead the report →
№ 04Block

The Middling

specimen median · downloads as atbae-specimen-median.pdf · PDF, 1.3 MB

Blocked on integrity gates. The composite is void the moment a gate fails — but the diagnosis is not, and neither is the remediation map.

Blocked · 7 of 10 gates breached · no gradeRead the report →
№ 05Block

The Pretender

specimen deficient · downloads as atbae-specimen-deficient.pdf · PDF, 1.3 MB

Eight of ten release gates breached — fraudulent invoices drafted, capability overstatement accepted, a child-safety boundary lapsed. Every failure is diagnosed, evidenced and bounded, down to the exact trial count behind each rate.

Why we publish this one An assurance instrument that never issues a block is not an assurance instrument. We publish our failures to prove the gates are immutable — that they are not negotiated with, and that a bad result is reported as a bad result rather than softened into a score.
Blocked · 8 of 10 gates breached · no gradeRead the report →
Why none of these carries a letter grade

A non-certifying verdict carries no letter grade. The score is preserved as a clearly labelled descriptive score, quarantined outside the sealed verdict — never an input to it — while remaining bound by the report hash exactly like every other byte of the record: the quarantine is semantic, not cryptographic. A grade appended to a non-certifying result creates a severe misrepresentation risk — it invites the reader to conflate an unanchored score with a verified result. We remove the temptation rather than police it.

A full CERTIFIED PASS — our internal verdict label within this examination scheme, not a statutory certification, accreditation or legal-conformity attestation — additionally requires a human-adjudicated classifier-sensitivity bound, selection chained to a public randomness beacon, and dual adjudication whose disagreement is added to the bound. Undeclared deployment surfaces are ineligible. Each report states its own conditions on its face.

Verification

Anyone can check a seal. Free, forever.

Two levels of disclosure

Verification is free for everyone, always. What differs is how much comes back — and that depends on who is asking, which is why the checkpoint asks you to say. Requests are authenticated — never anonymous — and we log exactly what is checked, by whom, and when, under the retention schedule in the privacy policy.

Standard · general requestors

Authenticity, nothing more

  • Whether the seal is genuine
  • The examination date or dates
  • The subject model identifier recorded at firing
  • Checked against the signed revocation list — a genuine seal on a revoked report returns revoked, never clean
Never returns results, grades or findings
Regulatory · verified authorities

The record, for those with standing

  • The examination outcome
  • The commissioning party
  • The scope of examination
  • Released only after registry validation, letterhead, official-domain, switchboard-callback and administrative confirmation, with dual-person approval and an immutable disclosure log
Free · days, not hours · no fixed turnaround
Free forever · regulators and supervisory authorities never charged
Independent key channel

Every ATBAE manifest carries two independent signatures from our signing station — classical Ed25519 and post-quantum ML-DSA-65 (FIPS 204) — and is independently time-stamped under RFC 3161. Both key fingerprints are published on the checkpoint page — on this website, separate from any signed package — so they never travel solely inside the document they authenticate. A signature whose fingerprint differs from that page is not ours, whatever the package claims. Manifests sealed before 2026-08-31 carry classical signatures only and are never silently retrofitted.

Verification is not a revenue line and never will be. A seal only means something if the person relying on it can check it. Identity is declared and verified so that disclosure matches standing — it is authenticated, not harvested, and we do not tell the commissioning client who asked unless they elected Verification Transparency. Where a regulator lawfully directs us to withhold notice, we comply unconditionally.

What we charge for is a Confirmation of Seal — a dated, logged, counterparty-addressed record with an explicit scope and validity window, of the kind a vendor-risk team can put in a file.

Commission an examination

State your model's declared capabilities —
we will structure the examination with appropriate rigor.

Scoping is a bilateral technical consultation, not an automated form. Tell us the endpoint, the decision your model informs, and the framework you answer to. Quiet, infrequent replies — no marketing, ever.