Documentation

Performance

The sub-50 ms verify claim, measured by an executable benchmark and gated in CI.

POST /v1/verify is deterministic: the same input always gets the same verdict, with no third-party calls in the path. The full request covers authentication, optional idempotency deduplication, parallel signal reads (enrichment data, per-tenant rate counters, signature checks), the scoring function, and issuing a signed verdict. Degraded paths still return within budget: they fail to challenge.

How it's measured

An executable benchmark boots the real API service in its full production configuration: the same runtime, storage, and signing paths your traffic gets, and drives it over local HTTP:

StepWhat runs
Benchmark50 warm-up rounds, then 1,000 samples per scenario
RuntimeReal edge runtime in production configuration
MachineThe artifact records what it ran on: CPU model, core count, OS, memory, and sample count
CI gatep50 < 60 ms, p95 < 120 ms, and p99 < 120 ms on shared CI runners (stricter 10 ms / 50 ms locally), enforced in CI
ArtifactWrites end-to-end and worker-execution p50, p95, p99, p99.9, mean, and worst observed sample

Scenarios

ScenarioWhat the 1,000 samples exercise
e2eFull local HTTP verify: authentication, idempotency dedup, rate counters, scoring, verdict
e2e worker executionThe same request timed within the service, excluding the local HTTP harness
e2eChallengeSame, with a farm-shaped payload (headless browser, sparse headers) that deterministically returns challenge
e2eIdempotentSame, plus an Idempotency-Key so dedup sits on the critical path
e2eSignedSame, with a real signing key (JWKS-served verdicts)
pureScoringThe decision logic alone, in-memory: the algorithmic floor

The e2e scenarios also run a 200-request concurrent burst to report throughput (requests/sec) at concurrency 8.

Honest reading of the numbers

These measurements capture the whole pipeline in the local runtime, where storage and counters run locally in the same process: no data-center I/O is included. In production you add real network round trips to that storage (tens of milliseconds on cold misses; typically single-digit on warm reads), plus your own client-to-edge path.

The benchmark drives the service over local HTTP. The end-to-end measurement includes the Node-to-service harness; the worker-execution measurement is stamped inside the service and excludes that harness. Neither includes a real client-to-edge hop. The machine the benchmark ran on (CPU, cores, OS, memory) is recorded in the artifact with every run, so any number can be traced back to where it was produced: a developer machine or a CI runner, never a claim about a production cluster.

Each scenario reports the tail as p99.9 in addition to p50/p95/p99, plus the worst observed sample: the headline metrics are p50, p95, and p99 (each with comfortable margin); p99.9 and maximum are reported as supplemental due to small sample size per run (∼1 in 1,000 observations per run; 5,000 total per scenario).

Farm-shaped traffic that returns decision: "challenge" stays within the same budget. For outages: when storage is unavailable, verify short-circuits to a degraded challenge verdict (never allow, never a retry), a path exercised by the test suite and bounded by the same budget because it returns before scoring.