Documentation

Everything to integrate, configure and self-host.

From your first decision to a production deployment — the full path, in depth.

Beta

Verdict is currently in beta (v0.7.0, pre-1.0). It's functional and self-hostable, but APIs, schemas, and defaults may still change between releases — pin a version, review the changelog before upgrading, and evaluate carefully before relying on it in production. It's open source (Apache-2.0); issues and contributions are welcome.

Introduction

Verdict is an open-source fraud & risk decisioning engine. You send it an event; it evaluates the event against rules you can read and returns a verdict — allow, challenge, review, or deny — in a single synchronous call, along with a 0–100 score and the exact reasons behind it.

It is deliberately a decisioning layer, not a data vendor. You bring the signals you already have — the transaction, the device, the IP, whatever your enrichment provides — and Verdict turns them into a consistent, auditable, versioned decision. That is the half most teams end up rebuilding badly: the rules, the scoring, the policy, the human review loop, and the record of why. Verdict is that half, self-hosted and transparent.

Under the hood it is a modular monolith (NestJS + TypeScript): bounded contexts — ingest, feature store, rules, scoring, decision, lists, cases, feedback, analytics, graph, anomaly, auth and API keys — that talk only over typed ports and events, so any one can be extracted into its own service when you outgrow it. Two surfaces sit on top: the engine API (documented here) and the operator dashboard (review queue, configuration, analytics, graph explorer, keys).

event → features + graph + anomaly → rules → score 0–100 → policy bands → verdict

Quickstart

You need Docker with the Compose plugin (docker compose version). One command brings up Postgres, the engine, and the dashboard, wired together over the compose network — no CORS, nothing else to configure:

# 1 · bring up postgres + engine + dashboard
git clone https://github.com/verdict-engine/verdict-engine && cd verdict-engine
AUTH_SECRET=$(openssl rand -hex 32) docker compose up --build

#   db         postgres 16 · durable users, cases, verdicts, keys, graph
#   engine     http://localhost:4000   (POST /v1/decisions, /health)
#   dashboard  http://localhost:3000   (review queue, config, analytics)

# 2 · open the dashboard, create the admin account (first user),
#     then Keys → Create key, and copy it (shown once).

# 3 · send your first decision (next section).

The first person to register in the dashboard becomes the admin; public sign-up then closes and further operators are added by an admin under Team. Create a service API key under Keys — it is shown once and stored only as a SHA-256 hash — and you are ready to send decisions. Confirm the engine is up with curl localhost:4000/health.

Integration

Integrating Verdict is five steps: get a key, call the decision API on the path you want to protect, act on the verdict, feed outcomes back so it learns, then tune and go live. Each step below is self-contained — you can be live with steps 1–3 in an afternoon and layer in 4–5 later.

  1. Create a service API key

    In the dashboard, Keys → Create key. It's shown once — store it as a secret (env var / secret manager). Every call to the decision API sends it as X-API-Key.

  2. Call the decision API on your critical path

    Right before you commit the action you want to protect — authorize a charge, complete a login, release a payout — send the event to POST /v1/decisions and wait for the verdict (a single synchronous call). Send every identifier you have; absent fields simply mean the rules that read them don't fire. Backfilling or scoring in bulk? Use POST /v1/decisions/batch (up to 100 events at once) or, to keep it off your response path, POST /v1/decisions/async (202 + an id you poll at GET /v1/decisions/{id} or receive by webhook).

  3. Act on the verdict

    Branch on verdict: proceed on allow, step the user up (3-D Secure / OTP) on challenge, hold for an analyst on review (a case opens automatically), and block on deny. The reasons tell you exactly why.

  4. Feed outcomes back

    Close the loop so the engine learns: analysts resolve review cases in the dashboard, and your PSP posts chargebacks to POST /v1/labels/chargeback. Those labels sharpen the adaptive model and power backtesting.

  5. Tune, then go to production

    Author rules and policies in Configure, backtest a change against your labeled history before publishing, and deploy with a real AUTH_SECRET and Postgres (see Deployment).

POST /v1/decisions is a machine-to-machine endpoint — it takes a service API key, not an operator login. Here it is end to end:

curl
curl -X POST http://localhost:4000/v1/decisions \
  -H 'content-type: application/json' \
  -H 'x-api-key: vk_live_…' \
  -H 'idempotency-key: order_1837' \
  -d '{
    "type": "card.authorize",
    "amount": 3500, "currency": "USD",
    "subject": { "userId": "usr_3f9a", "deviceId": "dev_2b1", "ip": "203.0.113.7", "channel": "visa" },
    "instrument": { "kind": "card", "bin": "411111", "threeDS": false }
  }'
Node / TypeScript
const res = await fetch(VERDICT_API_URL + "/v1/decisions", {
  method: "POST",
  headers: {
    "content-type": "application/json",
    "x-api-key": process.env.VERDICT_API_KEY,
    "idempotency-key": order.id,       // retries return the same decision
    "x-correlation-id": traceId,        // threaded through logs + events
  },
  body: JSON.stringify({
    type: "card.authorize",
    amount: order.amount, currency: order.currency,
    subject: { userId: order.userId, deviceId, ip: req.ip, channel: "visa" },
    instrument: { kind: "card", bin: card.bin, threeDS: order.threeDS },
  }),
});

if (res.status === 429) return retryAfter(res.headers.get("retry-after"));
const { verdict, score, reasons } = await res.json();
switch (verdict) {
  case "deny":      return block();
  case "challenge": return stepUp();          // 3-D Secure / OTP
  case "review":    return hold(order.id);     // a case opens for an analyst
  default:          return proceed();          // "allow"
}

Two headers make integration robust. Pass an Idempotency-Key (an order id works well) so a retry returns the original decision instead of scoring twice; without one it defaults to the event id. Pass an X-Correlation-Id to thread the request through the engine's logs and internal events for tracing.

The endpoint is rate-limited per API key (default 600/min): over budget it returns 429 with a Retry-After header, and every response carries X-RateLimit-Limit and X-RateLimit-Remaining. A malformed event returns 400 with a coded reason; a transient internal failure returns the policy's on_error verdict rather than an error, so you always get a decision. The response is fully explainable:

{
  "id": "vd_0mu4…", "eventId": "evt_0mu4…",
  "verdict": "review", "score": 55,
  "reasons": [
    { "tag": "takeover", "points": 33 },
    { "tag": "no_3ds",   "points": 22 }
  ],
  "policyId": "pol_card_authorize", "policyVersion": "v0.4.0",
  "decidedAt": "2026-09-17T10:22:00.000Z"
}

Every field of the request and response is defined in the API reference. Prefer a client library? The official SDKs wrap all of this — typed calls, retries and timeouts included.

SDKs

Official client libraries wrap the API-key data plane so you don't hand-write HTTP: typed events and verdicts, a per-request timeout, and automatic retry with backoff on 429/5xx/network errors (honoring Retry-After). They cover the integrator endpoints only — operator and admin actions (auth, cases, rules, config) stay in the dashboard.

TypeScript / JavaScript@verdict/sdkAvailable
Dart / Flutterverdict_sdkAvailable
Pythonverdict-sdkPlanned
React Native@verdict/react-nativePlanned
1 · Install
# TypeScript / JavaScript (Node 18+, Deno, Bun, edge)
npm i @verdict/sdk

# Flutter / Dart — in pubspec.yaml:
#   dependencies:
#     verdict_sdk: ^0.1.0
2 · Score an event

Construct a client once with your service key, then call decide on your critical path. The idempotencyKey (an order id works well) makes a retry return the original decision instead of scoring twice.

TypeScript
import { VerdictClient } from "@verdict/sdk";

const verdict = new VerdictClient({ apiKey: process.env.VERDICT_API_KEY });

const decision = await verdict.decide({
  type: "card.authorize",
  amount: order.amount, currency: order.currency,
  subject: { userId: order.userId, ip: req.ip, fingerprint },
}, { idempotencyKey: order.id });        // retries return the same decision

switch (decision.verdict) {
  case "deny":      return block(decision.reasons);
  case "challenge": return stepUp();
  case "review":    return hold(order.id);
  default:          return proceed();
}
Flutter / Dart
import 'package:verdict_sdk/verdict.dart';

final verdict = VerdictClient(apiKey: apiKey);

final decision = await verdict.decide(const VerdictEvent(
  type: VerdictEventType.cardAuthorize,
  amount: 4900, currency: 'ETB',
  subject: Subject(userId: 'usr_3f9a', ip: '196.188.120.4', fingerprint: 'fp_9c1e'),
));

if (decision.verdict == Verdict.deny) block(decision.reasons);
3 · Handle failure — decide how to fail

A deny should block; a transport failure is a product decision. Each SDK throws a typed error/exception so you branch on the kind — rate limit, auth, or an engine outage — and choose fail-open (allow, favor availability) or fail-closed (challenge/deny, favor safety) per event type, mirroring the policy's on_error. The client also surfaces the rate-limit budget from the last response.

TypeScript
import {
  VerdictClient,
  VerdictRateLimitError, VerdictAuthError, VerdictConnectionError,
} from "@verdict/sdk";

const verdict = new VerdictClient({
  apiKey: process.env.VERDICT_API_KEY,
  baseUrl: "https://verdict.internal",  // default http://localhost:4000
  timeoutMs: 10_000,                     // per request
  maxRetries: 2,                         // 429 / 5xx / network, backoff + Retry-After
});

try {
  const decision = await verdict.decide(event, { idempotencyKey: order.id });
  // …act on decision.verdict…
} catch (err) {
  if (err instanceof VerdictRateLimitError)   await backoff(err.retryAfterMs);
  else if (err instanceof VerdictAuthError)   rotateKey();       // 401 / 403
  else if (err instanceof VerdictConnectionError) failOpen();    // outage — your call
  else throw err;                                                // 400 / 5xx
}

verdict.rateLimit;   // { limit, remaining } from the last response's headers
Flutter / Dart
try {
  final decision = await verdict.decide(event, idempotencyKey: order.id);
  if (decision.verdict == Verdict.deny) block(decision.reasons);
} on VerdictRateLimitException catch (e) {
  await Future<void>.delayed(e.retryAfter ?? const Duration(seconds: 1));
} on VerdictConnectionException {
  failOpen();   // engine outage — availability vs safety is your call
}

// Batch results are a sealed type — switch exhaustively
for (final item in await verdict.decideBatch(events)) {
  switch (item) {
    case BatchOk(:final decision): act(decision);
    case BatchError(:final code):  log(code);
  }
}
4 · Batch, async & feedback

Score in bulk, move scoring off your response path, and feed outcomes back — all typed. Poll getDecision for an async verdict, or (better at volume) subscribe a webhook to verdict.reached.v1.

// Batch — up to 100, order preserved; a bad event fails only its own entry
for (const r of await verdict.decideBatch(events)) {
  if (r.ok) act(r.decision);
  else      log(r.error.code);
}

// Async — 202 now, verdict later (poll, or subscribe a webhook to verdict.reached.v1)
const { id } = await verdict.decideAsync({ id: "evt_9f2a", ...event });
const decision = await verdict.getDecision(id);  // throws VerdictNotFoundError while pending

// Close the loop — record a chargeback as a fraud label (needs the "labels" scope)
await verdict.recordChargeback(order.eventId);
Method → endpoint
decide(event, opts?)POST /v1/decisionsScore one event synchronously → Decision.
decideBatch(events)POST /v1/decisions/batchUp to 100 events; per-entry ok/error, order preserved.
decideAsync(event)POST /v1/decisions/async202 + event id; verdict computed off the response path.
getDecision(id)GET /v1/decisions/:idFetch an async verdict; not-found error while pending.
recordChargeback(id)POST /v1/labels/chargebackRecord a chargeback as a fraud label (labels scope).

Keep service keys server-side. A key shipped in a mobile or browser binary can be extracted, so score from a trusted backend or proxy through your server — the SDKs target that integrator context, not untrusted clients. Full method docs, options and error tables are in each package's README (@verdict/sdk and verdict_sdk); every field is in the API reference.

Best practices

Everything above gets you a verdict. These are the habits that make an integration robust in production — how to call the engine, how to fail, and how to keep decisions getting better over time. None are required to start; adopt them as you harden.

Calling the engine
  • Call it from your backend, never a browser or mobile client. The service key is a secret — keep it in an env var / secret manager, server-side. The decision and label endpoints authenticate with X-API-Key; every dashboard, admin and read endpoint uses a short-lived operator bearer token instead. Client code should hold neither. Rotate keys under Keys and revoke a leaked one instantly.
  • Put the call on the critical path, with a timeout. Send the event right before the irreversible action and wait for the verdict. Set a tight client timeout (≈1–2 s) and decide your fallback per surface: on a network failure or timeout, let a login through but hold a payout. The engine already degrades its own faults to the policy's on_error verdict — but a failure between you and the engine is yours to handle.
  • Always send an Idempotency-Key. Use a natural id — the order id, the login-attempt id. Retries then return the same verdict instead of scoring twice, so a network retry is always safe.
  • Send every identifier you have. userId, deviceId, ip and phone feed the entity graph and the anomaly baseline; a field you omit simply means the rules that read it don't fire. More context, better decisions.
  • Respect rate limits. Back off on 429 using the Retry-After header, and watch X-RateLimit-Remaining so you tune concurrency before you hit it.
  • Thread a correlation id. Pass X-Correlation-Id so one request is traceable across your logs, the engine's logs, and its events.
Acting on the verdict
  • Handle all four verdicts, not just allow/deny. Wire challenge to a real step-up (3-D Secure / OTP) and review to a hold — collapsing them to a block is where good users get lost. Keep the reasons with the transaction for support and compliance.
  • Treat a -1 score as “short-circuited”, not low-risk. A watchlist hit or a degraded decision reports -1 with an explanatory tag — branch on the verdict, never on the raw score.
Getting better over time
  • Close the loop — it's the whole point. Resolve review cases in the dashboard and post chargebacks to POST /v1/labels/chargeback. Chargebacks are the only signal for fraud you allowed and never reviewed, so this is what makes the adaptive scorer and backtests trustworthy.
  • Backtest before you publish, and roll out gradually. Replay any rule or policy change against your labeled history first. When you go live, start lenient — or run in shadow mode, logging the verdict without enforcing it — then tighten the bands as the data confirms them.
  • Audit what you send. Send a card bin only, never a full PAN. The engine masks phone numbers and IPs in its own logs, and the API activity log (GET /v1/activity, or the dashboard's Activity tab) lets you see exactly what each call sent and what it got back.
  • Monitor the mix, not just uptime. Watch verdict distribution, false-positive rate, p99 latency and 429s on the Analytics tab (or scrape GET /metrics). Wire outcomes onward two ways, for two audiences: alert notifications ping a person (a Slack channel on a deny spike or a dead-letter), while webhooks feed a system — the event goes to your own services so they can react automatically (update a risk store, kick off downstream automation, land it in your warehouse). Same events, different consumer.

Develop against live, interactive docs. The engine serves Swagger UI at http://localhost:4000/docs — log in with an operator account right on the page and call any endpoint against your own instance, with request and response schemas for every field.

Use cases

The same event → verdict shape covers very different risk surfaces. Here are four common ones — the event you'd send, and how you'd act on the answer.

Card payment authorization

Protect checkout — score a card charge before you authorize it.

POST /v1/decisions
{ "type": "card.authorize",
  "amount": 4200, "currency": "USD",
  "subject": { "userId": "usr_9", "deviceId": "dev_3", "ip": "…" },
  "instrument": { "kind": "card", "bin": "411111", "threeDS": false } }

allow → charge · challenge → trigger 3-D Secure · review → hold & queue · deny → decline. Signals: velocity, no-3DS, device takeover, amount.

Account login / takeover

Catch credential stuffing and account takeover at sign-in.

POST /v1/decisions
{ "type": "account.login",
  "subject": { "userId": "usr_9", "deviceId": "dev_new", "ip": "203.0.113.7", "channel": "web" } }

allow → sign in · challenge → send an OTP · deny → block. Signals: login velocity, new device, shared-IP ring.

Wallet withdrawal / payout

Guard money leaving the platform — where a miss is expensive, so fail closed.

POST /v1/decisions
{ "type": "wallet.withdraw",
  "amount": 9000, "currency": "USD",
  "subject": { "userId": "usr_9", "deviceId": "dev_3" } }

review / deny hold or block a suspicious payout. The policy's on_error is fail_closed here. Signals: amount anomaly (vs the user's own history), dormant-then-spike, velocity.

Signup & promo abuse

Stop one actor farming bonuses across many fake accounts.

POST /v1/decisions
{ "type": "account.login",
  "subject": { "userId": "usr_new", "deviceId": "dev_shared", "ip": "198.51.100.4" } }

The entity graph links accounts that share a device or IP; a large graph.ringSize flags the ring → review or deny. Explore any entity's cluster in the dashboard's Graph tab.

LLM & MCP integration

Verdict ships an official Model Context Protocol server (verdict-mcp), so any LLM client — Claude Code, Claude Desktop, Cursor — can drive the engine as tools. Integrate, operate and explore in plain language, no glue code:

  • “Score a $4,200 card authorization for user usr_9 on a new device.”
  • “What's our false-positive rate this week?”
  • “Show me the ring around device dev_shared.”
  • “Backtest moving the review band to 35–69 for card.authorize.”

The server exposes tools for decide, recent_decisions, analytics_summary, lookup_entity, list_rules, model_weights and backtest_policy. Build it (npm install && npm run build in verdict-mcp/), then point your client at it:

{
  "mcpServers": {
    "verdict": {
      "command": "node",
      "args": ["/path/to/verdict-mcp/dist/index.js"],
      "env": {
        "VERDICT_API_URL": "http://localhost:4000",
        "VERDICT_API_KEY": "vk_live_…",   // for the decide tool
        "VERDICT_TOKEN":   "eyJ…"          // operator token, for reads/admin
      }
    }
  }
}

Credentials are scoped by what you pass: omit VERDICT_API_KEY for a read-only client, or omit VERDICT_TOKEN for a decisions-only one — a tool without its credential refuses rather than acting.

Prefer to wire it yourself with an LLM? Point your coding agent at /llms.txt — a self-contained integration brief (auth, the event schema, verdict handling, every endpoint, the signal namespace) written for code assistants. Drop it into your repo or hand it to your agent and it can scaffold the integration against the real contract; the full field-by-field API reference is there too.

Events & fields

Every channel is mapped into one normalized event, so the engine reasons about payments, logins and withdrawals the same way. Supported type values: card.authorize · card.capture · card.refund · payment.authorize · wallet.withdraw · account.login · order.place. Adding a channel is a field mapping, not a core change.

Send as much as you have — every field is a potential signal, and absent fields simply mean the rules that read them do not fire. Identifiers you send (userId, deviceId, ip) are what the entity graph links, and amount feeds each user's anomaly baseline.

Request fields
typeenum (required)Event type — selects the ruleset and policy.
amount / currencynumber / string(3)Value + ISO-4217 code for money-bearing events. Feeds amount and anomaly signals.
subject.userIdstring (required)The acting user — the primary graph node and anomaly key.
subject.deviceIdstringStable device/client id — velocity, first-seen and device-linking signals.
subject.fingerprintstringDevice fingerprint hash — reuse across users and fingerprint↔device mismatch (cloning/spoofing), independent of deviceId.
subject.ipstringClient IP — resolved to a location for geo signals, and linked in the graph (shared-IP / ring). Never placed in URLs or logs.
subject.phonestringMSISDN for mobile-money / telecom rails — a graph entity; many accounts on one phone flags a SIM box.
subject.channelstringOrigin rail, e.g. visa, telebirr, web.
instrument.kind"card" | "wallet" | "bank"Payment instrument type.
instrument.binstring(6–8)Card issuer BIN only — never a full PAN.
instrument.issuerCountrystring(2)ISO-3166 issuer country.
instrument.threeDSbooleanWhether the transaction carried a 3-D Secure result.
attributesobjectExtra validated scalars, e.g. { geoMismatch: true }. Readable in rules as attr.<key>.

Signals & rules

Rules read a flat namespace of signals derived from the event and the engine's own state. Some come straight off the event; others are computed at decision time by the feature store (velocity), the entity graph, and the anomaly baseline. Any of these can appear in a rule condition:

Signal namespaces
event.*type · channel · amount · currencyStraight from the event.
instrument.*kind · bin · issuerCountry · threeDSPayment instrument metadata.
velocity.*attemptsLast2m · attemptsLast24h · amountLast1hRolling counters from the feature store.
device.*firstSeen · usersOnDevice · usersOnFingerprint · fingerprintFirstSeen · fingerprintDeviceMismatchDevice & fingerprint recency, sharing, and cloning/spoofing.
geo.*country · distanceKm · countryChanged · impossibleTravel · ipSimMismatchIP-resolved location, distance from the user's last location, and travel faster than a jet.
graph.*ringSize · usersOnDevice · usersOnIp · usersOnPhone · devicesOnUserEntity-graph link counts and cluster size — incl. accounts sharing a phone (SIM box).
anomaly.*amountZScore · amountMean · samplesDeviation from this user's own history.
attr.*any scalar you sendYour custom attributes.

A rule proposes: when its condition matches it adds weight under a tag — it never decides on its own. Rules are authored in a small DSL, versioned in git, hot-reloaded, and diffable in review. That separation (rules propose, scoring weighs, policy decides) is what keeps every verdict attributable.

rule r_takeover {
  when  device.firstSeen == true && geo.ipSimMismatch == true
  then  score += 33, tag "takeover"
}

rule r_ring {
  when  graph.ringSize > 6           // shared devices/IPs cluster
  then  score += 26, tag "ring"
}

rule r_amount_anomaly {
  when  anomaly.amountZScore > 3     // 3σ from this user's own baseline
  then  score += 22, tag "amount_spike"
}

Verdicts & scoring

Scoring sums the weights of every matched rule, per tag, into a single 0–100 number. A policy then maps score bands to one of four verdicts. How you act on each is up to your integration:

allowLet it through.
challengeStep the user up (3-D Secure / OTP) before proceeding.
reviewHold and let an analyst decide — a case opens automatically.
denyBlock it.

Every verdict carries the score and the reasons (tag + points) that produced it — so a decision is always auditable, and a support or compliance question has a concrete answer. A watchlist or degraded short-circuit reports a score of -1 with an explanatory tag. Because decisions are written to an append-only verdict log, any of them can be replayed or backtested later.

Three scorers, one seam. The default is the transparent hand-weighted sum. Set SCORER=learned for the adaptive model that learns each tag's fraud rate from your labels, or SCORER=ml for a trained logistic-regression model that consumes the whole feature vector — velocity, device fingerprint, geolocation, graph and anomaly signals, plus the rules' own score — and returns a calibrated risk with the per-feature terms as the reasons. It is trained offline on a synthetic dataset (npm run train:model); inference is a single dot-product plus a sigmoid, so it adds nothing to decision latency. All three implement the same port, so switching is one environment variable — nothing else moves. The ML weights are hot-swappable and can load from a local file(MODEL_PATH), an HTTPS URL (MODEL_URL), or S3-compatible storage (MODEL_S3_* — AWS S3, MinIO, R2, Spaces), refreshed on a timer — validated against the feature vector, with the bundled weights as a safe fallback — so you roll out a retrained model without a redeploy. See the active model and its per-feature weights at GET /v1/config/model.

Scaling the ML pipeline. Serving and training scale independently by design. Serving is a fixed-cost dot-product held in the image — no model server, no network hop, no GPU — so it adds microseconds and scales horizontally with engine replicas, which each load and validate the same weights. The practical ceiling at very high volume is the feature reads (velocity, graph), not the model. Training is fully decoupled: the engine never trains at runtime and only consumes a validated JSON weights file, so the bundled gradient-descent trainer is just a reference — you can train on real labels at any scale in an external batch job (Python, a GBM, a feature store) and publish the weights to MODEL_S3_* on whatever cadence you retrain. Because labels arrive as events (analyst resolutions and chargebacks), the same loop that improves the adaptive scorer is the training set for the ML model.

Configuration

Rules, scoring and policies are configuration you own — versioned like code, editable live in the dashboard. Rules add weighted tags (previous section). Scoring sums them. A policy maps score bands to verdicts, with an on_error mode (fail-open / fail-closed) that decides the outcome when a dependency is degraded — on purpose, never silently.

policy card_authorize {
  bands {
    allow      0..24
    challenge  25..44   → step-up (3DS)
    review     45..69   → queue "risk-ops"
    deny       70..100
  }
  on_error  fail_open   // degrade to allow, never silently block
}

Admins edit a policy's bands and publish a new version; every version is retained, so you can roll back in one click and simulate a score before shipping. Bands must cover 0–100 with no gaps or overlaps — the engine validates this on publish and refuses an invalid policy. All of this is live in the dashboard's Configure tab, or over the config API.

The operator dashboard

Everything below is driven from a self-hosted operator console (a separate Next.js app that talks to the engine's API). Analysts work the review queue and resolve cases; admins configure rules, policies, scoring, retention, keys and webhooks — all server-side, nothing the client can bypass.

Review queue — cases the engine sent to review, ready to triage, assign and resolve.
Case detail — risk score, the event/verdict ids, and the resolution + audit trail.
Configure → Scoring model — the active scorer, model provenance and per-feature weights.
Configure → Storage — document-store size plus server-disk health per mount.
Analytics — decision and outcome rollups, projected from the verdict log.

Operating

A review verdict opens a case in the dashboard queue. An analyst assigns it to themselves, works it, and resolves it as fraud, legit, or inconclusive. Every action is recorded on an append-only audit trail, and the analyst is always taken from the authenticated token, never the request body.

Configuration changes are audited too. Every operator mutation — a rate-limit or retention tweak, a published policy, a new alert channel or API key — is written to a config-change audit log (GET /v1/audit) with the actor, the action, the time and the result; secret-ish fields (URLs, tokens) are redacted before storage. So “who changed this, and when?” always has an answer.

Resolving a case records a label — ground truth about that event. Chargebacks post the same kind of label from your PSP via POST /v1/labels/chargeback, which crucially covers the frauds you allowed and never sent to review. Both feed the same learning loop.

The Analytics tab shows verdict mix, score distribution, top signals, decisions per day, label counts (including chargebacks), and the false-positive rate — a read-model maintained by a projector that consumes events, so it never touches the decision path.

The Activity tab is the API interaction log: every scored call, newest first, pairing the exact request payload with the verdict it produced (GET /v1/activity). It's your audit trail for “what did we send and what did we get back” — with phone numbers and IPs masked and card data BIN-only, so the log itself is safe to keep.

Webhooks push those same events to your systems in real time. Register an endpoint in the dashboard (admin → Webhooks) for verdict.reached, case.resolved or label.recorded; each delivery is a signed POST — verify the X-Verdict-Signature header (HMAC-SHA256 of the body) before trusting it. Deliveries are durable: retried with backoff and dead-lettered after repeated failure, off the decision path — inspect and re-drive them under Webhooks.

Alert notifications send those events to a person, not a system. Add a Slack incoming-webhook, a Telegram bot (URL https://api.telegram.org/bot<token>/sendMessage plus a chat id), or a generic HTTPS channel (POST /v1/notifications), subscribe it to the events you care about, and — for verdict.reached — set a minimum severity so you only get pinged on, say, deny. Built-in alert.anomaly (a decision whose amount is far from the user's own baseline — tune the z-score threshold under Configure) and alert.dead_letter (a delivery that finally gave up) events mean unusual activity and operational failures reach you too. Delivery is durable — each alert is queued, retried with backoff and dead-lettered if it never lands (inspect and re-drive via GET/POST /v1/notifications/deliveries), all off the request path — and throttled per channel (set throttlePerMin) so an alert storm can't bury an operator or rate-limit your Slack.

Intelligence

Four capabilities turn that feedback into better decisions. All are optional and observable.

Adaptive scoring. Set SCORER=learned and the scorer weights each signal by how often it has ridden a fraud label in your own resolved cases and chargebacks — a tag keeps its hand weight until it has enough labels to trust, so a fresh deploy behaves exactly like the weighted scorer and drifts toward your data. Inspect the learned weights any time at GET /v1/model (or the Analytics page), even while the weighted scorer is live.

Rule & policy backtesting. Before you ship a change, replay it against your labeled history and see how it would have performed — precision, recall, and the deltas in fraud caught, false positives, and verdict flips versus the live policy. It is read-only; nothing is published.

# would this policy change catch more fraud without more false positives?
curl -X POST http://localhost:4000/v1/backtest \
  -H 'authorization: Bearer <admin-token>' \
  -H 'content-type: application/json' \
  -d '{ "eventType": "card.authorize",
        "policy": { "bands": [ … candidate bands … ] } }'

# → { "labeled": 312,
#     "candidate": { "precision": 0.83, "recall": 0.9, … },
#     "delta": { "fraudCaught": +20, "falsePositives": -15, "flips": 48 } }

Entity graph. Users, devices and IPs are linked as they co-occur on events. Rules can read graph.ringSize and the link counts to flag rings a single event can't reveal, and analysts can explore any entity's cluster at GET /v1/graph/:kind/:id or the dashboard's Graph tab.

Anomaly detection. The engine keeps a per-user spending baseline and exposes anomaly.amountZScore — how many standard deviations a charge sits from that user's own history — so a rule can flag "this doesn't look like this account" without a fixed threshold.

Data governance

Verdict stores the identifiers you send and the decisions it makes about them. It expects a card bin only and never a full PAN. Personal data lives in four places: the entity graph, the per-user anomaly baseline, the replay log (which keeps whole events for backtesting), and the activity log (which keeps the masked request per decision).

Right to erasure. An admin can remove a user across all four with one call. The append-only verdict log is intentionally out of scope — it holds only an event id, verdict and tags — and should be aged out with a retention window instead.

# right-to-erasure: wipe a user's graph identity, baseline and replay samples
curl -X POST http://localhost:4000/v1/privacy/erase \
  -H 'authorization: Bearer <admin-token>' \
  -H 'content-type: application/json' \
  -d '{ "userId": "usr_3f9a" }'
# → { "userId": "usr_3f9a", "replaySamplesRemoved": 12, "activityEntriesRemoved": 12,
#     "graph": true, "baseline": true }

Postgres is the system of record; back it up on your normal schedule. The dashboard holds no data of its own.

Deployment

Bring up Postgres, the engine, and the dashboard with one command:

# from verdict-engine/
AUTH_SECRET=$(openssl rand -hex 32) docker compose up --build

There is no fallback for AUTH_SECRET — the stack refuses to start until you supply a real one (placeholders and repeated-character filler are rejected too). Generate it once and keep it stable across restarts, or existing operator tokens stop verifying.

What runs where
engine:4000 · privateThe API and the interactive Swagger at /docs. Speaks plain HTTP — keep the port private and put a TLS proxy in front. This is the only service that must reach Postgres and Redis.
dashboard:3000 · internalThe operator UI (review queue, config, analytics, keys). Reaches the engine over VERDICT_API_URL; expose it only to your team.
postgres / redisnot publishedBacking services on the compose network — no host ports. Postgres is the system of record; Redis is optional (shared idempotency/velocity/rate-limits).

Docs vs. your deployment. This documentation site and its Swagger preview are the public project site — they are not part of the stack you run. Your own live, try-it-out API docs are served by the engine itself at /docs on port 4000, behind your proxy. Nothing you deploy runs on port 3001.

Environment
AUTH_SECRETrequired in prodSigns operator tokens. Generate with: openssl rand -hex 32
DATABASE_URLrequired in prodpostgres:// connection string. Durable users, cases, verdicts, keys, graph.
DATABASE_POOL_MAXdefault 20Max Postgres connections per engine replica. Raise it if decisions queue under load.
REDIS_URLmulti-replicaredis:// connection. Shares idempotency, velocity and rate limits across replicas — required to run more than one engine replica.
PORTdefault 4000Engine HTTP port.
VERDICT_API_URLdashboardWhere the dashboard reaches the engine (e.g. http://engine:4000).
PERSISTENCEoptionalSet to `memory` to explicitly accept a non-persistent deploy (demos only).
SCORERoptional`weighted` (default), `learned` (adaptive), or `ml` (trained model). Set MODEL_URL to hot-load ML weights from a cloud bucket.
RATE_LIMIT_DECISIONS_PER_MINdefault 600Per-API-key budget for POST /v1/decisions before a 429.
RATE_LIMIT_LOGIN_PER_MINdefault 10Per-IP budget for login — a brute-force brake.
TRUST_PROXYdefault offProxy hops to trust for the real client IP (per-IP limits). Default ignores X-Forwarded-For so it can't be spoofed; set 1 behind a single LB.
AUTH_SECRET_PREVIOUSrotationPrevious signing secret(s), comma-separated, kept valid while you rotate AUTH_SECRET — see the security section.

The engine validates its environment at startup and refuses to boot if it is misconfigured — so a bad deploy fails loudly instead of coming up insecure or without a database:

✗ verdict-engine cannot start — fix these environment variables:
    • AUTH_SECRET is required in production — it signs operator tokens.
    • DATABASE_URL is required in production for durable persistence.

Persistence is Postgres, durable across restarts; the in-memory store is for dev and tests only. Put a TLS-terminating reverse proxy in front and keep the engine port private — it speaks plain HTTP. Health for probes: GET /health. The engine emits one structured JSON log line per operational event (a degraded decision logs engine.degraded with a scrubbed cause) — ship these to your log stack. Metrics are exposed at GET /metrics in Prometheus format (decision latency and outcomes, 429s, outbox and webhook delivery counters and queue depth, retention rows pruned) — keep that endpoint on the private port.

Running at scale? Set REDIS_URL and run several engine replicas behind a load balancer — idempotency, velocity and rate limits are shared in Redis, so a retry or a burst is enforced consistently across nodes. The transactional outbox is durable (Postgres) and delivers at-least-once with backoff and a dead-letter, so an event survives a crash. The full guide — scaling limits, sizing, key rotation and a security checklist — is in DEPLOYMENT.md.

Requirements & capacity

Verdict is light to run. The engine is an IO-bound Node process and the dashboard is stateless — Postgres is the component you size for growth. A small setup comfortably handles a startup's traffic; the numbers below are a starting point, not a ceiling.

Recommended servers (to start)
Engine (per replica)1–2 vCPU · 1 GBStateless behind Postgres; mostly waiting on the database. Scale out once the Redis adapters land (see below).
Dashboard0.5 vCPU · 512 MBStateless Next.js. One instance is plenty for an ops team.
PostgreSQL 162 vCPU · 4 GB · SSDThe system of record. Give it fast disk, backups, and room to grow with your history.

Measured throughput. A single engine replica sustained ~300 decisions/second in our benchmark — and that was a dev laptop running Postgres over Docker Desktop, a deliberately pessimistic setup for database round-trip latency. Latency was ~20 ms for a lone request, and p50 ~70 ms · p99 ~140 ms under moderate concurrency, with zero errors up to 64 requests in flight. On a tuned Linux host with Postgres on a low-latency network, expect materially higher throughput and lower latency.

What sets the ceiling. Each decision makes several database round-trips — velocity, the entity graph, the anomaly baseline, the policy, then the log and outbox writes — so per-node throughput is bound by database latency, not CPU. The two levers are (1) keeping Postgres close and sizing the pool (DATABASE_POOL_MAX, default 20), and (2) setting REDIS_URL so the hot velocity reads and idempotency go to Redis — which also lets you run several engine replicas behind a load balancer, sharing one budget. Without it, those are per-replica, so run a single one. Postgres and the dashboard scale out either way.

Sizing storage. Budget roughly three rows per decision — the append-only verdict log, the replay sample (which holds the whole event, for backtesting), and a scoring sample — plus a graph node per new entity. That is on the order of a few GB per million decisions.

Retention & pruning. A built-in job keeps those append-only collections from growing without limit: on a timer it deletes rows older than each collection's window — verdict log, activity log, replay samples, idempotency keys and dead-lettered events. Defaults are 365 / 90 / 90 / 7 / 30 days; a window of 0 keeps a collection forever. Tune them per deployment via RETENTION_* env vars, or at runtime on the dashboard's Configure tab (GET/PUT /v1/config/retention); POST /v1/config/retention/run forces a sweep. Counts are exported as verdict_retention_pruned_total. For disk management, GET /v1/config/storage reports the store's size on disk and per-collection row counts (also sampled into the verdict_storage_* metrics each sweep) — so you can see what's growing; set RETENTION_VACUUM=true to reclaim freed space promptly.

Server disk health. For Docker deployments where data sits on a volume, the same GET /v1/config/storage (and the dashboard's Storage panel) also reports the filesystem health of the disks that data lives on — total, free and used, per mount — plus a per-service breakdown of what's consuming the disk (the document store, and a disk-backed model file). Point DISK_HEALTH_PATHS at the volume mounts you want watched; the engine can only see filesystems mounted into its own container, so to watch the Postgres volume from here, bind-mount it (read-only) into the engine — otherwise monitor the database container's volume with node_exporter or cAdvisor. The values are exported as verdict_disk_total_bytes, verdict_disk_free_bytes, verdict_disk_used_ratio and verdict_disk_component_bytes for alerting.

Security

  • API keys are random (vk_live_…), SHA-256-hashed at rest, shown once, and revocable instantly. They can be scoped (decisions / labels) and given an expiry; a request outside a key's scope is rejected with 403.
  • Operator auth uses scrypt-hashed passwords and signed bearer tokens (12-hour expiry) carrying issuer, audience, a key id and a unique token id, with roles (admin / analyst). The first user bootstraps as admin; public registration then closes.
  • Session revocation: POST /v1/auth/logout revokes the current token immediately (before it expires), /logout-all ends every session for the user, and /refresh slides a session — so a leaked or stale token can be cut off, not just waited out.
  • Key rotation: each token records the key that signed it (kid), so AUTH_SECRET rotates with zero forced logouts — set the old secret as AUTH_SECRET_PREVIOUS until its tokens age out. Placeholder/weak secrets are rejected at boot.
  • Rate limiting guards the decision endpoint (per key) and login (per IP), returning 429 + Retry-After. Per-IP limits use the proxy-aware client IP (TRUST_PROXY), never a spoofable header.
  • Data: BIN only, never a full PAN; the verdict log is append-only; phone/IP are masked in the activity log; personal data is erasable per user (see Data governance).
  • Ownership is enforced server-side — a case action uses the analyst from the token, never an id from the body.
  • Webhooks are HMAC-signed (X-Verdict-Signature); deliveries are durable with retry, backoff and a dead-letter you can inspect and re-drive.
  • Payment callbacks and chargebacks are machine endpoints — key them, and re-verify amount and status server-side before fulfilling.
  • Supply chain: CI runs npm audit (high+ severity) alongside type-check and tests on every change.