Documentation

The outcome loop

How verify, feedback, and the weekly Agent Farming Ledger close the loop between decisions and outcomes.

Same loop as the Introduction: verify returns a decision, you take an action, feedback reports what happened and informs future decisions. Free-credit farming and the Agent Farming Ledger are applications of that loop.

Keep the paths separate. Request path: verify → allow / challenge / deny. Outcome path: account + eventId → outcome → feedback. The detector does not wait on whether the signup burned credits or converted; labels arrive when the business fact is known, which is also why the economics stay explainable.

Every action gets a verdict, every verdict gets a record, and every record eventually gets an outcome: credit burn, conversion, or chargeback. Those outcomes are labeled feedback for three uses:

UseWhat happens
AttributionJoin the outcome to the original verify on the stored eventId only
EvaluationAgent Farming Ledger, weekly exports, drift monitor, offline backtests
Scoring calibrationWeekly human-in-the-loop review: one documented threshold or decisionCosts change (not autonomous model training in the pilot)
StepWhat happensWhere
1 · Decideverify returns allow / challenge / deny and an eventIdYour edge
2 · RecordYou persist eventId on the account rowYour database
3 · LabelWhen the outcome lands, feedback joins it on that eventIdYour queue to Chitmark
4 · CalibrateWeekly review uses labels to adjust allow/challenge/deny boundaries via decisionCostsJoint review
5 · ReportThe weekly Agent Farming Ledger shows what was blocked, what slipped, and what it costChitmark

The mechanics

Security invariant: feedback cannot create attribution. It can only attach an outcome to an existing verify eventId for your tenant. Never invent a join from email, IP, or other subject fields.

verify returns an eventId with every decision. Persist it on the account row. When the outcome lands, feedback joins only on that stored id. The outcome-label connectors guide covers the patterns: read the id off the account row, and if there is no stored id, skip and fix the signup path.

Append-only observations (not latest-wins): one eventId may carry several labels. Exact duplicates collapse to one feedbackId. Compatible labels coexist. Conflicting pairs (e.g. converted then abuse_confirmed) quarantine the new row for audit and exclude it from effective ledger totals without deleting prior evidence. HTTP { ok: true } still acknowledges quarantined ingest.

The edge enforces ownership on every label: an event id that does not belong to your tenant returns 404 unknown_event_id (the invariant working, not a soft miss). Duplicate deliveries are safe when bodies match exactly on eventId, outcome, value, unit, and observedAt (same warehouse feedbackId, no double-counted value), or when you send an Idempotency-Key. If the ledger is temporarily unavailable, feedback returns 503 and you retry with backoff: the loop degrades loudly, never silently.

verify
   │
   └── eventId ───────────────┐
                              ↓
account / request ───────→ feedback
                              │
                              ↓
                         outcome label(s)

evt_123
 ├─ converted          (accepted)
 └─ abuse_confirmed    (quarantined if it conflicts)

From outcomes to the next decision

In the pilot, this is a weekly human-in-the-loop review, not autonomous learning: each week, jointly inspect a small set of challenged and allowed cohorts, compare 72-hour burn and conversion, and make one documented threshold or rule change via context.decisionCosts. That is scoring calibration from labeled feedback, not online model training on the verify path. Labels also feed evaluation (ledger, drift, backtests) after attribution on eventId. A tenant whose traffic converts keeps its boundary; a tenant whose traffic burns credits tightens it. Delayed labels are handled explicitly: chargebacks arrive weeks after the decision, so the loop treats them as late-arriving evidence, and the drift monitor checks whether late labels change what enforcement should have done.

Verify is a three-way decision, not a single risk threshold. Uncertainty sits in the middle; decisionCosts moves both edges. confidence on the Verdict is separate: a model confidence measure for the returned decision, not a calibrated probability of abuse and not the threshold those weights move.

The boundaries are a documented cost trade-off, not a hidden model. context.decisionCosts takes relative, dimensionless weights (not dollars): falseAllow = abusive traffic allowed, falseChallenge = legitimate traffic challenged when allow would have been fine, falseDeny = legitimate traffic denied. Defaults weight those 1:1:3. Until a real partner result supports a self-improving claim, describe this as calibration.

                    uncertainty
                         ↓
allow  ←──────────  challenge  ──────────→  deny
         ↑                         ↑
   allow/challenge           challenge/deny
      boundary                  boundary

falseAllow ↑       → challenge earlier (abuse-averse)
falseChallenge ↑   → let more of the gray zone through
falseDeny ↑        → reserve deny for confirmed abuse

Not:  risk < t → allow, else block
Yes:  price the three error classes; both boundaries move

The weekly Agent Farming Ledger

The weekly export is the readout for that review: it joins each decision back to the outcome labels that arrived for it, straight from the event ledger:

ExportWhat it answers
WeeklyWhat was blocked, what slipped, and what it cost, in dollars (from the stored verdict at decide time, joined to later outcomes)
Outcome estimatesWhat the unlabeled traffic probably became, before the labels arrive
Drift monitorWhether late labels (chargebacks) change how past stored verdicts look in the ledger

The weekly export is the ledger of record for the loop: defense you can show a board, not a score you have to take on faith. Bring challenge completion and downstream conversion by challenge type to the same table.

Isolation and the network effect

Today the loop is per-tenant by design: your outcomes tune your decisions, and your event data is never mixed with another tenant's. Ownership is enforced on every label, and everything a tenant writes is scoped to that tenant.

The ledger is built as a neutral record, and that is the property a cross-customer intelligence layer would build on: as labeled outcomes accumulate across tenants, patterns in economically abusive agent behavior become observable in aggregate. That layer is the long-horizon direction, not a feature today: nothing about your traffic is shared with other tenants.