Documentation
The outcome loop
How verify, feedback, and the weekly Agent Farming Ledger close the loop between decisions and outcomes.
Same loop as the Introduction: verify returns a decision, you take an action, feedback reports what happened and informs future decisions. Free-credit farming and the Agent Farming Ledger are applications of that loop.
Keep the paths separate. Request path: verify → allow / challenge / deny. Outcome path: account + eventId → outcome → feedback. The detector does not wait on whether the signup burned credits or converted; labels arrive when the business fact is known, which is also why the economics stay explainable.
Every action gets a verdict, every verdict gets a record, and every record eventually gets an outcome: credit burn, conversion, or chargeback. Those outcomes are labeled feedback for three uses:
| Use | What happens |
|---|---|
| Attribution | Join the outcome to the original verify on the stored eventId only |
| Evaluation | Agent Farming Ledger, weekly exports, drift monitor, offline backtests |
| Scoring calibration | Weekly human-in-the-loop review: one documented threshold or decisionCosts change (not autonomous model training in the pilot) |
| Step | What happens | Where |
|---|---|---|
| 1 · Decide | verify returns allow / challenge / deny and an eventId | Your edge |
| 2 · Record | You persist eventId on the account row | Your database |
| 3 · Label | When the outcome lands, feedback joins it on that eventId | Your queue to Chitmark |
| 4 · Calibrate | Weekly review uses labels to adjust allow/challenge/deny boundaries via decisionCosts | Joint review |
| 5 · Report | The weekly Agent Farming Ledger shows what was blocked, what slipped, and what it cost | Chitmark |
The mechanics
Security invariant: feedback cannot create attribution. It can only attach an outcome to an existing verify eventId for your tenant. Never invent a join from email, IP, or other subject fields.
verify returns an eventId with every decision. Persist it on the account row. When the outcome lands, feedback joins only on that stored id. The outcome-label connectors guide covers the patterns: read the id off the account row, and if there is no stored id, skip and fix the signup path.
Append-only observations (not latest-wins): one eventId may carry several labels. Exact duplicates collapse to one feedbackId. Compatible labels coexist. Conflicting pairs (e.g. converted then abuse_confirmed) quarantine the new row for audit and exclude it from effective ledger totals without deleting prior evidence. HTTP { ok: true } still acknowledges quarantined ingest.
The edge enforces ownership on every label: an event id that does not belong to your tenant returns 404 unknown_event_id (the invariant working, not a soft miss). Duplicate deliveries are safe when bodies match exactly on eventId, outcome, value, unit, and observedAt (same warehouse feedbackId, no double-counted value), or when you send an Idempotency-Key. If the ledger is temporarily unavailable, feedback returns 503 and you retry with backoff: the loop degrades loudly, never silently.
verify
│
└── eventId ───────────────┐
↓
account / request ───────→ feedback
│
↓
outcome label(s)
evt_123
├─ converted (accepted)
└─ abuse_confirmed (quarantined if it conflicts)From outcomes to the next decision
In the pilot, this is a weekly human-in-the-loop review, not autonomous learning: each week, jointly inspect a small set of challenged and allowed cohorts, compare 72-hour burn and conversion, and make one documented threshold or rule change via context.decisionCosts. That is scoring calibration from labeled feedback, not online model training on the verify path. Labels also feed evaluation (ledger, drift, backtests) after attribution on eventId. A tenant whose traffic converts keeps its boundary; a tenant whose traffic burns credits tightens it. Delayed labels are handled explicitly: chargebacks arrive weeks after the decision, so the loop treats them as late-arriving evidence, and the drift monitor checks whether late labels change what enforcement should have done.
Verify is a three-way decision, not a single risk threshold. Uncertainty sits in the middle; decisionCosts moves both edges. confidence on the Verdict is separate: a model confidence measure for the returned decision, not a calibrated probability of abuse and not the threshold those weights move.
The boundaries are a documented cost trade-off, not a hidden model. context.decisionCosts takes relative, dimensionless weights (not dollars): falseAllow = abusive traffic allowed, falseChallenge = legitimate traffic challenged when allow would have been fine, falseDeny = legitimate traffic denied. Defaults weight those 1:1:3. Until a real partner result supports a self-improving claim, describe this as calibration.
uncertainty
↓
allow ←────────── challenge ──────────→ deny
↑ ↑
allow/challenge challenge/deny
boundary boundary
falseAllow ↑ → challenge earlier (abuse-averse)
falseChallenge ↑ → let more of the gray zone through
falseDeny ↑ → reserve deny for confirmed abuse
Not: risk < t → allow, else block
Yes: price the three error classes; both boundaries moveThe weekly Agent Farming Ledger
The weekly export is the readout for that review: it joins each decision back to the outcome labels that arrived for it, straight from the event ledger:
| Export | What it answers |
|---|---|
| Weekly | What was blocked, what slipped, and what it cost, in dollars (from the stored verdict at decide time, joined to later outcomes) |
| Outcome estimates | What the unlabeled traffic probably became, before the labels arrive |
| Drift monitor | Whether late labels (chargebacks) change how past stored verdicts look in the ledger |
The weekly export is the ledger of record for the loop: defense you can show a board, not a score you have to take on faith. Bring challenge completion and downstream conversion by challenge type to the same table.
Isolation and the network effect
Today the loop is per-tenant by design: your outcomes tune your decisions, and your event data is never mixed with another tenant's. Ownership is enforced on every label, and everything a tenant writes is scoped to that tenant.
The ledger is built as a neutral record, and that is the property a cross-customer intelligence layer would build on: as labeled outcomes accumulate across tenants, patterns in economically abusive agent behavior become observable in aggregate. That layer is the long-horizon direction, not a feature today: nothing about your traffic is shared with other tenants.