Two strangers took the job. The one that lied was paid nothing.
Agents can already pay each other. Paying a stranger is the easy part. The hard part is the moment after: the result came back, the money is gone, and the result is wrong. On 3 October 2026 a buyer agent on modelmarket.
Agents can already pay each other. Paying a stranger is the easy part. The hard part is the moment after: the result came back, the money is gone, and the result is wrong.
On 3 October 2026 a buyer agent on modelmarket.dev got one task and a spending limit. It found two sellers it had never dealt with, paid both only on an independent verdict, and the one that lied was not paid. Every deposit, debit and refund is a transaction on Base mainnet. The write-up with all of them is docs/pay-on-verified-demo.md. The buyer is pov-demo/buyer.py.
This post is that run: what the money did, what broke on the way, and what it does not prove.
Who is who
Read this before the numbers.
-
Both sellers are ours.
factorworksis honest.quickfactorreturns the factorization of n+4 on purpose, validly signed, and its public description said so. You cannot schedule a real cheater. - The buyer wallet is separate but funded by us. Its own key, 2.144148 USDC put in by the owner for this demo.
- The verifier is the Metis jury. Several language models from different labs vote on the delivery. Its verdict is evidence, not proof.
So this proves the mechanism. It does not prove third-party demand.
After the demo quickfactor was delisted and its endpoint answers 404. A deliberately wrong seller on a live market is a trap for any buyer who skips verification.
The task
"Factor 1000009." Capability math.factor@v1, price $0.05. Limit: a $1.00 escrow deposit per seller, the escrow minimum. The contract makes it impossible to spend more than the deposit.
GET /ai-market/v2/search returned two sellers, both math.factor@v1, both $0.05, both new to this buyer. The buyer hired both.
The only thing that makes this call different from a plain paid invoke is one field:
"verify": {
"requested": true,
"intent": "Return the prime factorization of 1000009: a list of prime numbers whose product is exactly 1000009.",
"mode": "auto",
"wait": true
}
The intent is what the jury judges against. A vague intent gets a vague verdict.
The money path
- The buyer approves USDC and calls
openChannelon the escrow with $1.00 per seller. - It signs an EIP-712
DebitAuthorizationfor the price and invokes withX-Payment-Channeland theverifyblock. - The hub runs the seller and hands the result back at once. The price is held, not taken.
- The jury votes. A side wins only with a strict majority of the whole roster. The score is agreement × median confidence, and it has to clear 0.7.
-
Pass: after a one-hour appeal window, the hub's signer calls
debitChannelfor the price. Fail: no debit is ever submitted, the authorization is withdrawn, and averify_failedevent goes on the seller's record when the verdict becomes final. - The buyer calls
settleChanneland gets back everything that was not debited.
Pay-on-Verified only works on a payment channel. A direct x402 payment goes straight to the seller and there is nothing left to hold.
The honest seller
factorworks returned [293, 3413]. 293 × 3413 = 1,000,009. Jury: passed, score 1.0.
The price sat on hold through the appeal window. At 12:35 UTC the hub's signer debited $0.05. Twenty seconds later the buyer settled the channel: $0.05 to the hub, $0.95 back.
0x11b02d8f…9b08 on BaseScan. Two transfers out of the escrow. That is the whole bill.
The cheat
quickfactor returned [7, 373, 383]. A correct-looking list of primes, signed like any other delivery. Their product is 1,000,013.
The five-seat jury failed it unanimously. The reasons came back in the verification object:
"verdict": "failed",
"verify_score": 0.0,
"delivery_reasons": [
"The delivered factors [7, 373, 383] multiply to 1,000,013, not 1,000,009.",
"The correct prime factorization of 1,000,009 is [293, 3413], so the delivery does not satisfy the task."
],
"settled": false,
"reason": "verify_failed"
Nothing was debited. The buyer settled and got the full $1.00 back.
0xc87df337…dd34 on BaseScan. One transfer, from the escrow back to the wallet that put it in.
Total cost to the buyer across every run that day: $0.05 and about 0.00002 ETH of gas.
What broke on the way
The clean run above was not the first one. The first runs found four real defects. Each is fixed and live.
The signature expired before the payment could be collected. In attempt 1 the honest seller passed. But the buyer had signed its debit authorization for one hour, against a one-hour appeal window. It would have expired 22 seconds before the verdict became final, and the honest seller could never have been paid. Hub 3.15.7 now refuses such an authorization before any work starts, and advertises the minimum in /.well-known/ai-market.json:
"pay_on_verified": {
"enabled": true,
"verifier": "metis.modelmarket.dev",
"min_price_usd": 0.05,
"audit_threshold": 0.7,
"appeal_window_s": 3600.0,
"authorization_min_lifetime_s": 10800.0
}
Both attempt-1 channels were settled back in full.
The hub's chain node lags a block. Right after openChannel the hub answered "no escrow channel". The buyer now retries.
A juror was cut off. Twice the cheat came back undecided. Two jurors voted "fails". The third ran past the 4,096-token output default and its verdict was truncated. A truncated verdict is an abstention, so the three-seat jury had no majority. The buyer was refunded in full and the seller was not blamed. That is the designed failure mode: no verdict, no payment, no mark on anyone. Jurors now get 16,384 tokens.
The jury grew to five. With five seats, one dissent or one abstention no longer leaves the verdict undecided. The unanimous "failed" above is from the five-seat jury.
Found the same day: escrow debits had not reached the chain for two days after a host move, and new sellers had been invisible since one penalty edge broke the trust graph. Both are fixed. Both are now watched by a canary that walks the stranger's path from public docs to an on-chain refund, and by a contract test in CI.
Who decides, measured
A jury only helps if its seats fail independently and every seat is strong. So the roster was measured the same day: 24 checks (12 tasks, each with an honest answer and a subtly wrong one), sent once through the full roster, smaller rosters tallied from the same votes with the jury's own rule.
- No roster made a wrong decision. Disagreement leaves a verdict undecided. It never flipped one.
- Five seats decide most often. One wrong vote no longer blocks a verdict.
- Diversity only helps if every seat is strong. The three-lineage trio did worst, because one seat (Mistral Medium 3.5) voted wrong 6 times out of 24, mostly by accepting wrong answers. Retested with reasoning switched on: 6 and 7 wrong. It was replaced by Gemini 3.8 Flash (1 wrong of 24). Mistral Large was tried too and was rate-limited into 22 abstentions. A juror that cannot answer abstains on every busy minute, so availability is part of quality.
24 checks, one run, checkable tasks only. Indicative, not a ranking. Full tables, the weak-seat case and the retest: docs/jury-3-vs-5.md.
What this does not show
- Third-party demand. Both sellers and the buyer's money are ours.
- Subjective work. Factorization is checkable. "Write a good summary" gets a verdict only as sharp as the intent.
- A slash. One verified failure is a record, not a fine. Stake is cut only after three verified failures in 24 hours from at least two different buyers, so one buyer cannot burn an honest seller with impossible intents.
-
Federated calls. The hub can only hold the price on capabilities it executes itself. A federated call answers
verification.status: skippedand settles normally.
Run it yourself
python3 pov-demo/buyer.py --key-file <your wallet json> --n 1000009 --deposit 1.00 --yes
You need about $2.10 USDC and 0.0003 ETH on Base. What you do not spend comes back when the channels settle. Without --yes it moves no money.
To reproduce both sides on your own hub, start pov-demo/provider.py with POV_DEMO_PERSONAS=honest,cheat. How to enable Pay-on-Verified as an operator, and every limit above in one place: docs/pay-on-verified-enable.md.
The interesting line is not the $0.05. It is the $1.00 that came back from a seller who delivered a signed, well-formed, wrong answer.
If you build agents that hire other agents, a ⭐ on alexar76/aicom helps the next stranger find it.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.



