Dev.to AI 🤖 Ai 👁 0 📖 4 min read

We Put Qwen 3.8 Max in Charge of Three Onchain Trading Agents. Here's What Happened.

How Alibaba's flagship reasoning model became the brain of Contex Arena — an AI trading arena live on Monad testnet. The arena Contex Arena is simple to explain and hard to run: three AI trading agents — Dege

How Alibaba's flagship reasoning model became the brain of Contex Arena — an AI trading arena live on Monad testnet.

The arena

Contex Arena is simple to explain and hard to run: three AI trading agents — Degen Dan (aggressive momentum), The Professor (disciplined mean-reversion), Whale (patient giant) — trade a synthetic asset against each other, live and onchain, on Monad testnet. Every trade is a real transaction. Spectators bet MON on which agent finishes the round with the highest portfolio value; winners split the pool, parimutuel style.

The agents are fully autonomous. Each tick, an agent reads the market state onchain (price, its own positions, its rivals' portfolios, time left in the round), thinks, and outputs a single machine-parseable decision:

{"action": "BUY", "amount": 12.5, "reason": "Price 4% below average, momentum turning up."}

That decision becomes an agentBuy / agentSell transaction seconds later.

The brain is the whole game. A dumb brain makes random trades; a flaky brain breaks the JSON contract and the round stalls; a slow brain misses the round cadence. We'd been running on a failover chain of fast models — Nebius's Nemotron Lightning first, then Atria, then NVIDIA NIM. It worked. But we wanted to see what a real reasoning model would do with the same job.

Why Qwen 3.8 Max

The Metropolis hackathon's Alibaba Cloud bounty asks builders to push Qwen 3.8 Max into "genuinely agentic territory" — not chatbots answering prompts, but agents planning, using tools, and executing multi-step work. An autonomous trader that reads onchain state, reasons about it, and fires transactions is about as agentic as it gets: every decision has financial consequences, recorded forever on a public ledger.

Qwen 3.8 Max is Alibaba's flagship reasoning model, and it speaks the OpenAI-compatible chat completions dialect. That made the integration almost boring — in the best way. Our agent runner already talks to every provider through raw POST /chat/completions behind a provider failover chain, so adding Qwen took about fifteen lines of code: a new entry at the head of the chain, model qwen3.8-max, pointed at the QwenCloud international endpoint (maas.qwencloudapi.com).

The one gotcha: thinking eats tokens

There was exactly one surprise, and it's the kind you only learn from a reasoning model in production. Qwen 3.8 Max does its reasoning in a thinking trace, and on some endpoints thinking can't be disabled — the trace burns output tokens against your max_tokens budget. Our old budget was 300 tokens, tuned for a fast non-reasoning model that returns bare JSON. Hand 300 tokens to a reasoning model and the thinking trace eats the budget before the JSON is complete. We'd fought (and fixed) the same failure mode with Nemotron Lightning before, so we knew the shape of the fix.

The fix: give Qwen a 4,000-token budget and let the parser do its job. Our parser doesn't trust the model to be tidy — it strips <think> traces, extracts every balanced {...} candidate from the response, and takes the first one that parses and validates. If the model answers in prose instead of JSON, we nudge it once with its own bad answer as context before failing over to the next provider. Defense in depth, because onchain agents don't get a human in the loop.

Live results

We pointed one agent at Qwen first, watched the logs, then promoted it to primary for all three fighters. After a 30-minute live session with all three agents running on Qwen — 4 full rounds settled onchain — here are the numbers from the QwenCloud pay-as-you-go console:

  • 66 requests, 66 successful — 100% success rate, zero JSON format retries. Not one response needed the retry nudge; every decision parsed clean on the first attempt.
  • 64 of 66 decisions answered directly by Qwen. Twice, a request exceeded our 45-second client timeout and the chain failed over to Nebius automatically — the arena never stalled, no round was missed. That's the failover doing exactly what it was built for.
  • 14.3s average latency — inside our 45-second per-decision timeout and the round cadence. (The two timed-out requests pull the average up; the typical decision is faster.) A reasoning model is slower than a non-reasoning model; the architecture absorbs it.
  • 55.7K tokens total — roughly 844 tokens per full decision cycle. Reasoning doesn't have to be expensive.
  • Personality-faithful decisions, measurably. Same prompts, three brains, three behaviors over the session: Degen Dan went 13 buys / 5 sells / 3 holds ("ape half the stack into the breakout"); the Professor went 12 buys / 1 sell / 10 holds, repeatedly declining trades with "not the 3% dip required for a buy"; Whale went 5 buys / 17 holds ("2.6% move is noise below the 8% dislocation threshold; no strike warranted"). The reasoning shows its work, and the work matches the character.

What Qwen brought to the project

Three things, honestly:

1. Decision quality you can watch. The jump from a fast small model to a flagship reasoning model isn't subtle when the output is money on a leaderboard. The agents' reason fields — one short sentence per trade — went from plausible-sounding to actually grounded in the market state they were shown.

2. Format reliability. Zero parse failures across 66 decisions. For an agent whose entire interface to the world is a JSON object that becomes a blockchain transaction, that's the difference between "demo" and "product."

3. A real agentic story. The bounty asked for Qwen doing real work, not answering prompts. Our agents plan (persona + market analysis), use tools (onchain reads via RPC, transaction submission), and execute multi-step loops (decide → re-validate the round is still active → send the trade → repeat). Qwen 3.8 Max sits at the center of that loop, settling decisions without human intervention.

Try it

Contex Arena is live on Monad testnet at contexarena.xyz — connect a wallet, grab testnet MON from the faucet, bet on a fighter, and watch three Qwen-powered brains try to out-trade each other. Every trade, bet, and payout is onchain and verifiable.

Watch the 85-second demo on YouTube.

Built for the Monad Metropolis hackathon — Trust, Identity & AI Infrastructure track.

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.