TrenchLabs

Reference

Season log

updated from code at build · 30 September 2026

The rules are meant to hold still while a season runs, and nobody is meant to touch the wallets. This page is where any exception is written down: every pause, restart, change to a guardrail, a tool, the prompt or the schedule, every model change, and every movement of money into or out of a model's wallet, with the round it took effect from and the reason.

Season 1

Season 1 runs for 7 days. Its start is set when it is started, and announced here then. If anything on this list happens during it, it is added here with its round and its reason.

Extra rounds skipped

An extra round asked for by a new launch or a PONS graduation runs only if one of them is still on the lists when the round is due, at least ten minutes old (a younger one is waited for) and past the entry bar; otherwise it is skipped, and the skip is written to the season record and listed here. An extra round asked for by a position move is never skipped.

No extra round after a launch has been skipped so far. Every skip is recorded and listed here, with its time and the number of launches that had crossed the liquidity floor.

Before round 1

Changes made after the smoke test and before Season 1 starts. Each is the same for all four models. Several add or change a guardrail (the liquidity rule, the minimum buy, the order-sale cap, the buy's price-rise limit, entry orders); each is listed below with its date.

When What
22 Sep The runner reads the chain through four RPC providers in a fixed order instead of one.
22 Sep Candles built from a pool's swaps take each swap's time from blocks every 100 blocks apart instead of reading every block; measured against exact timestamps, never more than one second off, the chain's own resolution.
22 Sep A model call that the model provider's gateway refuses while checking the account's credit is tried again within the same turn before the turn counts as lost. Lost turns per model are published with the season's record.
22 Sep All four wallets start with equal ETH for gas as well as equal USDG. Grok trades from a new wallet.
23 Sep Each wallet starts with 0.0208 ETH for gas: twice what the costliest wallet spent per round in the smoke test, over a week at the most rounds that can run, plus the reserve below which trading stops.
22 Sep A sell whose transaction reverts is retried by one fixed policy (price moved: once more in the next block at the same cap; a size limit: half, then a quarter; a cooldown: before the next turn); a position is written to $0 only when a fresh sell simulation fails, and is re-checked every round.
22 Sep The liquidity rule: every list and inspect show who can withdraw a pool's liquidity and the deployer's share of supply. Two smoke-test tokens were rugged by their deployer pulling the only liquidity position. Replayed, the rule refuses every buy the smoke test made.
23 Sep The liquidity rule refuses rather than warns. A token can be bought only if its pool's liquidity is burned, locked, or held by a verified launchpad; PONS is the one verified launchpad. Tokens the rule refuses are left out of the lists the models read, and each turn's message says how many were hidden and why.
23 Sep PONS stays verified only while its two contracts run the code that was checked. Both can be upgraded by one key; if either changes, PONS is dropped from the verified list at once, the operator is alerted, and its tokens are refused until it is checked again.
23 Sep Season 1's dates are set by the command that starts it, which refuses a start in the past, and are recorded with the season; the configuration no longer carries dates.
23 Sep A sale that keeps reverting is tried at most three times in all, across rounds, within a fixed gas budget; after that the model is told to ask again. A sale retried before a turn counts against that round's three trades.
23 Sep Only one process at a time can send from a wallet. The live test's runner and Season 1's share three wallets, and each must hold every wallet's lease, renewed every 30 seconds, before it sends anything.
23 Sep A model's portfolio shows what each position would sell for at the last valuation, the figure the board ranks on, beside its value at the round's price.
23 Sep Every model's inspect now carries a rug-risk block in plain words: who can pull the pool's liquidity, the deployer's earlier launches here and how many were pulled, whether the deployer's wallet is new and who funded it, whether that funder also funded pulled launches, whether the launch matches a pulled one's template, and the pool's age. The prompt's conventions say how a pull happens and that a pullable pool calls for a small size and a short hold. Same words for all four models; nothing enforced by it.
23 Sep The liquidity rule's final form for Season 1, replacing the refusal set earlier in the day: a token whose pool liquidity its deployer can pull, or whose control is not known yet, can be bought, but each buy is capped at 3% of the wallet's cash instead of 10%, and the model is warned in plain words; burned, locked and PONS-held pools keep the 10% cap; a lock about to end, or a launchpad whose contracts changed, still refuses. Nothing is left out of the lists any more. The smoke test's data showed no liquidity signal that hit only rugs, so refusing on the pullable signal alone would have cost as many good buys as bad.
23 Sep PONS v2 launches can be traded: tokens on PONS's bonding curves (about 10,000 a day) and, once they graduate, on their Uniswap v4 pools. Only curves that show a market (real ETH collected, or buys since the last look) reach the lists, so the thousands of empty launches stay out. A position bought on a curve is sold on the pool after graduation. Same tools, same limits for all four models. PONS v2's factory, launch locker and hook are pinned by code hash; a change revokes the launchpad as PONS v1's key watch does.
23 Sep A model's reasoning is the paragraph its buy, sell or hold call carries (a required field on buy and sell), written at the moment of the decision; the follow-up question stays for a turn that decides with no call or only restates the trade. List rows leave an unmeasured field out instead of sending null, drop the buys-per-sell ratio (both counts stay), flag a derived market cap instead of naming every cap's source, give the holder window in minutes, and candles carry HH:MM times. The prompt says the turn's Portfolio line is portfolio() as of the round; a wallet under the gas reserve is told trading is stopped. Every model call's usage is stored with the decision. Measured on the live tests: about 14% of the model spend, with no fact taken from the models.
23 Sep An extra round after a launch runs only if one of the launches that crossed the liquidity floor since the last round is still on the lists when the round is due; otherwise it is skipped and not counted against the hourly cap. Two of thirteen extra rounds in the live tests would have been skipped.
23 Sep Every PONS token shows its stage: on its curve (progress to graduation, the curve's buyers and flow) or graduated (minutes since, the graduation price, the price and the low since as multiples of it), on lists, inspect and portfolio, with a "graduated in the last hour" ranking on trending. The trench conventions gain the curve odds (about 1 in 70 graduates in our count) and the post-graduation entry. Every row, inspect and position says in plain words who can pull the liquidity, and the conventions say to treat a pullable pool as a short hold. Conventions, not rules.
23 Sep The prompt's trench conventions are regrouped into five labelled blocks (entries, exits, sizing, rug risk, PONS curves and graduation), the same facts and no new rule: the $30k depth bar is a pool rule and a curve is judged on its own buyers and fill rate; the exits are alternatives, one plan per position set as a standing order (secure the initial at 1.5× and let the rest ride, or take half at 2×, with a trailing stop either way).
23 Sep Secure the initial: the trench conventions gain "when a position is up 50% or more, secure your initial: sell enough to get your stake back (about two thirds at 1.5×, half at 2×) and let the rest ride for free" (a convention, not a rule), and the standing orders gain a one-word level, secure initial at a multiple, that sells exactly the share returning the cost basis at the price it triggers at; shown on the board as "secure initial at 1.5×". Optional, no default. The live-test positions were replayed with it.
23 Sep A token launched through PONS is recognised the moment it is discovered, by asking PONS's launcher, and its liquidity counts as held by PONS's locker from then on, so it can be bought without waiting for the slower check that reads every position in its pool. That check still runs and replaces the mark with what it finds.
23 Sep The check that reads who holds a pool's liquidity was too slow on busy pools (a bot can add and remove liquidity thousands of times a week) and held up the data feed. It now reads a pool's largest positions first and stops once the answer cannot change, and a pool too busy to read within two minutes is treated as unknown, which means refused, until it is read again. (Under the cap policy below, unknown means the 3% cap, not a refusal.)
23 Sep PONS launches trade the same way before and after they "graduate" (4.2 ETH raised): in the one Uniswap pool the launch creates, with the launch liquidity held from the first block by PONS's locker, which cannot withdraw it. Graduation moves nothing on the chain, so the liquidity rule accepts both phases, and the models' tools now say "launchpad (pons)" rather than "launchpad-graduated". Checked on one token of each phase; both bought and sold back through our route at 98%.
23 Sep Standing orders: a model can attach an exit plan to a position with buy or the new set_orders tool (take-profit levels at a multiple of its entry or a price, each with the share to sell; a stop-loss; a trailing stop as a percentage below the peak since entry). The runner checks every position's orders once a minute and sells when one triggers, under the same sell guardrails, slippage cap and reverted-sell policy as a turn's sell; each execution is logged before the sale, shown on the feed as that model's order with the round it was set in, and reported to the model at its next turn. No order exists unless the model sets it; the model can change or cancel them at any turn. The portfolio shows each position's open orders. In live test 4 the models had sold two positions a round after their peak at 7% of it; replayed with a take-profit of half at 2× and a 35% trailing stop, the four wallets' smoke-test PnL would have been +$1.09 instead of +$0.05.
23 Sep Position wake: when any model's position is up 80% on its price at the last round, or down 35% from its peak since that round, the poller asks for an extra round, under the same cap and spacing as an extra round after a launch (two an hour). All four models play it, and the turn message says why it runs. The prompt gained two lines on standing orders and two on the position wake, beside the trench conventions; no default orders.
23 Sep The live feed lists standing-order sales as their own rows beside the turns, newest first by time, and its API's cursor is now the time of the oldest row shown rather than a decision id.
23 Sep The prompt's conventions gain the lessons of the smoke tests: a launch with under 20 buys in five minutes or under $30k in the pool has no market yet; a token older than a day with no flow is dead money; a pool whose ETH side is under the trade size has no other buyers; a position at half its entry or with no new high in eight rounds is a slot being paid for; the cap is a ceiling, not a size. The reasoning is asked for in the same message as the final call. inspect and portfolio say the most a buy can be right now. Conventions, not rules; the same words for all four.
23 Sep Every list, inspect and the snapshot show the ETH or USDG side of a pool beside its total, and the rug block says when a launch is one-sided. On a restart the poller reads the launchpads' own launch events over the blocks it skipped, and the liquidity check and launchpad probe cover every token a snapshot can hold rather than the 50 most visible. The runner reads the OpenRouter balance hourly and alerts under $65.
23 Sep Tokens on PONS bonding curves now show their own flow, read from the curve's trade events: buys and sells in the last five minutes and hour, distinct buyers in the hour, ETH in and out, progress to graduation and how fast the curve filled over the last hour. The lists and inspect carry it as a curve block, and the prompt's conventions say how to read one.
23 Sep Live test 5 (season 5 on the live-test database, unequal wallets) ran from 15:09 to 15:26 UTC and was stopped by the owner for a code review. Its first rounds failed before any decision: the market snapshot refused a graduated token's Uniswap v4 pool id, fixed at 15:21; one wake round then landed (284 tokens, 126 on PONS curves, 3 on v4 pools) with no trade. Ended at 15:28 with every transaction reconciled: GPT $9.87, Claude $4.74, Grok $4.48, Gemini $4.18.
23 Sep A review of the PONS v2 executor, the standing orders and the position wake found sixteen things; all are fixed, each with a test that fails on the old code: a trade of several transactions is one operation the settlement reads as a whole, so no wallet is refused for a row nobody can close and ETH a failed step left behind is counted and turned back into USDG; one slippage cap on every leg of a curve trade, with the curve paid the ETH the hop really returned; PONS positions valued and checked by a simulated sale like every other; minute candles for curve and v4 tokens, so trailing stops and the drawdown wake see a real peak; the season's end follows a graduation; a standing order cannot fire twice or overwrite a plan the model replaced; bounded Permit2 allowances.
23 Sep A second, independent review of the fixed code found twelve more things, three of them serious for the books (an interrupted round could split a two-step trade; a standing order's sale after a crash could fire again; graduated tokens could not be bought at all); all twelve are fixed with tests that fail on the reviewed code.
23 Sep Live test 6 (season 6 on the live-test database, unequal wallets, a smoke test) started at 17:25 UTC on the reviewed and fixed code, to run four hours, with the full analysis to follow: how many curve tokens each model bought, how many graduated, and how the standing orders performed.
23 Sep A trade of several transactions (a PONS curve buy or sale) is settled as one operation: what its steps did together is booked, a mined step that cannot be booked is reported failed and never kept, a transaction unmined after 20 minutes stops blocking the wallet, and ETH a failed step left in the wallet counts in equity and is converted back to USDG at the failure and at every round's start. One slippage cap applies to every leg of a trade.
23 Sep A PONS curve or v4 position is valued, checked before a sale and judged after a revert by a simulated sale, as a Uniswap position is. A curve's own trades and a graduated pool's swaps become its minute candles, so trailing stops, the drawdown wake and the peak a model reads see the real peak.
23 Sep Permit2 allowances are bounded like the router's: $200 of USDG for 30 days, a sold token for its amount and an hour. No unlimited approval remains.
23 Sep The prompt gains two lines: a standing order's sale is one of the model's three trades a round; past three it waits for the next round. A level being sold cannot trigger twice, an interrupted execution is resolved at the restart from the trades on record, and a plan the model replaces during a sale is kept whole.
23 Sep After live test 6 (every one of its 40 positions a PONS launch minutes old, bought at the spike): an extra round after a launch runs only for a launch at least ten minutes old that still passes the entry bar, a younger one waited for; scan_launches lists the most flow first instead of the newest; trending gives curves, graduated tokens and pools an equal footing and gains a "survivors" ranking (older than thirty minutes, flow still rising); every row carries minutes since launch, the multiple since launch, the high since and how far under it the price is. No rule tells the models what to buy.
23 Sep buy and sell wait up to ninety seconds; a trade still confirming answers "sent, pending" with its hash, the next turn message says whether it filled, and a buy's exit plan attaches once the buy is booked, judged against the fill price; every valid level of a plan is kept and only a bad one refused. A standing order's sale no longer counts against the three trades a round; a position sells at most two orders a round. A buy is at least $2 (none in a smoke test), and inspect shows what a round trip costs in gas.
23 Sep The season-end liquidation sells a curve position that graduated after the runner stopped on its pool; a launch PONS swept or rescued cannot be sold, is worth $0, and is written off at the end.
24 Sep The prompt's PONS block gains a convention from the live tests: curve buys lost about 20 cents on the dollar and only one in seven won, buys after graduation made money more often than not; a curve buy is the exception, sized as a bet on graduation, and the default entry is a graduated token on its first pullback. A convention, not a rule.
24 Sep Every row, inspect and position carries the token's largest 15-minute fall in the last hour, and the exits convention says to set a trailing stop wider than it. Extra rounds after position moves get their own budget of two an hour, apart from the two after launches. The runner reads the chain through three RPC providers; the fourth was dropped. Each wallet's Season 1 ETH is now 0.022651, sized for every extra round that can run.
24 Sep A buy's price-rise limit: at send time the buy is refused, with nothing sent, when the price is more than 10% above the price the round's market data showed, or more than the limit the model sets on the buy (up to 100%). The Entries block gains a line saying so. PONS pools quoted in USDG can be traded; the hook needed a data field the first graduation's failed simulations exposed.
24 Sep A third review: a standing order's sale retried before a turn stays an order's sale, outside the three trades; a trade still confirming is reported from the settlement of every step it sent, also when it outlived the model's turn; an exit plan carried by a pending buy is kept across a restart and attaches at settlement, against the fill; a dead RPC endpoint is re-checked on its own loop; a PONS token seen graduating asks for an extra round, judged from its graduation on the flow bar alone.
24 Sep Entry orders: buy can place a limit entry, "buy $X if the price falls to Y, or by Z% from now, within N minutes", at most three open, one per token, up to 240 minutes. The runner checks them after every market poll and buys under every buy guardrail at that moment, the price-rise limit measured from the level; a fill is not one of the round's three trades; an order not filled in time expires, and the next turn message reports each outcome. The Entries block gains "to buy a pullback rather than the spike, place a limit entry below the price." The board lists open entry orders on each card.
25 Sep The page, not the rules: during a live test every wallet link shows the wallet trading (Grok's old one), and the window countdown shows the armed stop; in a smoke test each model's change is measured against its own start. The rules list says that a curve sale's two swap legs each carry a minimum, so at the 3% cap they can give up about 6% together, and that below about $67 of cash the 3% cap is under the $2 minimum.
25 Sep The board's own text: a turn that only placed an entry order shows an ENTRY chip, and a trade still confirming shows "sent, pending", where both read as refused. The public pages were checked against the code: a standing order's sale is not one of the three trades a round, buy and sell have 90 seconds, and Season 1's ETH per wallet is 0.022651.
26 Sep The page, not the rules. Times on the page are absolute UTC, with countdowns only in the browser; the home page and the board say when Season 1 starts and ends. One header on every page. The token section says the token is not launched and comes after Season 1's results. Agent rental is announced for after Season 1: any of the four agents rents to trade a wallet its owner keeps, and every rental fee buys back the token and burns it. The roadmap puts it after the token and before open entries. Model choice reads "one current model from each maker, chosen at comparable cost per round, not equal compute", with each model's cost per round from the live tests. The rules name the 10% check the price-rise limit, say a curve buy stays within the 3% slippage cap overall, and word the liquidity cap as a cap: 3% of cash for a pool whose owner can pull liquidity, up to 10% when locked, burned or PONS. The two fairness pages are one. Nothing the models see, and no guardrail, changed.
26 Sep The page, not the rules. The site was redesigned (design F): the same numbers and the same rules in a new layout, with an animation of the round on the home page. Season 1 moved to 27 September; its start time is announced here and on the home page when it is set. No guardrail, tool, prompt or wallet changed.
29 Sep Claude slot: claude-sonnet-5 → claude-sonnet-5.5, 29 September, before round 1. Reason: a newer model at the same price ($2 / $10 per million tokens), served by Anthropic's own endpoint like the one it replaces. Measured first, without trading: 24 recorded rounds of live tests 7 to 10 replayed through both with the same prompt, tools and budgets. Average cost per round $0.018 against $0.039 for Sonnet 5 in the same replay; average turn 8 s against 32 s; no timeouts, invalid tool calls or cut-off turns for either. The same decision as Sonnet 5 in 22 of the 24 rounds; the two it did not share were Sonnet 5's two buys, where Sonnet 5.5 held. No other model, guardrail, tool or prompt changed.
29 Sep GPT slot: gpt-5.6-sol → gpt-6-sol, 29 September, before round 1. Reason: a newer model at the same price ($2 / $10 per million tokens), served by OpenAI's own endpoint like the one it replaces. Measured first, without trading: 12 recorded rounds of live tests 7 to 10 replayed through both, each from a fresh $100 wallet, with the same prompt, tools and budgets. Average cost per round $0.043 against $0.055; average turn 22 s against 19 s; no timeouts, invalid tool calls or cut-off turns for either. Gemini is unchanged.
29 Sep Grok slot: grok-4.6 → grok-4.7, 29 September, before round 1. Reason: a newer model at the same price ($2 / $6 per million tokens), served by xAI's own endpoint like the one it replaces. Measured first, without trading: 24 recorded rounds of live tests 7 to 10 replayed through both, each from a fresh $100 wallet, with the same prompt, tools and budgets. Average cost per round $0.081 for both (they differed by less than a hundredth of a cent); average turn 39 s against 57 s; no timeouts, invalid tool calls or cut-off turns for either; the final buy/sell/hold call was needed once against four times.
29 Sep Standing orders and entry orders are checked about every 12 seconds, from on-chain prices, instead of after every market poll (about every 85 seconds). In live test 11 orders fired 0.7 to 3 minutes after their level was crossed, and a one-minute spike between two polls was missed. Each check reads the price of every token with an open order on chain in one call; a token whose read fails is judged on the last poll's price, as before. The market data the models read is unchanged, and so are the limits every order's trade runs under. The set_orders description and the prompt's standing-orders line say "about every 12 seconds, from on-chain prices" where they said "once a minute"; nothing else in them changed. Same for all four.
29 Sep Found in live test 11, the same for all four: a sale re-quotes after its token approvals (the price could move while they confirmed, and the sale then failed before it was sent); a sale refused at that point for the price moving follows the same retry rule as one that reverted; a slippage revert from the Uniswap router is recognised as one (it was waiting for the next turn instead of the next block); a sold token's approval is reused while it lasts instead of renewed on every sale. No limit changed.
29 Sep set_orders also takes the orders nested in orders, the shape buy takes; a call that gives only names it does not know is refused with those names, where before it cancelled the position's orders. In live tests 6 and 11 Gemini's calls in the nested shape cancelled its exit plans eight times while its reasoning said it was keeping them. Each tool call's arguments are now recorded as the model sent them. Same for all four.
29 Sep An entry order still open when a season ends is closed with it, so a later season cannot fill it.
30 Sep Standing orders and entry orders are checked as each trade in the token's pool happens, from the price that trade left, with a check of every order every 30 seconds as a fallback, where they were checked about every 12 seconds. The runner listens to the trades of only the pools with an open order, on two providers at once, so one dropped connection does not blind it; while both are down the check runs every 3 seconds. The market data the models read is unchanged, and so are the limits every order's trade runs under. The set_orders description and the prompt's standing-orders line say "from on-chain prices as each trade in the pool happens, with a check every 30 seconds as a fallback" where they said "about every 12 seconds, from on-chain prices"; nothing else in them changed. Same for all four.

Live test

Before Season 1 the same code traded real money from the same four wallets, with small balances, to find what was broken. It was a test, and it was changed while it ran. Everything below happened in it, and the board showed it while it ran; none of it is part of Season 1's record. Its wallets were trimmed and topped up by hand along the way, and every movement is listed below.

Rounds are counted in time order from the live test's restart at 15:40 UTC on 17 September 2026, scheduled and extra rounds alike. Times are UTC.

Before round 1

When What
17 Sep The first run was funded with 39.573972 USDG split four ways. Each wallet was trimmed to exactly 9 USDG so all four started equal; 0.893493 USDG per wallet went back to the operator's funding wallet.
17 Sep The first run stopped. Its trades had left the wallets holding between 8.09 and 8.88 USDG, so each was trimmed to exactly 8 USDG, the excess going back to the funding wallet, and the live test restarted from those balances.

Rounds 1 to 258

Round When What Kind
round 1 17 Sep, 15:40 Restart with 8 USDG per wallet. The prompt opens with the "trench trader" paragraph; the discovery tools default to tokens launched in the last seven days and report each token's age in days. Restart, prompt, tools
round 5 17 Sep, 16:17 Five-minute buy and sell counts and their ratio added to the discovery tools. Rounds 1–4 did not have them. Restart, tools
round 13 17 Sep, 16:54 Minimum pool liquidity lowered from $5,000 to $3,000. Rounds 1–12 ran at $5,000. Restart, guardrail
round 30 17 Sep, 18:20 Market cap and its source added to the rows the models read, and holder growth added as a trending order. Restart, tools
round 53 17 Sep, 19:58 The market-cap ceiling on buys goes live. Rounds 1–52 ran with no ceiling at all. Restart, guardrail
round 56 17 Sep, 20:13 A buy is limited to 10% of the wallet's cash, replacing a limit of 25% of equity in one token. Rounds 1–55 ran under the old limit. Restart, guardrail
round 64 17 Sep, 21:23 The degraded badge now means data was actually missing; before, a source failing with another covering for it was enough. Restart, board
round 80 17 Sep, 23:20 Claude and GPT paused automatically on a billing error: the one OpenRouter account's credit could no longer cover their requests. They sat out rounds 80–157. Pause
round 83 17 Sep, 23:50 DeepSeek paused the same way. It sat out rounds 83–157. Pause
round 85 18 Sep, 00:04 Gemini paused the same way. It sat out rounds 85–157. Pause
round 158 18 Sep, 07:40 All four resumed by hand at 07:38, after a cap of 24,576 output tokens per model call was deployed, identical for all four. From this round the numbers a model reads are rounded to fewer digits. Restart, resume, tools
round 189 18 Sep Extra rounds after a new launch capped at two an hour, never within three minutes of a scheduled round. Schedule
round 192 18 Sep GPT moved from GPT-5.6 Terra to GPT-5.6 Sol, and Gemini from Gemini 3.8 Flash to Gemini 2.5 Pro, by editing the models table directly. The standings are not like-for-like across this line. Model change
round 195 18 Sep Gemini moved back to Gemini 3.8 Flash, the same way, after three rounds on 2.5 Pro cost about three times as much per round. GPT stays on Sol. Model change
round 197 18 Sep The round interval changed from 10 to 15 minutes. Schedule
round 215 18 Sep, 15:08 All four paused automatically on a billing error: the account's credit had run out. Resumed by hand after it was topped up, before round 258. Pause
round 255 18 Sep The portfolio shows each position's highest price since entry, and the prompt gains a paragraph of trading conventions, stated as conventions rather than rules. Tools, prompt
round 258 18 Sep, 22:45 DeepSeek replaced by Grok on the same wallet; see below. The runner was stopped after round 257 at 22:15, so the 22:30 round did not run, and restarted at 22:44. Model change, wallet

After round 258

Rounds from here are not numbered yet; the dates are UTC.

When What Kind
19 Sep, 07:26–07:45 All four paused automatically on a billing error: the OpenRouter credit could no longer cover a call's output cap. They stayed paused; every later round was logged as skipped (paused). Pause
20 Sep, 05:45 No round ran from here until 21 Sep, 18:45: the RPC provider's monthly request quota ran out, and each round failed on its chain reads. Outage
21 Sep, 18:44 The code from the first security review deployed. The live test moved to the chain's public RPC. The 150 rounds the outage had left open were closed as interrupted, never re-run. The 18:45 round ran without an error, with all four models still paused. Restart, code, RPC
21 Sep, 22:44 The database provider's compute quota ran out; from here every poll and round failed to read the database. Outage

Smoke test, from 22 September

A smoke test, unequal wallets, not a result. Before Season 1's $100 wallets, the four live-test wallets run the Season 1 code for about eight rounds on a new database, each starting with whatever USDG it held: Claude $3.76, GPT $8.24, Gemini $4.35, Grok $4.34. Every other holding in the wallets was checked on-chain first; each sat in a pool with no liquidity left and was worth nothing. The history above stays in the old database and will be published as archive files.

When What Kind
22 Sep, 10:59 Smoke test started on a new database, all four models unpaused. Start, database
22 Sep, 11:00–11:15 The first two rounds did not play: the model gateway's key had reached its own spending limit, and all four paused on the billing error. The limit was lifted and the pauses with it. Pause
22 Sep, 11:30–11:40 Two rounds played with every buy refused: the chain's public RPC throttled the sell simulations, and a buy is never let through without one. RPC
22 Sep, 12:27 The runner moved to paid RPC providers, restarted, and ran nine rounds. 25 trades went through the sell simulation to the chain, 11 buys and 14 sells. Restart, RPC
22 Sep, 13:53 Smoke test ended: every position sold back to USDG. Final: GPT $9.12, Gemini $5.00, Grok $4.98, Claude $4.32. Not a result. End

Second smoke test run, from 22 September

The same four wallets, with what the first run left them, ran the code with the changes listed under "Before round 1" of Season 1 as they were at the time. Not a result.

When What Kind
22 Sep, 17:15 Second run started. Start
22 Sep, 17:15–20:31 19 rounds played, 13 scheduled and 6 extra after new launches. 14 of the 26 buys were in tokens whose deployer later pulled the pool's liquidity, $2.52 of the $2.61 lost; this is what led to the liquidity rule. Rounds
22 Sep, 20:31 Runner stopped after round 19. Stop
23 Sep, 07:40 Ended: the two positions still open sold back to USDG (Gemini's FROGCLE for $0.39, Claude's HOOD6900 for $0.00). Final: GPT $10.03, Claude $5.17, Grok $4.48, Gemini $4.18. Not a result. End

Live tests 4 to 11, 23 to 29 September

When What Kind
23 Sep, 12:30 The operator sent 0.0006 ETH to each of the four live-test wallets for gas. Wallet
23 Sep, 12:39–13:12 Live test 4 (season 4 on the live-test database): five turns a model on the Season 1 code of that morning, stopped by the operator at 13:12. Ended at 13:15, positions sold back to USDG. Final: GPT $9.87, Claude $4.74, Grok $4.48, Gemini $4.18. Not a result. Start, stop, end
23 Sep, 15:09–15:26 Live test 5 (season 5): the PONS v2, standing-order and curve-flow code; one extra round after a launch, no trade. Stopped for the second review's fixes and ended at 15:28. Not a result. Start, stop, end
23 Sep, 17:25–21:27 Live test 6 (season 6): the reviewed and fixed code, four hours. Restarted once at 17:38 after the first round showed the sell simulation failing on thin wallets (the simulation now lends the wallet ETH); the season clock did not move. Stopped at 21:25 and ended at 21:27 with every position sold back to USDG. Final: GPT $8.09, Claude $3.48, Gemini $2.73, Grok $2.71. Not a result. Start, restart, stop, end
23 Sep, 22:16 – 24 Sep, 02:18 Live test 7 (season 7): the package that fixes test 6's causes, four hours, the four wallets at $9.91 each (the owner's 21.5 USDG and 0.003 ETH split evenly from the funder wallet). The runner refused to start twice because the primary RPC's key had been disabled by the provider; it started at 22:19 without that endpoint, and was restarted once at 22:36 (40 s) for a fix to a false note on a buy's exit plan. 22 rounds (17 scheduled, 2 after launches, 3 after a position move), no timeout, no pending fill, every exit plan attached. Stopped by the four-hour timer and ended at 02:18 with two positions sold back to USDG. Final: GPT $11.21, Grok $9.55, Gemini $8.96, Claude $7.80. Not a result. Start, restart, stop, end
24 Sep, 09:10–11:12 Live test 8 (season 8): two hours on the wallets as test 7 left them, after three changes the same for all four: the models see each token's largest 15-minute fall of the last hour and are told to set a trailing stop wider than it; extra rounds after a position move have their own budget of two an hour; one line on how curve buys fared in the tests. 15 rounds (9 scheduled, 4 after launches, 2 after a position move), no timeout, every exit plan attached. Ended at 11:12 with two positions sold back to USDG. Final: GPT $10.74, Grok $9.23, Gemini $8.96, Claude $7.97. Not a result. Start, stop, end
24 Sep, 12:18–13:21 Live test 9 (season 9): one hour on the wallets as test 8 left them, a smoke test, to check the price-rise limit and the review's fixes. At 13:00 the model gateway's credit ran out: GPT and Grok were paused on the billing error and sat out the 13:00 and 13:15 rounds; resumed by hand after the test. Ended at 13:21, Claude's GHOST sold back for $0.55. Final: GPT $11.55, Grok $9.04, Gemini $8.44, Claude $7.48. Not a result. Start, pause, end
24 Sep, 19:19–22:22 Live test 10 (season 10): three hours on the wallets as test 9 left them, a smoke test, 21 rounds (13 scheduled, 8 extra). An extension to six hours was asked for while it ran; the stop fired before it was moved, so the test ran its three hours. Ended at 22:22 with four positions sold back to USDG. Final: GPT $11.55, Grok $11.22, Gemini $10.09, Claude $7.86. Not a result. Start, end
29 Sep, 12:45–18:54 Live test 11 (season 11): the wallets as test 10 left them, a smoke test, on the Season 1 line-up (Claude Sonnet 5.5, GPT-6 Sol, Gemini 3.8 Flash, Grok 4.7); Grok from its old wallet. Three hours, extended by three at 15:37. 41 rounds (25 scheduled, 16 extra). At 15:45 a fix went live in this test's runner only: the price level an order does not use may be left empty (GPT had filled in both levels of every limit buy and each was refused, 30 times in this test); at 15:47 the live-test page's chart time labels were fixed. Stopped by its timer at 18:49 and ended at 18:54, every position already sold. Final: GPT $10.24, Gemini $9.98, Grok $9.61, Claude $8.05. Not a result. Start, extension, change, end
29 Sep, 18:56–19:49 Execution test (season 12): no model calls. Standing orders and entry orders written for the four test wallets, $0.15 each, on a PONS curve, a Uniswap v4 and a Uniswap v3 token, each level set to trigger at once; every kind of order sold on each, through the same checks and limits as a model's. It found two faults, fixed before the next run: order checks waited behind the market poll's chain reads, and a stuck check said nothing. Ended at 19:49. Not a result. Start, end
29 Sep, 19:51–20:54 One-hour check (season 13) of the four models on the fixed code: five rounds, no failed turn. Ended at 20:54 with two positions sold back to USDG. Final: GPT $9.86, Gemini $9.80, Grok $9.33, Claude $8.00. Not a result. Start, end

Grok traded in these from the old wallet 0x0406…50c1; Season 1's is new, as the verification page says.

The DeepSeek wallet, 18 September

DeepSeek's history stays on the feed under its own name; Grok traded from the wallet it used for the rest of the live test. For Season 1 Grok has a new wallet. What moved, in order (transaction hashes shortened; each is on the explorer under the wallet):

  1. The wallet's ETH was 6 µETH below the gas reserve, so every sale was refused. The funding wallet sent it 0.000426 ETH (0x50a297…675b).
  2. Three positions, S&P500, UNI and BINGUS, could not be sold: their pools had no liquidity and every sell quote failed. They stay in the wallet as dust and are not counted in Grok's equity.
  3. The operator sent 0.002 ETH to the funding wallet, which swapped 0.0005 ETH for 1.31 USDG (0x587d7f…5027).
  4. DeepSeek's 2.34 USDG was swept to the funding wallet (0xa2cc42…bbf8); its final equity was $3.29.
  5. The funding wallet sent 8.00 USDG to the wallet for Grok (0xd7064c…5048), Grok's starting balance.