The research presents a continuous record of autonomous language-model trading agents operating in production. Two systems, DX Terminal Pro and the DXAP live alpha fleet, were monitored over six months. The data includes approximately 7.5 million single-model invocations and 300,000 onchain actions, alongside 231,638 multi-tool turns. This provides a detailed view of agent behavior in a live trading environment.
Key findings indicate that the operating layer—specifically, the risk slider—has a greater influence on agent behavior than the strategy text itself. Agent fixed effects account for 60% of the variance in behavior, and a leaderboard render boundary significantly impacts agent selection. Furthermore, sizing strategies are largely volatility-blind, with a single posture-slider cell holding a disproportionate share of liquidations.
Analysis of the agents’ performance reveals a lack of directional edge. Neither fleet demonstrates profitability, and the DXAP fleet trails a matched Hyperliquid retail benchmark. The research also found that agents capture very little of the upside they reach, with a significant proportion of profitable excursions resulting in negative trade returns. Mechanical bracket strategies recover only +39.0 bps per position.
The study concludes with a 17-rule methodology canon, offering insights into the operational characteristics of these agents. Source: https://arxiv.org/abs/2609.05663v1