Roni Rechter

← Projects

LLM Trading Arena

Status: in progress

An arena in which several language models, each given a different trading strategy, operate against the same market data and are scored continuously against each other.

The point is not to find a model that makes money. It is to watch what different models do with the same autonomous mandate, under a scoring function that cannot be argued with.

Why this exists

Most evaluation of language models happens where being wrong is cheap — a benchmark, a graded answer, a task with a correct response someone already knows. That measures capability under conditions the model's author selected.

Trading is a poor place to make money and an unusually good place to observe behaviour. It is continuous, adversarial, immediately scored, and impossible to revise after the fact. A model that is confidently wrong finds out, in public, on a schedule it does not control.

What I am actually looking for is the behavioural stuff that benchmarks do not surface. Whether a model's confidence tracks its accuracy or drifts away from it. What it does when the data is ambiguous rather than merely hard — hesitate, or invent conviction. Whether it recognises a regime it has not seen before, or applies the previous regime's logic to it with no signal that anything changed. Whether it escalates, and whether it escalates for the right reasons.

These are the same questions as the billing work, asked somewhere the feedback is fast and the excuses are unavailable.

What is not here

No results, no returns, no equity curve. Publishing performance from a short run would be noise presented as signal, and the whole reason this is interesting is that the scoring is honest.

If it produces something worth reading — including a stretch where the answer was that this approach does not work — it will be here, with the losing periods shown at the same size as the rest.