The slow model was right more often. The fast one made more money.
Arne Kellmann ·
Two of the three models in this experiment gave better calls than the third. The third earned three times as much per call. It was not smarter. It answered in a quarter of a second.
This is a replay, not a live account, and there are no fees in the numbers below. Read to the end before you conclude anything. The point is not a trading strategy. The point is a measurement of what latency is worth when the world keeps moving while a model thinks.
Who is behind it. Jev is the first model of TypeSafe AI, a San Francisco lab that left stealth on 15 September 2026 with a $40 million seed round led by DCVC. The founders are Diogo Almeida, formerly a researcher at OpenAI, with Erik Gafni and Sasha Sheng. On X: @typesafeai and @CompleteSkeptic. I have no relationship with the company; the API key is a normal early-access key and I paid for every call in this post.
The setup
- The day: 15 September 2026, BTC/USD spot on Kraken. 132,800 trades with millisecond timestamps, pulled from the public API. Range of the day 4.5 percent, one of the wilder days of the month.
- The moments: 446 points where something was happening, a five-second move of at least 8 basis points or a burst of trades at four times the day’s rate, at least 20 seconds apart. Plus 60 quiet points as a control. 506 decisions.
- The state: the same JSON for every model at every point. Returns over the last 1, 5, 15, 30, 120 and 300 seconds in basis points, trade count and volume of the last five seconds against the day’s average, buy share, range of the last 30 seconds. Nothing from the future.
- The question: buy, sell or hold for the next 30 seconds, where buy means “more likely to rise than fall by more than 2 bps”.
- The deciders: Jev 1.13.0 asked as one Choice plus one Noul. GPT-6 Astra and GPT-5.4 nano asked for the same answer as JSON, reasoning effort on the lowest setting the API accepts, so each runs as fast as it can.
- The fill: a decision made at time t is filled at the first real trade after t plus that model’s measured latency. Exit at t plus 30 seconds. Every latency is the real round trip from Frankfurt, measured per call.
All three ran in parallel on every point, so no model saw a different market.
Latency
| Model | p50 | p90 | Mean |
|---|---|---|---|
| Jev 1.13.0 | 252 ms | 538 ms | 330 ms |
| GPT-5.4 nano | 1,574 ms | 2,289 ms | 1,739 ms |
| GPT-6 Astra | 2,609 ms | 4,526 ms | 2,987 ms |
Ten times. That was expected. What it does to the money was not.
Who was right
Hit rate is the share of directional calls where the price was on the called side 30 seconds later.
| Model | Directional calls | Hit rate at 30 s |
|---|---|---|
| Jev | 408 | 54.2 % |
| GPT-5.4 nano | 414 | 58.5 % |
| GPT-6 Astra | 295 | 59.7 % |
Astra is the most accurate and the most cautious: it said hold on 210 of 506 points. Jev is the most trigger-happy and the least accurate. If you scored this like a quiz, Astra wins.
Who made money
Now the number that matters. Take only the 258 points where all three models made the same directional call. Same decision, same direction, same moment. The only difference is when each one’s order reaches the book.
| Enter at | Gross result per call |
|---|---|
| Jev’s fill, ~250 ms | 2.45 bps |
| Fixed 300 ms | 2.42 bps |
| Fixed 1 s | 1.11 bps |
| nano’s fill, ~1.6 s | 0.96 bps |
| Fixed 2 s | 0.82 bps |
| Astra’s fill, ~2.6 s | 0.80 bps |
| Fixed 5 s | 0.47 bps |
| Fixed 10 s | 0.38 bps |
Two thirds of the edge is gone by the time Astra’s answer arrives. On 30 percent of those agreed calls the price had already moved a full basis point in the called direction before Astra’s fill. The market does not wait for the reasoning to finish.
Scored on their own calls at their own latency, per directional call: Jev 1.47 bps, nano 1.04 bps, Astra 0.88 bps. The least accurate model, filled first, comes out ahead. Not because it saw more. Because it was there.
Confidence did its job
Jev returns a confidence with every Choice. I did nothing with it during the run. Afterwards:
| Jev confidence | Calls | Hit rate | Per call at own fill |
|---|---|---|---|
| under 0.5 | 232 | 48.3 % | 0.43 bps |
| 0.5 and above | 176 | 62.3 % | 2.84 bps |
Below 0.5 it is a coin flip, and it says so. Above 0.5 it is as accurate as Astra and still filled in a quarter of a second. Act only on the confident half and the per-call result at Jev’s latency is 2.84 bps. The same calls filled at Astra’s latency: 1.00 bps. That is the whole article in one row.
The potential win, in numbers
The tables above are per call. Here is the day.
On the 258 trades all three models agreed on, filled at each model’s own latency and closed 30 seconds later:
| Filled at | Gross result for the day | On $10k per trade | On $100k per trade | On $1M per trade |
|---|---|---|---|---|
| Jev, ~250 ms | 631 bps | $631 | $6,310 | $63,100 |
| nano, ~1.6 s | 248 bps | $248 | $2,480 | $24,800 |
| Astra, ~2.6 s | 207 bps | $207 | $2,070 | $20,700 |
The latency premium, Jev’s fill against Astra’s on identical trades, is 424 bps for the day: 1.64 bps per call, standard error 0.20, bootstrap 95 percent interval 1.25 to 2.03. Jev’s fill was the better one on 77 percent of the calls. That is not noise. It is the same decision, reaching the book two seconds earlier, 258 times.
Now the fees, because they decide whether any of this is money:
| Fee per side | Jev fill, net day | Astra fill, net day | Premium |
|---|---|---|---|
| 0 bps (maker rebate, internalised flow) | +631 | +207 | +424 |
| 1 bps | +115 | −309 | +424 |
| 2 bps | −401 | −825 | +424 |
| 10 bps (Kraken high-volume taker) | −4,529 | −4,953 | +424 |
| 26 bps (Kraken base taker) | −12,785 | −13,209 | +424 |
Two things fall out. First, at retail fees the strategy is a way to lose money quickly, and being fast only slows the loss. Second, the premium itself does not care about fees: it is 424 bps whichever row you are on, because both sides trade the same 258 times. Speed is worth the same whether the desk is profitable or not. What speed cannot do is make a bad edge good.
Where it is money: a desk that already trades at near-zero cost, maker-only, rebated, or internalised. For that desk, on this day, on a $100k clip, the difference between a 250 ms decider and a 2.6 second decider was $4,240. On $1M, $42,400. In a day. From the same signal. Restricting to Jev’s confident half, 176 calls, the day is 500 bps at Jev’s fill and 175 bps at Astra’s latency: the confident subset keeps most of the edge and cuts the trade count by a third, which is what a fee-paying desk needs.
Largest drawdown of the Jev curve during the day: 107 bps. Not a smooth line. A day.
What it cost
| Model | Tokens | Price |
|---|---|---|
| Jev | 320,826 in, output free | $0.013 |
| GPT-5.4 nano | 190,889 | $0.04 |
| GPT-6 Astra | 125,061 in, 18,905 out | $2.20 |
Five hundred decisions for a cent, or for two dollars. At ten decisions a second, one is a business and the other is a bill.
Now the cold water
- No fees. Kraken’s taker fee is 26 bps per side at the base tier. Every number above is single digits. On this day, at this frequency, with market orders, all three models lose money, and Jev loses it fastest because it trades most. This is a measurement of latency, not a strategy.
- The fill is optimistic. First trade after the timestamp, unlimited size, no queue, no impact. Real fills are worse, and worse for the slow model too.
- One day, one market. 258 agreed calls is enough to see the shape and not enough to bank on it. A quiet day would show less; the control points showed almost nothing.
- Latency is from my desk. Colocated, all three would be faster, and the ratio would probably widen, not shrink.
None of that changes the shape of the entry table. Whatever your edge is, it decays by the second, and a model that needs three of them starts from a worse price.
Where this applies
Trading is the example because it has a price every millisecond. The shape is general. Any decision where the world moves while you decide has an entry table like the one above:
- a bid in a real-time ad auction, where the slot is gone in 100 ms,
- a fraud check on a payment, where the customer abandons at three seconds,
- a route for a support message while the sender is still typing,
- a rate-limit or a block on a request that is already being served.
For those, “the more accurate model” is the wrong question. The question is: accurate enough at the latency the decision lives in. A quarter of a second and a confidence number you can threshold covers more of them than people think.
The replay script is 150 lines of Node: fetch the day, cut it into seconds, build the state, ask three APIs, look up the fill. If you want to run it on your own day, the Jev docs are the only unusual dependency.