Can a foundation model forecast crypto?
We ran the open-source Kronos candlestick model through a full year of Bitcoin, Ether and Solana. It has real skill at 5 to 15 minutes, and that skill is about twenty times too small to pay the fees.
Key findings
- Kronos-small predicts the direction of the next 5 minutes of Bitcoin better than chance: a rank correlation (IC) of +0.086 over 1,456 forecasts (p = 0.001) and a 53.3% hit rate.
- The skill is gone by 30 minutes. Only 5 of 37 run × horizon cells were significant at the 5% level, close to the 1.9 that chance alone would give, and one of the five pointed the wrong way.
- The average 5-minute Bitcoin move in the test year was 0.094%. A spot round trip costs 0.2% with maker orders and 0.5% with taker orders. 0 of 43 simulated trading cells made money after fees.
- Beyond a few minutes the model mostly predicts a reversal of its input window. Its 4-hour error was 2.8 times larger than a “no change” forecast. Fine-tuning and the four times larger base model did not fix this.
What Kronos is
Kronos is an open-source foundation model for financial candlesticks (arXiv 2508.02739, MIT licence). It turns open, high, low, close and volume bars into tokens and generates future bars one step at a time, the way a language model generates text. We tested Kronos-small (24.7 million parameters) and Kronos-base (102.3 million), both with a 512-bar context window, and a copy of Kronos-small that we fine-tuned on 5-minute crypto bars.
We wanted to know whether its forecasts are good enough to trade after the fees of a European spot exchange.
The test
- Test year: 1 September 2025 to 31 August 2026, Binance spot prices for BTC, ETH and SOL on 5-minute, 15-minute and 1-hour candles.
- Walk-forward origins: every 6 hours (1,456 forecast origins) or every 8 hours (1,092). At each origin the model saw only the bars before that moment: 448 five-minute bars (about 37 hours) or 488 hourly bars. It was never refit on test data.
- Sampling: 20 independent future paths per origin (temperature 1.0, top-p 0.9), up to 48 steps ahead.
- Scoring: the information coefficient (Spearman correlation between the mean of the 20 paths and the realised return), the direction hit rate, the error against a “no change” forecast and the coverage of the 10–90% band of the paths.
- Trading simulation: go long or short when at least 70% of the paths predict a move larger than the round-trip fee, at 0.2% and at 0.5% per round trip.
- Scale: eight runs, about six hours on one consumer graphics card.
Kronos forecast skill fades after 15 minutes
Rank correlation (IC) between Kronos's mean forecast and the realised return, by forecast horizon, over the Sep 2025 – Aug 2026 walk-forward test year.
Show data
| Forecast horizon (log scale) | BTC, 5-min candles, Kronos-small | BTC, 5-min candles, Kronos-base | BTC, 5-min candles, fine-tuned small | BTC, 1-hour candles, Kronos-small | ETH, 1-hour candles, Kronos-small |
|---|---|---|---|---|---|
| 5 | 0.09 | 0.07 | 0.02 | – | – |
| 15 | 0.06 | 0.06 | 0.02 | – | – |
| 30 | 0.02 | 0.02 | 0.01 | – | – |
| 60 | 0.02 | 0.01 | 0.02 | -0.0047 | -0.02 |
| 120 | 0.03 | – | 0.01 | – | – |
| 240 | 0.01 | – | -0.01 | – | – |
| 180 | – | – | – | 0.05 | 0.0024 |
| 360 | – | – | – | 0.04 | 0.02 |
| 720 | – | – | – | 0.02 | 0.04 |
| 1440 | – | – | – | 0.0053 | 0.03 |
Skill that fades in minutes
On Bitcoin 5-minute candles the IC was +0.086 at 5 minutes and +0.060 at 15 minutes. From 30 minutes on it fell to between +0.01 and +0.03, with p-values of 0.32 or more. The larger Kronos-base did no better (+0.072 and +0.061). On 1-hour candles there was no meaningful skill for Bitcoin, Ether or Solana at any horizon.
Part of the 5-minute skill is a known effect. The previous 5-minute return on its own has an IC of −0.067, because very short moves tend to reverse. After removing that effect, Kronos keeps an IC of +0.059 (p = 0.024), so the model does learn something beyond the simple reversal.
The moves are smaller than the fees
Where Kronos has skill, the moves are smaller than the fees
Average absolute BTC price move over each horizon at the 1,456 test origins (Sep 2025 – Aug 2026), against Bybit EU spot round-trip fees.
Show data · values in %
| Mean |BTC return| | |
|---|---|
| 5 min | 0.09 |
| 15 min | 0.15 |
| 30 min | 0.21 |
| 1 h | 0.28 |
| 2 h | 0.42 |
| 4 h | 0.64 |
| 6 h | 0.77 |
| 12 h | 1.09 |
| 24 h | 1.61 |
At 5 minutes the average absolute Bitcoin move was 0.094%. With an IC of about 0.06 and a typical move of 0.12%, the expected edge is roughly 0.01% per trade. The cheapest round trip costs 0.2%, twenty times more. The simulation agrees: none of the 43 run × horizon cells had a positive average after costs. At 0.2% per round trip the average simulated trade lost 0.23% and only 38% of trades won.
Longer forecasts mostly reverse the input
Beyond a few minutes one pattern dominated the forecasts: the model predicted that the recent trend in its input window would reverse. The rank correlation between the 4-hour forecast and the trend of the 37-hour lookback was −0.84. As a result the 4-hour forecast error (1.81% mean absolute error) was 2.8 times larger than simply forecasting no change (0.64%). The 10–90% band of the sampled paths covered 57% of outcomes at 5 minutes and only 14% at 4 hours, where a calibrated model would cover 80%.
We tried three fixes.
- Fine-tuning Kronos-small on BTC, ETH and SOL 5-minute bars up to July 2025 lowered the validation loss by 1.9%. The test-year IC at 5 minutes dropped to +0.022 (p = 0.47).
- The larger base model gave the same picture as the small one.
- Removing the bias: we measured the dependence on the lookback trend in the first half of the test year and subtracted it in the second half. What remained had an IC of −0.013 at 4 hours (p = 0.72).
What it is good for
The spread of the 20 sampled paths does predict how large the next move will be: a volatility IC of 0.28 to 0.42 up to one hour. On 1-hour candles that was slightly better than plain historical volatility (Bitcoin 0.31 against 0.26), on 5-minute candles slightly worse (0.36 against 0.38). Kronos is a reasonable volatility input for a larger model. It is not a trading signal on its own.
A small-sample trap
On the 364 origins that the original run shared with the fine-tuned run, the original model showed an IC of about +0.10 at one to two hours, with p ≈ 0.05. It looked promising. On all 1,456 origins the same horizons gave +0.018 and +0.026. A few hundred forecasts are not enough to trust a correlation of 0.1.
Caveats
- The test year was a bear market: Bitcoin −27%, Ether −44%, Solana −49%.
- The p-values treat forecast origins as independent. Horizons longer than the 6–8 hour spacing overlap, so their p-values are optimistic. That makes the lack of long-horizon skill more convincing, not less.
- We used 20 samples per origin and the default sampling settings. Other settings might sharpen the mean forecast. They would not change how small a 5-minute move is compared with the fee.
- The simulation includes short trades, which need margin on a spot exchange.
Method details for every study on this site are on the methods page.