June 17, 2026
Featured Insights
Do AIs Make Good Traders, and Do They Make Good Traders Better?
Two years ago, we put that conjecture to the test. We called it "The Crystal Ball Challenge." We staked 120 finance-trained adults with $50 each and handed them the front page of the Wall Street Journal - one day before publication, with any mention of market moves blacked out. In effect, we gave them what every trader dreams of: a working crystal ball. They could go long or short the S&P 500 and 30-year Treasury bonds, with leverage if desired. For example, shown Wednesday's front page - reporting on Tuesday's events - they placed their trades at Monday's close and were closed out at Tuesday's close, once the news had played out in the market. Each player got 15 trading opportunities, one front page per year from 2008 to 2022.2
The result: our paid players didn't do too well! On average, they broke even and about 1/6 went bust. They weren't great at inferring market direction from the crystal ball, but they were particularly bad at position-sizing. We also invited five very successful macro-traders to play the game and they did pretty well - more on that shortly.3
The Crystal Ball game has been hugely popular since we made it freely available to play on our website. About 60,000 people have given it a whirl, and they fared substantially worse on average than our paid players - not surprising, given they were staked with $1 million in play money rather than real greenbacks.
Many of our readers have asked: why should humans have all the fun? So we invited four leading LLMs - Claude, ChatGPT, Gemini, and Grok - to try their hand at our game. We even went a step further and updated the game so you can compete against the AIs mano-a-machina, no tears. Everyone trades simultaneously - sealed bid - and the AIs see the same headlines you do, reasoning from first principles without knowing the market returns. There's also a leaderboard where players can see their own and the AIs performance. Want to give it a try before reading on? Play the Crystal Ball Game.
How'd the AIs do?
Below are the results for the four AIs over ten rounds of play, with a starting wealth of $1 million. In terms of average ending wealth, Claude does best, followed by ChatGPT. Grok just about breaks even, while Gemini on average loses a considerable amount of starting wealth.
| Model | Avg ending wealth | Hit ratio | Sharpe ratio (per day) | Mean stock bet | Mean bond bet | Mean daily return | Std Dev of daily return |
| Claude | $2,594,258 | 60.7% | 0.31 | 10.8x | 7.8x | 10.48% | 33.3% |
| Gemini | $492,357 | 42.3% | 0.02 | 11.7x | 11.2x | 1.00% | 38.5% |
| Grok | $968,181 | 48.8% | 0.12 | 11.9x | 11.3x | 4.46% | 34.8% |
| ChatGPT | $1,472,347 | 54.3% | 0.21 | 6.9x | 5.6x | 3.90% | 19.4% |
Click for full-size
The AIs took way too much risk
Whether the AIs made money, as Claude and ChatGPT did, isn't the end of the story. In order to assess how well they played, we need to assess their performance relative to the objective we provided before they played the game. Our instructions to the AIs were:
Please play the game as if you are a typical middle-aged, wealthy investor in the United States. The starting bankroll represents 100% of your financial wealth, you have pension income covering your subsistence needs, and there are no taxes.4
A high-functioning LLM could reasonably have inferred from this that their objective should not be to maximize expected gain, but rather their risk-adjusted profit, assuming a typical investor's degree of risk-aversion.
We believe all the AIs were too aggressive in their position-sizing. One perspective on this is to note that the US stock market has moved by over 5% on 23 days and by over 9% on seven days since the year 2000. Given average position sizing in stocks of 7x to 12x across the AIs, we think they were taking too much risk of a catastrophic loss of capital, given none of them had (or could reasonably expect to have) super high hit ratios.
A more coherent way of thinking about their bet-sizing is to refer to the Merton share rule-of-thumb for position sizing, which says that the optimal amount of volatility to stomach is equal to your Sharpe ratio.5 divided by your coefficient of risk-aversion. We told the AIs to act as though they were a typical middle-aged, wealthy investor, so they should have realized that such an investor would have a coefficient of risk-aversion in the range of 2 - 4. If the AI estimated its Sharpe ratio even as high as Claude's at 0.3, then the Merton share would say it should have been sizing its positions to have a daily standard deviation of returns of 7.5% to 15%6 Instead, the AIs ran daily standard deviations of 20% to 40%, which is just way too high given even an optimistic expectation on their part for how well they'd do.7
Even though Gemini and Grok both have positive Sharpe ratios of 0.02 and 0.12, they both lose money on average, and Gemini loses a lot!8 This is a result of betting too aggressively, using about 12x leverage on their stock trades and 11x leverage on their bond trades, generating a standard deviation of daily returns of 25% - 40%. If Grok reduced its bet size by 60% across the board, it would have wound up with $1.27 million on average, and if Gemini bet 90% less than it bet, it would have eked out a small profit ending with $1.005 million on average.
Interestingly, all four of the AIs knew about Kelly betting. We asked the AIs to play our coin flip game, and all four gave excellent answers as to their approach to the game, grounded in the Kelly Criterion, which can be thought of as a special case of the Merton share. It's interesting that they understood how to take risk on a biased coin-flip, but when it came to betting on stocks and bonds, they didn't seem able to make the connection between the two problems.
How did the AIs do against humans?
Overall, we'd say that Claude and ChatGPT did a better job than most people at guessing the direction of markets, but as we discussed already, all the AIs bet too aggressively in their basic model versions. The many people who have played the game on our website exhibited a wide range of position-sizing approaches - many took too much risk, while quite a few took very little risk, which may have been consistent with their lack of confidence in correctly reading the crystal ball.
The table below reports how often AIs beat - i.e. finished with more wealth than - human players, measured over about 200 sessions.9
| Model | % of sessions AIs beat humans |
| Claude | 76% |
| Gemini | 43% |
| Grok | 51% |
| ChatGPT | 63% |
Click for full-size
In our initial experiment, we invited five professional macro traders - four men and one woman - to play the game. This was quite a select group of traders: head of trading at a top-five US bank, founder of a top-ten macro hedge fund, senior trader at a top-ten macro fund, former senior government bond trader at top-three US primary dealer, and former senior Jane Street trader. They did about as well as Claude and ChatGPT in guessing the direction of markets. Their average ending wealth was $2.3 million. Position-sizing varied quite a bit among these five expert traders, but the majority of them sized their bets sensibly assuming their realized return and risk metrics were close to what they reasonably expected ex ante.
Did people play differently when flying solo versus playing head-to-head against the AIs?
We speculated that players would play meaningfully differently when filled with competitive fire, intent on showing our robot overlords that they can't beat us at everything (yet). The high-level data doesn't obviously suggest this: when playing the AIs, about 50% of players were profitable and 20% went bust, quite similar to the solo data from the first Crystal Ball game.
When it comes to bet sizing however, we see more differences. Head-to-head players were more aggressive, using an average gross leverage of 30x vs 25x when playing solo. When playing the AIs, players went all the way to the wall almost twice as often as well, maxing their leverage on 10% of plays, vs 6% for solo play.
Maybe the red haze of AI rage just makes us forget about optimal bet-sizing. Or perhaps when trying to beat the machine, we just focus on trying to squeeze out the most return we can each round, without thinking about the effect that's having on our compound return over many rounds.
Using top-of-the-line versions of the AI models
We also ran the highest reasoning versions of Claude, Gemini and ChatGPT.10 The results are in the table below. These AIs generated Sharpe ratios substantially higher than their more work-a-day siblings. Claude significantly reduced its position sizing, while Gemini and ChatGPT both increased their risk-taking to levels we deem too aggressive. Note that Gemini still is losing money despite increasing its Sharpe ratio to 0.16, a result of it taking over 50% risk on average on each trading day.
An approximation for the compound return an investor should expect to earn is equal to the average return each period minus one half the variance of those returns. So, for Gemini in "smart mode," it has an expected compound return of \(0.083 - .533^2/2 \approx -6%\). This negative expected compound return means that the median wealth outcome for Gemini is a loss of 60% of starting capital, which is pretty close to the 47% loss of wealth Gemini experienced.11
| Model | Avg ending wealth | Hit ratio | Sharpe ratio (per day) | Mean stock bet | Mean bond bet | Mean daily return | Std Dev of daily return |
| Claude | $3,335,220 | 67.9% | 0.50 | 5.2x | 3.8x | 9.75% | 19.5% |
| Gemini | $530,322 | 55.3% | 0.16 | 14.6x | 17.6x | 8.30% | 53.3% |
| ChatGPT | $2,755,492 | 62.6% | 0.34 | 6.1x | 5.3x | 9.15% | 27.0% |
Click for full-size
Can we help the AIs on position-sizing by making them read our book?
We tried this out by instructing the high-reasoning versions of the AIs to "read" our book, The Missing Billionaires: A Guide to Better Financial Decisions, and then play the game.12 As you can see from the table below, this helped them quite a bit, bringing their position-sizing to a more appropriate level, though Claude and ChatGPT may have gone a bit too far in reducing their sizing relative to the Merton share. Perhaps our book made the AIs a little too conservative!
Note that Gemini and Grok are in the money, and Gemini arguably got the most out of reading our book. Not surprisingly given our book is focused on investment sizing and not investment selection, it didn't help them with their prediction skills, which remained about the same.
| Model | Avg ending wealth | Hit ratio | Sharpe ratio (per day) | Mean stock bet | Mean bond bet | Mean daily return | Std Dev of daily return |
| Claude | $1,365,531 | 65.3% | 0.47 | 1.7x | 1.1x | 2.20% | 4.8% |
| Gemini | $1,297,241 | 57.9% | 0.22 | 3.3x | 3.7x | 2.15% | 9.9% |
| Grok | $1,117,116 | 46.8% | 0.14 | 2.2x | 1.9x | 0.90% | 6.3% |
| ChatGPT | $1,199,774 | 61.5% | 0.32 | 1.1x | 1.0x | 1.30% | 4.1% |
Click for full-size
Did the AIs play fair and square?
A few questions probably on your mind:
- How could we stop the AIs from cheating? Or, more precisely, how could we stop them from figuring out the actual date associated with the front page and then seeing how much stocks and bonds moved, and betting with 100% accuracy?
Each day's news is provided to the AIs with context prompts to generate bets using general knowledge and reasoning only, with no recourse to market data. The internal reasoning scroll is monitored for compliance,13 as is the AIs ex-post self-summary of its reasoning. - Which version of each AI did we use?
The base models used were claude-sonnet-4.6, gpt-5, gemini-2.5-flash, and grok-4-fast-reasoning. The high-reasoning models used were claude-opus-4.7, gpt-5 (reasoning_effort=high), gemini-2.5-pro, and grok-4. Each model is used with the default temperature. - What did the AIs actually see?
The AIs were provided with OCR-generated text from each of the WSJ pages. We suspect this provides a modest advantage to the humans, who have more context about the overall page, what text was redacted, etc.
Connecting the dots
Investing involves two decisions: choosing what to invest in, and deciding how much to invest in what you've chosen. Our Crystal Ball experiment suggests that AIs - particularly Claude and ChatGPT - are roughly as good as expert humans on the what decision, at least when it comes to connecting macro-economic news to near-term market direction. That's a genuinely impressive result.
On the how much question, however, the AIs fell short, both in absolute terms and relative to the seasoned, expert traders in our original experiment. The AIs all took far too much risk for the objective we gave them. The irony is that they all knew about Kelly betting and the Merton share in the abstract - they just couldn't apply it when the problem was dressed up in market rather than coin-flip clothing.
In recent months, AIs have been tested on a range of financial challenges, from stock-picking to complex, multi-variable personal finance decisions. Our impression of this broad swath of experimentation is that AIs aren't doing great. The stock picking result is not that surprising, given its zero-sum nature: one investor's alpha comes at another's expense, and any edge that's freely and widely available is quickly arbitraged away. As Bloomberg's Matt Levine observes: "If ChatGPT knew what stocks would go up, it would be managing Renaissance."
But most personal financial decisions - how much to save, how to allocate between stocks and bonds, when to convert a Roth IRA - are not zero-sum. We can all make good decisions simultaneously. There's no adversary to adapt and neutralize the AIs advice. This is where we expect AIs to improve most, and most durably.
The result of the AIs reading our Missing Billionaires book points in this direction: with more specific guidance on risk-taking, the AIs calibrated their position sizing considerably better. Perhaps they overcorrected slightly - our book may have made them a touch too conservative - but the responsiveness is encouraging. A more systematic approach to teaching AIs the principles of good financial decision-making seems likely to pay off.
The deeper promise may be in human-AI collaboration. AIs may prove most valuable not as autonomous traders but as a counterweight to our very human behavioral biases such as overconfidence, recency bias, and the tendency to bet too big on views that feel certain but aren't.
Appendix: AI Bet-Sizing Detailed Statistics
Stock Bet Size by Question For Base AI Models
| Model | Stat | Q1 | Q2 | Q3 | Q4 | Q5 | Q6 | Q7 | Q8 | Q9 | Q10 | Q11 | Q12 | Q13 | Q14 | Q15 |
| Actual | Ret | 2.70% | -1.25% | -3.51% | -0.93% | 0.62% | 1.03% | 1.96% | -0.56% | -1.90% | 0.13% | 0.95% | -0.75% | -0.56% | -0.22% | -2.79% |
| Claude | Avg | -11 | 13 | -19 | -8 | 5 | 5 | 9 | 10 | -9 | 13 | 11 | -5 | 18 | -2 | -14 |
| Min | -12 | 12 | -20 | -8 | 5 | 5 | 8 | 8 | -12 | 12 | 8 | -15 | 15 | -8 | -20 | |
| Max | -8 | 15 | -15 | -8 | 5 | 5 | 15 | 15 | -8 | 18 | 12 | 12 | 20 | 5 | -12 | |
| Gemini | Avg | -14 | 18 | -14 | -12 | 4 | -4 | -8 | 12 | -8 | 12 | 9 | 14 | 14 | -14 | -4 |
| Min | -15 | 10 | -15 | -15 | 0 | -10 | -15 | 10 | -10 | 10 | 5 | 10 | 10 | -15 | -15 | |
| Max | -10 | 20 | -10 | 0 | 10 | 5 | 15 | 15 | 0 | 15 | 15 | 15 | 15 | -10 | 10 | |
| Grok | Avg | -15 | 15 | -13 | -6 | -1 | -10 | 5 | 13 | -14 | 10 | 13 | 18 | 16 | -10 | -7 |
| Min | -20 | 8 | -20 | -15 | -10 | -15 | -10 | 8 | -20 | 5 | 8 | 10 | 10 | -15 | -15 | |
| Max | -12 | 25 | 3 | 5 | 5 | -5 | 10 | 20 | -10 | 15 | 25 | 25 | 25 | -5 | 8 | |
| ChatGPT | Avg | 2 | 8 | -8 | -8 | 5 | -4 | 6 | 6 | -6 | 7 | 7 | 8 | 9 | -7 | 5 |
| Min | -10 | 8 | -10 | -8 | 4 | -6 | 6 | 6 | -6 | 6 | 6 | 6 | 8 | -8 | -6 | |
| Max | 6 | 12 | -8 | -8 | 6 | 4 | 6 | 8 | -5 | 8 | 8 | 8 | 12 | -6 | 8 |
Click for full-size
Bond Bet Size by Question For Base AI Models
| Model | Stat | Q1 | Q2 | Q3 | Q4 | Q5 | Q6 | Q7 | Q8 | Q9 | Q10 | Q11 | Q12 | Q13 | Q14 | Q15 |
| Actual | Ret | -0.81% | 1.32% | 2.74% | 1.58% | -0.77% | -0.77% | -1.02% | 1.06% | 0.41% | 0.60% | -0.20% | 1.07% | 1.44% | 1.09% | -0.88% |
| Claude | Avg | 8 | -8 | 14 | 5 | 1 | 1 | 6 | -6 | 12 | -8 | -7 | 12 | 5 | 1 | -10 |
| Min | 5 | -10 | 10 | 5 | -3 | -3 | 5 | -10 | 10 | -10 | -8 | -5 | -10 | -5 | -15 | |
| Max | 10 | -8 | 15 | 5 | 3 | 3 | 10 | -5 | 15 | -8 | -5 | 20 | 10 | 5 | -8 | |
| Gemini | Avg | 14 | -9 | 20 | 15 | 2 | 4 | 16 | -8 | 7 | -10 | 1 | -10 | 19 | -8 | 14 |
| Min | 5 | -15 | 15 | 0 | -5 | 0 | -10 | -10 | 0 | -10 | -10 | -10 | 15 | -10 | 10 | |
| Max | 20 | 15 | 20 | 20 | 10 | 10 | 20 | -5 | 15 | -10 | 10 | -5 | 20 | -5 | 20 | |
| Grok | Avg | 15 | -7 | 16 | 16 | 2 | 12 | 2 | -7 | 16 | -10 | -8 | -4 | 19 | -3 | 6 |
| Min | 10 | -15 | 8 | 10 | -3 | 5 | -15 | -15 | 12 | -15 | -20 | -15 | 10 | -15 | -10 | |
| Max | 20 | 5 | 25 | 25 | 12 | 20 | 15 | -3 | 20 | -5 | -4 | 15 | 30 | 10 | 20 | |
| ChatGPT | Avg | 6 | -6 | 9 | 10 | -3 | 5 | -1 | -4 | 5 | -4 | -4 | -6 | 4 | 6 | 6 |
| Min | 4 | -8 | 6 | 10 | -4 | 3 | -5 | -6 | 4 | -5 | -5 | -6 | -5 | 4 | 5 | |
| Max | 8 | -5 | 10 | 10 | -2 | 6 | 5 | -4 | 6 | -4 | -3 | -4 | 10 | 8 | 8 |
Click for full-size
Crystal ball riddles (answers in footnotes):14
• Where does a crystal ball do its shopping?15
• What is a crystal ball's least favorite drug?16
The content on this page is being provided as general market commentary and for educational purposes only. It does not constitute any form of investment advice, or recommendation to buy or sell any securities or adopt any investment strategy mentioned herein. Any investment strategies and investment results discussed herein are for illustration purposes only in the context of the commentary, and do not reflect actual or hypothetical Elm Wealth strategies or results, or an offer to provide such strategies or results.
This content is intended only to provide observations and views of the author(s) at the time of writing, both of which are subject to change at any time without prior notice. The information contained in the commentaries is derived from sources deemed by Elm Wealth to be reliable, but its accuracy and completeness cannot be guaranteed. This material does not have regard to specific investment objectives, financial situation and the particular needs of any specific reader. Any views regarding future prospects may or may not be realized. Past performance is no guarantee of future results.
- Thanks to our stable of traders for their diligent testing: Nolan Lenaghan, Frankie Agrest, Brandon Labbe, Mike Fothergill, and Steven Schneider. Elm's Crystal Ball Trading Challenge is for entertainment and educational purposes only, and nothing in it should be considered investment advice, a solicitation, or a recommendation to use AI for trading. Performance by any AI model in a simulated, hypothetical exercise is not indicative of its ability to generate returns in real markets.
- Trading days drawn one per year from the past 15 years, randomly selected from employment reports, Fed announcements, and other days, all drawn from the top half of days by market volatility. The days are chosen at random and not intended to be misleading or trick the players.
- Here's the original article: When a Crystal Ball Isn't Enough to Make You Rich.
- We did not tell the AIs our randomized methodology, described above, for how the 15 days were chosen.
- The ratio of excess return to volatility.
- i.e. a range of 0.3/4 - 0.3/2.
- The AIs might have expected they'd do better at the game than they did, which would justify taking more risk than our analysis based on their realized success, but we don't think their position-sizing could be supported by any reasonable expectation.
- You may be wondering why Gemini and Grok both have positive Sharpe ratios, despite having Hit Ratios saying they're guessing the direction of stocks and bonds incorrectly over 50% of the time. The explanation is that they're making more money on the trades they get right, either because they're betting bigger on those or the market moves more in their favor when they're guessing right.
- Fewer for ChatGPT, due to some rate-limiting we hit.
- A high-reasoning mode for Grok wasn't available.
- While you might have thought that Gemini would wind up with its expected rather than median wealth outcome, note that while it got to play the game ten times, it played it quite similarly all ten times and so it's not surprising that it wound up with an average outcome close to the median.
- For Grok, we used their standard model here too.
- With the exception of Grok.
- With thanks to Belle Osvath, who stumped Victor with these two as a guest on her Smarter Planner podcast where they talked about this AI Crystal Ball experiment.
- Sears.
- Crack.