Results — can the AI read the news and call the stock?
Over 13 trading days and 3,873 consensus trades, the panel is right 48.8% of the time and makes +0.01% per trade — it does not beat a coin flip, but that is NOT distinguishable from luck yet (p=0.13).
Does it work?
Can it beat a coin flip?
Right 48.8% of the time — it needs to clear 50%. At 3,873 trades this is not statistically distinguishable from 50/50 (p=0.13).
Would it have made money?
About +0.01% per trade, equal-weight. Cumulatively over the 13 days it drifts to +0.07% — noise, no trend.
Is it better at winners or losers?
Its “bad news / short” calls did better (51.8%, +0.04%) than its “good news / long” calls (47.5%, -0.01%).
Which setup is best?
Which model is best?
The paid anchor gpt-4.1-nano leads. Full table:
| trades | hit rate | return/trade | |
|---|---|---|---|
| gpt-4.1-nano | 2842 | 49.5% | +0.10% |
| gpt-4o-mini | 3911 | 48.6% | -0.01% |
| consensus | 3873 | 48.8% | +0.01% |
Does the 3-model consensus beat a single model?
Barely — consensus lands mid-pack; gpt-4.1-nano alone edges it. Combining didn’t help here.
When all three agree, is it more reliable?
Only a hair: unanimous 49.2% (2,525 trades) vs split 48.0% (1,348).
When and where does it work?
Does it work better at the open or the close?
The morning wins. Open trades (R1) hit 53.0% (+0.19%); afternoon trades (R2) only 47.0% (-0.07%). The clearest signal on the page.
Big companies or small?
Mixed, and small samples per bucket:
| trades | hit rate | return/trade | |
|---|---|---|---|
| Mega | 401 | 50.6% | -0.09% |
| Large | 1238 | 49.4% | -0.02% |
| Mid | 850 | 49.3% | -0.01% |
| Small | 586 | 48.5% | +0.12% |
| Micro | 448 | 47.1% | -0.03% |
| Unknown | 350 | 46.0% | +0.11% |
Do the law-firm “investor alert” spam PRs drag it down?
A little. Dropping the 227 flagged solicitations lifts the average from +0.01% to +0.04% per trade.
The one bet we’re actually testing
Frozen rule (pre-registered 2026-07-02): Short large-cap (Mega/Large) bad-news calls, morning round — panel consensus
We locked this the instant the search suggested it, and we do not retune it. The in-sample number is what generated the rule, so on its own it means nothing. The out-of-sample number — fresh days the rule has never seen — is the only one that counts.
in-sample (cherry-picked — ignore)60.7%
OUT-OF-SAMPLE — the real test—
Could we really have done it?
Could we actually have traded these in time?
These numbers already use only the 3,873 trades whose decision existed before the entry — the honest, no-look-ahead set. Including the late ones would add ~118 more.
Anything still unsettled?
247 trades from the last collected day are waiting on the next day’s close — shown as “pending,” never counted as a result.
Look at the trades yourself
Let me look at the actual trades.
Open the full list — search any ticker, sort by biggest win/loss, filter to the morning trades or the unanimous calls, and click any trade to see the news + each model’s words + the price math, with links out to the article and the stock.
Can I trust the numbers?
What is this based on, and can I trust it?
4,187 settled common-stock trades across 13 days, priced on split/dividend-adjusted daily bars. Round membership is the news’s publish time, and every trade’s decision is verified to pre-date its entry — no look-ahead.