
From DowJo — where this research becomes practice. Train your judgment before you risk your money.
Course Research · Episode 13
Can AI Become a Better Investor Than You?
The evidence-led takeaway
Markets pay everyone for carrying risk. They pay extra only for being right about something the price does not already know.
Jo reads this note before answering. Sign in to ask Jo with this evidence attached — you come straight back here.
In brief
What this research concludes
Markets pay everyone for carrying risk. They pay extra only for being right about something the price does not already know.
The judgment skill this hands over
When someone says AI will out-invest you, ask four things: better at what (one task, or the whole chain of goal, reading, forecasting, sizing and trading); tested on what (a past it had already read); what happens when everyone has it; and better than whom (you, a professional, or an index fund you could already own).
What we investigated
The question, and who it looks at
This is episode 13 of Investing in the AI Era, a 14-episode course. The film tells the story; this note is the research behind it, kept inspectable.
Who and what this looks at
How we tested it
Researched to be refuted, not confirmed
Candidate claims went to an independent pass instructed to refute them against primary sources. Survivors became evidence; casualties became the refused list below.
The method, in the research team's own words
Pre-registered five competing hypotheses with falsifiers and committed them (a2e387e) before any research. One wave of six bounded primary-source lanes, one agent each, no retries (run wf_2d4e3e5a-5bd): track record, forecasting vs returns, overfitting/crowding/decay/regime, execution and market structure, contemporary AI in real functions (frontier), and the baseline (index funds, behaviour, human + machine). 128 claims, 53 famous examples tested, 63 bounded not-founds. The producer then spent six direct primary checks (Virtu S-1, M6 paper, M6 team file reproduction, Lopez-Lira & Tang v6, McLean & Pontiff abstract, SPIVA). The sentences to be spoken were refuted by four independent agents in one pass and adjudicated before any TTS. Verdict vocabulary: SUPPORTED / CORRECTED / MISLEADING / UNSUPPORTED / DO-NOT-NARRATE.
Pre-registered before any data was pulled
The firms, metrics and window for the quantitative work were fixed in writing before a single number was requested, so the result could not be cherry-picked after the fact. Every pre-registered case is reported regardless of direction. The full pre-registration is in the research desk.
What the evidence says
33 claims survived refutation
Each carries its source, the date the thing happened and the date it was said — two different facts — and its limits, stated by the research team rather than left for you to discover.
- FactPrimary sourceF01
In its 2014 Form S-1, Virtu Financial reported only one losing trading day in the period 1 January 2009 to 31 December 2013, a total of 1,238 trading days.
Source: Virtu Financial, Inc. — Form S-1 registration statement (opens sec.gov)
Event 1 Jan 2009 · Published Mar 2014
Limits: Measured on daily Adjusted Net Trading Income, a non-GAAP measure; 1 January 2009 to 31 December 2013; includes Madison Tyler before July 2011. The S-1/A of 27 March 2014 extended the window to 1,278 days and the 2015 prospectus to 1,485, each with one losing day; the 2015 prospectus adds that the firm profitably exited only 49% of its positions in 2014.
- FactPrimary sourceF02
Virtu described its business as avoiding long or short positions in favour of earning small bid/ask spreads on large trading volumes across thousands of securities and other financial instruments.
Source: Virtu Financial, Inc. — Form S-1 registration statement (opens sec.gov)
Event 2014 · Published Mar 2014
Limits: The filing says more than 10,000 securities and other instruments; 'thousands' is conservative.
- FactPrimary sourceF03
Virtu stated it does not engage in the types of principal investing and predictive, momentum and signal trading in which many other broker-dealers and trading firms engage, and described its trading as non-directional, non-speculative and market neutral.
Source: Virtu Financial, Inc. — Form S-1 registration statement (opens sec.gov)
Event 2014 · Published Mar 2014
Limits: Self-description. The same filing says a market maker's success depends on posting the best prices and responding to market data, so the film says 'not betting on which way prices would go', not 'not forecasting anything' (N05 MISLEADING in refutation).
- FactResearchF04
The M6 competition was a live forecasting competition lasting twelve months from February 2022, in which participants forecast the rankings of 50 US stocks and 50 international ETFs and submitted notional portfolio weights on them, scored on real prices, for $300,000 in prizes.
Event Feb 2022 · Published Oct 2023
Limits: Investment decisions were competition submissions scored on information ratio, not real money under management.
- FactResearchF05
M6 scored forecasts by ranked probability score (RPS) and investment decisions by information ratio (IR), a measure of return per unit of risk.
Event 20 2022 · Published Oct 2023
Limits: IR here is the competition's own risk-adjusted return measure.
- FactResearchF06
Of the 162 teams on the M6 Global leaderboard (163 rows including the organisers' benchmark entry), 38 forecast more accurately than the benchmark, 47 constructed better portfolios, and 11 did both.
Event 20 2022 · Published Oct 2023
Limits: The paper's performance sentence says 163 teams because it counts the benchmark row; its H2 analysis states 162 with the benchmark excluded. Verified in the organisers' file (id 32cdcc24, RPS 0.16, IR 0.45347). Portfolios were notional.
- FactPrimary sourceF07
Before launching the competition the M6 organisers stated ten hypotheses, including Hypothesis No.3: that there would be a weak link between teams' forecasting accuracy and their risk-adjusted investment returns.
Event Feb 2022 · Published 2023
Limits: The repository readme describes "the ten hypotheses of the M6 made before launching the competition".
- FactResearchF08
Across the 138 M6 Global teams whose forecasts were not identical to the benchmark, the organisers found no connection between forecasting accuracy and investment performance (r = 0.04).
Event 20 2022 · Published Oct 2023
Limits: Reproduced: Pearson 0.0448, Spearman 0.034. The paper's subgroup result (r = 0.7 among the top 20% by overall rank) is not used: with RPS lower-is-better it reflects selection on a combined rank, not a link among the best teams. The abstract's own phrase is 'limited connection'.
Show the remaining 25 pieces of evidence
- FactResearchF09
In M6, the top-performing forecasting teams constructed relatively inefficient portfolios on average, while the top-performing investment teams submitted forecasts of various accuracy levels.
Event 20 2022 · Published Oct 2023
Limits: Not narrated: the reading rests on correlation coefficients whose sign is ambiguous with a lower-is-better RPS (refutation R1). Replaced in the film by F32.
- FactResearchF10
The M6 organisers found most participants were measurably overconfident: they assumed much more investment risk than was justified by the accuracy of their forecasts.
Event 20 2022 · Published Oct 2023
Limits: Measured with the organisers' investment risk model; participants were competition entrants, not retail investors.
- FactResearchF11
Using headlines published after the model's knowledge cutoff, Lopez-Lira and Tang report that GPT-4 scores significantly predict the subsequent drift in stock returns, especially for small stocks and negative news.
Event Oct 2021 · Published 28 Oct 2025
Limits: Working paper, not peer-reviewed as of v6. The ~90% hit rate applies to the non-tradable initial reaction and is not narrated.
- FactResearchF12
Before trading costs, the annualised Sharpe ratio of Lopez-Lira and Tang's equal-weighted GPT-4 overnight-news long-short strategy declined in each of four subperiods, from 6.54 in 2021Q4 to 3.68 in 2022, 2.33 in 2023 and 1.22 over January-May 2024, which the authors describe as suggestive evidence consistent with LLM adoption improving market efficiency.
Event Oct 2021 · Published 28 Oct 2025
Limits: Figure 8 uses the Figure 3 strategy, 'Without Transaction Costs'. One backtest sample split into unequal periods; confidence intervals not inspected; the intraday-news strategy shows 'less clear evidence of a decline'. The paper is now published in the Journal of Financial Economics (vol. 184, 2026); figures were checked on arXiv v6.
- FactResearchF13
At a transaction cost of 20 basis points per round trip, Lopez-Lira and Tang's strategy is unprofitable over their sample period (October 2021 to May 2024); at 10 basis points it still earns over 100% cumulatively.
Event Oct 2021 · Published 28 Oct 2025
Limits: Cumulative, not risk-adjusted; costs modelled as fixed round-trip basis points at auctions.
- FactResearchF14
Across 97 variables shown to predict cross-sectional stock returns, McLean and Pontiff found portfolio returns 26% lower out of sample and 58% lower post-publication, and conclude that investors learn about mispricing from academic publications.
Event through 2013 · Published Feb 2016
Limits: Average declines; the out-of-sample decline is an upper bound on data mining, and 32% is attributed to publication-informed trading. Jacobs and Muller (JFE 2020) find a reliable post-publication decline only in the United States.
- FactResearchF15
A 2025 working paper by Lopez-Lira, Tang and Zhu finds that LLMs recall exact economic and financial values from before their training cutoff and that instructions to respect historical boundaries do not prevent it.
Event 2025 · Published 15 Dec 2025
Limits: Working paper; model coverage per the paper.
- FactFrontier · Sep 2026Primary sourceF16
The arXiv version of Kim, Muhn and Nikolaev's "Financial Statement Analysis with Large Language Models", which claimed GPT-4 outperformed financial analysts at predicting earnings changes, was withdrawn on 2025-02-20 after a co-author identified inconsistencies in the data and analyses.
Event 20 Feb 2025 · Published 20 Feb 2025
Limits: Still withdrawn on arXiv as of 13 September 2026; the authors' pages say 'temporarily withdrawn'. A different Kim and Nikolaev paper was retracted by the Journal of Accounting Research in 2026 — not narrated.
- FactResearchF17
FINSABER (Li et al., KDD 2026) re-tested two previously reported LLM trading agents (FinMem and FinAgent) over 2004-2024 and 100+ symbols with transaction costs and found previously reported LLM advantages deteriorate significantly.
Event 2004/2024 · Published 26 Jun 2026
Limits: Backtest, not live money; timing strategies, not portfolio construction.
- FactFrontier · Sep 2026Primary sourceF18
On 2026-07-16 the Forecasting Research Institute reported that AI models have likely reached parity with superforecasters on ForecastBench, with several models statistically indistinguishable from superforecaster-level accuracy; its superforecaster predictions were last elicited in 2024 and a fresh round is planned for fall 2026.
Event Mar 2026 · Published 4 Mar 2026
Limits: World-event forecasting, not investment returns. The comparison relies on a statistical extrapolation of 2024 human forecasts; some top AI entries are crowd-adjusted. Earlier (2026-03-04): superforecasters 70.6 vs best AI 67.9 on the Brier Index.
- FactResearchF19
Cao, Jiang, Wang and Yang (Journal of Financial Economics, 2024) report that an AI analyst trained on corporate disclosures, industry trends and macroeconomic indicators surpasses most analysts in stock return predictions, by a narrow margin per prediction.
Event 2021/2024 · Published 1 Oct 2024
Limits: Abstract verified by the producer. Margin: 53.7% of target-price predictions in NBER w28800 (2001-2016); 54.5% (2001-2018) for the published version per a University of Maryland summary, not spoken. A machine-learning ensemble, not a generative model.
- FactFrontier · Sep 2026Primary sourceF20
In AIMA's September 2025 survey of 150 hedge fund managers running about $788 billion, 95% said they use generative AI in their work.
Event 2025 · Published 16 Sep 2025
Limits: Self-reported adoption, not results; "use" includes administrative and research tasks.
- FactResearchF21
Cao et al. find the AI analyst wins when information is transparent but voluminous, while human analysts win where institutional knowledge is crucial, such as intangible assets and financial distress.
Event 2021/2024 · Published 1 Oct 2024
- FactResearchF22
Cao et al. find that combining human analysts with the AI adds significant value and substantially reduces extreme errors.
Source: Cao, Jiang, Wang, Yang — Journal of Financial Economics 160 (2024) (opens repository.lsu.edu)
Event 2021/2024 · Published 1 Oct 2024
Limits: Size of the gains not extracted.
- FactResearchF23
Cao et al. find analysts caught up with the machine after alternative data arrived only where their employers had built AI capabilities.
Source: Cao, Jiang, Wang, Yang — Journal of Financial Economics 160 (2024) (opens repository.lsu.edu)
Event 2021/2024 · Published 1 Oct 2024
Limits: 'Only' rests on statistical significance: improvement concentrates in, and is only significant for, brokerages that built AI capability.
- FactFrontier · Sep 2026JournalismF24
S&P Dow Jones Indices' SPIVA U.S. Scorecard reports that 79% of all active large-cap U.S. equity funds underperformed the S&P 500 in 2025.
Event 2025 · Published Mar 2026
Limits: The primary page blocked scripted access; the sentence was recovered from the search index of the primary page and from two outlets quoting it.
- FactResearchF25
Rossi and Utkus (Journal of Financial Economics, 2024), studying previously self-directed Vanguard investors who joined its hybrid Personal Advisor Services in 2015-2017, find that the service increased indexing and reduced home bias, the number of assets held, and fees.
Event 20 2015 · Published Sep 2020
Limits: Co-author Stephen Utkus is a Vanguard employee. The refuter read the published figures (expense ratios 23 to 10 basis points; average cash 19% to 2%, median 4%, measured the month before sign-up); the producer could not reach the published text directly, so no figure is spoken or shown. The 2020 working paper's figures (19 to 9 basis points; cash down 19 points) are superseded.
- FactFrontier · Sep 2026Primary sourceF26
Morningstar's 2026 Mind the Gap study estimates the average dollar in US mutual funds and ETFs earned about 1.2 percentage points a year less than the funds themselves over the ten years to 2025.
Source: Morningstar, Mind the Gap (US, 2026 edition) (opens morningstar.com)
Event 20 2016 · Published 6 Aug 2026
Limits: An IRR on pooled flows; cannot separate discretionary mistiming from the timing of regular saving.
- FactResearchF27
Fulkerson, Jordan, Riley and Yan (Financial Analysts Journal, 2026), using Morningstar's 2015-2024 sample, estimate that poor timing costs mutual fund investors about 0.10% a year.
Event 20 2015 · Published 12 May 2026
Limits: Only the abstract and summary were read in the harvest.
- Historical analogyPrimary sourceF28
In its April 1999 report on the 1998 crisis, the President's Working Group on Financial Markets concluded that the size, persistence and pervasiveness of the widening of risk spreads confounded the risk management models employed by LTCM and other participants.
Source: President's Working Group on Financial Markets (April 1999) (opens cftc.gov)
Event Aug 1998 · Published 28 Apr 1999
Limits: A pre-AI model-driven fund; used only as historical evidence that similar models can be wrong together.
- FactFrontier · Sep 2026Primary sourceF29
The Bank of England's Financial Policy Committee describes correlated AI-driven trading as a potential future risk, and the Financial Stability Board's October 2025 monitoring report says there is still little empirical evidence that AI-driven market correlations affect market outcomes.
Published 9 Apr 2025
Limits: A statement of concern, not a measurement.
- FactFrontier · Sep 2026Primary sourceF30
ForecastBench's operators extrapolated AI-superforecaster parity to November 2026 on 2026-01-29 and to May 2027 on 2026-03-04.
Event Jan 2026 · Published 29 Jan 2026
Limits: Not narrated: part of the six-month move was a change of scoring rule that FRI called a modelling artefact, and by July 2026 FRI judged parity already likely (F18).
- FactPrimary sourceF32
Of the 38 M6 Global teams that forecast more accurately than the benchmark, only 11 also constructed a better portfolio than the benchmark.
Event 20 2022 · Published 2023
Limits: Computed by the producer from the organisers' file (RPS < 0.16 and IR > 0.45347); matches the paper's 'just 11 teams report better IR and RPS scores than the benchmark'.
- FactResearchF33
The M6 organisers found that a limited number of teams actually chose to exploit their forecasts to define the weights of their investments.
Event 20 2022 · Published Oct 2023
Limits: Verified by the producer in the paper text. The organisers add that accurate forecasts, properly risk-sized, can improve investment decisions.
- InterpretationFrontier · Sep 2026Primary sourceF31
Man Group describes its AlphaGPT signal-research system in detail but discloses neither how many AI-generated signals trade live nor any returns attributable to them.
Source: Man Group, 'What AI Can (and Can't Yet) Do for Alpha' (opens man.com)
Event 2025 · Published 13 Nov 2025
Limits: One checkable example; the broader generalisation that best-placed firms mostly do not publish AI-attributable returns was UNSUPPORTED and is not spoken.
3 further audited series are recorded but not drawn
M6 Global teams: RPS and IR
Source: M6DATA
Lopez-Lira & Tang overnight-news strategy annualised Sharpe
Source: LLT
McLean & Pontiff predictor returns indexed
Source: MP
What surprised us
Where the simple story did not survive
Logged by the research lanes as they worked, before anything was written. The full log is in the research desk.
- 1
AIEQ, the pioneering AI stock ETF, holds only about $120M after nearly nine years, and over the latest year trailed the AI index it tracks by 0.84 points, close to its 0.75% fee. (Lane: L1-track-record)
- 2
The SEC's first 'AI washing' cases were not about AI that failed; they found firms had made false or misleading claims about using AI at all. (Lane: L1-track-record)
- 3
In 2020, the same year Medallion reportedly gained 76%, Renaissance's funds open to outside investors lost about 22.6% and 33.6%. (Lane: L1-track-record)
- 4
Medallion's legendary return of over 60% a year is an arithmetic, pre-fee figure from a book; a 2024 reconstruction puts the compounded pre-fee return probably under 35%. (Lane: L1-track-record)
- 5
Bridgewater's AI fund reportedly matched, rather than beat, its flagship in H1 2026 (8.1% each), and neither Bridgewater nor the press has disclosed how much of its decision-making is automated. (Lane: L1-track-record)
The strongest case against this
The evidence against our own conclusion
Carried at full strength, before the conclusion — not as a footnote.
- F18
- F19
- F20
- F14
Statement: The machines are not standing still.
Concession: The best AI has likely drawn level with the best human forecasters (FRI, July 2026), an AI analyst beat most human analysts narrowly, published edges do not vanish on average, and adoption is near-universal. The film's reply is built from the last of these: what everyone uses, the price learns.
History, under test
The strongest historical case
What the past licenses — and, stated just as plainly, what it does not.
M6 forecasting competition, February 2022 to February 2023
- F04
- F05
- F06
- F07
- F08
- F10
Lesson: Being right about what will happen and being paid for it are two jobs, rarely joined; most teams sized their bets beyond their accuracy.
What would change our mind
The conclusion is wrong if…
A widely available AI tool that keeps beating the index after costs and for the risk it took, on years it had never read about, while more and more people use it.
What remains uncertain
What we still do not know
Stated by the research team, in full, rather than smoothed over.
- Whether similar AI models make firms wrong together: described as a potential risk by the Bank of England and the FSB; no measurement found (F29).
- Which organisations have a real machine edge: this record found no disclosed AI-attributable returns to test; the generalisation that the best-placed firms do not publish was UNSUPPORTED in refutation and is not spoken (F31; one checkable example, Man Group, is).
- How long the forecasting gap stays open, and whether parity would show up in returns (F18, F30, F08).
- How much of the investor return gap is bad timing: 1.2 points (Morningstar) vs 0.10% (FAJ 2026) on overlapping data (F26, F27).
- Whether the Lopez-Lira & Tang decline is caused by adoption: the authors call it suggestive (F12).
- H3 (organisations win) for returns: open, as pre-registered, because of non-disclosure.
- Objectives, mandates and career risk: no harvested source verified (Cremers & Petajisto could not be opened).
What we refused to publish
20 claims we would not say — and why
The do-not-narrate list. Some are popular; some are true but unproven; some are simply not this note’s to make. Each refusal is enforced in production, not just recorded.
- UnverifiableX01
“Renaissance Medallion as proof that machines or AI beat human investors; its "66% a year" return.”
Verdict: UNVERIFIABLE — record undisclosed
Why: Closed to outsiders since 1993; figures from a journalist's book and anonymous investors; the headline is arithmetic and gross of fees, and a 2024 reconstruction puts the compounded pre-fee return probably under 35%. Renaissance's outside-investor funds lost ~22.6% and ~33.6% in 2020 per a third-party scoreboard (L1-09..13).
- UnverifiableX02
“An AI-run fund or ETF has beaten the market (e.g. AIEQ since 2017).”
Verdict: UNVERIFIABLE — no primary benchmark comparison opened
Why: The issuer reports 10.08%/yr at NAV since inception but benchmarks against its own AI index; the fund changed structure; no audited autonomous AI record was found (L1-02..05, L5-18, L5-21).
- RefutedX03
“GPT-4 beats human analysts at financial statement analysis (Kim, Muhn, Nikolaev).”
Verdict: REFUTED as established evidence — withdrawn
Why: arXiv version withdrawn 2025-02-20 over data inconsistencies (L2-10, L5-07). Only the withdrawal may be narrated.
- RefutedX04
“ChatGPT predicts stock moves with ~90% accuracy.”
Verdict: REFUTED as a tradable claim
Why: The ~90% is a portfolio-day hit rate for the non-tradable initial reaction (L2-11).
- RefutedX05
“Most / half of all trading is algorithmic or high-frequency (as an SEC measurement).”
Verdict: REFUTED as sourced
Why: The SEC's 2010 "50% or higher" line footnotes press articles; ESMA measured 24-43% of value traded and 58-76% of orders in EU equities, May 2013, depending on definition (L4-14, L4-15).
- RefutedX06
“DALBAR's investor behaviour gap.”
Verdict: REFUTED as a measure of bad timing
Why: Compares dollar-weighted investor returns with a time-weighted index, conflating the sequence of market returns with behaviour (L6-09, L6-10).
Show the remaining 14 refused claims
- RefutedX07
“The 2010 Flash Crash or Knight Capital 2012 as failures of intelligent or AI trading systems.”
Verdict: REFUTED
Why: A volume-only execution algorithm "without regard to price or time"; a deployment error reactivating code unused since 2003 (L4-01..08).
- RefutedX08
“LTCM as an AI or machine-learning failure.”
Verdict: REFUTED
Why: A leverage and liquidity failure of a model-driven relative-value fund in 1998 (L3-11..13).
- RefutedX09
“Regulators have found AI is already causing herding in markets; the SEC's 2023 predictive-analytics proposal was a herding rule.”
Verdict: REFUTED
Why: BoE and FSB describe potential risk; the SEC proposal addressed conflicts of interest and was withdrawn on 2025-06-12 (L3-16..21).
- Partly unverifiableX10
“Centaur (human + engine) chess shows human-plus-machine teams will dominate investing.”
Verdict: PARTIALLY UNVERIFIABLE — untested analogy
Why: The 2005 result is real but a single event; later engine comparisons move toward parity; markets adapt and chess does not (L6-21, L6-22).
- RefutedX11
“Alpha Arena proved LLMs can trade profitably with real money.”
Verdict: REFUTED as evidence of skill
Why: Six models, $10,000 each, ~16 days, one market path, four of six lost (L5-08, L5-09).
- UnverifiableX12
“Bridgewater's AI-driven fund returns (e.g. 11.3%/yr, 8.1% in H1 2026, 11.9% in 2025).”
Verdict: UNVERIFIABLE — anonymous sources
Why: Reported by Reuters from anonymous sources; the 11.9% figure was not found in the opened article (L1-17..19).
- UnverifiableX13
“NBIM's AI saves ~$100m a year / will save $400m; AI made NBIM 20% more productive.”
Verdict: UNVERIFIABLE
Why: The primary statement says "several hundred million kroner" in trading costs for a project using AI among other tools; the 20% is an internal self-report survey (L4-21, L4-22, L5-14).
- Partly unverifiableX14
“"Bad timing costs fund investors ~15% of their funds' returns" as settled.”
Verdict: PARTIALLY UNVERIFIABLE — disputed
Why: Morningstar's 2026 edition says ~12%; an FAJ 2026 paper on the same sample puts the timing cost near 0.10%/yr (L6-04..08). Narrated only as a dispute.
- UnverifiableX15
“Machine learning doubles the performance of stock-picking strategies (Gu, Kelly, Xiu).”
Verdict: UNVERIFIABLE — figures not verified
Why: Exact Sharpe figures not verified; Avramov et al. show profits concentrate in hard-to-arbitrage stocks and weaken after costs (L2-06..09).
- RefutedX16
“High Active Share reliably predicts outperformance.”
Verdict: REFUTED — disputed
Why: Frazzini, Friedman and Pomorski (FAJ 2016) find it correlates with benchmark returns rather than predicting fund returns (L6-20).
- Partly unverifiableX17
“Sentient's AI hedge fund closure as evidence that AI funds fail.”
Verdict: PARTIALLY UNVERIFIABLE — thin and anecdotal
Why: Anonymous Bloomberg sourcing; one closure is not a failure rate (L1-08).
- UnverifiableX18
“The value factor is dead; ML models predicted the March 2020 crash.”
Verdict: UNVERIFIABLE
Why: Value's 55% drawdown is documented but its death is disputed; the only COVID-prediction hit was unopened and retrospective (L3-14, L3-15).
- UnresolvableX19
“FinanceBench's 81% failure rate (2023) as a description of current models.”
Verdict: UNRESOLVABLE — dated benchmark; MUST BE STATED AS HISTORICAL
Why: A 2023 configuration on a 150-question sample (L5-01). Not narrated: benchmark scores are not investment evidence (founder brief).
- UnresolvableX20
“Quantum computing as part of the AI investing question.”
Verdict: UNRESOLVABLE IN THIS EPISODE — coverage item not carried
Why: Listed ADJACENT in the coverage matrix; the founder's 2026-09-13 brief defines the episode and does not ask for it. Named in the founder review.
What to watch next
Dated material, and what would make it stale
ForecastBench AI vs superforecaster gap
4 Mar 2026Status: best AI 67.9 vs superforecasters 70.6 (Brier Index); parity projected May 2027
What changes: S06 line one and N39; N56
SPIVA U.S. Scorecard
Mar 2026Status: Year-End 2025: 79% of active large-cap funds lagged the S&P 500
What changes: N49 and S08 line one at Year-End 2026
AIMA generative AI adoption survey
16 Sep 2025Status: 95% of 150 managers use generative AI
What changes: N41
Kim, Muhn & Nikolaev paper status
20 Feb 2025Status: arXiv version withdrawn
What changes: N36 and the S05 card if a corrected version is published
Morningstar Mind the Gap
6 Aug 2026Status: ~1.2 points a year
What changes: N52 at the next edition
Evidence & sources
20 sources, by tier
Tier 1 is primary and authoritative — filings, regulators, official statistics. Journalism and books are attributed ingredients, never proof by reputation.
- Primary sourceVirtu Financial, Inc. — Form S-1 registration statement (opens sec.gov)supports F01, F02, F03
- ResearchMakridakis, Spiliotis, Hollyman, Petropoulos, Swanson & Gaba — The M6 forecasting competition: Bridging the gap between forecasting and investment decisions (International Journal of Forecasting; arXiv 2310.13357) (opens arxiv.org)supports F04, F05, F06, F08, F09, F10, F33
- Primary sourceM6 organisers — Mcompetitions/M6-methods, "IJF paper/performance_vs_answers_unique.csv" (sha256 2d69575c457cbf5b977746327c6468bfbe367dd8e199e885eab73668cb2dab36) (opens github.com)supports F07, F32
- ResearchLopez-Lira & Tang — Can ChatGPT Forecast Stock Price Movements? Return Predictability and Large Language Models (arXiv 2304.07619v6) (opens arxiv.org)supports F11, F12, F13
- ResearchMcLean & Pontiff — Does Academic Research Destroy Stock Return Predictability? Journal of Finance 71(1), 5-32 (abstract via Crossref) (opens api.crossref.org)supports F14
- ResearchLopez-Lira, Tang, Zhu — The Memorization Problem: Can We Trust LLMs' Economic Forecasts? (arXiv 2504.14765v2) (opens arxiv.org)supports F15
- Primary sourcearXiv record 2407.17866v3 — Kim, Muhn, Nikolaev, Financial Statement Analysis with Large Language Models (withdrawal notice by Valeri Nikolaev) (opens arxiv.org)supports F16
- ResearchLi, Kim, Cucuringu, Ma, Can LLM-based Financial Investing Strategies Outperform the Market in Long Run? (FINSABER; arXiv 2505.07078; KDD 2026) (opens arxiv.org)supports F17
- Primary sourceForecasting Research Institute — Making Forecasting Scores Easier to Interpret: Introducing the Brier Index (opens forecastingresearch.substack.com)supports F18
- ResearchCao, Jiang, Wang, Yang — From Man vs. Machine to Man + Machine: The art and AI of stock analyses, Journal of Financial Economics 160 (2024) (opens repository.lsu.edu)supports F19, F21
- Primary sourceAIMA press release, 'Front-office Gen AI adoption shifts from if to when for leading fund managers' (opens aima.org)supports F20
- ResearchCao, Jiang, Wang, Yang — Journal of Financial Economics 160 (2024) (opens repository.lsu.edu)supports F22, F23
- JournalismTKer (Sam Ro), reporting SPIVA U.S. Scorecard Year-End 2025; corroborated by InvestmentNews (2026-03-04) (opens tker.co)supports F24
- ResearchRossi & Utkus, 'Who Benefits from Robo-advising? Evidence from Machine Learning' (working paper draft hosted by GFLEC) (opens gflec.org)supports F25
- Primary sourceMorningstar, Mind the Gap (US, 2026 edition) (opens morningstar.com)supports F26
- ResearchFulkerson, Jordan, Riley, Yan, 'Bad Timing Does Not Cost Investors 15% of Their Funds' Returns', Financial Analysts Journal 82(3) (opens rpc.cfainstitute.org)supports F27
- Primary sourcePresident's Working Group on Financial Markets (April 1999) (opens cftc.gov)supports F28
- Primary sourceBank of England, Financial Stability in Focus: Artificial intelligence in the financial system (April 2025) (opens bankofengland.co.uk)supports F29
- Primary sourceForecasting Research Institute (ForecastBench operators) — LLMs Are Closing the Gap on Human Superforecasters (opens forecastingresearch.substack.com)supports F30
- Primary sourceMan Group, 'What AI Can (and Can't Yet) Do for Alpha' (opens man.com)supports F31
Complete evidence depth
The research desk
Everything the research evaluated before it was distilled — including what it rejected, and why.
The research desk
Before the evidence above was distilled, the research evaluated 128 candidate claims across 0 lanes, rejected 20 with a recorded reason, and logged 38 surprises. A rejected claim with a reason is the most reusable thing research produces — the desk keeps all of them.
- laneRenaissance Medallion as proof that machines or AI beat human investors; its "66% a year" return. (Id: X01; Verdict: UNVERIFIABLE — record undisclosed; Reason: Closed to outsiders since 1993; figures from a journalist's book and anonymous investors; the…
- laneAn AI-run fund or ETF has beaten the market (e.g. AIEQ since 2017). (Id: X02; Verdict: UNVERIFIABLE — no primary benchmark comparison opened; Reason: The issuer reports 10.08%/yr at NAV since inception but benchmarks against its own AI index; the fund…
Save this research and keep following the story — you come straight back to this desk.
Community
Talk it through
Discuss this research
Members only · in 📡 Intelligence Watch
Challenge the argument, bring evidence, or ask what would change our mind. This note is discussed in 📡 Intelligence Watch, a DowJo room where members compare reads. Filings, news events, and what they actually mean — the discussion room for the intelligence feed.
Read and reply, then keep following the story — you come straight back to this note.
Education-first ground rules apply in every room: mechanisms and evidence, never calls.
Now train it
Take it further
Reading is where judgment starts. Practice is where DowJo measures it — and remembers what to train next.
Train it
Make five honest calls on today's market
Commit before the outcome exists, then see what happened. Practice is where this research gets measured.
Train itTrain it
Think you can spot it before the reveal?
Train it on a chart: an unseen setup, revealed bar by bar, graded on your process.
Sign in to train this setupLearn the concepts underneath
Join the discussion
Challenge the argument, bring evidence, or ask what would change our mind — in the open, with other members.
MentorAsk Jo what could invalidate this thesis.
Jo reads this note before answering. Sign in to ask, with this evidence attached — you come straight back here.
Ask Jo with this evidence attachedProvenance — where this note comes from
This page is a deterministic projection of canonical research artifacts. It adds presentation and discovery; it never adds, removes or softens a finding. Question taken from the title; thesis from the packet. Projected 21 Sep 2026 by dowjo-research-projector v1.1.0 from origin commit c83c8797ab79.
| Artifact | Path (origin-relative) | Identity | Version |
|---|---|---|---|
| Episode registry | course/episodes.json | sha256 75b65b1618fc… | — |
| Evidence packet | course/research/AI13-evidence-packet.json | sha256 eb797367cbd9… | compiled 2026-09-13 · rev 1 |
| Research shortlist | course/research/AI13-research-shortlist.json | sha256 cda1a543d991… | — |
| Approved script | course/scripts/AI13-what-the-price-already-knows.json | sha256 4fabd85fbbec… | v draft2 |
| Episode manifest | course/manifest/AI13.manifest.json | sha256 88ea06c0f331… | — |
| Approval record | renders/approved/AI13-FINAL.json | sha256 2b333bfe49a8… | compiled 2026-09-13 |
| Pre-registration | course/research/AI13-lanes/PR0-thesis-preregistration.md | sha256 237adcb65b70… | — |
| Poster still | qa/stills-AI13/RESEARCH-POSTER.png | sha256 307d99fc6e7e… | — |
Approved master (AI13-final02-corrected.mp4): sha256 c131855e476787e9fc9d37dae5c053468f0c70555d78dd2ce80403dd3fedcdb0
Commercial disclosure: none. No affiliate relationship, review copy, sponsorship or publisher relationship applies to this note.
For education only. Not financial advice. No buy, sell or hold recommendations, ever.


