DowJoResearch
A data-centre corridor between rows of lit server racks with one person walking away down the aisle; the title reads: One losing trading day in 1,238. It said it did not engage in predictive trading.
Episode 13 · Investing in the AI Era · 10 min film
You may have seen the film. This is the research behind it — every claim, every refusal, kept inspectable.

From DowJo — where this research becomes practice. Train your judgment before you risk your money.

Course Research · Episode 13

Can AI Become a Better Investor Than You?

The evidence-led takeaway

Markets pay everyone for carrying risk. They pay extra only for being right about something the price does not already know.

Make five honest calls on today's market Jo portraitAsk Jo about this research

Jo reads this note before answering. Sign in to ask Jo with this evidence attached — you come straight back here.

Researched Current33 pieces of evidence20 claims refused20 sources

In brief

What this research concludes

Markets pay everyone for carrying risk. They pay extra only for being right about something the price does not already know.

The judgment skill this hands over

When someone says AI will out-invest you, ask four things: better at what (one task, or the whole chain of goal, reading, forecasting, sizing and trading); tested on what (a past it had already read); what happens when everyone has it; and better than whom (you, a professional, or an index fund you could already own).

What we investigated

The question, and who it looks at

This is episode 13 of Investing in the AI Era, a 14-episode course. The film tells the story; this note is the research behind it, kept inspectable.

Who and what this looks at

Virtu FinancialMan GroupVanguardMorningstarS&P Dow Jones IndicesBank of EnglandForecasting Research InstituteAlternative Investment Management AssociationSpyros MakridakisAlejandro Lopez-LiraSean CaoR. David McLean and Jeffrey PontiffAlberto Rossi and Stephen Utkus

How we tested it

Researched to be refuted, not confirmed

Candidate claims went to an independent pass instructed to refute them against primary sources. Survivors became evidence; casualties became the refused list below.

The method, in the research team's own words

Pre-registered five competing hypotheses with falsifiers and committed them (a2e387e) before any research. One wave of six bounded primary-source lanes, one agent each, no retries (run wf_2d4e3e5a-5bd): track record, forecasting vs returns, overfitting/crowding/decay/regime, execution and market structure, contemporary AI in real functions (frontier), and the baseline (index funds, behaviour, human + machine). 128 claims, 53 famous examples tested, 63 bounded not-founds. The producer then spent six direct primary checks (Virtu S-1, M6 paper, M6 team file reproduction, Lopez-Lira & Tang v6, McLean & Pontiff abstract, SPIVA). The sentences to be spoken were refuted by four independent agents in one pass and adjudicated before any TTS. Verdict vocabulary: SUPPORTED / CORRECTED / MISLEADING / UNSUPPORTED / DO-NOT-NARRATE.

Pre-registered before any data was pulled

The firms, metrics and window for the quantitative work were fixed in writing before a single number was requested, so the result could not be cherry-picked after the fact. Every pre-registered case is reported regardless of direction. The full pre-registration is in the research desk.

What the evidence says

33 claims survived refutation

Each carries its source, the date the thing happened and the date it was said — two different facts — and its limits, stated by the research team rather than left for you to discover.

Show the remaining 25 pieces of evidence
3 further audited series are recorded but not drawn
  • M6 Global teams: RPS and IR

    Source: M6DATA

  • Lopez-Lira & Tang overnight-news strategy annualised Sharpe

    Source: LLT

  • McLean & Pontiff predictor returns indexed

    Source: MP

What surprised us

Where the simple story did not survive

Logged by the research lanes as they worked, before anything was written. The full log is in the research desk.

  1. 1

    AIEQ, the pioneering AI stock ETF, holds only about $120M after nearly nine years, and over the latest year trailed the AI index it tracks by 0.84 points, close to its 0.75% fee. (Lane: L1-track-record)

  2. 2

    The SEC's first 'AI washing' cases were not about AI that failed; they found firms had made false or misleading claims about using AI at all. (Lane: L1-track-record)

  3. 3

    In 2020, the same year Medallion reportedly gained 76%, Renaissance's funds open to outside investors lost about 22.6% and 33.6%. (Lane: L1-track-record)

  4. 4

    Medallion's legendary return of over 60% a year is an arithmetic, pre-fee figure from a book; a 2024 reconstruction puts the compounded pre-fee return probably under 35%. (Lane: L1-track-record)

  5. 5

    Bridgewater's AI fund reportedly matched, rather than beat, its flagship in H1 2026 (8.1% each), and neither Bridgewater nor the press has disclosed how much of its decision-making is automated. (Lane: L1-track-record)

The strongest case against this

The evidence against our own conclusion

Carried at full strength, before the conclusion — not as a footnote.

  • F18
  • F19
  • F20
  • F14

Statement: The machines are not standing still.

Concession: The best AI has likely drawn level with the best human forecasters (FRI, July 2026), an AI analyst beat most human analysts narrowly, published edges do not vanish on average, and adoption is near-universal. The film's reply is built from the last of these: what everyone uses, the price learns.

History, under test

The strongest historical case

What the past licenses — and, stated just as plainly, what it does not.

M6 forecasting competition, February 2022 to February 2023

  • F04
  • F05
  • F06
  • F07
  • F08
  • F10

Lesson: Being right about what will happen and being paid for it are two jobs, rarely joined; most teams sized their bets beyond their accuracy.

What would change our mind

The conclusion is wrong if…

A widely available AI tool that keeps beating the index after costs and for the risk it took, on years it had never read about, while more and more people use it.

What remains uncertain

What we still do not know

Stated by the research team, in full, rather than smoothed over.

  • Whether similar AI models make firms wrong together: described as a potential risk by the Bank of England and the FSB; no measurement found (F29).
  • Which organisations have a real machine edge: this record found no disclosed AI-attributable returns to test; the generalisation that the best-placed firms do not publish was UNSUPPORTED in refutation and is not spoken (F31; one checkable example, Man Group, is).
  • How long the forecasting gap stays open, and whether parity would show up in returns (F18, F30, F08).
  • How much of the investor return gap is bad timing: 1.2 points (Morningstar) vs 0.10% (FAJ 2026) on overlapping data (F26, F27).
  • Whether the Lopez-Lira & Tang decline is caused by adoption: the authors call it suggestive (F12).
  • H3 (organisations win) for returns: open, as pre-registered, because of non-disclosure.
  • Objectives, mandates and career risk: no harvested source verified (Cremers & Petajisto could not be opened).

What we refused to publish

20 claims we would not say — and why

The do-not-narrate list. Some are popular; some are true but unproven; some are simply not this note’s to make. Each refusal is enforced in production, not just recorded.

6 Unverifiable9 Refuted3 Partly unverifiable2 Unresolvable
  1. UnverifiableX01

    “Renaissance Medallion as proof that machines or AI beat human investors; its "66% a year" return.”

    Verdict: UNVERIFIABLE — record undisclosed

    Why: Closed to outsiders since 1993; figures from a journalist's book and anonymous investors; the headline is arithmetic and gross of fees, and a 2024 reconstruction puts the compounded pre-fee return probably under 35%. Renaissance's outside-investor funds lost ~22.6% and ~33.6% in 2020 per a third-party scoreboard (L1-09..13).

  2. UnverifiableX02

    “An AI-run fund or ETF has beaten the market (e.g. AIEQ since 2017).”

    Verdict: UNVERIFIABLE — no primary benchmark comparison opened

    Why: The issuer reports 10.08%/yr at NAV since inception but benchmarks against its own AI index; the fund changed structure; no audited autonomous AI record was found (L1-02..05, L5-18, L5-21).

  3. RefutedX03

    “GPT-4 beats human analysts at financial statement analysis (Kim, Muhn, Nikolaev).”

    Verdict: REFUTED as established evidence — withdrawn

    Why: arXiv version withdrawn 2025-02-20 over data inconsistencies (L2-10, L5-07). Only the withdrawal may be narrated.

  4. RefutedX04

    “ChatGPT predicts stock moves with ~90% accuracy.”

    Verdict: REFUTED as a tradable claim

    Why: The ~90% is a portfolio-day hit rate for the non-tradable initial reaction (L2-11).

  5. RefutedX05

    “Most / half of all trading is algorithmic or high-frequency (as an SEC measurement).”

    Verdict: REFUTED as sourced

    Why: The SEC's 2010 "50% or higher" line footnotes press articles; ESMA measured 24-43% of value traded and 58-76% of orders in EU equities, May 2013, depending on definition (L4-14, L4-15).

  6. RefutedX06

    “DALBAR's investor behaviour gap.”

    Verdict: REFUTED as a measure of bad timing

    Why: Compares dollar-weighted investor returns with a time-weighted index, conflating the sequence of market returns with behaviour (L6-09, L6-10).

Show the remaining 14 refused claims
  1. RefutedX07

    “The 2010 Flash Crash or Knight Capital 2012 as failures of intelligent or AI trading systems.”

    Verdict: REFUTED

    Why: A volume-only execution algorithm "without regard to price or time"; a deployment error reactivating code unused since 2003 (L4-01..08).

  2. RefutedX08

    “LTCM as an AI or machine-learning failure.”

    Verdict: REFUTED

    Why: A leverage and liquidity failure of a model-driven relative-value fund in 1998 (L3-11..13).

  3. RefutedX09

    “Regulators have found AI is already causing herding in markets; the SEC's 2023 predictive-analytics proposal was a herding rule.”

    Verdict: REFUTED

    Why: BoE and FSB describe potential risk; the SEC proposal addressed conflicts of interest and was withdrawn on 2025-06-12 (L3-16..21).

  4. Partly unverifiableX10

    “Centaur (human + engine) chess shows human-plus-machine teams will dominate investing.”

    Verdict: PARTIALLY UNVERIFIABLE — untested analogy

    Why: The 2005 result is real but a single event; later engine comparisons move toward parity; markets adapt and chess does not (L6-21, L6-22).

  5. RefutedX11

    “Alpha Arena proved LLMs can trade profitably with real money.”

    Verdict: REFUTED as evidence of skill

    Why: Six models, $10,000 each, ~16 days, one market path, four of six lost (L5-08, L5-09).

  6. UnverifiableX12

    “Bridgewater's AI-driven fund returns (e.g. 11.3%/yr, 8.1% in H1 2026, 11.9% in 2025).”

    Verdict: UNVERIFIABLE — anonymous sources

    Why: Reported by Reuters from anonymous sources; the 11.9% figure was not found in the opened article (L1-17..19).

  7. UnverifiableX13

    “NBIM's AI saves ~$100m a year / will save $400m; AI made NBIM 20% more productive.”

    Verdict: UNVERIFIABLE

    Why: The primary statement says "several hundred million kroner" in trading costs for a project using AI among other tools; the 20% is an internal self-report survey (L4-21, L4-22, L5-14).

  8. Partly unverifiableX14

    “"Bad timing costs fund investors ~15% of their funds' returns" as settled.”

    Verdict: PARTIALLY UNVERIFIABLE — disputed

    Why: Morningstar's 2026 edition says ~12%; an FAJ 2026 paper on the same sample puts the timing cost near 0.10%/yr (L6-04..08). Narrated only as a dispute.

  9. UnverifiableX15

    “Machine learning doubles the performance of stock-picking strategies (Gu, Kelly, Xiu).”

    Verdict: UNVERIFIABLE — figures not verified

    Why: Exact Sharpe figures not verified; Avramov et al. show profits concentrate in hard-to-arbitrage stocks and weaken after costs (L2-06..09).

  10. RefutedX16

    “High Active Share reliably predicts outperformance.”

    Verdict: REFUTED — disputed

    Why: Frazzini, Friedman and Pomorski (FAJ 2016) find it correlates with benchmark returns rather than predicting fund returns (L6-20).

  11. Partly unverifiableX17

    “Sentient's AI hedge fund closure as evidence that AI funds fail.”

    Verdict: PARTIALLY UNVERIFIABLE — thin and anecdotal

    Why: Anonymous Bloomberg sourcing; one closure is not a failure rate (L1-08).

  12. UnverifiableX18

    “The value factor is dead; ML models predicted the March 2020 crash.”

    Verdict: UNVERIFIABLE

    Why: Value's 55% drawdown is documented but its death is disputed; the only COVID-prediction hit was unopened and retrospective (L3-14, L3-15).

  13. UnresolvableX19

    “FinanceBench's 81% failure rate (2023) as a description of current models.”

    Verdict: UNRESOLVABLE — dated benchmark; MUST BE STATED AS HISTORICAL

    Why: A 2023 configuration on a 150-question sample (L5-01). Not narrated: benchmark scores are not investment evidence (founder brief).

  14. UnresolvableX20

    “Quantum computing as part of the AI investing question.”

    Verdict: UNRESOLVABLE IN THIS EPISODE — coverage item not carried

    Why: Listed ADJACENT in the coverage matrix; the founder's 2026-09-13 brief defines the episode and does not ask for it. Named in the founder review.

What to watch next

Dated material, and what would make it stale

  • ForecastBench AI vs superforecaster gap

    4 Mar 2026

    Status: best AI 67.9 vs superforecasters 70.6 (Brier Index); parity projected May 2027

    What changes: S06 line one and N39; N56

  • SPIVA U.S. Scorecard

    Mar 2026

    Status: Year-End 2025: 79% of active large-cap funds lagged the S&P 500

    What changes: N49 and S08 line one at Year-End 2026

  • AIMA generative AI adoption survey

    16 Sep 2025

    Status: 95% of 150 managers use generative AI

    What changes: N41

  • Kim, Muhn & Nikolaev paper status

    20 Feb 2025

    Status: arXiv version withdrawn

    What changes: N36 and the S05 card if a corrected version is published

  • Morningstar Mind the Gap

    6 Aug 2026

    Status: ~1.2 points a year

    What changes: N52 at the next edition

Evidence & sources

20 sources, by tier

Tier 1 is primary and authoritative — filings, regulators, official statistics. Journalism and books are attributed ingredients, never proof by reputation.

  1. Primary sourceVirtu Financial, Inc. — Form S-1 registration statement (opens sec.gov)supports F01, F02, F03
  2. ResearchMakridakis, Spiliotis, Hollyman, Petropoulos, Swanson & Gaba — The M6 forecasting competition: Bridging the gap between forecasting and investment decisions (International Journal of Forecasting; arXiv 2310.13357) (opens arxiv.org)supports F04, F05, F06, F08, F09, F10, F33
  3. Primary sourceM6 organisers — Mcompetitions/M6-methods, "IJF paper/performance_vs_answers_unique.csv" (sha256 2d69575c457cbf5b977746327c6468bfbe367dd8e199e885eab73668cb2dab36) (opens github.com)supports F07, F32
  4. ResearchLopez-Lira & Tang — Can ChatGPT Forecast Stock Price Movements? Return Predictability and Large Language Models (arXiv 2304.07619v6) (opens arxiv.org)supports F11, F12, F13
  5. ResearchMcLean & Pontiff — Does Academic Research Destroy Stock Return Predictability? Journal of Finance 71(1), 5-32 (abstract via Crossref) (opens api.crossref.org)supports F14
  6. ResearchLopez-Lira, Tang, Zhu — The Memorization Problem: Can We Trust LLMs' Economic Forecasts? (arXiv 2504.14765v2) (opens arxiv.org)supports F15
  7. Primary sourcearXiv record 2407.17866v3 — Kim, Muhn, Nikolaev, Financial Statement Analysis with Large Language Models (withdrawal notice by Valeri Nikolaev) (opens arxiv.org)supports F16
  8. ResearchLi, Kim, Cucuringu, Ma, Can LLM-based Financial Investing Strategies Outperform the Market in Long Run? (FINSABER; arXiv 2505.07078; KDD 2026) (opens arxiv.org)supports F17
  9. Primary sourceForecasting Research Institute — Making Forecasting Scores Easier to Interpret: Introducing the Brier Index (opens forecastingresearch.substack.com)supports F18
  10. ResearchCao, Jiang, Wang, Yang — From Man vs. Machine to Man + Machine: The art and AI of stock analyses, Journal of Financial Economics 160 (2024) (opens repository.lsu.edu)supports F19, F21
  11. Primary sourceAIMA press release, 'Front-office Gen AI adoption shifts from if to when for leading fund managers' (opens aima.org)supports F20
  12. ResearchCao, Jiang, Wang, Yang — Journal of Financial Economics 160 (2024) (opens repository.lsu.edu)supports F22, F23
  13. JournalismTKer (Sam Ro), reporting SPIVA U.S. Scorecard Year-End 2025; corroborated by InvestmentNews (2026-03-04) (opens tker.co)supports F24
  14. ResearchRossi & Utkus, 'Who Benefits from Robo-advising? Evidence from Machine Learning' (working paper draft hosted by GFLEC) (opens gflec.org)supports F25
  15. Primary sourceMorningstar, Mind the Gap (US, 2026 edition) (opens morningstar.com)supports F26
  16. ResearchFulkerson, Jordan, Riley, Yan, 'Bad Timing Does Not Cost Investors 15% of Their Funds' Returns', Financial Analysts Journal 82(3) (opens rpc.cfainstitute.org)supports F27
  17. Primary sourcePresident's Working Group on Financial Markets (April 1999) (opens cftc.gov)supports F28
  18. Primary sourceBank of England, Financial Stability in Focus: Artificial intelligence in the financial system (April 2025) (opens bankofengland.co.uk)supports F29
  19. Primary sourceForecasting Research Institute (ForecastBench operators) — LLMs Are Closing the Gap on Human Superforecasters (opens forecastingresearch.substack.com)supports F30
  20. Primary sourceMan Group, 'What AI Can (and Can't Yet) Do for Alpha' (opens man.com)supports F31

Complete evidence depth

The research desk

Everything the research evaluated before it was distilled — including what it rejected, and why.

The research desk

Before the evidence above was distilled, the research evaluated 128 candidate claims across 0 lanes, rejected 20 with a recorded reason, and logged 38 surprises. A rejected claim with a reason is the most reusable thing research produces — the desk keeps all of them.

91 Do not narrate19 Corrected16 Supported2 Misleading
  • laneRenaissance Medallion as proof that machines or AI beat human investors; its "66% a year" return. (Id: X01; Verdict: UNVERIFIABLE — record undisclosed; Reason: Closed to outsiders since 1993; figures from a journalist's book and anonymous investors; the…
  • laneAn AI-run fund or ETF has beaten the market (e.g. AIEQ since 2017). (Id: X02; Verdict: UNVERIFIABLE — no primary benchmark comparison opened; Reason: The issuer reports 10.08%/yr at NAV since inception but benchmarks against its own AI index; the fund…

Save this research and keep following the story — you come straight back to this desk.

Community

Talk it through

Discuss this research

Members only · in 📡 Intelligence Watch

Open in Community

Challenge the argument, bring evidence, or ask what would change our mind. This note is discussed in 📡 Intelligence Watch, a DowJo room where members compare reads. Filings, news events, and what they actually mean — the discussion room for the intelligence feed.

Read and reply, then keep following the story — you come straight back to this note.

Education-first ground rules apply in every room: mechanisms and evidence, never calls.

Now train it

Take it further

Reading is where judgment starts. Practice is where DowJo measures it — and remembers what to train next.

Provenance — where this note comes from

This page is a deterministic projection of canonical research artifacts. It adds presentation and discovery; it never adds, removes or softens a finding. Question taken from the title; thesis from the packet. Projected 21 Sep 2026 by dowjo-research-projector v1.1.0 from origin commit c83c8797ab79.

ArtifactPath (origin-relative)IdentityVersion
Episode registrycourse/episodes.jsonsha256 75b65b1618fc…—
Evidence packetcourse/research/AI13-evidence-packet.jsonsha256 eb797367cbd9…compiled 2026-09-13 · rev 1
Research shortlistcourse/research/AI13-research-shortlist.jsonsha256 cda1a543d991…—
Approved scriptcourse/scripts/AI13-what-the-price-already-knows.jsonsha256 4fabd85fbbec…v draft2
Episode manifestcourse/manifest/AI13.manifest.jsonsha256 88ea06c0f331…—
Approval recordrenders/approved/AI13-FINAL.jsonsha256 2b333bfe49a8…compiled 2026-09-13
Pre-registrationcourse/research/AI13-lanes/PR0-thesis-preregistration.mdsha256 237adcb65b70…—
Poster stillqa/stills-AI13/RESEARCH-POSTER.pngsha256 307d99fc6e7e…—

Approved master (AI13-final02-corrected.mp4): sha256 c131855e476787e9fc9d37dae5c053468f0c70555d78dd2ce80403dd3fedcdb0

Commercial disclosure: none. No affiliate relationship, review copy, sponsorship or publisher relationship applies to this note.

For education only. Not financial advice. No buy, sell or hold recommendations, ever.