The provocation is deliberate. AI is not the problem. Uncontrolled research throughput is.
Generative and agentic systems can read more papers, propose more features, write more tests and explore more combinations than a conventional research team. That is useful. It also multiplies the opportunities to mistake noise for signal.
Generating hypotheses
Research capacity was constrained by analyst time, coding effort and data preparation.
Rejecting false discoveries
AI reduces the cost of another test. It does not reduce the statistical burden created by testing it.
Governed rejection
Durable research systems must make ideas easy to reject and difficult to promote.
When the number of hypotheses rises faster than validation quality, the expected output is not necessarily more alpha. It can be a larger library of persuasive backtests with weaker economic meaning.
More tests can manufacture confidence—even when every candidate is noise.
Consider a deliberately simplified experiment. Every candidate strategy has no true edge. Each is tested at a 5% false-positive threshold and the tests are assumed independent. The chance of finding at least one apparently significant result rises quickly with the number of trials.
Illustration only. Real research tests are often correlated. The chart is not an estimate of the probability that a specific strategy is overfit.
published return factors
Harvey, Liu & Zhu documented a large factor literature and argued that conventional significance thresholds are too lenient under multiple testing.
predictors studied
McLean & Pontiff examined a broad sample of published return predictors across in-sample, out-of-sample and post-publication periods.
out-of-sample returns
Average portfolio returns were 26% lower outside the original sample than in the published in-sample evidence.
post-publication returns
Average returns were 58% lower after publication, consistent with statistical bias and investor learning.
False confidence rarely enters through one dramatic error. It accumulates through small research choices.
Hypothesis multiplication
Features, transformations, thresholds and horizons proliferate faster than the economic thesis becomes clearer.
Point-in-time failure
Revised data, current index membership, late fundamentals or misaligned timestamps leak future information into the past.
Repeated holdout use
An “out-of-sample” period stops being untouched once researchers repeatedly inspect it and redesign the model around the result.
Implementation fiction
Frictionless fills, unlimited liquidity, stable borrow and ideal contract rolls create returns that cannot survive production.
Narrative after results
A causal story written after the strongest backtest is selected can convert selection luck into apparent conviction.
Research monoculture
Shared models and data sources can generate different-looking signals that ultimately express the same factor or regime.
The scarce asset is no longer another idea. It is a research process capable of saying “no.”
A governed agentic architecture should separate idea generation from promotion authority. Agents may accelerate collection, coding and challenge tests; they should not silently redefine the hypothesis, consume the final holdout or move a model into production.
Register
Fix the economic thesis, expected sign, horizon and rejection criteria before testing.
Reconstruct
Make point-in-time data, timestamps, universe and transformation lineage reproducible.
Challenge
Use placebos, subperiods, perturbations, factor overlap and alternative definitions to attack the result.
Validate
Open a genuinely untouched sample once, under a pre-committed decision rule.
Implement
Model turnover, slippage, liquidity, borrow, contract rolls and capacity explicitly.
Approve
Keep promotion, limits, rollback and accountability with responsible people.
Architecture, not an AI claim.
Tensor is building its research operating system to accelerate investigation without allowing research agents to promote their own conclusions.
Separated authority
Discovery, challenge, validation and human approval are distinct stages with different permissions.
Evidence before promotion
Point-in-time reconstruction, untouched validation, realistic costs and overlap checks precede escalation.
Complete research lineage
Data vintages, tested variants, assumptions and rejection reasons are preserved—not only the winner.
Human accountability
Humans retain ownership of limits, deployment, exceptions, rollback and ongoing decay monitoring.
Do not ask whether a manager uses AI. Ask what AI is allowed to change—and what evidence survives.
- How many related hypotheses and parameter variants were tested before the presented result was selected?
- Can the exact point-in-time dataset and universe be reproduced for every historical decision?
- What period remains genuinely untouched, and who controls access to it?
- Which cost, liquidity, borrow and capacity assumptions separate research from implementation?
- How is incremental information measured against conventional factors and existing exposures?
- Who can approve, modify, suspend or retire a model—and is that decision auditable?
AI will make analytical intelligence more abundant. That makes disciplined rejection more valuable—not less. The manager with the most ideas may not win. The manager with the clearest evidence about why most ideas were rejected may be better positioned to build durable systematic research.Discuss our research architecture
