Engineering deep dive · · 16 min read

Teaching a 27B Model to Write Trading Alphas: 101 Formulas, 12 Rewards and One Unseen Year

An LLM writes one-line trading formulas in the language of WorldQuant's 101 Formulaic Alphas; a verifier turns each into a dollar-neutral long-short book on 46 US stocks and scores it on 12 reward channels. The DSL contract, alpha vs. beta and five layers that keep the model off the market's direction (holdout beta 0.09), every reward channel including a JEPA world-model stress test with a correlation-geometry objective, the reward-hack log, and train vs. holdout results: mean holdout score −1.37 → +1.80, equal-weight book +16.2% over an unseen year, with limits.

Qwen3.8-27B is trained with multi-reward RL to write one-line trading formulas in the language of WorldQuant’s 101 Formulaic Alphas. A deterministic verifier turns each formula into a daily dollar-neutral long-short book on 46 US stocks and scores it on 12 reward channels. On 50 prompts and a year of market data that training never saw, the mean score of its formulas went from −1.37 to +1.80, and an equal-weight book of them returned +16.2% with a market beta of 0.09.

This post documents the system: the formula language, how a formula becomes a portfolio, the difference between alpha and beta and how we keep the model from learning the market, each reward channel and the hack that created it, and the train and holdout results. The training recipe (GDPO, CISPO, REPO-R) and its experiment ladder are in Multi-Reward RL, Part 3, which described this gym with the domain hidden.

Results of the best run (1790958861) on the 50 unseen prompts and the unseen year, before and after training.
MetricStep 0 (untrained)Step 500 (shipped)
Mean holdout score per formula−1.37+1.80
Valid formulas45 / 5049 / 50
Formulas with holdout score above zero4 / 4547 / 49
Equal-weight book: 1-year return after costs−0.5%+16.2%
Equal-weight book: Sharpe−0.212.58
Book beta to the market—0.09 (correlation 0.19)

0. System at a glance

Configuration of the run reported in this post.
PolicyQwen3.8-27B, NF4 weights + rank-32 LoRA, no visible reasoning
ActionOne formula in <alpha>…</alpha>, at most 128 tokens
VerifierDeterministic: parse → daily dollar-neutral book → costs → 12 reward channels
Universe46 US large caps, daily OHLCV from June 2012
Train window2012-06-29 → 2025-08-20 (3,304 sessions)
Holdout2025-08-21 → 2026-08-21 (252 sessions), never used for training or selection
RL recipeGDPO over 12 channels, CISPO loss, REPO-R, 12 answers per prompt, 600 steps, one H100
Checkpoint ruleBest train-window CPCV of the 50 probe answers (step 500)

1. The formula language

Source. In 2016 Zura Kakushadze published 101 trading alphas from WorldQuant, with the firm’s permission, as “explicit formulas – that are also computer code”. Eighty were in production at the time. They use daily open, high, low, close, volume, returns and vwap, have average holding periods of 0.6 to 6.4 days and an average pairwise correlation of 15.9%. The shortest:

Alpha#101: ((close - open) / ((high - low) + .001))

If a stock closed well above its open relative to its daily range, go long the next day: the paper calls it a delay-1 momentum alpha. Our DSL uses the same vocabulary, with a fixed contract:

The DSL contract. Every operator looks only backward in time.
GroupElementMeaning
Fieldsopen high low close volumedaily prices and share volume
returnsclose-to-close return
vwapproxy: (high + low + close) / 3
Cross-sectionalrank(x)rank of x across the 46 stocks on each day, in [0, 1]
scale(x)rescale so that sum(|x|) = 1
Time seriesdelay(x, d)value d sessions ago
ts_delta(x, d)x[t] − x[t−d]
ts_mean, ts_std, ts_sum, ts_max, ts_min (x, d)rolling statistics over d sessions
ts_rank(x, d)share of the window below today's value
ts_corr, ts_cov (x, y, d)rolling correlation and covariance
ts_decay_linear, ts_ema, ts_skew, ts_kurt, ts_slope, …weighted averages and shape statistics
Element-wise+ − * /, abs, log, sign, power, min, maxper stock, per day
Limits1 ≤ d ≤ 252total nested lookback ≤ 252 sessions
≤ 256 DSL tokens, depth ≤ 16the model's answer is also capped at 128 tokens
Safetyrecursive-descent parsernothing the model writes is executed as code

2. From formula to book

market data up to close(t)
  → evaluate formula          one value per stock
  → rank across stocks        values in [0, 1]
  → subtract the mean         weights sum to 0      (long $ = short $)
  → divide by sum(|w|)        gross exposure = 1    (0.5 long, 0.5 short)
  → fill at open(t+1)         earn open(t+1) → open(t+2), minus costs
A formula is a ranking; the ranking is a portfolioEvery evening the formula scores each stock. Ranks are centred and scaled, so longs and shorts are equal in dollars.1 · formula value2 · cross-sectional rank3 · weight in the bookKO0.80.0−27.8%XOM1.10.2−16.7%JPM1.90.4−5.6%MSFT2.40.6+5.6%AAPL3.00.8+16.7%NVDA4.21.0+27.8%short 50%long 50%Illustration with six names and made-up values; the real book ranks 46 stocks. Orders fill at the next day's open.
The same three steps for six stocks with made-up values. Only the order of the formula's values matters.
Cost model. Holdout costs are deliberately harsher than training costs.
StageCosts
Training3.5 bps per unit of turnover: 2 fee + 1 slippage + 0.5 half-spread, plus a small fixed impact
HoldoutPer-stock square-root market impact for a $1M book, capacity penalty, 50 bps a year to borrow the shorts

3. Alpha vs. beta: keeping the model off the market’s direction

Any daily return series can be split into a part that follows the market and a part that does not:

rbook(t)=α+β⋅rmarket(t)+ε(t)r_{\text{book}}(t) = \alpha + \beta \cdot r_{\text{market}}(t) + \varepsilon(t)

Beta is market exposure: a book with β = 1 gains 1% when the market gains 1%. Beta is cheap; an index fund sells it for a few basis points. Alpha is what remains when the market’s move is taken out: return from ranking stocks correctly against each other. A model that “learns to trade” by holding the market would look skilled in a rising year and fail in a falling one, so the gym is built to pay for alpha only.

When everything falls together, a long-short book barely moves$100 long in the top-ranked stocks, $100 short in the bottom-ranked ones. Profit = long move − short move.The whole market falls 4%longs −4%, shorts −4%−4%long side−4%short sidebook P&L$0.00Longs fall 1%, shorts fall 5%the ranking was right−1%long side−5%short sidebook P&L+$4Dollar-neutral is not fully market-neutral: a book can still lean on size, momentum or volatility. A separate reward taxes that.
With equal dollars long and short, a market-wide move cancels; only the spread between the two sides is earned.
Five layers between the model and the market’s direction.
LayerMechanismWhat it blocks
1. ConstructionRanks are centred and scaled: long dollars always equal short dollars. Measured net exposure on the holdout: 0.00.Direct market exposure
2. Factor reward (R4)Daily P&L is regressed on the market, size, 12-1 momentum, 5-day reversal and low volatility; the Sharpe those factors explain is penalised.Indirect exposure, e.g. long high-beta stocks, short low-beta ones
3. Prompt constraintsBriefs forbid known proxies, e.g. “do NOT lean on long-horizon price sums — they proxy market beta”.Beta by design
4. Motif guard (R10)Pure price-level or price-momentum formulas with no predictive power lose their positive reward.Slow tilts that ride a trend
5. Holdout checkBeta and correlation of the shipped book against the market on the unseen year.Anything the first four missed

Measured on the holdout year. The equal-weight market of the same 46 stocks rose +27.3%. The book returned +16.2% with β = 0.09 and a correlation of 0.19, so the market explains only about 2–3 points of its return. On the 7 days when the market fell more than 1.5% (on average −1.9%), the book lost on average 0.3%. In a strongly rising year a long-only index made more money; the book made less but independently of the market, with a higher Sharpe (2.58 vs. 1.95) and half the drawdown (−4.3% vs. −8.6%).

The book earned its return without riding the marketHoldout year, Aug 2025 – Aug 2026. Book = equal-weight step-500 formulas after realistic costs. Market = the same 46 stocks, equal weight.$0.90$1.00$1.10$1.20$1.30Growth of $1market +27.3%book +16.2%Aug 2025Aug 2026Daily returnsβ = 1β = 0.09market daily return →book ↑Beta 0.09, correlation 0.19. Book Sharpe 2.58 vs. market 1.95; worst drawdown −4.3% vs. −8.6%.
Left: growth of $1 for the book and for the market over the unseen year. Right: each dot is one day; the book's slope against the market is 0.09, close to flat.

4. The task

The task contract.
InputA design brief: mechanism, horizon, input rule, structural hint, forbidden pattern, risk preference
OutputOne formula in <alpha>…</alpha>
Group12 answers per brief, ranked against each other on every channel
Prompts500 generated briefs → 226 kept: those the untrained model solved in some but not all of 8 tries

A real brief from the training set:

Market: US-listed stocks trading on weekday sessions. Design an alpha whose PRIMARY mechanism is: mixed-signal extremes (rolling max or min of a price-volume composite). Target horizon: roughly 168-day windows. Inputs: must use exactly one ts_corr. Structure: take the sign of a time-series relationship. Constraint: do NOT lean on long-horizon price sums — they proxy market beta. Risk preference: make it robust to a volatility-regime change.

Briefs the model always or never solves give no learning signal. Dropping them was worth +0.47 on holdout in Part 3.

5. Twelve reward channels

GDPO standardises each channel within the group of 12 answers, multiplies by its priority and sums, so no channel can dominate through its numeric range. Two hard rules follow: an invalid answer never receives a positive advantage, and each group stays zero-sum.

Twelve reward channels, each answering one questionGDPO priority of each channel in the best run. Each channel is ranked within the group of 12 answers before weighting.Does it make money out of sample?cpcv3.0CPCV net SharpeIs it overfit or fragile?pbo1.0in-sample vs out-of-sample rankjepa_stress0.4survives simulated marketsfactor_neutrality0.4not just a known factorCan you actually trade it?turnover1.0pays for churning > 50%/daymin_turnover0.5pays for sitting stillDoes it predict?ic0.2weights vs next-day returnsIs it new?novelty_pnl0.4P&L unlike the archivenovelty_syntax0.2formula unlike the archiveIs it clean?motif_guard0.3no lazy one-motif formulascomplexity0.2> 24 nodes costsformat1.0inactive in no-think modeProfit carries the most weight, but more than half of the total priority guards against ways of looking profitable without being so.
Channels grouped by the question they answer, with their GDPO priorities in the best run.
The verifier matrix. Formulas are the raw channel values before GDPO standardisation.
IDChannelPriorityComputesBlocks
R1cpcv3.0mean − 0.5·std of net Sharpe over 15 CPCV splitsluck; fit to one period
R2pbo1.0penalty when a formula ranks far better in-sample than out of sample within its group (≥ −0.48)overfitting
R3jepa_stress0.4score drop in 4 JEPA-generated markets, log-compressed (≥ −1)fragility to regime change
R4factor_neutrality0.4−0.4 · clip(raw Sharpe − residual Sharpe, 0, 3) after 5 known factorsmarket beta and known factors in disguise
R5turnover1.0−2 · max(0, turnover − 0.5)books that pay their profit away in costs
R6min_turnover0.5penalty below 6% a day; invalid below 2% a daystatic and empty books
R7ic0.20.2 · clip(20 · IC, −1, 1); at most 0.05 while CPCV ≤ 0profit without prediction
R8novelty_pnl0.40.4 · (1 − max |corr| of daily P&L vs. a 1,024-entry archive); duplicates penalisedclones of known books
R9novelty_syntax0.2bonus for an unseen canonical form (windows bucketed) × (1 − P&L corr); −0.3 per duplicateone template with jittered windows
R10motif_guard0.3bare one-motif or price-only formula with IC < 0.005 loses its positive rewardlazy single-motif formulas
R11complexity0.2−0.02 per expression node beyond 24bloated formulas
R12format1.0+0.05 for a reasoning block; always 0 in this run(inactive without reasoning)

R1 · CPCV core score

The train window is cut into 6 blocks; every pair of blocks is one test slice (15 slices), with 10 days purged and 5 embargoed at the edges. For each slice the annualised net Sharpe is computed:

CPCV=mean⁡(SR1,…,SR15)−0.5⋅std⁡(SR1,…,SR15)\text{CPCV} = \operatorname{mean}(\text{SR}_1,\ldots,\text{SR}_{15}) - 0.5 \cdot \operatorname{std}(\text{SR}_1,\ldots,\text{SR}_{15})

Example: Sharpe 0.8 on every slice scores 0.8. Sharpe 3.0 on three slices and −0.2 on twelve averages 0.44 but scores −0.20. Missing tag: −1.0. Unparseable formula or lookback over 252: −0.8.

CPCV: fifteen different “futures” cut from one historyTrain 2012–2025 is cut into 6 blocks; every pair of blocks is one test. The formula has nothing to fit, so each split is a fresh exam.block 1block 2block 3block 4block 5block 6split 1split 2split 3split 4split 5split 6split 7split 8split 9split 10split 11split 12split 13split 14split 15test daysin-samplepurge / embargoscore =mean − 0.5 × stdof 15 net SharpesA formula that shines on 3 splits and dies on 12 gets a low score: the std term punishes luck. Purge 10 days, embargo 5.
Fifteen test slices from six blocks. The formula has nothing to fit, so each slice is an independent exam; the grey blocks are the in-sample part used by R2.

R5, R6 · Turnover band

Trade too little or too much and the score dropsTurnover = share of the book traded per day. Sum of the two turnover channels (raw values, before GDPO weighting).0.00-0.25-0.50-0.75-1.001%2%5%10%20%50%100%invalid< 2%/dayno penalty6% to 50% a day−0.30 at 2%churn taxBelow 2% a day the whole answer counts as invalid: a book that never trades can't be judged on fresh data.
Free between 6% and 50% of the book a day; penalised on both sides; invalid below 2%.

Example, a real answer of the untrained model:

scale(ts_rank(ts_delta(volume, 23), 47) - ts_rank(ts_delta(volume, 23), 47))

The expression subtracts a value from itself, so every weight is zero every day: no trades, no costs, no losses. Under high costs this beats any honest formula. Its turnover is 0, below the 2% floor, so it is invalid.

R3 · JEPA stress test

A JEPA world model is trained only on the train window (2012 to August 2025) and then frozen. It reads a window of the 46-stock panel with 25% of the stock-day cells masked in spans of 8 days, encodes one latent per stock, and predicts the latent of the next window. The target is an exponential-moving-average copy of its own encoder, so the model learns to predict market states rather than raw prices.

A second target comes from Observable Matrix Dynamics (OMD; Halperin, 2026), a dynamical-systems view of the market: it follows the trajectory of a fixed-size matrix built from rolling stock correlations and of its spectrum, and finds that the matrix’s effective dimension collapses in crises such as 2008 and 2020. We use the same idea as an auxiliary loss: the predicted latent must also recover the geometry of the next window’s correlation matrix CC of the 46 stocks.

OMD-style geometry targets of the JEPA (loss weight 0.1). Computed on training data only.
TargetComputed asWhat it captures
Mean correlationaverage off-diagonal entry of CChow much stocks move together
Market-mode shareλmax⁡/tr⁡C\lambda_{\max} / \operatorname{tr} Cweight of the single market factor
Effective factor fraction(∑λi)2/∑λi2(\sum \lambda_i)^2 / \sum \lambda_i^2 divided by 46effective dimension of the market; collapses in crises
Residual factor fractionthe same without the market modediversity left beyond the market
Projector driftrotation of the top-3 eigenvector subspace from the past to the next windowrotation between sectors and themes
Leading spectrumshares of the top-5 eigenvaluesshape of the dominant modes

At reward time the latent is shifted along four channels — dispersion between stocks, trend flip, correlation and volatility regime — and decoded into complete price and volume panels. The formula is re-scored in each world with the same CPCV and costs, and the average drop becomes the penalty −min⁡(1, 0.4⋅log⁡(1+drop‾))-\min(1,\ 0.4 \cdot \log(1 + \overline{\text{drop}})). A fifth channel, a crash, was removed because it ranked formulas worse than chance (AUC 0.43).

The stress test: a frozen world model invents markets the formula has never seenJEPA trained only on 2012 – Aug 2025, then frozen. Each formula is re-scored in four decoded worlds; a large drop costs reward.Past window46 stocks × days25% of cells maskedEncoderone latent per stock→ market statePredictor→ predicted latentof the next windowTarget 1: future latentEMA copy of the encoderTarget 2: geometry (OMD)mean correlationmarket-mode shareeffective # of factorsrotation of top-3 modesAt reward timedispersiontrend flipcorrelationvol regimeshift the latent along one channelDecoder→ a full price/volumepanel for 46 stocksRe-scoresame formula, sameCPCV and costspenalty = −min(1, 0.4 · log(1 + mean drop of the score across the 4 worlds))The geometry target makes the model learn how the market's correlation structure moves, not only prices. A crash channel was dropped: it ranked formulas worse than chance.
Training (top) and use (bottom) of the stress test. The geometry target teaches the world model how the market's correlation structure moves, not only how prices move.

Other channels in one line each

  • R2 stability: first in-sample and tenth out of sample within the group is penalised; the penalty grows with the drop.
  • R4 factor neutrality: a formula whose Sharpe falls from 1.5 to 0.3 after removing the five factors pays −0.4 × 1.2 = −0.48.
  • R7 IC: the daily correlation of yesterday’s weights with today’s returns; an IC of 0.05 earns the full bonus.
  • R8, R9 novelty: ts_mean(close, 20) and ts_mean(close, 21) count as the same idea; new code that trades an old book earns nothing.
  • R10 motif guard, R11 complexity, R12 format: hygiene; R12 is inactive because this run answers without reasoning.

6. Reward hack log

The first version had a profit score and a format check. Each row below is a way the model found to score well without a real edge, and the change that closed it.

Reward and protocol changes, July – August 2026, from commit history and run reports.
DateObservedChange
Jul 4All answers collapsed onto one formulaSyntax novelty (R9)
Jul 5Different code, same bookP&L novelty with an archive seeded by Alpha101-style formulas (R8); smoke test: 2 → 5 distinct valid formulas
Jul 6Static books: slow price-level tilts with IC 0.000 took the top rewards (0.55–0.63)IC reward (R7); not enough alone (0.63 vs. 0.34 for momentum), so a minimum-turnover penalty the same day: 0.63 → 0.25
Jul 7High-CPCV books trading 0.2–0.6% of the book a day still wonHard floor: below 2% a day = invalid (R6)
Jul 7–8Formulas tuned to the exact scored historyJEPA stress test (R3): static books −12 offline, honest formulas ≈ 0; next day log-compressed, as 24% of groups had ≥ 2 formulas at the cap
Jul 969–83% of answers were one momentum template with jittered windows; every arm failed a residual-Sharpe checkWindow bucketing (5/21/63/126/252) before novelty; factor-neutrality penalty (R4)
Jul 1063–83% of invalid answers still got positive advantageValidity as a hard constraint on the advantage
Jul 23Checkpoint choice could leak the holdoutSelection by train-window CPCV only; CPCV floor on the advantage
Jul 26Stress crash scenario ranked formulas worse than chance (AUC 0.43 vs. 0.55–0.64); the turnover band charged top formulas −0.43Crash scenario removed; band moved to 3–6% a day (top decile now −0.09)
Aug 9Signal and fill on the same close: a subtle look-aheadFill at the next open, earn open-to-open returns

Every row was found by reading the top-scoring formulas. A rising reward curve was the first symptom of each hack, not evidence of learning.

7. Train vs. holdout

The best recipe from Part 3: CISPO + REPO-R, LoRA, 226 prompts, 600 steps. Part 3 shows the ladder of experiments that led to it, from DAPO at −0.47 to this recipe at +1.80.

Every checkpoint on the 50 unseen prompts (greedy). Train CPCV is on 2012 – Aug 2025 and selects the checkpoint; the other columns are on the unseen year.
StepTrain CPCVHoldout scoreValidFormula familiesBook returnBook Sharpe
0 (untrained)−0.98−1.3745 / 5045−0.5%−0.21
100−0.36−0.4747 / 5047+2.6%0.95
200−0.15−0.1349 / 5049+4.3%1.36
300+0.29+0.8548 / 5041+10.9%2.14
400+0.23+0.5149 / 5045+8.3%2.00
500 (shipped)+0.82+1.8049 / 5031+16.2%2.58
600+0.69+1.5848 / 5034+15.2%2.56
The model improved on years it trained on, and on the year it never sawBest run (CISPO + REPO-R, LoRA, 226 prompts). Left: training core score. Right: the same 50 unseen prompts scored every 100 steps.−1.0−0.50.0+0.5+1.0step 0step 200step 400step 600Training core score (CPCV, 10-step mean)−1.00.0+1.0+2.0step 0step 200step 400step 600Checkpoints on 50 unseen promptsshipped: +1.80base −1.37holdout yeartrain window (selector)Step 500 was picked by its train-window score (0.82), not by holdout. It also had the best holdout: +1.80 vs −1.37 for the untrained model.
Left: the training core score the model optimises. Right: the same 50 unseen prompts scored on the train window (grey, the selector) and on the unseen year (orange).

Step 500 has the highest train-window CPCV (0.82) and was shipped on that basis; it also has the highest holdout score. Holdout and train-window scores move together but not in lockstep: step 300 is better on holdout than step 400, and step 600 is slightly worse than step 500.

Every answer to the 50 unseen prompts, before and after trainingHoldout score of each valid formula over Aug 2025 – Aug 2026 (CPCV-style Sharpe after realistic costs). One dot per formula.−4−20+2untrained model45 valid · mean −1.37after 500 RL steps49 valid · mean +1.8047 of 49 trained formulas scored above zero on the unseen year, against 4 of 45 before training. All 49 use trading volume.
Every valid answer to the 50 unseen prompts, scored on the unseen year. Dashed lines: means.
Answers of the step-500 model to unseen prompts. Train CPCV and turnover on 2012 – Aug 2025; Sharpe and return on Aug 2025 – Aug 2026 after realistic costs.
FormulaTrain CPCVTurnover / dayHoldout SharpeHoldout return
rank(ts_max(volume / vwap, 11) / ts_mean(close, 21))0.953.1%3.13+19.3%
rank(ts_max(volume / vwap, 34) / ts_std(returns, 34))0.852.8%2.97+17.8%
rank(ts_std(volume / (high - low), 34))0.822.3%2.96+18.3%
rank(ts_std(volume, 8) / ts_std(close, 8))0.5914.9%2.70+15.7%
rank(ts_std(volume, 17) / ts_mean(volume, 17) + ts_std(volume, 17) / ts_mean(volume, 17))−0.4723.2%−1.09−7.6%

Convergence. All 49 trained answers use volume, and most rank stocks by the variability of their trading volume relative to price level or price volatility. Distinct formula families fell from 45 to 31. We have not established why this signal worked over this year; it is one idea expressed many ways, so 49 formulas are far fewer than 49 independent bets.

An equal-weight book of the model's formulas, over the unseen yearAll valid answers to the 50 unseen prompts, equal weight, realistic costs, Aug 2025 – Aug 2026.0%5%10%15%−0.5%untrainedSharpe -0.21+2.6%step 100Sharpe 0.95+4.3%step 200Sharpe 1.36+10.9%step 300Sharpe 2.14+8.3%step 400Sharpe 2.00+16.2%step 500Sharpe 2.58+15.2%step 600Sharpe 2.56Shipped checkpoint (step 500): +16.2% over the year, Sharpe 2.58. One year, one universe, one seed: a signal, not a track record.
One equal-weight book per checkpoint: all valid answers, realistic costs, compounded over the unseen year.

8. Selection: which checkpoint, which formulas

Checkpoint. Chosen by train-window CPCV of the 50 probe answers; the holdout is computed once, afterwards (rule in place since July 23).

Formulas for a book. Ranking by CPCV and taking the top 50 is the obvious choice and a poor one. On our formula archives that book had a median member CPCV of 0.71–0.75 but behaved like 1.8–2.6 independent bets (pairwise P&L correlation 0.50–0.63), with the worst drawdown of the selectors tested. A correlation-capped selector:

sort candidates by CPCV, descending
book = []
for f in candidates:
    if turnover(f) < 0.02: skip
    if max |corr(pnl(f), pnl(g))| >= 0.5 for any g in book: skip   # redundant bet
    book.append(f)

On two archives it reached a Sharpe of 2.24 and 2.00 against 1.85 and 1.39 for reward-ranked picks, with about 20 effective bets instead of 6–12 and half the drawdown, although its members’ median CPCV was only 0.38. The books in this post use neither selector: they average every valid answer with equal weight.

9. Limits

Our evaluation protocol marks a single-year result like this as no claim: evidence that the system learns, not evidence of a durable edge.

  • Survivorship. The 46 stocks are today’s large caps followed back to 2012; companies that shrank, merged or failed are absent.
  • One year, one seed. The holdout is 252 sessions from one market regime; the run is a single seed.
  • A gap in the probe. 15 of the 49 trained answers trade less than 2% of the book a day on the train window. Training would mark them invalid; the holdout probe scores CPCV only and does not apply the floor.
  • Neutrality. The book is dollar-neutral with a measured β of 0.09, but not sector-neutral, and our vwap is a proxy.

Why no model predicts the market with certainty

A market price is a projection of the real world: what companies sell and spend, what central banks and governments decide, wars, weather, new technologies, and the expectations of millions of people about all of it. Price and volume data are a shadow of that process. When the world changes in a way the past does not contain, the shadow changes too, and a formula learned from yesterday’s shadow cannot know it in advance. In principle, a computer that encoded every fundamental change in the world and every process driving it could forecast prices with high confidence. In practice the future is not in the data, and markets also react to the forecasts made about them. An alpha is therefore a small statistical edge that decays, and the only fair test is the one used here: out of sample, after costs, with the expectation that it will eventually stop working.

Takeaways

  1. The verifier is most of the work. Most channels exist to close a specific way of looking profitable without an edge; each was found by reading top-scoring outputs.
  2. Separate alpha from beta explicitly. Dollar-neutral construction, a factor reward and a measured holdout beta together kept the book’s market exposure at β = 0.09.
  3. Diversity must be selected for. RL converged on one family; a correlation cap, not the best scores, makes a book of independent bets.

Data and references

The JSON export contains the reward definitions, the JEPA stress configuration, the train curve, every checkpoint’s holdout and book results, the daily book and market returns on the holdout, and all 100 formulas from the untrained and step-500 models with train and holdout metrics. Recipe and experiment ladder: Multi-Reward RL, Part 3. Reward-design principles: the verifier design playbook.

Kakushadze, Z. (2016), 101 Formulaic Alphas, Wilmott Magazine 84, 72–80. Halperin, I. (2026), Observable Matrix Dynamics of Stocks, arXiv:2607.19005. Run 1790958861: Qwen3.8-27B, NF4 with a rank-32 LoRA, CISPO + REPO-R, GDPO, 600 steps, one H100. Holdout score per formula: CPCV-style Sharpe on the holdout year after realistic costs, minus an excess-turnover penalty, clipped to ±5, averaged over valid greedy answers to 50 unseen prompts. Market: equal-weight open-to-open return of the same 46 stocks.