Qwen3.8-27B is trained with multi-reward RL to write one-line trading formulas in the language of WorldQuant’s 101 Formulaic Alphas. A deterministic verifier turns each formula into a daily dollar-neutral long-short book on 46 US stocks and scores it on 12 reward channels. On 50 prompts and a year of market data that training never saw, the mean score of its formulas went from −1.37 to +1.80, and an equal-weight book of them returned +16.2% with a market beta of 0.09.
This post documents the system: the formula language, how a formula becomes a portfolio, the difference between alpha and beta and how we keep the model from learning the market, each reward channel and the hack that created it, and the train and holdout results. The training recipe (GDPO, CISPO, REPO-R) and its experiment ladder are in Multi-Reward RL, Part 3, which described this gym with the domain hidden.
| Metric | Step 0 (untrained) | Step 500 (shipped) |
|---|---|---|
| Mean holdout score per formula | −1.37 | +1.80 |
| Valid formulas | 45 / 50 | 49 / 50 |
| Formulas with holdout score above zero | 4 / 45 | 47 / 49 |
| Equal-weight book: 1-year return after costs | −0.5% | +16.2% |
| Equal-weight book: Sharpe | −0.21 | 2.58 |
| Book beta to the market | — | 0.09 (correlation 0.19) |
0. System at a glance
| Policy | Qwen3.8-27B, NF4 weights + rank-32 LoRA, no visible reasoning |
|---|---|
| Action | One formula in <alpha>…</alpha>, at most 128 tokens |
| Verifier | Deterministic: parse → daily dollar-neutral book → costs → 12 reward channels |
| Universe | 46 US large caps, daily OHLCV from June 2012 |
| Train window | 2012-06-29 → 2025-08-20 (3,304 sessions) |
| Holdout | 2025-08-21 → 2026-08-21 (252 sessions), never used for training or selection |
| RL recipe | GDPO over 12 channels, CISPO loss, REPO-R, 12 answers per prompt, 600 steps, one H100 |
| Checkpoint rule | Best train-window CPCV of the 50 probe answers (step 500) |
1. The formula language
Source. In 2016 Zura Kakushadze published 101 trading alphas from WorldQuant, with the firm’s permission, as “explicit formulas – that are also computer code”. Eighty were in production at the time. They use daily open, high, low, close, volume, returns and vwap, have average holding periods of 0.6 to 6.4 days and an average pairwise correlation of 15.9%. The shortest:
Alpha#101: ((close - open) / ((high - low) + .001))If a stock closed well above its open relative to its daily range, go long the next day: the paper calls it a delay-1 momentum alpha. Our DSL uses the same vocabulary, with a fixed contract:
| Group | Element | Meaning |
|---|---|---|
| Fields | open high low close volume | daily prices and share volume |
returns | close-to-close return | |
vwap | proxy: (high + low + close) / 3 | |
| Cross-sectional | rank(x) | rank of x across the 46 stocks on each day, in [0, 1] |
scale(x) | rescale so that sum(|x|) = 1 | |
| Time series | delay(x, d) | value d sessions ago |
ts_delta(x, d) | x[t] − x[t−d] | |
ts_mean, ts_std, ts_sum, ts_max, ts_min (x, d) | rolling statistics over d sessions | |
ts_rank(x, d) | share of the window below today's value | |
ts_corr, ts_cov (x, y, d) | rolling correlation and covariance | |
ts_decay_linear, ts_ema, ts_skew, ts_kurt, ts_slope, … | weighted averages and shape statistics | |
| Element-wise | + − * /, abs, log, sign, power, min, max | per stock, per day |
| Limits | 1 ≤ d ≤ 252 | total nested lookback ≤ 252 sessions |
≤ 256 DSL tokens, depth ≤ 16 | the model's answer is also capped at 128 tokens | |
| Safety | recursive-descent parser | nothing the model writes is executed as code |
2. From formula to book
market data up to close(t)
→ evaluate formula one value per stock
→ rank across stocks values in [0, 1]
→ subtract the mean weights sum to 0 (long $ = short $)
→ divide by sum(|w|) gross exposure = 1 (0.5 long, 0.5 short)
→ fill at open(t+1) earn open(t+1) → open(t+2), minus costs| Stage | Costs |
|---|---|
| Training | 3.5 bps per unit of turnover: 2 fee + 1 slippage + 0.5 half-spread, plus a small fixed impact |
| Holdout | Per-stock square-root market impact for a $1M book, capacity penalty, 50 bps a year to borrow the shorts |
3. Alpha vs. beta: keeping the model off the market’s direction
Any daily return series can be split into a part that follows the market and a part that does not:
Beta is market exposure: a book with β = 1 gains 1% when the market gains 1%. Beta is cheap; an index fund sells it for a few basis points. Alpha is what remains when the market’s move is taken out: return from ranking stocks correctly against each other. A model that “learns to trade” by holding the market would look skilled in a rising year and fail in a falling one, so the gym is built to pay for alpha only.
| Layer | Mechanism | What it blocks |
|---|---|---|
| 1. Construction | Ranks are centred and scaled: long dollars always equal short dollars. Measured net exposure on the holdout: 0.00. | Direct market exposure |
| 2. Factor reward (R4) | Daily P&L is regressed on the market, size, 12-1 momentum, 5-day reversal and low volatility; the Sharpe those factors explain is penalised. | Indirect exposure, e.g. long high-beta stocks, short low-beta ones |
| 3. Prompt constraints | Briefs forbid known proxies, e.g. “do NOT lean on long-horizon price sums — they proxy market beta”. | Beta by design |
| 4. Motif guard (R10) | Pure price-level or price-momentum formulas with no predictive power lose their positive reward. | Slow tilts that ride a trend |
| 5. Holdout check | Beta and correlation of the shipped book against the market on the unseen year. | Anything the first four missed |
Measured on the holdout year. The equal-weight market of the same 46 stocks rose +27.3%. The book returned +16.2% with β = 0.09 and a correlation of 0.19, so the market explains only about 2–3 points of its return. On the 7 days when the market fell more than 1.5% (on average −1.9%), the book lost on average 0.3%. In a strongly rising year a long-only index made more money; the book made less but independently of the market, with a higher Sharpe (2.58 vs. 1.95) and half the drawdown (−4.3% vs. −8.6%).
4. The task
| Input | A design brief: mechanism, horizon, input rule, structural hint, forbidden pattern, risk preference |
|---|---|
| Output | One formula in <alpha>…</alpha> |
| Group | 12 answers per brief, ranked against each other on every channel |
| Prompts | 500 generated briefs → 226 kept: those the untrained model solved in some but not all of 8 tries |
A real brief from the training set:
Market: US-listed stocks trading on weekday sessions. Design an alpha whose PRIMARY mechanism is: mixed-signal extremes (rolling max or min of a price-volume composite). Target horizon: roughly 168-day windows. Inputs: must use exactly one ts_corr. Structure: take the sign of a time-series relationship. Constraint: do NOT lean on long-horizon price sums — they proxy market beta. Risk preference: make it robust to a volatility-regime change.
Briefs the model always or never solves give no learning signal. Dropping them was worth +0.47 on holdout in Part 3.
5. Twelve reward channels
GDPO standardises each channel within the group of 12 answers, multiplies by its priority and sums, so no channel can dominate through its numeric range. Two hard rules follow: an invalid answer never receives a positive advantage, and each group stays zero-sum.
| ID | Channel | Priority | Computes | Blocks |
|---|---|---|---|---|
| R1 | cpcv | 3.0 | mean − 0.5·std of net Sharpe over 15 CPCV splits | luck; fit to one period |
| R2 | pbo | 1.0 | penalty when a formula ranks far better in-sample than out of sample within its group (≥ −0.48) | overfitting |
| R3 | jepa_stress | 0.4 | score drop in 4 JEPA-generated markets, log-compressed (≥ −1) | fragility to regime change |
| R4 | factor_neutrality | 0.4 | −0.4 · clip(raw Sharpe − residual Sharpe, 0, 3) after 5 known factors | market beta and known factors in disguise |
| R5 | turnover | 1.0 | −2 · max(0, turnover − 0.5) | books that pay their profit away in costs |
| R6 | min_turnover | 0.5 | penalty below 6% a day; invalid below 2% a day | static and empty books |
| R7 | ic | 0.2 | 0.2 · clip(20 · IC, −1, 1); at most 0.05 while CPCV ≤ 0 | profit without prediction |
| R8 | novelty_pnl | 0.4 | 0.4 · (1 − max |corr| of daily P&L vs. a 1,024-entry archive); duplicates penalised | clones of known books |
| R9 | novelty_syntax | 0.2 | bonus for an unseen canonical form (windows bucketed) × (1 − P&L corr); −0.3 per duplicate | one template with jittered windows |
| R10 | motif_guard | 0.3 | bare one-motif or price-only formula with IC < 0.005 loses its positive reward | lazy single-motif formulas |
| R11 | complexity | 0.2 | −0.02 per expression node beyond 24 | bloated formulas |
| R12 | format | 1.0 | +0.05 for a reasoning block; always 0 in this run | (inactive without reasoning) |
R1 · CPCV core score
The train window is cut into 6 blocks; every pair of blocks is one test slice (15 slices), with 10 days purged and 5 embargoed at the edges. For each slice the annualised net Sharpe is computed:
Example: Sharpe 0.8 on every slice scores 0.8. Sharpe 3.0 on three slices and −0.2 on twelve averages 0.44 but scores −0.20. Missing tag: −1.0. Unparseable formula or lookback over 252: −0.8.
R5, R6 · Turnover band
Example, a real answer of the untrained model:
scale(ts_rank(ts_delta(volume, 23), 47) - ts_rank(ts_delta(volume, 23), 47))The expression subtracts a value from itself, so every weight is zero every day: no trades, no costs, no losses. Under high costs this beats any honest formula. Its turnover is 0, below the 2% floor, so it is invalid.
R3 · JEPA stress test
A JEPA world model is trained only on the train window (2012 to August 2025) and then frozen. It reads a window of the 46-stock panel with 25% of the stock-day cells masked in spans of 8 days, encodes one latent per stock, and predicts the latent of the next window. The target is an exponential-moving-average copy of its own encoder, so the model learns to predict market states rather than raw prices.
A second target comes from Observable Matrix Dynamics (OMD; Halperin, 2026), a dynamical-systems view of the market: it follows the trajectory of a fixed-size matrix built from rolling stock correlations and of its spectrum, and finds that the matrix’s effective dimension collapses in crises such as 2008 and 2020. We use the same idea as an auxiliary loss: the predicted latent must also recover the geometry of the next window’s correlation matrix of the 46 stocks.
| Target | Computed as | What it captures |
|---|---|---|
| Mean correlation | average off-diagonal entry of | how much stocks move together |
| Market-mode share | weight of the single market factor | |
| Effective factor fraction | divided by 46 | effective dimension of the market; collapses in crises |
| Residual factor fraction | the same without the market mode | diversity left beyond the market |
| Projector drift | rotation of the top-3 eigenvector subspace from the past to the next window | rotation between sectors and themes |
| Leading spectrum | shares of the top-5 eigenvalues | shape of the dominant modes |
At reward time the latent is shifted along four channels — dispersion between stocks, trend flip, correlation and volatility regime — and decoded into complete price and volume panels. The formula is re-scored in each world with the same CPCV and costs, and the average drop becomes the penalty . A fifth channel, a crash, was removed because it ranked formulas worse than chance (AUC 0.43).
Other channels in one line each
- R2 stability: first in-sample and tenth out of sample within the group is penalised; the penalty grows with the drop.
- R4 factor neutrality: a formula whose Sharpe falls from 1.5 to 0.3 after removing the five factors pays −0.4 × 1.2 = −0.48.
- R7 IC: the daily correlation of yesterday’s weights with today’s returns; an IC of 0.05 earns the full bonus.
- R8, R9 novelty:
ts_mean(close, 20)andts_mean(close, 21)count as the same idea; new code that trades an old book earns nothing. - R10 motif guard, R11 complexity, R12 format: hygiene; R12 is inactive because this run answers without reasoning.
6. Reward hack log
The first version had a profit score and a format check. Each row below is a way the model found to score well without a real edge, and the change that closed it.
| Date | Observed | Change |
|---|---|---|
| Jul 4 | All answers collapsed onto one formula | Syntax novelty (R9) |
| Jul 5 | Different code, same book | P&L novelty with an archive seeded by Alpha101-style formulas (R8); smoke test: 2 → 5 distinct valid formulas |
| Jul 6 | Static books: slow price-level tilts with IC 0.000 took the top rewards (0.55–0.63) | IC reward (R7); not enough alone (0.63 vs. 0.34 for momentum), so a minimum-turnover penalty the same day: 0.63 → 0.25 |
| Jul 7 | High-CPCV books trading 0.2–0.6% of the book a day still won | Hard floor: below 2% a day = invalid (R6) |
| Jul 7–8 | Formulas tuned to the exact scored history | JEPA stress test (R3): static books −12 offline, honest formulas ≈ 0; next day log-compressed, as 24% of groups had ≥ 2 formulas at the cap |
| Jul 9 | 69–83% of answers were one momentum template with jittered windows; every arm failed a residual-Sharpe check | Window bucketing (5/21/63/126/252) before novelty; factor-neutrality penalty (R4) |
| Jul 10 | 63–83% of invalid answers still got positive advantage | Validity as a hard constraint on the advantage |
| Jul 23 | Checkpoint choice could leak the holdout | Selection by train-window CPCV only; CPCV floor on the advantage |
| Jul 26 | Stress crash scenario ranked formulas worse than chance (AUC 0.43 vs. 0.55–0.64); the turnover band charged top formulas −0.43 | Crash scenario removed; band moved to 3–6% a day (top decile now −0.09) |
| Aug 9 | Signal and fill on the same close: a subtle look-ahead | Fill at the next open, earn open-to-open returns |
Every row was found by reading the top-scoring formulas. A rising reward curve was the first symptom of each hack, not evidence of learning.
7. Train vs. holdout
The best recipe from Part 3: CISPO + REPO-R, LoRA, 226 prompts, 600 steps. Part 3 shows the ladder of experiments that led to it, from DAPO at −0.47 to this recipe at +1.80.
| Step | Train CPCV | Holdout score | Valid | Formula families | Book return | Book Sharpe |
|---|---|---|---|---|---|---|
| 0 (untrained) | −0.98 | −1.37 | 45 / 50 | 45 | −0.5% | −0.21 |
| 100 | −0.36 | −0.47 | 47 / 50 | 47 | +2.6% | 0.95 |
| 200 | −0.15 | −0.13 | 49 / 50 | 49 | +4.3% | 1.36 |
| 300 | +0.29 | +0.85 | 48 / 50 | 41 | +10.9% | 2.14 |
| 400 | +0.23 | +0.51 | 49 / 50 | 45 | +8.3% | 2.00 |
| 500 (shipped) | +0.82 | +1.80 | 49 / 50 | 31 | +16.2% | 2.58 |
| 600 | +0.69 | +1.58 | 48 / 50 | 34 | +15.2% | 2.56 |
Step 500 has the highest train-window CPCV (0.82) and was shipped on that basis; it also has the highest holdout score. Holdout and train-window scores move together but not in lockstep: step 300 is better on holdout than step 400, and step 600 is slightly worse than step 500.
| Formula | Train CPCV | Turnover / day | Holdout Sharpe | Holdout return |
|---|---|---|---|---|
rank(ts_max(volume / vwap, 11) / ts_mean(close, 21)) | 0.95 | 3.1% | 3.13 | +19.3% |
rank(ts_max(volume / vwap, 34) / ts_std(returns, 34)) | 0.85 | 2.8% | 2.97 | +17.8% |
rank(ts_std(volume / (high - low), 34)) | 0.82 | 2.3% | 2.96 | +18.3% |
rank(ts_std(volume, 8) / ts_std(close, 8)) | 0.59 | 14.9% | 2.70 | +15.7% |
rank(ts_std(volume, 17) / ts_mean(volume, 17) + ts_std(volume, 17) / ts_mean(volume, 17)) | −0.47 | 23.2% | −1.09 | −7.6% |
Convergence. All 49 trained answers use volume, and most rank stocks by the variability of their trading volume relative to price level or price volatility. Distinct formula families fell from 45 to 31. We have not established why this signal worked over this year; it is one idea expressed many ways, so 49 formulas are far fewer than 49 independent bets.
8. Selection: which checkpoint, which formulas
Checkpoint. Chosen by train-window CPCV of the 50 probe answers; the holdout is computed once, afterwards (rule in place since July 23).
Formulas for a book. Ranking by CPCV and taking the top 50 is the obvious choice and a poor one. On our formula archives that book had a median member CPCV of 0.71–0.75 but behaved like 1.8–2.6 independent bets (pairwise P&L correlation 0.50–0.63), with the worst drawdown of the selectors tested. A correlation-capped selector:
sort candidates by CPCV, descending
book = []
for f in candidates:
if turnover(f) < 0.02: skip
if max |corr(pnl(f), pnl(g))| >= 0.5 for any g in book: skip # redundant bet
book.append(f)On two archives it reached a Sharpe of 2.24 and 2.00 against 1.85 and 1.39 for reward-ranked picks, with about 20 effective bets instead of 6–12 and half the drawdown, although its members’ median CPCV was only 0.38. The books in this post use neither selector: they average every valid answer with equal weight.
9. Limits
Our evaluation protocol marks a single-year result like this as no claim: evidence that the system learns, not evidence of a durable edge.
- Survivorship. The 46 stocks are today’s large caps followed back to 2012; companies that shrank, merged or failed are absent.
- One year, one seed. The holdout is 252 sessions from one market regime; the run is a single seed.
- A gap in the probe. 15 of the 49 trained answers trade less than 2% of the book a day on the train window. Training would mark them invalid; the holdout probe scores CPCV only and does not apply the floor.
- Neutrality. The book is dollar-neutral with a measured β of 0.09, but not sector-neutral, and our
vwapis a proxy.
Why no model predicts the market with certainty
A market price is a projection of the real world: what companies sell and spend, what central banks and governments decide, wars, weather, new technologies, and the expectations of millions of people about all of it. Price and volume data are a shadow of that process. When the world changes in a way the past does not contain, the shadow changes too, and a formula learned from yesterday’s shadow cannot know it in advance. In principle, a computer that encoded every fundamental change in the world and every process driving it could forecast prices with high confidence. In practice the future is not in the data, and markets also react to the forecasts made about them. An alpha is therefore a small statistical edge that decays, and the only fair test is the one used here: out of sample, after costs, with the expectation that it will eventually stop working.
Takeaways
- The verifier is most of the work. Most channels exist to close a specific way of looking profitable without an edge; each was found by reading top-scoring outputs.
- Separate alpha from beta explicitly. Dollar-neutral construction, a factor reward and a measured holdout beta together kept the book’s market exposure at β = 0.09.
- Diversity must be selected for. RL converged on one family; a correlation cap, not the best scores, makes a book of independent bets.
Data and references
The JSON export contains the reward definitions, the JEPA stress configuration, the train curve, every checkpoint’s holdout and book results, the daily book and market returns on the holdout, and all 100 formulas from the untrained and step-500 models with train and holdout metrics. Recipe and experiment ladder: Multi-Reward RL, Part 3. Reward-design principles: the verifier design playbook.
Kakushadze, Z. (2016), 101 Formulaic Alphas, Wilmott Magazine 84, 72–80. Halperin, I. (2026), Observable Matrix Dynamics of Stocks, arXiv:2607.19005. Run 1790958861: Qwen3.8-27B, NF4 with a rank-32 LoRA, CISPO + REPO-R, GDPO, 600 steps, one H100. Holdout score per formula: CPCV-style Sharpe on the holdout year after realistic costs, minus an excess-turnover penalty, clipped to ±5, averaged over valid greedy answers to 50 unseen prompts. Market: equal-weight open-to-open return of the same 46 stocks.