Trading Journals and Performance Metrics
Traders who do not keep a journal are running a business without accounting records. They have impressions about their performance — typically optimistic ones, filtered through the confirmation bias and selective memory documented in the psychology course — but no data. Without data, there is no feedback loop for improvement. Every pattern of error that repeats is costing real money while remaining invisible to introspection. A trading journal transforms pattern recognition from the domain of fallible human memory into the domain of verifiable statistical evidence.
1. What the Journal Must Record
A trading journal is only as useful as the data it contains. Recording only P&L is insufficient; the journal must capture enough context to enable root-cause analysis of both winning and losing trades. The minimum required fields for every trade:
| Field | Why it matters |
|---|---|
| Date, ticker, direction | Basic identification; enables filtering by symbol, sector, time period |
| Setup type | Enables comparison of performance by setup (breakout vs pullback vs mean reversion) |
| Entry price, stop, target | Enables calculation of planned R:R at entry; comparison with actual execution |
| Exit price, exit reason | Distinguishes planned exits (target/stop hit) from emotional exits |
| Position size, dollar P&L | Required for expectancy and all performance metric calculations |
| Planned vs actual R:R | Reveals if targets are being cut short relative to initial plan |
| Market conditions | Broad market trend, VIX level; enables regime-specific performance analysis |
| Execution quality rating | Self-assessment 1–5: was the trade executed per the plan? Enables separating edge from execution quality |
The post-trade note is the most frequently skipped but analytically critical component. Within one hour of closing every trade, write 2–4 sentences: what the thesis was, whether the exit was planned or emotional, what you would do differently, and any pattern you noticed about your execution quality during the trade. This structured reflection creates the raw material for periodic performance reviews. Without these notes, reviewing old trades months later requires re-reading chart patterns without the psychological context that drove the decisions.
2. Expectancy: The Foundation Metric
Expectancy is the single most important performance metric because it directly measures the statistical value of each trade taken — the average dollar return per dollar risked across all trades in the sample. A positive expectancy confirms that the strategy has edge; a negative expectancy confirms it does not, regardless of recent P&L.
Formula: Expectancy = (Win Rate × Average Win) − (Loss Rate × Average Loss)
Worked example. Analysis of 150 trades from the journal: Win Rate = 48% (72 wins / 150 trades). Average win = $420. Loss Rate = 52% (78 losses). Average loss = $280. Expectancy = (0.48 × $420) − (0.52 × $280) = $201.60 − $145.60 = +$56.00 per trade. This means that on average, every trade taken generates $56 in profit — the strategy has genuine edge. If the trader takes 200 trades per year, the expected annual profit is $11,200, independent of the psychological experience of individual winning and losing streaks.
Expectancy also enables diagnosis of specific problems. A strategy with positive expectancy but negative actual P&L reveals an execution problem: the theoretical edge is present but not being captured. Common causes: cutting winners early (reducing average win below the plan), holding losers too long (increasing average loss above the plan), or taking trades that don’t meet the strategy criteria (reducing win rate below baseline). The trading plan from Course 35 defines the expected metrics; the journal compares actual performance against them.
3. Profit Factor, Sharpe Ratio, and Risk-Adjusted Returns
Profit Factor is the ratio of gross profits to gross losses across all trades: PF = Total Winning P&L ÷ Total Losing P&L. A profit factor above 1.0 means the strategy is profitable; above 1.5 is solid; above 2.0 is excellent and indicates a strong edge. The profit factor has an advantage over raw P&L as a performance measure because it is normalised by loss magnitude — a strategy that produces $100,000 in gains with $80,000 in losses (PF = 1.25) is generating more genuine alpha per dollar risked than one that produces $100,000 in gains with $50,000 in losses using twice the position size (PF = 2.0 but at double the risk).
Sharpe Ratio is the gold standard for risk-adjusted return measurement in institutional finance. Formula: Sharpe = (Strategy Return − Risk-Free Rate) ÷ Standard Deviation of Returns. A Sharpe above 1.0 indicates that the strategy is generating more than one unit of excess return per unit of volatility — acceptable. Above 1.5 is good; above 2.0 is excellent and rarely sustained. The Sharpe ratio’s primary weakness for active equity traders is its assumption that volatility is symmetric — it penalises upside volatility (large wins) the same as downside volatility (large losses), which is analytically incorrect for strategies where the distribution of outcomes is positively skewed.
The Sortino Ratio addresses this limitation by replacing the standard deviation denominator with the downside deviation (standard deviation of negative returns only): Sortino = (Strategy Return − Risk-Free Rate) ÷ Downside Deviation. For strategies with positive skew — small, frequent losses and occasional large wins — the Sortino ratio provides a more accurate measure of risk-adjusted performance than the Sharpe. Record both in the periodic performance review from Course 35’s monthly review cycle.
4. Maximum Drawdown and Calmar Ratio
Maximum drawdown (MDD) measures the largest peak-to-trough equity decline experienced during the strategy’s history. Formula: MDD = (Trough Equity − Peak Equity) ÷ Peak Equity, expressed as a negative percentage. A strategy with an all-time peak equity of $30,000 that subsequently declined to $24,000 has a maximum drawdown of −20%. Maximum drawdown is the primary risk metric for evaluating whether a strategy is survivable in live trading: a trader who cannot psychologically tolerate a 20% drawdown should not be running a strategy with a historical MDD of 20%, because doing so guarantees abandonment at the worst moment.
The relationship between maximum drawdown and required recovery is non-linear and frequently underestimated: recovering from a 20% drawdown requires a 25% gain; a 33% drawdown requires 50%; a 50% drawdown requires 100%. This asymmetry reinforces the capital preservation emphasis of the risk management framework: every point of drawdown prevention saves disproportionately more than its face value in required recovery effort.
The Calmar Ratio normalises return by maximum drawdown: Calmar = Annualised Return ÷ Maximum Drawdown (absolute value). A strategy returning 18% annually with a 12% maximum drawdown has a Calmar of 1.5 — generating 1.5 units of annual return per unit of worst-case drawdown experienced. Comparing Calmar ratios across strategies allows selection of the highest-return-per-unit-drawdown option among candidates, which is the correct optimisation target for a trader who must psychologically sustain the strategy through difficult periods.
5. Segmentation Analysis: Finding Patterns in the Data
Aggregate performance metrics are useful for strategy evaluation but insufficient for strategy improvement. The analytical value of the journal emerges from segmentation — filtering trades by specific characteristics to identify where the strategy performs well and where it loses money.
The most productive segmentation analyses for equity traders:
- By setup type. If your trend-following entries in bullish EMA stacks produce 2.2 expectancy but your mean-reversion entries produce −0.4 expectancy, immediately eliminate mean-reversion entries and concentrate on the profitable setup type. Most traders unknowingly carry a subset of losing setup types that drag down an otherwise profitable strategy.
- By market regime. Compare performance in bullish (SPY above 50 SMA) vs bearish/neutral regimes. Most equity trend-following strategies perform significantly better in bullish broad market environments. If your strategy loses money when SPY is below its 50 SMA, implement a rule to stop taking new long positions in that regime.
- By holding duration. Compare trades held <3 days vs 3–10 days vs >10 days. Many traders discover that their profitable setups require a specific holding period window to develop fully, and that cutting them short or holding them past the window systematically degrades returns.
- By execution quality rating. Trades where you rated your execution 4–5/5 should show significantly better outcomes than trades rated 1–2/5. If they don’t, the execution quality rating system needs refinement. If they do, this provides the empirical motivation to reject low-quality-execution setups in real time.
- By sector. Compare performance in technology vs financials vs healthcare. Some traders discover strong sector biases in their edge that suggest concentrating on familiar sectors and avoiding others.
6. Distinguishing Skill from Variance
A critical but frequently overlooked analytical challenge: how many trades are required before a performance result can be attributed to skill rather than variance? The answer is uncomfortable: far more trades than most retail traders accumulate before drawing conclusions. With a 55% win rate and 1:1.5 average R:R (a positive-expectancy strategy), the 95% confidence interval for the true win rate in a 50-trade sample spans approximately 41%–69% — an enormous range that includes outcomes consistent with both genuinely positive and genuinely negative expectancy strategies.
The practical standard: require a minimum of 100 completed trades before drawing meaningful conclusions about strategy edge, and 200–300 before making significant strategy modifications based on performance data. This is why paper trading evaluation periods should be substantial, and why the patience required to accumulate sufficient live trading data before abandoning or modifying a strategy is itself a skill. Short-sequence performance is dominated by variance; long-sequence performance reveals genuine edge.
Use the journal data to calculate the standard error of the win rate: SE = √(p × (1−p) / n), where p is the observed win rate and n is the number of trades. Two standard errors above and below the observed win rate gives an approximate 95% confidence interval. As n increases, the confidence interval narrows and the performance estimate becomes more reliable. Until the interval excludes 50% (the break-even win rate for a 1:1 R:R strategy), no statistically valid conclusion about edge can be drawn.
Key Takeaways
| Metric | Formula / benchmark |
|---|---|
| Expectancy | (Win% × Avg Win) − (Loss% × Avg Loss). Must be positive for viable strategy. |
| Profit Factor | Total wins ÷ Total losses. >1.5 solid; >2.0 excellent. |
| Sharpe / Sortino | Excess return ÷ volatility (Sharpe) or downside deviation (Sortino). >1.5 target. |
| Max Drawdown | Largest peak-to-trough equity decline. Must be psychologically tolerable; asymmetric recovery cost. |
| Calmar Ratio | Annual return ÷ Max Drawdown. Optimise for highest return per unit drawdown. |
| Minimum sample | 100 trades minimum before conclusions; 200–300 before strategy modifications. Variance dominates small samples. |
- Stock P&L Calculator — calculate accurate net P&L for each trade entry in your journal, including commissions on both sides.