OIH vs XOP: does extreme 10-day oil-services underperformance predict next-10-day OIH outperformance?
What happens after an oil-services basket gets crushed relative to its upstream peers for ten straight days? For OIH versus XOP, the mean-reversion story says extreme underperformance is a washout that sets up a rebound — that service names get repriced before they catch up. Over roughly three years of daily data, that story fails.
Across 609 eligible trading days, only 19 independent events pushed the 10-day OIH lag past a rolling 90th-percentile threshold. In the next ten days after those events, OIH returned 0.93 percentage points less than XOP on average — worse than the +0.08% baseline. With a t-stat of -1.15 and a p-value of 0.26, that tilt is indistinguishable from noise. The full methodology and statistical breakdown are in the analysis below.
For OIH over the past ~3 years, when its 10-day total return underperforms XOP's by more than the trailing 90th-percentile spread, does OIH outperform XOP over the next 10 trading days? Thesis: extreme oil-services underperformance versus upstream E&Ps marks a positioning washout that mean-reverts as service capacity gets repriced.
How this was measured
Daily closes were built from OIH_df and XOP_df minute bars. For each trading day t, the 10-day total-return differential was computed as XOP_ret10 - OIH_ret10, so positive values mean OIH underperformed XOP over the trailing 10 trading days. The trailing 90th percentile of that underperformance spread was estimated using a rolling 252-day lookback lagged by one day, so the trigger is known at close t-1. A trigger day is one where OIH's trailing 10-day return underperforms XOP's by more than that trailing 90th-percentile spread. Future outcome is the next-10-trading-day excess return of OIH over XOP. Because 10-day returns overlap, triggers were deduplicated to require at least 10 trading days between independent event anchors. Results are compared against non-trigger days using Welch's two-sample t-test.
The key numbers
Reading the numbers
The signal fired 19 independent times; after those days OIH lagged XOP by 0.93% on average over the next 10 days, versus a 0.08% gain otherwise. With p=0.26, that gap is not distinguishable from random chance, so the washout edge isn't demonstrated.
The charts
The wiggly line is how much XOP has beaten OIH over the trailing 10 days, and the smoother line is the moving 90th-percentile trigger, which stayed between roughly 3% and 5.5% (average 3.7%). The signal is meant to fire only when the wiggly line jumps above that trigger; that happens rarely, and mostly the line sits below it. This also helps explain why only 19 independent trigger events remained after overlapping 10-day windows were cleaned up.
This histogram shows the 10-day OIH-minus-XOP returns that followed each independent trigger. The average is -0.93%, the worst outcome is about -11%, and the best is about +5.9%. Only about 37% of the 19 trigger days ended with OIH actually beating XOP. If the mean-reversion thesis were working, the distribution would lean positive; instead it leans negative.
This bar chart directly compares trigger days against all non-trigger days. Trigger days average -0.93% forward excess return, while non-trigger days average +0.08%, so the signal underperforms the baseline by about 1.0 percentage point. The t-statistic of -1.15 and p-value of 0.26 say that gap could easily be noise. For the thesis to be supported, the trigger bar would need to be positive and clearly above the non-trigger bar; instead it is below zero.
Most recent independent trigger days
| anchor_date | under_spread_10d | threshold_p90 | fwd_excess_10d |
|---|---|---|---|
| 2024-02-12 | 0.0582 | 0.037 | -0.0067 |
| 2024-04-11 | 0.041 | 0.0341 | -0.0242 |
| 2024-04-29 | 0.0417 | 0.036 | 0.0366 |
| 2024-06-03 | 0.0467 | 0.0356 | 0.0262 |
| 2024-08-08 | 0.0348 | 0.0345 | -0.0133 |
| 2024-10-07 | 0.0381 | 0.0366 | -0.0143 |
| 2024-11-20 | 0.0615 | 0.0342 | 0.0332 |
| 2025-02-19 | 0.0372 | 0.0318 | 0.0189 |
| 2025-03-19 | 0.0466 | 0.0334 | 0.0022 |
| 2025-04-08 | 0.0371 | 0.0348 | -0.0139 |
| 2025-04-28 | 0.0393 | 0.0334 | -0.0015 |
| 2025-05-13 | 0.0498 | 0.0334 | -0.0103 |
| 2025-06-24 | 0.0438 | 0.0371 | 0.0594 |
| 2025-11-12 | 0.0398 | 0.034 | 0.0025 |
| 2026-03-02 | 0.0487 | 0.0349 | -0.1098 |
| 2026-03-16 | 0.1098 | 0.0466 | -0.0352 |
| 2026-06-09 | 0.0569 | 0.0398 | -0.0399 |
| 2026-06-26 | 0.073 | 0.0398 | -0.0389 |
| 2026-07-14 | 0.0556 | 0.0534 | -0.0484 |
The takeaway
The short answer is no: after a 10-day stretch where OIH badly lagged XOP (past the rolling 90th-percentile gap), OIH did not bounce back over the next 10 days — it actually did slightly worse than usual. Across 609 eligible trading days there were only 19 independent trigger episodes, and those were followed by an average -0.93% OIH-minus-XOP return versus +0.08% on normal days, a negative edge of about -1.01%. Fewer than 4 in 10 triggers were followed by OIH outperformance (37%), below the 47% baseline. With a t-stat of -1.15 and p=0.26, this is not a real signal — it's statistically indistinguishable from a coin flip, and the thin trigger count means even the apparent negative tilt could easily be noise. The washout/mean-reversion thesis isn't supported here; at best this is a null result, and the data lean mildly the other way. Practical takeaway: don't treat extreme oil-services underperformance as a buy signal against upstream names. The evidence over this window is too weak to act on, and any edge is at best unproven.
The fine print
- 10-day returns overlap; even with 10-day spacing between triggers, residual autocorrelation can inflate the effective sample size.
- The 90th-percentile trigger uses a rolling 252-day lookback, so results depend on the lookback window and the prevailing volatility regime.
- Returns are price returns only — ETF distributions and expense ratios are not included.
- Only 19 independent events from one pair over ~3 years; the result may be regime-specific and doesn't generalize.