Why Does Decision-Focused Learning Often Fail to Beat MSE?
Decision-focused learning (DFL) trains a predictor through the downstream objective it feeds — routing cost, portfolio return, knapsack value — instead of plain prediction error. The intuition is that a decision-aware loss should teach the model something MSE can't. This paper asks when that is actually possible.
One direction, two losses
A loss can only change learning by changing the parameter update. Through the chain rule, every gradient passes through the predictor Jacobian. If that Jacobian is rank one, nonzero per-example gradients from any loss lie on the same line — the task loss can rescale the step, not redirect it. A conditional spectral bound extends this to near-collinearity, and a batch-subspace characterization with counterexamples shows why these local statements imply neither common minimizers nor collinear batch updates.
A different loss need not provide an independent parameter-update direction.
What the experiments show
- Across 38 one-parameter equity configurations, DFL gains over MSE remain below 1.8%; a 385-parameter conditional predictor is also pointwise rank one.
- In validation-tuned shortest-path and knapsack experiments, full-capacity SPO+ reduces mean regret by 11.6% and 10.6%; only knapsack survives correction across eight comparisons.
- Holding expressivity fixed, invertible coordinate scaling lowers spectral effective rank and ordinary SGD gains; compensating for the scaling restores the original trajectories.
- Financial forward-target controls separate forecast accuracy from decision quality, and a matched neural comparison finds no aggregate DFL advantage in the tested architecture.
The practical takeaway
Before reaching for a decision-focused loss, check the geometry: if the predictor offers essentially one learning direction per example, the new loss has little room to help. Predictor geometry explains which directions are available — held-out decision quality remains the test of real benefit.
Frequently asked questions
What is decision-focused learning?
A training approach where a predictor is optimized for the quality of the downstream decision its predictions feed (for example, regret in an optimization problem) rather than for prediction error alone.
Why does decision-focused learning sometimes perform no better than MSE?
If the predictor's parameter Jacobian is rank one, gradients from the MSE loss and the task loss are collinear for each example, so the task loss cannot provide a new update direction.
How much does SPO+ improve over MSE?
In this paper's full-capacity experiments, SPO+ reduced mean regret by 11.6% on shortest path and 10.6% on knapsack, but only the knapsack gain survived Holm correction across eight comparisons.
Based on Jacobian Rank Collapse in Decision-Focused Learning (arXiv 2026, arXiv:2609.39261) by Aojie Yuan, Haiyue Zhang, Zijian Su. Written by Aojie (Justin) Yuan, USC Fortis Lab.