Overcoming Context Drift: Trust Guided Decision Transformer Uses Prediction Error and Conformal Prediction for Robust Offline RL
長序列決策救星:Trust Guided Decision Transformer 透過預測誤差與共形預測克服上下文漂移
Decision Transformers (DT) degrade on long rollouts as context drifts out of distribution. This paper shows that next-state prediction error acts as a reliable signal for this drift. To address this, the authors develop Trust Guided Decision Transformer (TGDT). TGDT evaluates several recent context suffixes, filters out those exceeding a calibrated error threshold using split conformal prediction, and then applies a frozen critic to choose the best action from the trusted candidates. Experiments on D4RL benchmarks demonstrate that TGDT reduces failure rates and outperforms vanilla DT and value-only selection.
Key points
Detecting Context Drift
Uses the model's own next-state prediction error to directly flag when the current context has drifted and become unreliable.
Trust-First Filtering
Reverses value-only selection by first eliminating untrusted contexts, ensuring the critic only selects actions from reliable historical states.
Conformal Calibration
Applies split conformal prediction calibrated against offline data to establish rigorous, non-parametric thresholds for acceptable error.
Mitigating Compounding Failures
Significantly reduces persistent high-error runs and improves cumulative returns on challenging D4RL control tasks.
How it works
Why it matters
A key bottleneck in offline RL is Decision Transformer failure over long horizons due to compounding extrapolation errors. TGDT offers a plug-and-play inference-time solution without retraining the base transformer model. This is highly valuable for deploying robust, safety-critical systems like autonomous vehicles and robotics where out-of-distribution drift must be actively managed.
Who it affects
- AI Developer
- AI Researcher
How to use it
- 1Decision filtering for autonomous driving under out-of-distribution road conditions
- 2Trajectory calibration in long-horizon robotic manipulation and navigation
- 3Anomaly detection and safeguarding during offline RL agent deployment
Limitations & caveats
- Relies heavily on a pre-trained frozen critic; sub-optimal critic quality will bound the final action selection performance.
- Evaluating multiple candidate context suffixes at each step introduces additional computational overhead during online inference.
Related
FurE: 10x Faster 3D Animal Fur Reconstruction Without Animal Datasets
FurE:免用動物毛髮資料集,實現 10 倍加速的 3D 動物毛髮重建技術
FurE is an efficient 3D animal fur reconstruction method that leverages a human-hair trained PCA decoder and Gaussian Frosting to achieve 10x faster, highly detailed, and editable groom reconstruction without animal datasets.
Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency
自我監督信心訓練:免於刻意「學會停止」即可提升 LLM 推理效率
Researchers found that training reasoning models to predict their own confidence at intermediate steps naturally reduces generated tokens by up to 25% at matched accuracy, without explicitly optimizing for length or stopping.
First-Order Stationarity of Reverse Diffusions: Bridging Optimization and Sampling
逆向擴散的一階駐點性:連結優化與生成採樣的數學機制
This research establishes a first-order optimization theory for diffusion models, proving that SDE-based reverse Langevin diffusions contract Fisher divergences exponentially under strongly convex noising.