Aivora
arXivAI ResearchAdvanced

Overcoming Context Drift: Trust Guided Decision Transformer Uses Prediction Error and Conformal Prediction for Robust Offline RL

長序列決策救星:Trust Guided Decision Transformer 透過預測誤差與共形預測克服上下文漂移

2 min read
Overcoming Context Drift: Trust Guided Decision Transformer Uses Prediction Error and Conformal Prediction for Robust Offline RL
The 30-second version

Decision Transformers (DT) degrade on long rollouts as context drifts out of distribution. This paper shows that next-state prediction error acts as a reliable signal for this drift. To address this, the authors develop Trust Guided Decision Transformer (TGDT). TGDT evaluates several recent context suffixes, filters out those exceeding a calibrated error threshold using split conformal prediction, and then applies a frozen critic to choose the best action from the trusted candidates. Experiments on D4RL benchmarks demonstrate that TGDT reduces failure rates and outperforms vanilla DT and value-only selection.

Key points

01

Detecting Context Drift

Uses the model's own next-state prediction error to directly flag when the current context has drifted and become unreliable.

02

Trust-First Filtering

Reverses value-only selection by first eliminating untrusted contexts, ensuring the critic only selects actions from reliable historical states.

03

Conformal Calibration

Applies split conformal prediction calibrated against offline data to establish rigorous, non-parametric thresholds for acceptable error.

04

Mitigating Compounding Failures

Significantly reduces persistent high-error runs and improves cumulative returns on challenging D4RL control tasks.

How it works

TGDT Decision Flow Architecture
Provide sequenceEvaluate multiple suffixesCompare with thresholdPass trusted candidates onlyOutput highest-value actionState & History InputGenerate SuffixCandidatesPredict Next StateErrorConformal Filter(Trust)Critic Action SelectionExecute Best Action

Why it matters

A key bottleneck in offline RL is Decision Transformer failure over long horizons due to compounding extrapolation errors. TGDT offers a plug-and-play inference-time solution without retraining the base transformer model. This is highly valuable for deploying robust, safety-critical systems like autonomous vehicles and robotics where out-of-distribution drift must be actively managed.

Who it affects

  • AI Developer
  • AI Researcher

How to use it

  1. 1Decision filtering for autonomous driving under out-of-distribution road conditions
  2. 2Trajectory calibration in long-horizon robotic manipulation and navigation
  3. 3Anomaly detection and safeguarding during offline RL agent deployment

Limitations & caveats

  • Relies heavily on a pre-trained frozen critic; sub-optimal critic quality will bound the final action selection performance.
  • Evaluating multiple candidate context suffixes at each step introduces additional computational overhead during online inference.

Related

FurE: 10x Faster 3D Animal Fur Reconstruction Without Animal Datasets
arXivAI Research

FurE: 10x Faster 3D Animal Fur Reconstruction Without Animal Datasets

FurE:免用動物毛髮資料集,實現 10 倍加速的 3D 動物毛髮重建技術

FurE is an efficient 3D animal fur reconstruction method that leverages a human-hair trained PCA decoder and Gaussian Frosting to achieve 10x faster, highly detailed, and editable groom reconstruction without animal datasets.

2 min read