REDACT: Robust Perceptive Locomotion
under Unseen Visual Corruption

A depth-conditioned locomotion framework that combines robust encoder design and consensus gating to retain useful terrain information under visual corruption unseen during training.

Natapat Kirdwichai  ·  Tobias Driskell-Poole  ·  Andrei Sontea  ·  Jadu Dash
Muhammad Burhan Hafez  ·  Danesh Tarapore

University of Southampton

Zero-shot real-world traversal. A policy trained in simulation traverses structured obstacles and forest terrain under unfamiliar visual conditions.

Abstract

Depth-conditioned locomotion policies have demonstrated impressive agile maneuvers, but can produce unpredictable actions when observations fall outside their training distribution. Occlusion, invalid returns, sensor noise, and visual distractors can shift deployment observations away from nominal simulated depth. While synthetic sensor augmentation targets specified degradations, it does not by itself define behavior under corruption families omitted from training. To address gaps in training-time coverage, we present REDACT (Retaining Evidence Despite Artifacts for Continued Traversal), a teacher–student framework combining an improved visual encoder architecture, persistent feature masking, and a novel consensus-gating algorithm to retain useful depth information under unmodeled corruption. The gate uses approximate conformal calibration on clean observations alone, requiring no prior knowledge of the corruption type. Trained on clean simulated depth, REDACT retains useful visual information under unseen corruption, supporting higher traversal success than existing parkour baselines. Evaluation of depth augmentation across corruption families further shows that REDACT improves robustness where augmentation coverage is missing. Real-world trials demonstrate zero-shot transfer to structured and forested environments with unfamiliar scene content.

Video overview

Method, controlled comparisons, and real-world trials.

Method

Teacher training, student learning, and consensus-gated deployment.

REDACT architecture: teacher training, depth-conditioned student training with persistent masks and cell prediction heads, and deployment with offline-calibrated consensus gating.

(a) Teacher training

Regularized multi-query attention structures the privileged teacher’s terrain representation to support learning from onboard depth.

(b) Student training

A residual encoder preserves spatial features. Cell predictions share a common target, while persistent masking prepares the recurrent policy to use partial visual evidence.

(c) Deployment

Clean-calibrated thresholds reject inconsistent features before pooling. Recurrent memory helps retain useful terrain information for continued control.

Depth encoder design

The encoder is designed to preserve useful terrain structure while limiting how visual corruption changes the representation supplied to the policy.

Learned downsampling
A deeper residual encoder replaces intermediate max pooling with strided convolutions. Max pooling selects local activation peaks, which can favour responses to artifacts. This is a concern when depth dropouts are represented at the 2 m clipping limit; learned downsampling provides an alternative to selecting those peaks.
Group normalization
Normalizing groups of intermediate feature channels regulates their scales, helping control variations in activation magnitude as observations change.
Spectral constraints
Capping weight spectral norms limits amplification through the constrained layers, helping control how input perturbations propagate toward the visual latent.
Retaining terrain information through shared predictions
Overlapping receptive fields give feature cells access to shared terrain context, while a common prediction target encourages each cell to contribute to the same terrain representation. Cells can nevertheless respond differently to visual corruption, allowing consensus gating to remove inconsistent contributions before they enter the recurrent policy. Persistent masking prepares the policy to use the remaining features, while recurrent memory helps retain useful terrain information from earlier observations.
Left: representative shallow ConvNet with flattening and fully connected layers. Centre: REDACT residual encoder retaining 4 by 6 spatial feature cells. Right: a stride-2 residual block with group normalization and SiLU activations.
Representative ConvNet (left), REDACT’s residual encoder (centre), and a stride-2 residual block (right). GN denotes group normalization.

Experiments

Comparisons in simulation and zero-shot transfer to real terrain.

Robustness beyond nominal observations

Trained without depth-image augmentation, REDACT maintains higher traversal success than REAL [1] and Extreme Parkour [2] across the tested corruption sweeps. The advantage also extends to pillar hurdles, where unfamiliar geometry changes the visual scene without changing the required maneuver.

Success-rate curves comparing REDACT, Extreme Parkour and REAL under flying pixels, Gaussian noise, Perlin dropout, vine occlusion and pillar hurdles. REDACT achieves higher success across the tested sweeps.
All three policies are trained without depth-image augmentation. Lines show means and shaded bands show 95% confidence intervals. In the pillar-hurdle sweep, obstacle heights vary rather than image-corruption severity; pillars are three times the hurdle height.

Generalization beyond augmentation coverage

Training on one corruption type does not ensure robustness to others. Augmentation transfers unevenly across the ConvNet baselines, while REDACT retains stronger performance on corruption types absent from training. Targeted augmentation can further improve REDACT, highlighting complementary benefits from encoder design and training coverage.

Paper table comparing ConvNet and REDACT with and without depth augmentation across nominal observations, flying pixels, Gaussian noise, Perlin dropout and vine occlusion. REDACT without augmentation reaches success rates of 96.0, 91.2, 93.0, 47.6 and 65.7 percent respectively.
All policies in this comparison share the same regularized teacher. Corruptions are tested at the maximum severity specified in the paper. SR denotes success rate; MXD measures progress through the terrain checkpoints. Select either figure to inspect it at full resolution.

Traversal demonstrations

Baseline comparisons, a controlled gating ablation, and real-world transfer.

01

REDACT and REAL under unseen observations

Both REDACT and REAL [1] were trained on clean simulated depth, without depth-image augmentation. Neither policy was trained on the Perlin dropout or flying-pixel noise shown here; the adjacent pillars were also absent from training.

Spatially structured depth dropout removes parts of the observed terrain.

REAL baseline [1]REDACT
02

Consensus gating

The same REDACT policy with gating disabled and enabled isolates its contribution under visual occlusion. With gating, recurrent memory can retain useful earlier terrain information while inconsistent current features are removed.

Gating disabledGating enabled
03

Real-world forest traversal

Zero-shot deployment introduces vegetation, unfamiliar scene content, and depth dropouts without additional training in the deployment environment.

LogLog with vegetation

Clips illustrate individual trials. Aggregate results, evaluation protocols, and failure cases are reported in the paper.

Implementation and training details

Architecture settings and the offline calibration procedure.

Architecture and calibration

Settings reported in the current manuscript. Calibration is approximate and is performed offline using clean observations.

SettingValue / procedure
Teacher attention queries4
Query/key dimension64
Teacher terrain latent dimension32
Student spatial feature grid4 × 6 cells
Finite scalar quantization5 levels per latent coordinate
Calibration observationsClean rollouts; frozen student; gating disabled
Calibration samplingOne frame per episode; approximately 3,500 episodes
Cell thresholdsEmpirical 99th percentile with linear interpolation
Gate fallbackIf fewer than 6 cells remain, retain the 18 lowest-scoring cells

References

  1. Jialong Liu, Dehan Shen, Yanbo Wen, Zeyu Jiang, and Changhao Chen. REAL: Robust Extreme Agility via Spatio-Temporal Policy Learning and Physics-Guided Filtering. arXiv:2603.17653, 2026.
  2. Xuxin Cheng, Kexin Shi, Ananye Agarwal, and Deepak Pathak. Extreme Parkour with Legged Robots. IEEE International Conference on Robotics and Automation (ICRA), pp. 11443–11450, 2024.
  3. Yohan Choi, Min-Jun Kim, Jin-Sung Kim, Yong-Jae Kim, and Youn-Hee Han. DAWN: Noise-Robust Quadruped Parkour via Depth-Denoising World Models. 2026.