arXiv Weekly Digest — Week 33, 2026

Fetched: 2026-08-10 | Categories: cs.RO, cs.LG, cs.HC, cs.CV | Papers: 20


ω-0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation

Authors: Zhe Li, Zhenzhe Zhang, Yangyang Wei et al. | Submitted: 2026-08-06 | arXiv: 2608.06375 Categories: cs.RO

Research Background: Household humanoid tasks demand tightly coordinated locomotion and manipulation simultaneously, yet current policies decompose these behaviors or focus on arm-only control, limiting whole-body deployment in unstructured environments.

Technical Approach: ω-0 introduces a latent predictive world-action model that jointly reasons over whole-body state and future outcomes. Given a language instruction and current observation, it generates latent predictions of future world states and decodes coordinated loco-manipulation actions end-to-end, trained on real-world humanoid data without decomposed subpolicies.

Key Takeaway: Unified latent world modeling enables humanoids to concurrently walk, balance, and manipulate objects under natural language instructions without task decomposition.


DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation

Authors: Junfeng Li, Junjie He, Zhide Zhong et al. | Submitted: 2026-08-06 | arXiv: 2608.06374 Categories: cs.RO

Research Background: Training a single generalist VLA policy across heterogeneous robot embodiments is an open challenge; existing methods underuse transferable dynamics priors and require laborious action-space normalization between embodiments.

Technical Approach: DyPES-VLA decouples learning into a shared dynamics prior module (capturing cross-embodiment physics and interaction patterns from diverse visual data) and embodiment-specific control heads (handling kinematic and action-space differences). The architecture eliminates manual preprocessing by learning the action format mapping jointly.

Key Takeaway: Disentangling shared dynamics from embodiment-specific control enables efficient cross-embodiment transfer without manual action-space alignment.


A Master-Slave Robot Manipulator for Needle-Based Teleoperation in MRI Chamber

Authors: Omar Curiel, Jing-Yuan Huang, Po-Chih Chen et al. | Submitted: 2026-08-06 | arXiv: 2608.06354 Categories: cs.RO

Research Background: Performing needle-based abdominal interventions inside MRI scanners requires MR-compatible robotic systems that can transmit motion and force without metallic or electronic components that would interfere with imaging.

Technical Approach: The system pairs a 2+1-DoF master controller with a slave manipulator connected via elastomeric fluid actuators that achieve high input impedance and low leakage. A complementary digital master controller adds multimodal control capability beyond conventional split-axis or mode-switchable hybrid interfaces.

Key Takeaway: Fluid-transmitted teleoperation with a hybrid digital-physical controller enables precise, MR-safe needle manipulation in interventional radiology workflows.


GeniWorld: A Generalizable Interactive World Model for Robotic Manipulation via Visual Actions

Authors: Chenghao Gu, Hanyang Yu, Jingbo Zhang et al. | Submitted: 2026-08-06 | arXiv: 2608.06332 Categories: cs.RO

Research Background: Generalist robot policies struggle in novel environments, and simulation-based scaling is expensive; action-conditioned world models offer a cheaper alternative but often suffer from poor controllability and out-of-distribution generalization.

Technical Approach: GeniWorld formulates robot interactions through visual actions (pixel-space action representations) that condition a generative world model to synthesize future frames. This visual action conditioning improves controllability and generalizes to unseen environments without requiring physical robot rollouts.

Key Takeaway: Visual-action conditioning in a generative world model achieves strong out-of-distribution generalization for robotic manipulation at a fraction of real-world data cost.


Robot Learning from Human Demonstrations: Handwritten Alphabet Trajectories and Human-Likeness Evaluation

Authors: Alperen Kenan, Paul Bremner, Manuel Giuliani | Submitted: 2026-08-06 | arXiv: 2608.06221 Categories: cs.RO, cs.HC, cs.LG

Research Background: Human-likeness in robot motion is a key trust factor in HRI, yet quantitatively evaluating whether learned motions resemble human dynamics remains an open problem, particularly for fine motor tasks like writing.

Technical Approach: The work presents an LfD framework for handwritten alphabet trajectories combining probabilistic movement primitives with a multi-metric human-likeness evaluation suite covering trajectory smoothness, timing variability, and perceptual ratings. A dataset of human handwriting demonstrations is collected and used to train and benchmark several LfD methods.

Key Takeaway: Probabilistic movement primitives reproduce human handwriting dynamics most faithfully, with the proposed multi-metric evaluation revealing aspects of human-likeness that single metrics miss.


Design and Evaluation of a Touchscreen-Based Teleoperation Interface for Robotic Manipulators

Authors: Juan José García Cárdenas, Alperen Kenan, Hamidreza Raei et al. | Submitted: 2026-08-06 | arXiv: 2608.06219 Categories: cs.RO, cs.HC

Research Background: Demanding teleoperation tasks in hazardous environments (e.g., nuclear swab sampling) require precise path and force tracking that conventional joystick interfaces struggle to provide intuitively.

Technical Approach: A touchscreen interface maps continuous finger movements to Cartesian end-effector commands, with pressure-sensitive force feedback cues. The design is validated through user studies measuring path accuracy, task completion time, and operator workload against a standard joystick baseline.

Key Takeaway: Touchscreen-based teleoperation significantly reduces operator workload and improves path accuracy for contact-rich surface tasks compared to joystick control.


VIDP: Variable Impedance Diffusion Policy for Compliant Robot Manipulation from Diverse Demonstrations

Authors: Hisham Khalil, Neil Fernandes, Thomas M. Kwok et al. | Submitted: 2026-08-06 | arXiv: 2608.06210 Categories: cs.RO

Research Background: Contact-rich manipulation benefits from variable impedance control, but compliance is a hidden variable in kinematic demonstration data, making it hard to learn from force-agnostic observations.

Technical Approach: VIDP augments a diffusion policy to jointly predict kinematic trajectories and variable impedance parameters (stiffness and damping) by inferring compliance from the statistical structure of diverse demonstrations rather than requiring explicit force labels. A consistency loss encourages coherent impedance profiles across similar contact phases.

Key Takeaway: Jointly diffusing trajectories and impedance profiles from kinematic data enables compliant, contact-robust manipulation without force-instrumented demonstrations.


ErgoSurf: Ergodic Control for the Coverage of Unknown Surfaces

Authors: Stefan Schneyer, Timo Bachmann, Maged Iskandar et al. | Submitted: 2026-08-06 | arXiv: 2608.06208 Categories: cs.RO

Research Background: Surface-centric tasks like inspection, cleaning, and polishing require full-coverage trajectories while maintaining stable contact, but ergodic control methods traditionally depend on known geometry or external sensing.

Technical Approach: ErgoSurf adapts ergodic trajectory optimization to incrementally estimate surface geometry from contact-force feedback alone, without cameras. The controller updates a spatial coverage distribution online as the robot touches unexplored surface regions, producing exploration-exploitation-balanced paths.

Key Takeaway: Contact-feedback-driven ergodic control achieves full-surface coverage on unknown geometries without vision, enabling robust task performance on arbitrary workpieces.


Prior-SG: Task and Prior Driven Region Segmentation for Scene Graphs in Arbitrarily-Structured Environments

Authors: Giorgio Tonetti, Laurent Kneip, Abel Gawel et al. | Submitted: 2026-08-06 | arXiv: 2608.06170 Categories: cs.RO, cs.CV

Research Background: Hierarchical 3D scene graphs support high-level spatial reasoning for mobile robots, but existing extraction methods rely on geometric heuristics (wall-separated rooms) that fail in open-plan or non-standard layouts.

Technical Approach: Prior-SG reformulates scene graph region segmentation as a probabilistic alignment problem: it fuses task-relevant priors (semantic object co-occurrence and navigational utility) with online visual observations into a probabilistic model that incrementally refines region boundaries as the robot explores.

Key Takeaway: Treating scene graph extraction as probabilistic alignment rather than rule-based clustering yields robust hierarchical representations even in environments without well-defined room boundaries.


Visual Grounding in Zero-Shot Vision-Language Control

Authors: J. de Curtò, Dayani Plasencia, Diego Sánchez et al. | Submitted: 2026-08-06 | arXiv: 2608.06154 Categories: cs.RO, cs.AI, cs.CV

Research Background: VLMs used as zero-shot robot controllers may appear effective due to simulator dynamics or conservative priors rather than genuine visual grounding, raising concerns about the validity of reported performance.

Technical Approach: The study applies an input-ablation battery to nine direct-action VLMs and six tool-use VLMs: blind images, repeated inputs, axis-reflected inputs, and non-visual baselines. Pipeline integrity checks ensure ablations isolate perceptual grounding from confounds, revealing the degree to which decisions depend on actual visual input.

Key Takeaway: Many zero-shot VLM controllers achieve competitive scores without genuine visual grounding, suggesting benchmark results may significantly overestimate real-world perceptual capabilities.


IcFuzz: Fuzzing Isaac Sim with Semantic Stage Guidance and Multi-level Mutation

Authors: Zhixiang Chen, Zhuangbin Chen, Ruoxi Jia et al. | Submitted: 2026-08-06 | arXiv: 2608.06088 Categories: cs.RO, cs.SE

Research Background: NVIDIA Isaac Sim is widely used for embodied AI development, but its complexity introduces software bugs that degrade simulation reliability; standard fuzzing struggles with the structured, semantic nature of simulator inputs.

Technical Approach: IcFuzz introduces semantic stage guidance to generate contextually valid simulation scenes and multi-level mutation (asset, scene, and physics parameter levels) to systematically explore the simulator’s input space. Coverage metrics track which physics behaviors and rendering paths have been exercised.

Key Takeaway: Semantic-aware multi-level fuzzing exposes previously unknown bugs in Isaac Sim that blind random mutation misses, improving confidence in simulation reliability for robot policy training.


Topometric Autonomous Vehicle Localization by Combining Visual Embeddings and Feed-Forward 3D Models

Authors: Eulogio Quemada-Torres, Alberto Jaenal, Francisco-Angel Moreno et al. | Submitted: 2026-08-06 | arXiv: 2608.06021 Categories: cs.RO

Research Background: Visual localization for autonomous vehicles must balance compact map representations with robustness to appearance changes and metric precision; visual place recognition alone lacks the metric accuracy needed for reliable navigation.

Technical Approach: The approach combines low-dimensional VPR image embeddings for coarse place retrieval with feed-forward 3D models that provide dense geometric correspondences for metric refinement. A topometric map stores both embeddings and local 3D structures, enabling efficient look-up followed by precise pose estimation.

Key Takeaway: Coupling VPR embeddings with lightweight feed-forward 3D models achieves map compactness and appearance robustness while recovering the metric precision needed for autonomous driving.


Adaptive-WAM: Quality-Guided Early-Exit Planning from Intermediate Video-Diffusion Features

Authors: Sining Ang, Yuguang Yang, Yan Wang | Submitted: 2026-08-06 | arXiv: 2608.06008 Categories: cs.RO

Research Background: World-action models for autonomous driving inherit the computational cost of full video generation even though deployment only requires ego trajectory prediction, creating a speed-accuracy trade-off.

Technical Approach: Adaptive-WAM analyzes how planning performance varies across video diffusion timesteps and transformer depth, finding that reliable decisions can be made from intermediate denoising features. A quality estimator gates early exit from the diffusion process when feature confidence is sufficient, skipping costly final denoising steps.

Key Takeaway: Early-exit from video diffusion at quality-guided intermediate features achieves near-full-model planning accuracy at a fraction of the inference cost.


Beyond Flat Policies: Hierarchical Post-Training for Embodied Agents in Robotic Manipulation

Authors: He Kong, Zengjue Chen, Qi Wang et al. | Submitted: 2026-08-06 | arXiv: 2608.05999 Categories: cs.RO

Research Background: VLA models fine-tuned as flat policies struggle with long-horizon manipulation because they cannot explicitly model task progression; hierarchical approaches rely on offline supervised learning and don’t generalize to novel task compositions.

Technical Approach: The work proposes hierarchical post-training that jointly optimizes a high-level task planner (predicting subtask sequences) and a low-level VLA executor through reinforcement signals. Online interaction data enables the hierarchy to adapt to unseen task compositions beyond the offline demonstration distribution.

Key Takeaway: Hierarchical post-training with online RL dramatically improves long-horizon manipulation performance and compositional generalization over flat VLA fine-tuning.


Observation-Grounded Self-Predictive Reinforcement Learning for Visual Continuous Control

Authors: Xinwei Liu, Junyuan Liang, Jianting Zhang et al. | Submitted: 2026-08-06 | arXiv: 2608.05989 Categories: cs.LG, cs.RO

Research Background: Learning robot control directly from pixels suffers from poor sample efficiency; both latent self-prediction and observation-space prediction methods still struggle on challenging visual control benchmarks.

Technical Approach: The method grounds latent self-prediction in actual observation space by enforcing consistency between predicted latent dynamics and decoded future observations. This observation-grounded auxiliary loss prevents representation collapse and provides a richer learning signal without requiring separate observation-space models.

Key Takeaway: Grounding latent self-prediction with observation-space consistency significantly improves sample efficiency on challenging visual continuous control tasks.


TRACE: Learned Proprioceptive Odometry for Legged Robots under Unreliable Contact Conditions

Authors: Taehyeon Kong, Woojin Kim, Jemin Hwangbo | Submitted: 2026-08-06 | arXiv: 2608.05975 Categories: cs.RO, cs.AI

Research Background: Accurate state estimation for legged robots is difficult on slippery or uneven terrain where contact assumptions break down; camera-dependent odometry is fragile to lighting, while purely inertial methods accumulate drift.

Technical Approach: TRACE (Tokenized Robust Attention for Contact-Aware Estimation) processes a history of IMU and joint measurements through a foot-aware cross-attention module that adaptively weights each foot’s contribution based on estimated contact reliability. It directly predicts relative displacement, rotation, and body-frame velocity end-to-end.

Key Takeaway: Foot-aware cross-attention over proprioceptive history yields robust odometry on legged robots even when contact conditions are unreliable, outperforming filter-based baselines.


SkillMemo: Expert-guided Skill Memory Framework for Compositional Embodied Manipulation

Authors: Changyuan Wang, Chubin Zhang, Zhenyu Wu et al. | Submitted: 2026-08-06 | arXiv: 2608.05970 Categories: cs.RO, cs.AI

Research Background: Diffusion policies and VLA models are constrained by the scarcity of large-scale trajectory data, limiting their compositional generalization to out-of-distribution task sequences not seen in training.

Technical Approach: SkillMemo constructs a skill memory bank from expert demonstrations by segmenting trajectories into reusable primitive skills. At inference, a retrieval mechanism selects and sequences relevant skill memories guided by the current task context, allowing novel task compositions to be assembled without retraining.

Key Takeaway: Retrieving and composing expert skill memories enables embodied agents to generalize to novel multi-step manipulation tasks without requiring additional trajectory data.


GAUGE: A Measurement-Grounded Benchmark for Physical Fidelity in Simulation Engines and Video World Models

Authors: Shuai Wang, Yaxin Feng, Xuekun Jiang et al. | Submitted: 2026-08-06 | arXiv: 2608.05948 Categories: cs.AI, cs.CV, cs.RO

Research Background: Physics engines and generative video world models are increasingly used for embodied AI training, but evaluating their physical fidelity relies on perceptual similarity or subjective ratings rather than grounded physical measurements.

Technical Approach: GAUGE provides a real-world-grounded diagnostic benchmark that tests specific physical principles (rigid body dynamics, friction, deformation) using calibrated real measurements as ground truth. Both simulation engines and video world models are evaluated jointly across a common set of physical scenarios.

Key Takeaway: Real-measurement-grounded benchmarking reveals systematic physical fidelity gaps in both simulation engines and video world models that perceptual metrics fail to detect.


Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models

Authors: Haodong Yan, Junfeng Li, Junjie He et al. | Submitted: 2026-08-06 | arXiv: 2608.05903 Categories: cs.CV, cs.RO

Research Background: World-action models built in VAE latent space are fragile to visual distribution shifts because the VAE optimizes pixel reconstruction rather than action-relevant features; semantic latent spaces improve robustness but sacrifice the generative pretraining advantages.

Technical Approach: Robust-WAM bridges both paradigms by pretraining in VAE space for rich dynamics priors, then aligning representations to a semantic latent space through a foresight distillation objective that transfers semantic structure without discarding generative knowledge.

Key Takeaway: Foresight distillation from generative to semantic latent space combines the training scalability of VAE-based pretraining with the visual robustness of semantic representations.


Shape-Aware Oriented Bounding Box (OBB) to Horizontal Bounding Box (HBB) Conversion

Authors: Badha Rathna Sabhapathy, Gotam Dahiya, Vishesh Vatsal | Submitted: 2026-08-06 | arXiv: 2608.05858 Categories: cs.CV, cs.RO

Research Background: Aerial and satellite object detection relies on oriented bounding boxes for tightly fitting spatially oriented objects, but downstream pipelines often require horizontal bounding boxes; naive conversion introduces excess background or crops detections.

Technical Approach: The paper introduces a shape-aware conversion algorithm that analytically minimizes background inclusion while preserving detection completeness. It uses object aspect ratio and rotation angle to adaptively project OBB vertices onto the horizontal axis, outperforming both tight-crop and loose-fit baselines.

Key Takeaway: Shape-aware OBB-to-HBB projection significantly reduces background noise in converted detections without losing object coverage, improving downstream aerial object detection pipelines.