KAI: A Kinematic-Aware Interface for Data-Efficient Articulated Object Manipulation
This paper introduces the Kinematic-Aware Articulation Interface (KAI), a representation that captures the kinematic structure of articulated objects, improving sample efficiency and generalization in robot manipulation.
The paper proposes KAI, a new representation for articulated object manipulation that embeds interpretable priors, improving sample efficiency and generalization.
Before reading this…
Applications
- →Robotics
- →Automation
To understand this paper, make sure you know these concepts first:
- Understanding of robotics and manipulationfind papers →
- Basic knowledge of machine learningfind papers →
Abstract
More Like ThisArticulated object manipulation requires an understanding of kinematic structure that is difficult and costly to learn from robot demonstrations alone. We introduce the Kinematic-Aware Articulation Interface (KAI), a structured intermediate representation that captures the kinematic structure of articulated objects. By embedding interpretable geometric and kinematic priors into policy learning, KAI provides a strong inductive bias aligned with the underlying structure of articulated motion. This design effectively improves sample efficiency, with gains particularly pronounced in low-data regimes: across six simulation tasks, our method achieves an average success rate of 82.9%, matching or surpassing baseline performance while using only half the demonstration data. Our method also exhibits robust generalization to unseen backgrounds and visual distractors, transferring from a single clean training environment to cluttered real-world scenes. KAI's action-agnostic design further enables co-training with human interaction videos to enhance real-world robustness: under diverse visual distractions, our method with video co-training achieves over 70% average success rate.