論文 / arXiv:2610.06843
Recursive Video In-Context Learning for Agentic Robot
PLAIN SUMMARY / やさしい要約
ロボットに作業を教える動画は便利ですが、長い動画を毎回全部見るのは大変です。そこでこの研究は、動画を「作業の流れ」から「細かな場面」へと段階的に分け、必要なところだけ見直せる形にします。たとえるなら、長い料理番組を最初から見返すのではなく、目次から「今必要な手順」だけ再生するイメージです。こうした仕組みで、2つのベンチマークで成功率が上がったと述べています。
AIが専門用語を使わずに書いた解説です。内容の正確さは、下の原文の要旨で確認してください。
ABSTRACT / 要旨(原文)
LLM agents that orchestrate frozen vision-language-action (VLA) policies improve across episodes through text memory, which records what the agent did but not how the task is done. A demonstration video shows it, but fits poorly into an agent's context. The full video slows every turn, fixed keyframes lose the contact detail that decides whether a grasp holds, and what the agent needs shifts from the task's structure while planning to the frames around each contact. We introduce Recursive Video In-Context Learning (RV-ICL), a training-free method that turns a demonstration into a hierarchy the agent navigates rather than a prompt it receives. The hierarchy is built from the sub-events of the demonstration, such as grasps and releases. Its levels grow finer, from keyframes of the whole task to phases, moments and short clips, and are exposed through read-only tools. The agent reads the coarse levels before planning. During execution it re-enters the hierarchy whenever a step needs more detail and loads only the clip of its current sub-goal. One demonstration per task is enough. Built on RPent, RV-ICL raises success from 92.6% to 96.5% on LIBERO-PRO and from 86.7% to 95.8% on LIBERO-Plus.
ここに表示しているのは原論文の要旨です。AIによる要約や解釈は含みません。
