{"id":18448,"date":"2026-09-05T21:27:09","date_gmt":"2026-09-05T21:27:09","guid":{"rendered":"https:\/\/techtrendfeed.com\/?p=18448"},"modified":"2026-09-05T21:27:09","modified_gmt":"2026-09-05T21:27:09","slug":"refactor-vla-unsupervised-library-studying-of-typed-motor-packages","status":"publish","type":"post","link":"https:\/\/techtrendfeed.com\/?p=18448","title":{"rendered":"REFACTOR-VLA: Unsupervised Library Studying of Typed Motor Packages"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div>\n<p>Most present vision-language-action (VLA) fashions\u2014similar to OpenVLA, \u03c00, RT-2, and RDT-1B\u2014are \u201cmonolithic.\u201d This implies they generate uncooked motor instructions or very quick sequences of actions, with out organizing behaviors into reusable, well-defined abstractions. Consequently, these fashions carry out poorly on long-horizon (multi-step) duties, and it\u2019s tough to interpret what they&#8217;ve realized. Current approaches for locating expertise typically keep away from the core drawback of deciding when two motion sequences are \u201cbehaviorally equal.\u201d For instance, AtomicVLA and AtomSkill group motion sequences by clustering their contrastive embeddings. In distinction, BLADE and LRLL depend on a big language mannequin (LLM) to evaluate whether or not two sequences are equal, however these LLMs should not calibrated to the robotic\u2019s personal dynamics. We introduce REFACTOR-VLA, a system that learns reusable expertise utilizing a \u201cwake\/sleep\u201d structure. Within the sleep section, the system clusters segments of motor packages utilizing a Behavioral-Equivalence Kernel (BEK). This BEK is predicated on the outcomes of rolling out actions in a realized latent world mannequin, M\u03c6. Within the wake section, the system generates typed lambda phrases (easy, structured packages) from a vocabulary impressed by the Hindley\u2013Milner sort system. These lambda phrases are then utilized by a library-conditioned rectified-flow motion decoder to provide actions. Solely abstractions that go each a Minimal Description Size (MDL) criterion and a return-preservation gate are accepted as expertise. To coach REFACTOR-VLA, we use a three-phase schedule: \u2022 Section A (World-model warmup): The latent world mannequin M\u03c6 is skilled. \u2022 Section B (Wake-phase coverage optimization): The coverage that makes use of the library of expertise is optimized. \u2022 Section C (Sleep-phase ability discovery): The system clusters motion fragments into reusable expertise. We evaluated REFACTOR-VLA on the complete LIBERO benchmark suite. Our outcomes present two major findings. First, merely rising the dimensions of the world mannequin\u2014from 188 million to 430 million parameters\u2014worsened efficiency on 4 out of 4 benchmark suites, disproving the concept that simply making the world mannequin greater at all times helps. Second, altering the coaching goal makes a giant distinction: including an auxiliary supervised contrastive loss (particularly, InfoNCE loss) throughout the world-model warmup (Section A) vastly improved the standard of ability clustering within the sleep section (Section C). We measured this utilizing Normalized Mutual Data (NMI) underneath n = 3 multi-seeding: \u2022 Object suite: 0.462 \u00b1 0.021 \u2022 Spatial suite: 0.867 \u00b1 0.025 \u2022 Purpose suite: 0.915 \u00b1 0.013 \u2022 LIBERO-10 suite: 0.754 \u00b1 0.010<\/p>\n<\/div>\n\n","protected":false},"excerpt":{"rendered":"<p>Most present vision-language-action (VLA) fashions\u2014similar to OpenVLA, \u03c00, RT-2, and RDT-1B\u2014are \u201cmonolithic.\u201d This implies they generate uncooked motor instructions or very quick sequences of actions, with out organizing behaviors into reusable, well-defined abstractions. Consequently, these fashions carry out poorly on long-horizon (multi-step) duties, and it\u2019s tough to interpret what they&#8217;ve realized. Current approaches for locating [&hellip;]<\/p>\n","protected":false},"author":2,"featured_media":18450,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[55],"tags":[136,3424,7968,1758,10448,9674,3934],"class_list":["post-18448","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-machine-learning","tag-learning","tag-library","tag-motor","tag-programs","tag-refactorvla","tag-typed","tag-unsupervised"],"_links":{"self":[{"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=\/wp\/v2\/posts\/18448","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=18448"}],"version-history":[{"count":1,"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=\/wp\/v2\/posts\/18448\/revisions"}],"predecessor-version":[{"id":18449,"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=\/wp\/v2\/posts\/18448\/revisions\/18449"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=\/wp\/v2\/media\/18450"}],"wp:attachment":[{"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=18448"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=18448"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=18448"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}