[{"data":1,"prerenderedAt":111},["ShallowReactive",2],{"\u002Fen\u002Fglossary\u002Fvla-model":3},{"id":4,"title":5,"alternateName":6,"body":7,"description":101,"extension":102,"keywords":103,"meta":104,"navigation":105,"path":106,"seo":107,"stem":108,"updated":109,"__hash__":110},"glossary\u002Fglossary\u002Fen\u002Fvla-model.md","VLA Model","VLA模型 \u002F Vision-Language-Action Model",{"type":8,"value":9,"toc":94},"minimark",[10,15,30,35,38,42,55,59,81],[11,12,14],"h1",{"id":13},"what-is-a-vla-model","What Is a VLA Model?",[16,17,18,19,23,24,29],"p",{},"A ",[20,21,22],"strong",{},"VLA model"," (Vision-Language-Action model) is an end-to-end robot foundation model: it takes camera images and a natural-language instruction (e.g. \"put the cup in the drawer\") as input and directly outputs robot actions — joint angles, end-effector poses, or gripper commands. By compressing perception, language understanding, and motion generation into one network, VLA has become one of the most watched policy architectures in ",[25,26,28],"a",{"href":27},"\u002Fen\u002Fglossary\u002Fembodied-ai","embodied AI",".",[31,32,34],"h2",{"id":33},"relationship-to-llms","Relationship to LLMs",[16,36,37],{},"VLA models are typically built on top of vision-language models (VLMs): they inherit the semantic and commonsense knowledge of large language models, then replace or extend the output head with action tokens or continuous actions. A useful mental model: an LLM predicts the next word; a VLA predicts the next action — it is the language-model family extended into the physical world.",[31,39,41],{"id":40},"why-real-robot-data-is-essential","Why Real-Robot Data Is Essential",[16,43,44,45,49,50,54],{},"The internet holds vast text and images, but almost no paired \"image + instruction → joint action\" data. VLA training therefore depends heavily on real-robot demonstrations, mostly collected via ",[25,46,48],{"href":47},"\u002Fen\u002Fglossary\u002Fteleoperation","teleoperation",". Simulation data can add scale, but must cross the ",[25,51,53],{"href":52},"\u002Fen\u002Fglossary\u002Fsim-to-real","Sim-to-Real"," gap. Data diversity — across tasks, scenes, and embodiments — often matters more for generalization than parameter count.",[31,56,58],{"id":57},"representative-work","Representative Work",[60,61,62,69,75],"ul",{},[63,64,65,68],"li",{},[20,66,67],{},"RT-2"," (Google DeepMind): co-trains a VLM with robot action tokens, demonstrating that web-scale knowledge transfers to manipulation;",[63,70,71,74],{},[20,72,73],{},"OpenVLA",": an open-source VLA trained on large open robot datasets, designed for community fine-tuning;",[63,76,77,80],{},[20,78,79],{},"π0"," (Physical Intelligence): uses flow matching to generate continuous actions, targeting a general cross-embodiment manipulation policy.",[16,82,83,84,88,89,93],{},"Deploying and fine-tuning VLA on real hardware requires a platform that executes high-rate action commands reliably — BXI's ",[25,85,87],{"href":86},"\u002Fen\u002Frobots\u002Fhumanoid-robot","humanoid robot"," and ",[25,90,92],{"href":91},"\u002Fen\u002Frobots\u002Frobotic-arms","dual-arm platform"," both expose ROS2 interfaces suited to VLA research.",{"title":95,"searchDepth":96,"depth":96,"links":97},"",2,[98,99,100],{"id":33,"depth":96,"text":34},{"id":40,"depth":96,"text":41},{"id":57,"depth":96,"text":58},"A Vision-Language-Action model maps camera observations and language instructions to robot actions for manipulation and embodied-AI tasks.","md","VLA model, vision language action model, robot foundation model, end-to-end robot policy, embodied AI model",{},true,"\u002Fglossary\u002Fen\u002Fvla-model",{"title":5,"description":101},"glossary\u002Fen\u002Fvla-model",null,"NFikZ8HZmgXboc8Yz0BrE7ZWSYyddapOZmEH4UV-Tt0",1785156467504]