In this sense, the first generation was highly significant as the starting point of robot learning; however, its input diversity and representational structure were fundamentally limited for achieving the robustness and safety required in real-world deployment. 1.2. 2nd Generation The defining characteristic of the second generation lies in constructing VLA models by combining the semantic representations of pre-trained vision language models with an action head or decoder that outputs robot actions.
← all excerpts
Sensing the Action: Rethinking Sensor Modalities and Multi-Modal Fusion in Vision-Language-Action Models for Robotic Manipulation.
1
—
—