# UnifoLM-VLA-0: A Vision-Language-Action (VLA) Framework under UnifoLM Family
Project Page | Models | Datasets
πEnglish | π¨π³δΈζ
| Spatial Semantic Enhancement | Manipulation Generalization |
|---|---|
| To address the requirements for instruction comprehension and spatial understanding in manipulation tasks, the model deeply integrates textual instructions with 2D/3D spatial details through continued pre-training, substantially strengthening its spatial perception and geometric understanding capabilities. | By leveraging full dynamics prediction data, the model achieves strong generalization across diverse manipulation tasks. In real-robot validation, it can complete 12 categories of complex manipulation tasks with high quality using only a single policy. |