Skip to content
Kernelia
All news
ResearchHugging Face Blog

SmolVLA: Efficient Vision-Language-Action Model trained on Lerobot Community Data

Hugging Face's team has introduced SmolVLA, an efficient vision-language-action model trained on Lerobot community data. This model combines text and vision understanding to perform actions.

Summary written by Kernelia from the original article by Hugging Face Blog. The story and its rights belong to its author.