Publication:

Distillation of Vision Language Action Models for Robust Robotic Manipulation

Loading...
Thumbnail Image

Files

MENG_AARON_THESIS.pdf (3.04 MB)

Date

2026-04-13

Journal Title

Journal ISSN

Volume Title

Publisher

Research Projects

Organizational Units

Journal Issue

Access Restrictions

Abstract

Vision Language Action models demonstrate strong generalization in robotic manipulation but remain impractical for deployment due to their computational cost. We investigate whether rollout based distillation can transfer the capabilities of a large VLA into compact, deployable policies using only action level supervision. On the LIBERO benchmark, we compare both a BC Transformer architecture and SmolVLA architecture student model trained on human demonstrations to the same model trained on trajectories generated by a teacher policy, π0.5. Distillation improves average success from 51.8% to 63.9%, and 84.25% to 86.5%, for each model respectively under equal dataset sizes. However, scaling the distilled dataset by 10× yields inconsistent improvements, indicating that distillation effectiveness is governed more by data distribution than dataset size. Under distribution shift, all models degrade, but training on teacher rollouts generated from perturbed initial states improves robustness, demonstrating that distillation can transfer not only task performance but also generalization behavior when the training distribution is sufficiently diverse. These results highlight that effective distillation in embodied settings depends not only on the quantity of supervision, but on the structure of the data and the capacity of the student model.

Description

Type of resource

Princeton University Senior Theses

Keywords

Location

Citation