Author: 김 영규

Posted in X-Review

[ICRA 2026] VITRA : Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos

안녕하세요, 이번주 X-review는 unscripted real-life human video를 VLA pretraining 데이터로 바꾸는 연구를 리뷰해보려고 합니다. 지난 ActiveMimic 리뷰에 이어 egocentric human video를 로봇 학습으로 끌어오는 결의…

Continue Reading
Posted in X-Review

[arXiv 2026] ActiveMimic: Egocentric Video Pretraining with Active Perception

안녕하세요, 이번주 리뷰는 egocentric video pretraining에 대한 연구입니다. 최근 egocentric human video의 pretraining을 다루는 연구들이 늘어나고, 대부분 cam의 움직임을 노이즈로 다루는데, 오히려 해당 부분을 살리는…

Continue Reading
Posted in X-Review

[arxiv 2025] Is Diversity All You Need for Scalable Robotic Manipulation?

안녕하세요, 이번에는 로봇 조작 학습에서 데이터 다양성이 정말 항상 좋은 것인지에 대해 다룬 연구를 리뷰해보려고 합니다. Agibot에서 진행한 연구이고, 저자들은 task diversity, multi-embodiment pre-training, expert…

Continue Reading
Posted in X-Review

[ICLR 2026 Workshop] World Action Models are Zero-shot Policies

안녕하세요 이번주는 WAM을 소개하려고 합니다. 최근 로봇 파운데이션 모델들의 연구에서 로봇 데이터의 teleoperation 의존도를 낮추는 연구와 기존 데이터를 통해서 3차원 현실에서 작동하기 위한 모델 구조,…

Continue Reading
Posted in X-Review

[ICML 2026] DECO: Decoupled Multimodal Diffusion Transformer for Bimanual Dexterous Manipulation with a Plugin Tactile Adapter

안녕하세요, 이번주 X-review에는 tactile 관련 연구를 가져왔습니다. 최근 제안서 작업한 과제 내용에 기존 pretrained VLA에 tactile 센싱 모듈을 추가하겠다는 내용을 적었는데, 이거 어떻게 하면 효과적으로…

Continue Reading
Posted in X-Review

GR00T : An Open Foundation Model for Generalist Humanoid Robots

안녕하세요, 이번주 X-review는 NVIDIA의 가장 간판 프로젝트 중 하나인 GR00T에 대해 작성하려고 합니다. 기존 로봇 파운데이션 모델들이 주로 단일 팔, 병렬 그리퍼, tabletop manipulation 중심으로…

Continue Reading
Posted in X-Review

[arXiv 2026] PokeVLA: Empowering Pocket-Sized Vision-Language-Action Model with Comprehensive World Knowledge Guidance

안녕하세요, 이번주는 작은 모델임에도 불구하고 대용량 학습 데이터로 학습한 큰 모델 대비 강인하고 성능 좋은 모델을 다룬 연구에 대해서 리뷰해보려고 합니다. 얼마 전 VLA-Adapter라는 연구도…

Continue Reading
Posted in X-Review

[arXiv 2026] Beyond Imitation: Reinforcement Learning-Based Sim-Real Co-Training for VLA Models

안녕하세요, 이번주는 RSS 2026에 submit된 Co-training 연구를 리뷰해보려고 합니다. 시뮬레이션 데이터는 현실 데이터와 함께 co-training되면서 low-cost로 VLA training을 풍부하게 해주는데, 대부분의 co-training 연구들은 SFT 방식으로…

Continue Reading
Posted in X-Review

[ICLR 2026] Emergent Dexterity via Diverse Resets and Large-Scale Reinforcement Learning

안녕하세요, 이번주는 Large-Scale RL에 대해 다루어보려고 합니다. RL을 통해 policy를 학습하게되면 너무 optimal한 행동에 fitting되고 여러 상황에 대응하기는 좀 힘들 뿐 만 아니라 reward shaping이…

Continue Reading
Posted in X-Review

[ICLR 2026] Self-Improving Vision-Language-Action Models with Data Generation via Residual RL

안녕하세요, 이번주 X-review는 data generator로써 RL을 활용하며 VLA에 대한 SFT를 진행하며, 제목처럼 self improving 하는 policy 학습법을 다룬 연구입니다. Recovery behavior를 위한 generalist 데이터셋 구성에…

Continue Reading