Posted in X-Review

[IROS 2026]Repurposing RGB-based Foundation Model for Depth Estimation on Thermal Images Using Hierarchical Supervision

오늘은 IROS2026에 붙은 열화상 RGB VFM에서의 representation 정렬관점에 연구입니다. 최근 IROS accept 논문을 보다가 제가 관심있는 영역의 논문이 나와 보게 되었습니다 Introduction 열화상 카메라는 모두가…

Continue Reading
Posted in Paper X-Review

[RA-L 2026]From Obstacles to Etiquette: Robot Social Navigation with VLM-Informed Path Selection

안녕하세요. 이번에 리뷰로 가져온 논문은 From Obstacles to Etiquette: Robot Social Navigation with VLM-Informed Path Selection이라는 논문입니다. 해당 논문의 핵심은 엄청 간단한데 플래닝 모듈이 여러…

Continue Reading
Posted in X-Review

[ICLR 2026] Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training

이번에 리뷰하려는 페이퍼는 텍스트로만 학습된 LLM이 이미지를 전혀 보지 않고도 시각적 사전 지식인 Visual Priors 를 이미 가지고 있다는 것을 밝힌 논문입니다. 1. 이미지 없이,…

Continue Reading
Posted in X-Review

[CVPR 2026] An Empirical Study on How Video-LLMs Answer Video Questions

1. Introduction 최근 Video-LLM은 단순히 이미지 한 장을 이해하는 것을 넘어, 여러 프레임으로 구성된 비디오의 내용을 이해하고 다양한 형태의 질문에 답할 수 있을 정도로 빠르게…

Continue Reading
Posted in Paper X-Review

[arXiv 2025] SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Abstract 안녕하세요, 강희승입니다. 오늘은 지난주에 리뷰한 SigLIP의 후속 연구로, Google DeepMind에서 발표한 SigLIP2를 리뷰해보려고 합니다. SigLIP의 경우, Memory efficiency, 모델의 확장성, 학습 과정에서 효율성 등의…

Continue Reading
Posted in X-Review

[AAAI 2026] Endowing Vision-Language Models with System 2 Thinking for Fine-Grained Visual Recognition

Abstract VLMs는 query 이미지에서 주요한 시각적 특징을 추출하는 능력이 뛰어나지만, fine-grained 시나리오에서는 미묘한 차이를 식별하는 데 어려움을 겪습니다. 저자들은 인간의 “System 1 & System 2”…

Continue Reading
Posted in Paper X-Review

[ICCV 2023] Sigmoid Loss for Language Image Pre-Training

Prologue 안녕하세요. 강희승입니다. 이번에는 SigLIP입니다. 기초교육 간 읽은 논문 중 최대한 X-review에 작성되지 않은 논문들을 리뷰해보려고 합니다. SigLIP도 최근 자주 쓰이는 Visual Backbone으로 알고 있는데,…

Continue Reading
Posted in X-Review

[arXiv 2026] Emotion-LLaMAv2 and MMEVerse – A New Framework and Benchmark for Multimodal Emotion Understanding

안녕하세요. 이번에는 Emotion-LLaMAv2 and MMEVerse: A New Framework and Benchmark for Multimodal Emotion Understanding 논문을 읽어봤습니다. Emotion-LLaMA라는 emotion reasoning 분야로 획을 그은 논문의 후속 논문인데요….

Continue Reading
Posted in X-Review

[CVPR 2025] VGGT: Visual Geometry Grounded Transformer

안녕하세요. 오늘 리뷰할 논문은 CVPR 2025에서 Best Paper Award를 받은 VGGT(Visual Geometry Grounded Transformer)입니다. Introduction 본 논문은 3D reconstruction task를 다루고 있습니다. 3D reconstruction은 기본적으로…

Continue Reading
Posted in X-Review

[CVPR 2026] Thinking Diffusion: Penalize and Guide Visual-Grounded Reasoning in Diffusion Multimodal Language Models

1. Introduction Qwen-VL이나 LLaVA와 같은 Vision-Language Model(VLM)은 이미지를 이해하고, 이미지에 대한 질문에 답하거나 복잡한 시각적 추론을 수행하는 등 다양한 vision-language task에서 좋은 성능을 보여주고 있습니다….

Continue Reading