Recent Posts

Posted in Paper X-Review

[ArXiv 2025] Towards a Science of Scaling Agent Systems

이번에 소개드릴 논문은 LLM 기반 멀티 에이전트 시스템의 확장성을 분석한 연구입니다. 기존 연구가 성능 향상에 초점을 맞췄던 것과 달리, 에이전트 간 협업이 어떤 조건에서 성능…

Continue Reading
Posted in Paper X-Review

[Arxiv 2026] AdaThinkV: Adaptive Thinking for Token-Efficient Video Reasoning

안녕하세요. 이번에는 AdaThinkV: Adaptive Thinking for Token-Efficient Video Reasoning 논문을 읽어보았습니다. 아카이브 논문이지만, 저자가 이번 ECCV에 1,7저자로 두편 붙인 이력이 있고, 1저자가 26년도 석사시작인 것…

Continue Reading
Posted in Paper X-Review

[arXiv 2026] LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation

안녕하세요. 이번에 리뷰로 가져온 논문은 LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation이라는 논문입니다. 해당 논문은 기존에 사전학습된 VLM의 spatial reasoning 능력을 실제 로봇의…

Continue Reading
Posted in X-Diary

Earth Rover Challenge 2026 후기

안녕하세요 이번 Earth Rover Challenge에 참여한 후기를 적어보고자 합니다.먼저 Earth Rover Challenge가 어떤 대회인지 소개드리겠습니다 위 로봇은 Earth Rover Mini+라는 로봇입니다. 가격은 대략 $399정도하며 전방/후방…

Continue Reading
Posted in X-Review

[Neurocomputing 2022] GSV-Cities: Toward Appropriate Supervised Visual Place Recognition

안녕하세요. 이번주에 GSV-Cities : Toward Appropriate Supervised Visual Place Recognition 논문을 읽고, 세미나 발표 대신 X-Review로 대신하게 되었습니다 Visual Place Recognition(VPR)은 주어진 query image가 어느…

Continue Reading
Posted in Paper X-Review

[CVPR 2026] M4-RAG: A Massive-Scale Multilingual Multi-Cultural Multimodal RAG

안녕하세요 오늘은 M4-RAG: A Massive-Scale Multilingual Multi-Cultural Multimodal RAG라는 논문을 읽어보았습니다이 논문은 multilingual + multimodal 환경에서 RAG가 실제로 얼마나 잘 작동하는지를 평가하기 위한 대규모 benchmark인…

Continue Reading
Posted in X-Review

[ECCV 2026] Open-Vocabulary Long-term Action Anticipation

안녕하세요. 이번 리뷰에서는 ECCV 2026에서 접한 Long-term Action Anticipation 논문을 가져왔습니다. Long-term Action Anticipation(LTA)은 관찰된 비디오를 바탕으로, 이후에 이어질 Z개의 행동으로 이루어진 시퀀스를 예측하는 task입니다….

Continue Reading
Posted in Paper X-Review

[CVPR 2026] Latent Implicit Visual Reasoning

안녕하세요. 이번에는 Latent Implicit Visual Reasoning (LIVR) 논문을 읽어보았습니다. 지난번에 리뷰한 Coconut이 LLM의 reasoning을 language space에서 continuous latent space로 옮겼다면, 이번 논문은 비슷한 문제의식을 Large…

Continue Reading
Posted in X-Review

[CVPR 2025] Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language Models

안녕하세요. 이번 주는 Image representation의 구조도 text representation과 같이 attribute, object, relation 과 같은 concept들을 분리하고 재조합할 수 있는 구조를 가지고 있는 지에 대해 연구한…

Continue Reading
Posted in Paper X-Review

[AAAI 2024] GMMFormer: Gaussian-Mixture-Model Based Transformer for EfficientPartially Relevant Video Retrieval

안녕하세요. 오늘의 리뷰도 PRVR 관련 논문을 리뷰하고자 합니다. 이번 논문은 최근 PRVR 연구들에서 베이스라인으로 많이 사용되는 방법론입니다. 그럼 바로 리뷰 시작하겠습니다. 1. Introduction 최근 비디오가…

Continue Reading