Author: 이 재윤

Posted in X-Diary

ICML 2026 참관기

안녕하세요, 오늘은 지난 주에 코엑스에서 열린 ICML 2026 참관 후기를 작성해보았습니다. 1만 5000명이 참여했고 6600여 편의 논문이 accept 되었다고 하는데, 그래서 그런지 사람이 굉장히 많았던…

Continue Reading
Posted in Paper X-Review

[ECCV 2024] PALM : Predicting Actions through Language Models

안녕하세요, 이번에 리뷰할 논문은 action anticipation 이라는 task를 다루는 논문입니다. 창의학기제 논문이 마무리되는대로 본 연구 주제로 넘어갈 예정이라 입문할 겸 해서 읽어보게 되었습니다. Action Anticipation…

Continue Reading
Posted in Paper X-Review

[ICML 2026] VideoBrain : Learning Adaptive Frame Sampling for Long Video Understanding

안녕하세요, 요즘 SAR만 파다 보니 루즈해지기도 해서 마침 ICML conference 참가 신청도 했겠다 어떤 논문들이 있는지 찾아보았는데, adaptive frame sampling이라는 말에 끌려 이 논문을 읽어보게…

Continue Reading
Posted in Paper X-Review

[CVPR 2026] SARMAE : Masked Autoencoder for SAR Representation Learning

안녕하세요, 이번에 리뷰할 논문은 SAR 이미지를 위한 자기주도 사전학습법을 제안한 논문입니다. 현재 창의학기제와 기업과제가 모두 SAR Object Detection이기 때문에 논문에서의 인사이트가 도움이 될 만한 부분이…

Continue Reading
Posted in Paper X-Review

[NIPS 2023] Scaling Open-Vocabulary Object Detection

안녕하세요, 이번에 리뷰할 논문은 Google Deepmind에서 2023년에 발표한 NIPS spotlight 논문입니다. 현재 저희 팀 과제에 투입되기 위한 팔로우업 중에 읽게 된 논문으로, detection 데이터셋이 제한적인…

Continue Reading
Posted in Paper X-Review

[ICLR 2024] CLIPSELF: VISION TRANSFORMER DISTILLS ITSELF FOR OPEN-VOCABULARY DENSE PREDICTION

안녕하세요, 오늘은 ICLR 2024 Spotlight 논문인 CLIPself를 리뷰해 보려고 합니다. object detection 논문인 만큼 아마 많은 분들이 흥미롭게 읽을 수 있는 논문이지 않으까 싶네요. CLIP이…

Continue Reading
Posted in Paper X-Review

[AAAI 2026] SM3Det: A Unified Model for Multi-Modal Remote Sensing Object Detection

안녕하세요, 오늘 리뷰할 논문은 AAAI 2026 Oral 논문인 SM3Det 입니다. LVU 논문 작업 이후 다시 저희 팀 기업 과제 팔로우업과 창의학기제를 겸해서 SAR Object Detection…

Continue Reading
Posted in Paper X-Review

[ICCV 2025]Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs

Video Large Language Models(Video-LMMs)는 시공간 토큰(spatiotemporal tokens)을 활용해서 강력한 비디오 이해 능력을 가지게 되었지만 토큰 개수가 많아질수록 연산량이 2차적으로 증가한다는 문제점을 가지고 있었습니다. 이에 저자들은…

Continue Reading
Posted in Paper X-Review

[CVPR 2025] Apollo: An Exploration of Video Understanding in Large Multimodal Models

안녕하세요, 3번째 x-review는 Apollo라는 논문입니다. (논문 기준) 현재까지 video-LLM 연구의 문제점을 짚고, 저자 자신들의 모델을 제안하는 구성이기 때문에 LVU task에 익숙하지 않으신 분들도 꽤(?) 재밌게…

Continue Reading
Posted in Paper X-Review

[arXiv 2025] WorldMM:Dynamic MultiModal Memory Agent for Long Video Understanding

안녕하세요, 두 번 째 x-review로 WorldMM을 가지고 왔습니다. 저희 논문 작업에서 벤치마크를 만들면, 그걸 테스트할 여러 LVU methods 중 하나가 WorldMM인데, 처음에 아키텍처를 봤을 때…

Continue Reading