Author: 박 성준

Posted in X-Diary

2026년 상반기 회고 – 박성준

안녕하세요. 박성준 연구원입니다. 벌써 2026년이 절반이 넘게 지나갔습니다. 연초에 세웠던 목표, 계획들을 돌아보면서 2026년 상반기 어떻게 지나갔는지를 회고하려합니다. 상반기 목표 먼저 상반기 목표는 크게 두가지…

Continue Reading
Posted in X-Review

[CVPR 2025] VGGT: Visual Geometry Grounded Transformer

안녕하세요. 오늘 리뷰할 논문은 CVPR 2025에서 Best Paper Award를 받은 VGGT(Visual Geometry Grounded Transformer)입니다. Introduction 본 논문은 3D reconstruction task를 다루고 있습니다. 3D reconstruction은 기본적으로…

Continue Reading
Posted in X-Diary

ICML 2026 참관기

안녕하세요. 서울 코엑스에서 열린 ICML을 다녀오며 느낀점을 참관기로 남겨보려합니다. ICML은 ML 분야 최상위 학회이고 서울 코엑스에서 열렸습니다. 이번 ICML 2026에는 2만4천에 달하는 논문들이 투고되었고 그…

Continue Reading
Posted in X-Review

[ ICML 2026 ] Think in Latent, Explain in Language: Self-Explainable Latent Reasoning

안녕하세요. 오늘은 ICML 2026에서 흥미로운 주제를 다룬 논문이 있어 소개하려 합니다. 요즘 latent reasoning 쪽을 흥미 있게 팔로우하고 있었는데, ICML에 관련 논문이 있어 읽어보고 리뷰하게…

Continue Reading
Posted in X-Review

[ ICLR 2024 ] ANTGPT: CAN LARGE LANGUAGE MODELS HELP LONG-TERM ACTION ANTICIPATION FROM VIDEOS?

안녕하세요. 오늘 리뷰할 논문은 ICLR 2024에 발표된 AntGPT입니다. AntGPT는 영상을 입력 받아 영상 이후에 나올 사람의 행동을 예측하는 long-term action anticipation(이하 LTA) 문제에 대규모 언어…

Continue Reading
Posted in X-Review

[arXiv 2026] Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video

안녕하세요. 오늘 리뷰할 논문은 Video-MME-v2입니다. Video-MME는 긴 비디오 이해 분야에서 가장 널리 활용되는 데이터셋입니다. 최근에 Video-MME 팀이 새로 데이터셋을 공개하여 해당 논문을 리뷰하려합니다. Introduction 최근…

Continue Reading
Posted in X-Review

[CVPR 2026] Learnability-Guided Diffusion for Dataset Distillation

안녕하세요, 박성준 연구원입니다. 최근 CVPR 2026에 accept된 논문들을 읽어보는 중에 흥미로운 주제를 발견하여 리뷰하고자합니다. 당분간은 CVPR 2026 논문들을 읽고 소개하려합니다. Before Review 리뷰할 논문이 다루는…

Continue Reading
Posted in X-Review

[CVPR 2026] Generative Video Compression with One-Dimensional Latent Representation

오늘 리뷰는 CVPR 2026에 게재된 Video Compression 논문입니다. Introduction 비디오 데이터의 증가로 인해서 낮은 비트레이트에서도 높은 품질을 유지하는 동시에 효율적으로 압축하는 기술이 점점 중요해지고 있습니다….

Continue Reading
Posted in X-Review

[ArXiv 2025] Active Video Perception: Iterative Evidence Seekingfor Agentic Long Video Understanding

안녕하세요, 오늘 리뷰할 논문은 Active Video Perception(AVP)입니다. Long Video Understanding 연구로 기존의 agentic 파이프라인의 단점을 보완한 연구입니다. Introduction 긴 비디오 이해(Long Video Understanding, LVU)는 대부분…

Continue Reading
Posted in X-Review

[NIPS2025] Vgent: Graph-based Retrieval-Reasoning-Augmented Generation For Long Video Understanding

안녕하세요. 박성준 연구원입니다. 오늘 리뷰할 논문은 LVU연구인 Vgent입니다. NIPS2025에서 spotlight로 선정된 연구입니다. Introduction 대규모 비디오 언어 모델(Large Video Language Model, LVLM)은 영상과 자연어를 동시에 다루며…

Continue Reading