Efficient Vision Transformers
Attention that scales down without giving up accuracy — partial attention, shuffled local self-attention, spatial decay matrices, and refined attention maps for recognition and dense prediction.
Open to collaboration
Postdoctoral Researcher UNIST Vision and Learning (UVL) Lab Ulsan, South Korea
My research connects seeing, describing, predicting, and acting. I care about systems that generalize past their training data — and that run at a cost someone can actually afford.
Attention that scales down without giving up accuracy — partial attention, shuffled local self-attention, spatial decay matrices, and refined attention maps for recognition and dense prediction.
Segmenting categories never seen at training time by grounding pixel-level representations in language rather than a closed label set.
Joint image-text representations, cross-modal fusion, and text-guided visual learning for retrieval, re-identification, and grounded recognition.
Adapting and evaluating MLLMs — GPT-4o, LLaVA, Qwen, LLaMA, MiniCPM — for visual reasoning tasks under practical latency and memory constraints.
LLM agents that plan and call tools — ReAct and code-agent loops, LangChain and LlamaIndex pipelines, and deployed voice agents.
Prediction in latent rather than pixel space — object-centric models that discover keypoints, boxes, and masks from video alone, learn action spaces without annotation, and trade extra loss terms for inductive bias.
Showing 74 of 74
No publications match those filters.
Efficient Multi-Scale Spatial Interactions for Visual Recognition Tasks pdf ↗
Refining Attention Maps for Vision Transformer pdf ↗
Hierarchical Vision Transformers with Shuffled Local Self-Attentions pdf ↗
Balancing Multiple Object Tracking Objectives based on Learned Weighting Factors pdf ↗
Awarded for IEEE Transactions on Industrial Informatics (2022) and Neurocomputing (2022)
I am always glad to hear about collaborations, student supervision, or reviewing invitations. Email is the fastest way to reach me.