Publications
In reverse chronological order. * denotes equal contribution. Also on Google Scholar.
2026
-
ECCV 2026Multiple Images Distract Large Multimodal Models via Attention FragmentationEuropean Conference on Computer Vision (ECCV), 2026While many tasks require reasoning across multiple images, open-source Large Multimodal Models (LMMs) remain unreliable in these settings. We analyze multi-image LMMs and identify a phenomenon we term attention fragmentation: in each image, tokens at similar background locations act as attention sinks, absorbing disproportionate attention. Causal masking further skews this effect, as earlier images accumulate more sink attention than later images. Using an entropy score over per-image attention, we find that visual focus remains highly dispersed across images rather than isolating key evidence. By applying Pinsker's inequality, we establish a theoretical bound showing that this high-entropy dispersion, combined with stronger early-image sinks, strictly reduces the usable non-sink attention available to earlier images, providing a mechanistic link to image order sensitivity. Motivated by this diagnosis, we propose Attention Remasking (AR), a post-training edit that blocks sink keys and opens a sparse set of cross-image links, routing the recovered attention to task-relevant tokens. AR improves accuracy and reduces order sensitivity across multi-image benchmarks, narrowing the gap between open-source LMMs and leading commercial models.
@inproceedings{qiao2026multiple, title = {Multiple Images Distract Large Multimodal Models via Attention Fragmentation}, author = {Qiao, Tingrui and Zhao, Di and Li, Yuzhuo and Pang, Bo and Walker, Caroline and Cunningham, Chris and Koh, Yun Sing}, booktitle = {European Conference on Computer Vision (ECCV)}, pages = {361--379}, publisher = {Springer Nature Switzerland}, year = {2026} } -
WACV 2026Overcoming Fine-Grained Visual Challenges in Animal Re-Identification via Semantic Feature AlignmentIEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2026Identifying individual animals at different points in space and time is vital for effective wildlife monitoring and biodiversity conservation. While existing computer vision methods have shown promise in re-identifying animals, their capability in Animal Re-Identification (Animal ReID) remains restricted by the inherent visual variations, specifically high intra- and low inter-identity variations. High intra-identity variations refer to high visual diversity within the same individual due to pose or form changes and occlusions, and low inter-identity variations refer to subtle visual differences between distinct individuals due to fine-grained appearances. To address these challenges, we propose the Clip-based Animal RE-identification (CARE) framework, which leverages the image-conditioned textual description generation and individual-level semantic feature alignment, mitigating the negative impacts of visual variations in Animal ReID. Crucially, we have packaged CARE into a stand-alone toolkit and piloted it with stakeholders, facilitating real-world wildlife monitoring for biodiversity conservation. Extensive experiments on benchmark and in-the-wild datasets further demonstrate that CARE consistently outperforms state-of-the-art methods, validating its effectiveness in Animal ReID.
@inproceedings{Wu_2026_WACV, title = {Overcoming Fine-Grained Visual Challenges in Animal Re-Identification via Semantic Feature Alignment}, author = {Wu, Yihao and Zhao, Di and Li, Yuzhuo and Alajas, Matthew and Glen, Alistair S. and Zhang, Jingfeng and Dobbie, Gillian and Wilson, Daniel and Koh, Yun Sing}, booktitle = {Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)}, month = {March}, pages = {371--381}, year = {2026} }
2025
-
ACM MM 2025MetaWild: A Multimodal Dataset for Animal Re-Identification with Environmental MetadataACM International Conference on Multimedia (ACM MM), 2025Identifying individual animals within large wildlife populations is essential for effective wildlife monitoring and conservation efforts. Recent advancements in computer vision have shown promise in animal re-identification (Animal ReID) by leveraging data from camera traps. However, existing Animal ReID datasets rely exclusively on visual data, overlooking environmental metadata that ecologists have identified as highly correlated with animal behavior and identity, such as temperature and circadian rhythms. Moreover, the emergence of multimodal models capable of jointly processing visual and textual data presents new opportunities for Animal ReID, but existing datasets fail to leverage these models' text-processing capabilities, limiting their full potential. Additionally, to facilitate the use of metadata in existing ReID methods, we propose the Meta-Feature Adapter (MFA), a lightweight module that can be incorporated into existing vision-language model (VLM)-based Animal ReID methods, allowing ReID models to leverage both environmental metadata and visual information to improve ReID performance. Experiments on MetaWild show that combining baseline ReID models with MFA to incorporate metadata consistently improves performance compared to using visual information alone, validating the effectiveness of incorporating metadata in re-identification. We hope that our proposed dataset can inspire further exploration of multimodal approaches for Animal ReID.
@inproceedings{li2025metawild, title = {MetaWild: A Multimodal Dataset for Animal Re-Identification with Environmental Metadata}, author = {Li, Yuzhuo and Zhao, Di and Qiao, Tingrui and Wu, Yihao and Pang, Bo and Koh, Yun Sing}, booktitle = {Proceedings of the 33rd ACM International Conference on Multimedia}, pages = {13009--13015}, year = {2025} } -
MSc ThesisImproving Domain Generalisation with Vision-Language IntegrationMSc Thesis, The University of Auckland, 2025@mastersthesis{li2025improving, title = {Improving Domain Generalisation with Vision-Language Integration}, author = {Li, Yuzhuo}, school = {The University of Auckland}, year = {2025} }