2026

SpaMEM: Benchmarking Dynamic Spatial Reasoning via Perception-Memory Integration in Embodied Environments

Chih-Ting Liao, X. Xiao, C. Meng, Z. Chen, Y. Qiao, W. Zhou, T. Wang, X. Zheng

European Conference on Computer Vision (ECCV) 2026

TODO: add a 1~2 sentence TLDR. A benchmark probing dynamic spatial reasoning of embodied vision-language models through perception-memory integration.

SpaMEM: Benchmarking Dynamic Spatial Reasoning via Perception-Memory Integration in Embodied Environments

Chih-Ting Liao, X. Xiao, C. Meng, Z. Chen, Y. Qiao, W. Zhou, T. Wang, X. Zheng

European Conference on Computer Vision (ECCV) 2026

TODO: add a 1~2 sentence TLDR. A benchmark probing dynamic spatial reasoning of embodied vision-language models through perception-memory integration.

Staying VIGILant: Mitigating Visual Laziness via Counterfactual Visual Alignment in MLLMs

X. Xiao, C. Liu, Chih-Ting Liao, Y. Zhang, Q. Lan, Y. Wei, L. Zhao, J. Wang, J. Gu

European Conference on Computer Vision (ECCV) 2026

TODO: add a 1~2 sentence TLDR of VIGILant here.

Staying VIGILant: Mitigating Visual Laziness via Counterfactual Visual Alignment in MLLMs

X. Xiao, C. Liu, Chih-Ting Liao, Y. Zhang, Q. Lan, Y. Wei, L. Zhao, J. Wang, J. Gu

European Conference on Computer Vision (ECCV) 2026

TODO: add a 1~2 sentence TLDR of VIGILant here.

Do LLMs Build World Models From Text? A Multilingual Diagnostic of Spatial Reasoning

Z. Pan, Chih-Ting Liao, C. Liu, X. Xiao, Y. Qiao, C. Meng, Z. Chen, Xin Cao

arXiv preprint 2026 Under review

TODO: add a 1~2 sentence TLDR. A multilingual diagnostic of whether LLMs build world models from text, focused on spatial reasoning.

Do LLMs Build World Models From Text? A Multilingual Diagnostic of Spatial Reasoning

Z. Pan, Chih-Ting Liao, C. Liu, X. Xiao, Y. Qiao, C. Meng, Z. Chen, Xin Cao

arXiv preprint 2026 Under review

TODO: add a 1~2 sentence TLDR. A multilingual diagnostic of whether LLMs build world models from text, focused on spatial reasoning.

Reasoning-VLA: A Fast and General Vision-Language-Action Reasoning Model for Autonomous Driving

D. Zhang, Z. Yuan, Z. Chen, Chih-Ting Liao, Y. Chen, F. Shen, Q. Zhou, Tat-Seng Chua

International Conference on Machine Learning (ICML) 2026

TODO: add a 1~2 sentence TLDR of Reasoning-VLA here.

Reasoning-VLA: A Fast and General Vision-Language-Action Reasoning Model for Autonomous Driving

D. Zhang, Z. Yuan, Z. Chen, Chih-Ting Liao, Y. Chen, F. Shen, Q. Zhou, Tat-Seng Chua

International Conference on Machine Learning (ICML) 2026

TODO: add a 1~2 sentence TLDR of Reasoning-VLA here.

2025

Adversarial Robustness for Unified Multi-Modal Encoders via Efficient Calibration

Chih-Ting Liao, Z. Chen, C. Meng, TY. Huang, Xin Cao, X. Zheng

arXiv preprint 2025 Under review

TODO: add a 1~2 sentence TLDR. Improving adversarial robustness of unified multi-modal encoders (e.g. ImageBind / UniBind) via efficient calibration.

Adversarial Robustness for Unified Multi-Modal Encoders via Efficient Calibration

Chih-Ting Liao, Z. Chen, C. Meng, TY. Huang, Xin Cao, X. Zheng

arXiv preprint 2025 Under review

TODO: add a 1~2 sentence TLDR. Improving adversarial robustness of unified multi-modal encoders (e.g. ImageBind / UniBind) via efficient calibration.