Sections
Text Area

PUBLICATIONS

Journal Article
Trustworthy Generative AI
The Curse of CoT: On the Limitations of Chain-of-Thought in In-Context Learning

Zheng, T., Chen, Y., Li, C., Li, C., Zong, Q., Shi, H., Xu, B., Song, Y., Wong, G. Y. & See, S., 2025, The Curse of CoT: On the Limitations of Chain-of-Thought in In-Context Learning. Transactions on Machine Learning Research. vol. 2025-November.

Journal Article
Trustworthy Generative AI
Threat-Aware UAV Dodging of Human-Thrown Projectiles with an RGB-D Camera

Zhang, Y., Fan, N., Zheng, H., Liang, J., Pan, Z., Chen, Q. & Lyu, X., 2025, Threat-aware UAV Dodging of Human-Thrown Projectiles with an RGB-D Camera. IEEE Robotics and Automation Letters. vol. 11, no. 2, p. 1178–1185, article no. 11283030.

Journal Article
Sustainability
PAL: Boosting Skin Lesion Segmentation via Probabilistic Attribute Learning

Yuan, Y., Wang, X., Li, J., Chen, G. & Heng, P. A., 2025, PAL: Boosting Skin Lesion Segmentation via Probabilistic Attribute Learning. IEEE Transactions on Medical Imaging. vol. 44, no. 12, p. 5183–5196, article no. 11078393.

Journal Article
Sustainability
Temporal-multimodal consistency alignment for Alzheimer’s cognitive assessment prediction

Yang, X., Dang, X., Cai, J., Li, J., Wang, X. & Heng, P. A., 2025, Temporal-Multimodal Consistency Alignment for Alzheimer's Cognitive Assessment Prediction. Medical Physics. vol. 52, no. 6, p. 5064–5080.

Conference Paper
Human-Robot Interaction
“AI Afterlife” as Digital Legacy: Perceptions, Expectations, and Concerns

Lei, Y., Ma, S., Sun, Y. & Ma, X., 2025, "AI Afterlife" as Digital Legacy: Perceptions, Expectations, and Concerns. CHI 2025 - Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems. Association for Computing Machinery (ACM), p. 1–18 (article no. 981).

Conference Paper
Human-Robot Interaction
Scaffolded Turns and Logical Conversations: Designing Humanized LLM-Powered Conversational Agents for Hospital Admission Interviews

Liu, D., Zhang, Y., Zhao, B., Ma, S., Shi, C. & Ma, X., 2025, Scaffolded Turns and Logical Conversations: Designing Humanized LLM-Powered Conversational Agents for Hospital Admission Interviews. CHI 2025 - Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems. Association for Computing Machinery (ACM), p. 1–23 (article no. 643).

Conference Paper
Robotics & Embodied AI
A Data-Driven Velocity Estimator for Autonomous Underwater Vehicles Experiencing Unmeasurable Flow and Wave Disturbance

Cai, J., Mayberry, S., Yin, H. & Zhang, F., 2025, A Data-Driven Velocity Estimator for Autonomous Underwater Vehicles Experiencing Unmeasurable Flow and Wave Disturbance. 2025 IEEE International Conference on Robotics and Automation, ICRA 2025. Institute of Electrical and Electronics Engineers (IEEE), p. 4138–4144.

Conference Paper
Robotics & Embodied AI
BEINGS: Bayesian Embodied Image-Goal Navigation with Gaussian Splatting

Meng, W., Wu, T., Yin, H. & Zhang, F., 2025, BEINGS: Bayesian Embodied Image-Goal Navigation with Gaussian Splatting. 2025 IEEE International Conference on Robotics and Automation, ICRA 2025. Institute of Electrical and Electronics Engineers (IEEE), p. 5252–5258.

Conference Paper
Robotics & Embodied AI
Design of a Formation Control System to Assist Human Operators in Flying a Swarm of Robotic Blimps

Wu, T., Fu, J., Meng, W., Cho, S., Zhan, H. & Zhang, F., 2025, Design of a Formation Control System to Assist Human Operators in Flying a Swarm of Robotic Blimps. 2025 IEEE International Conference on Robotics and Automation, ICRA 2025. Institute of Electrical and Electronics Engineers (IEEE), p. 8929–8935 (article no. 11128354).

Conference Paper
Robotics & Embodied AI
Design of a Gesture-Controlled Multi-Blimp System

Lyu, M., Wu, T. & Zhang, F., 2025, Design of a Gesture-Controlled Multi-Blimp System. Proceedings - 2025 10th International Conference on Automation, Control and Robotics Engineering, CACRE 2025. Institute of Electrical and Electronics Engineers (IEEE), p. 65–71 (article no. 11119595).

Conference Paper
Robotics & Embodied AI
QuietBlimp: A Human-Friendly Miniature Autonomous Noise-Mild Blimp for Indoor Environment

Wu, T., Lyu, M., Zhao, Y., Fu, J., Yang, Y., Tao, Q. & Zhang, F., 2025, QuietBlimp: A Human-Friendly Miniature Autonomous Noise-Mild Blimp for Indoor Environment. 2025 IEEE/ASME International Conference on Advanced Intelligent Mechatronics, AIM 2025. Institute of Electrical and Electronics Engineers (IEEE), article no. 11175635.

Conference Paper
Trustworthy Generative AI
ComparisonQA: Evaluating Factuality Robustness of LLMs Through Knowledge Frequency Control and Uncertainty

Zong, Q., Wang, Z., Ren, X., Zheng, T. & Song, Y., 2025, ComparisonQA: Evaluating Factuality Robustness of LLMs Through Knowledge Frequency Control and Uncertainty. Findings of the Association for Computational Linguistics: ACL 2025. Association for Computational Linguistics (ACL), p. 4101–4117.

Conference Paper
Trustworthy Generative AI
ConKE: Conceptualization-Augmented Knowledge Editing in Large Language Models for Commonsense Reasoning

Zhang, L., Wang, W., Fang, T. & Song, Y., 2025, ConKE: Conceptualization-Augmented Knowledge Editing in Large Language Models for Commonsense Reasoning. Findings of the Association for Computational Linguistics, ACL 2025. Association for Computational Linguistics (ACL), p. 627–635.

Conference Paper
Trustworthy Generative AI
EcomScriptBench: A Multi-task Benchmark for E-commerce Script Planning via Step-wise Intention-Driven Product Association

Wang, W., Cui, L., Liu, X., Nag, S., Xu, W., Luo, C., Sarwar, S. M., Li, Y., Gu, H., Liu, H., Yu, C., Bai, J., Gao, Y., Zhang, H., He, Q., Ji, S. & Song, Y., 2025, EcomScriptBench: A Multi-task Benchmark for E-commerce Script Planning via Step-wise Intention-Driven Product Association. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025. Association for Computational Linguistics (ACL), p. 1–22.

Conference Paper
Trustworthy Generative AI
KnowShiftQA: How Robust are RAG Systems when Textbook Knowledge Shifts in K-12 Education?

Zheng, T., Li, W., Bai, J., Wang, W. & Song, Y., 2025, KnowShiftQA: How Robust are RAG Systems when Textbook Knowledge Shifts in K-12 Education? Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), ACL 2025. Association for Computational Linguistics (ACL), p. 183–195.

Conference Paper
Trustworthy Generative AI
MARS: Benchmarking the Metaphysical Reasoning Abilities of Language Models with a Multi-task Evaluation Dataset

Wang, W. & Song, Y., 2025, MARS: Benchmarking the Metaphysical Reasoning Abilities of Language Models with a Multi-task Evaluation Dataset. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025. Association for Computational Linguistics (ACL), p. 1568–1596.

Conference Paper
Trustworthy Generative AI
PrivaCI-Bench: Evaluating Privacy with Contextual Integrity and Legal Compliance

Li, H., Hu, W., Jing, H., Chen, Y., Hu, Q., Han, S., Chu, T., Hu, P. & Song, Y., 2025, PrivaCI-Bench: Evaluating Privacy with Contextual Integrity and Legal Compliance. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics (ACL), p. 10544–10559.

Conference Paper
Trustworthy Generative AI
Revisiting Epistemic Markers in Confidence Estimation: Can Markers Accurately Reflect Large Language Models’ Uncertainty?

Liu, J., Zong, Q., Wang, W. & Song, Y., 2025, Revisiting Epistemic Markers in Confidence Estimation: Can Markers Accurately Reflect Large Language Models' Uncertainty? Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), ACL 2025. Association for Computational Linguistics (ACL), p. 206–221.

Conference Paper
Human-Centric Benchmarking
Large Language Model Tokens are Psychologically Salient

Haslett, D. A., Chan, A. B. & Hsiao, J. H.-W., 2025, Large Language Model Tokens are Psychologically Salient. Paper presented at The 47th Annual Meeting of the Cognitive Science Society (COGSCI2025), San Francisco, United States. p. 4819.

Conference Paper
Human-Centric Benchmarking
Whose Values Prevail? Bias in Large Language Model Value Alignment

Qi, R., Papyshev, G., Tsai, K. S., Chan, A. B. & Hsiao, J. H.-W., 2025, Whose Values Prevail? Bias in Large Language Model Value Alignment. Paper presented at 47th Annual Meeting of the Cognitive Science Society, San Francisco, United States. p. 665–672.

Conference Paper
Trustworthy Generative AI
Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement Learning

Hu, W., Li, H., Jing, H., Hu, Q., Zeng, Z., Han, S., Xu, H., Chu, T., Hu, P. & Song, Y., 2025, Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement Learning. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics (ACL), p. 865–883.

Conference Paper
Trustworthy Generative AI
INTEGROUND: On the Evaluation of Verification and Retrieval Planning in Integrative Grounding

Cheng, J., Zhuang, Q., Li, H., Chan, C., Liu, X., Qiu, L. & Song, Y., 2025, InteGround: on the Evaluation of Verification and Retrieval Planning in Integrative Grounding. EMNLP 2025 - 2025 Conference on Empirical Methods in Natural Language Processing, Findings of EMNLP 2025. Association for Computational Linguistics (ACL), p. 13587–13602.

Conference Paper
Trustworthy Generative AI
LogiDynamics: Unraveling the Dynamics of Inductive, Abductive and Deductive Logical Inferences in LLM Reasoning

Zheng, T., Cheng, J., Li, C., Shi, H., Wang, Z., Bai, J., Song, Y., Wong, G. Y. & See, S., 2025, LogiDynamics: Unraveling the Dynamics of Inductive, Abductive and Deductive Logical Inferences in LLM Reasoning. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics (ACL), p. 20710–20731.

Conference Paper
Trustworthy Generative AI
MCIP: Protecting MCP Safety via Model Contextual Integrity Protocol

Jing, H., Li, H., Hu, W., Hu, Q., Xu, H., Chu, T., Hu, P. & Song, Y., 2025, MCIP: Protecting MCP Safety via Model Contextual Integrity Protocol. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics (ACL), p. 1177–1194.

Conference Paper
Sustainability
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling

Chen, T., Xu, X., Liu, Z., Li, P., Song, X., Jaiswal, A. K., Zhang, F., Hu, J., Wang, Y., Chen, H., Diao, S., Liu, S., Li, Y., Yin, L. & Yang, C., 2025, GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling. Proceedings of the 39th Conference on Neural Information Processing Systems (NeurIPS 2025). article no. 21177.

Conference Paper
Trustworthy Generative AI
MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly

Wang, Z., Yu, W., Ren, X., Zhang, J., Zhao, Y., Saxena, R., Cheng, L., Wong, G., See, S., Minervini, P., Song, Y. & Steedman, M., 2025, MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly. Paper presented at The Thirty-Ninth Annual Conference on Neural Information Processing Systems (NeurIPS 2025), California, United States.

Conference Paper
Trustworthy Generative AI
ZEBRA: Towards Zero-Shot Cross-Subject Generalization for Universal Brain Visual Decoding

Wang, H., Lu, J., Li, H. & Li, X., 2025, ZEBRA: Towards Zero-Shot Cross-Subject Generalization for Universal Brain Visual Decoding. Proceedings of the 39th Conference on Neural Information Processing Systems (NeurIPS 2025). 2025.

Conference Paper
Human-Centric Benchmarking
Made in China, thinking in America: U.S. Values Persist in Chinese LLMs

Haslett, D., Huang, L. T.-L., Khalatbari, L., Hsiao, J. H.-W. & Chan, A. B., 2025, Made-in China, Thinking in America: U.S. Values Persist in Chinese LLMs. CogSci Asia–Pacific Meetup 2025.