Zhenwei Dai 戴振威
I am a Principal Applied Scientist at Microsoft M365 Copilot, where I lead LLM post-training and evaluation. My work focuses on improving instruction following, response quality, reliability, and inference efficiency for agents in everyday productivity workflows.
Previously, I worked across Amazon Ads, Amazon Search, and AWS AI Research and Education (AIRE), focusing on NLP model training and LLM post-training.
I earned my PhD at Rice University in August 2022, advised by Prof. Anshumali Shrivastava and Prof. Reinhard Heckel. My doctoral research focused on efficient algorithms for large-scale machine learning, randomized algorithms, and data mining.
Research interests
LLM post-training; reliable agents, tool use, and evaluation; efficient reasoning and adaptation; memory-efficient algorithms and probabilistic data structures.
MiniCorp: The Last Mile of the AI Agent Firm
Featured research project · 2026
Office Simulators for Enterprise AGI
MiniCorp is an office simulator where AI agents work together, make decisions, and respond to market feedback. We use this setting to study how an AI-run company learns over time.
A collaboration with Jingying Zeng and our coauthors.
Dec 2025 – Present
Principal Applied Scientist
Microsoft M365 Copilot · Redmond, WA
May 2025 – Nov 2025
Senior Applied Scientist
Amazon Ads · LLM Agent · Seattle, WA
May 2023 – Apr 2025
Senior Applied Scientist
Amazon Search · Palo Alto, CA
Sep 2022 – Apr 2023
Applied Scientist
AWS AI Research and Education · Santa Clara, CA
* denotes equal contribution.
MiniCorp: The Last Mile of the AI Agent Firm
Jingying Zeng, Zhenwei Dai, Jinning Li, Changho Shin, Dylan Zhang, Yuxuan Lu, Qi He, Dakuo Wang, Kai-Wei Chang.
arXiv preprint, October 2026
Firefly: Illuminating Verified Real-world Tool Call Data Generation
Yuxuan Lu, Ziyi Wang, Yingzhou Lu, Yisi Sang, Jiri Gesi, Xianfeng Tang, Yimeng Zhang, Zhenwei Dai, Hui Liu, Hanqing Lu, Chen Luo, Qi He, Benoit Dumoulin, Jing Huang, Dakuo Wang.
A Reward-Guided Dual-Phase Framework for Adaptive Inference-Time Reasoning
Yingqian Cui, Zhenwei Dai, Pengfei He, Bing He, Hui Liu, Zhan Shi, Xianfeng Tang, Jingying Zeng, Suhang Wang, Yue Xing, Jiliang Tang, Benoit Dumoulin.
TRAJECT-Bench: A Trajectory-Aware Benchmark for Evaluating Agentic Tool Use
Pengfei He, Zhenwei Dai, Bing He, Hui Liu, Xianfeng Tang, Hanqing Lu, Juanhui Li, Jiayuan Ding, Subhabrata Mukherjee, Suhang Wang, Yue Xing, Jiliang Tang, Benoit Dumoulin.
To Trust or Not to Trust: Attention-based Trust Management for LLM Multi-Agent Systems
Pengfei He, Zhenwei Dai, Xianfeng Tang, Yue Xing, Hui Liu, Jingying Zeng, Qiankun Peng, Shrivats Agrawal, Samarth Varshney, Suhang Wang, Jiliang Tang, Qi He.
ACL 2026, main conference; NeurIPS 2026, poster
Keeping an Eye on LLM Unlearning: The Hidden Risk and Remedy
Jie Ren, Zhenwei Dai, Xianfeng Tang, Yue Xing, Shenglai Zeng, Jingying Zeng, Qiankun Peng, Samarth Varshney, Suhang Wang, Qi He, Charu Aggarwal, Hui Liu.
View all selected publications →