Yifei Ming
Contact: alvinming5 [at] gmail [dot] com
Hi! I am a senior research scientist at . I work on research that advances Gemini’s agentic capabilities on open-ended, long-horizon, and economically valuable tasks, along with the ecosystem that supports them, such as agent harness, environment scaling, and self-evolution.
In the past, I worked on:
-
Post-training: training agents for tool use, long-context reasoning, deep research, and multi-turn interaction
-
Foundational research: understanding the capabilities and limits of LLM-based systems
-
Evaluation & inference scaling: building sandboxes, challenging benchmarks, and improving agentic workloads at test time
I am open to collaborations. We are also hiring self-motivated student researchers / research interns. If you are interested in the directions below, feel free to reach out.
Research directions
(1) Self-evolving agent. Agents can self-evolve by learning from past experience, but today’s trial-and-error process is largely heuristic-driven, time-consuming, prone to overfitting, and poorly supported by existing systems. What infrastructure and algorithms can support a continuously evolving agentic system?
(2) Long-horizon agent. Economically valuable tasks often require humans to work for a long time. When agents operate at that scale, they tend to fail in ways humans wouldn’t (e.g., poorly followed instructions, reward hacking, and context pollution), leading to saturating returns. How do we design scalable environments and agent harnesses that hold up over long horizons?
News
| 05/2026 |
Starting a new position as a Senior Research Scientist at |
|---|---|
| 05/2024 |
Starting a new position as a Research Scientist at |
| 05/2024 |
Defended my Ph.D. thesis on Reliable Foundation Models in the Open World |
| 06/2023 | Research intern at Microsoft Research working on spatial and mathematical reasoning for vision-language models. |
| 05/2022 | Research intern at Adobe working on multi-modal document understanding and robustness. |
Publications
-
ACLHard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier MathIn Annual Conference of the Association for Computational Linguistics (ACL) 2026
-
ACLDiscovering the Gems in Early Layers: Accelerating Long-Context LLMs with 1000x Input Token ReductionIn Annual Conference of the Association for Computational Linguistics (ACL) 2026
-
TMLRA Survey of Frontiers in LLM Reasoning: Inference Scaling, Learning to Reason, and Agentic SystemsIn Transactions on Machine Learning Research (TMLR) 2025
-
TMLRGeneralized Out-of-Distribution Detection and Beyond in Vision Language Model Era: A SurveyIn Transactions on Machine Learning Research (TMLR) 2025
-
NeurIPSIs A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language ModelsIn Neural Information Processing Systems (NeurIPS) 2024
-
ICMLUnderstanding Retrieval-Augmented Task Adaptation for Vision-Language ModelsIn International Conference on Machine Learning (ICML) 2024
-
ICLRProvable Out-of-Distribution Generalization in HypersphereIn International Conference on Learning Representations (ICLR) 2024
-
CPAL
Oral Domain Generalization via Nuclear Norm RegularizationIn Conference on Parsimony and Learning (CPAL) 2023 -
EMNLPA Critical Analysis of Document Out-of-Distribution DetectionIn Empirical Methods in Natural Language Processing (EMNLP Findings) 2023
-
IJCVHow Does Fine-Tuning Impact Out-of-Distribution Detection for Vision-Language Models?In International Journal of Computer Vision (IJCV) 2023
-
NeurIPSDomain Generalization with Nuclear Norm RegularizationIn Neural Information Processing Systems (NeurIPS’W) DistShift Workshop 2022
-
ICMLAre Vision Transformers Robust to Spurious Correlations?In International Conference on Machine Learning (ICML’W), SCIS Workshop 2022