Marvin Gao
AI Engineering & Applied Researcher
Founder & CEO of Parallight · Previously IBM, Alibaba & Ant Group
What I Do
Since 2019 I have been building the mechanism that turns engineers into AI engineers — so frontier AI is something people build, not only consume. I do this as a practicing researcher and engineer, not a commentator: publications in JMLR, TPAMI, ICRA and AAAI; production AI shipped at IBM, Alibaba and Ant Group; and 2,000+ engineers taught across 13 countries, over fifteen years and four technology waves.
- AI engineering & applied research — first-author robotics work on LLM-guided policy enhancement (GUIDES, ICRA 2026), red-teaming as a multi-agent game (TPAMI 2026), sim-to-real transfer with prompt learning (AAAI 2024), and MARLlib (JMLR 2023), a multi-agent RL library now widely used in the field. In industry: China Construction Bank's automated Q&A system, Alipay risk models that earned a national invention patent, and document automation that cut 70% of manual work for the world's largest small-batch PCB manufacturer.
- AI enablement — I still teach directly, and retention data — not intuition — sets the pedagogy. At Parallight I build agent systems that train engineers one-to-one: ~57,000 tutor turns across 997 cloud-sandbox sessions, with weekly-active learners flat at ~50% for eight consecutive weeks. Earlier I ran what was then China's largest online AI-education program for working adults.
- Entrepreneurship — three companies founded or operated: Parallight (profitable in its first three months), KeploreAI (enterprise R&D agents, pre-seed at a $27M valuation), and iLearning.zone ($300K net profit in year one, acquired by Huike Group).
News
Experience
Selected Publications
Published under my Chinese name, Minquan Gao · Google Scholar

GUIDES: Guidance Using Instructor-Distilled Embeddings for Pre-trained Robot Policy Enhancement
IEEE ICRA 2026Gives pre-trained robot policies semantic awareness without architectural redesign: instructor-distilled guidance embeddings fused into the policy's latent space. +10 points absolute task success for transformer policies; real UR5 grasp strike-zone engagement rose from 4.1% to 38.8%.

Red-teaming Language Models as a Multi-round, Multi-agent Game
IEEE TPAMI 2026A game-theoretic alternative to single-round heuristic attacks: red-team and blue-team LLMs co-evolve over multi-round play, producing diverse attackers and measurably safer defenders.

Prompt to Transfer: Sim-to-Real Transfer for Traffic-Signal Control with Prompt Learning
AAAI 2024Uses an LLM's world knowledge as prompt-based dynamics modeling to close the sim-to-real gap for traffic-signal control policies.

MARLlib: A Scalable and Efficient Multi-agent Reinforcement Learning Library
JMLR 2023A unified, reproducible framework and environment interface for multi-agent RL — I owned the system design and architecture, and the distributed implementation for the core algorithms, HAPPO and HATRPO among them. Now widely used in the field.
More About Me
I hold an M.S. in Computer Science from Johns Hopkins University (Interactive Computing & Robotics Lab), an M.S. in Computer Science & Software Engineering from Zhejiang University (Knowledge Graph Laboratory), and a B.S. in Computer Science from Lanzhou University.
My path has run through industry, research and founding in turn. I started as an engineer shipping AI where a wrong answer costs real money: at Ant Group my risk-classification models guarded Alipay transfers and earned a national invention patent; at IBM, as a Band-8 senior scientist, I put China Construction Bank's automated Q&A live in three cities and cut 70% of a manual step for the world's largest small-batch PCB manufacturer. When foundation models changed the field, I went back to research — Peking University, then Johns Hopkins — and the research has stayed active ever since: JMLR 2023, AAAI 2024, and in 2026 both a first-author ICRA paper and a TPAMI paper. It runs on real robot arms and live traffic grids, not benchmarks alone.
Teaching, for me, is an expression of social values rather than a profession. It began in 2011, when I took a volunteer team into Gansu's poorest counties to teach farmers and schoolchildren to use a computer — and was named Outstanding Volunteer of Gansu Province. I've had help from a lot of people throughout my journey, and I believe in giving it back.
Outside of work I head outdoors: I've trekked to Everest Base Camp, crossed the Qinghai–Tibet Plateau twice, and will take any excuse to visit another national park. I also love long drives — I've driven across the United States, coast to coast, twice. The pattern is the same as in my work: pick a far destination, then figure out the route.