Bowen Qu
Bowen Qu Brian

About Me

👋Hi, I’m Bowen(Brian) Qu, a graduate of Peking University and MTS at Moonshot.ai (Kimi, 月之暗面) Posttrain. My research interests include MLLM, TIR(Tool-Integrated Reasoning) and Agentic RL. The logo of this website is my lovely cat - Baka (巴卡 in Chinese)!

  • Experience: Fortunately, I have the honor to participate in some interesting MLLM research projects:
    • [2025.03 - Present] Moonshot.ai, Technical Staff (PostTrain). Cooking K3 & K2 series:
      • 🔹 Vision Agentic RL (K3 & K2.6 - Visual Agent, Frontier-Level Vision Agentic)
      • 🔹 Native Multimodal RL (K2.5 Report Chap2.2, ZeroVision ColdStart -> Vision-Centric RL)
      • 🔹 Chart Understanding and Chart-to-Code
    • [2024.04 - 2024.12] 01.ai & Rhymes.ai, Research Intern, Multimodal Team, supervised by Junnan Li, working closely with Dongxu Li and Haoning Wu.
      • 🏆 Core Contributor of Aria — an Open Multimodal Native MoE
    • [2024.02 - 2024.07] IDEA Research, Research Intern, working closely with Zhengzhuo Xu, Yiyan Qi and Chengjin Xu.
      • 🏆 Co-first Author of ChartMoE (ICLR2025 Oral): Mixture of Diversely Aligned Expert Connector for Chart Understanding
  • Status: I'm always eager to learn new insights and ideas. The potential of Agents are still under exploration. If you're in Beijing, let's grab coffee to discuss it. Also feel free to drop me an 📧 if there is a good fit!
Interests
  • Agentic RL
  • MLLM Reasoning, Knowledge and Perception
Education
  • M.S., Computer Science

    Peking University (PKU), 2022-2025

  • B.E., Electronic Engineering

    Huazhong University of Science and Technology (HUST), 2018-2022

🌔 Work at Moonshot.ai (Kimi)

Post-training on K3 Vision Agentic · K2.6 VTIR · K2.5 Vision Perception & Knowledge & Reasoning

2026.07 K3 RL scaling curves: Agentic Chart Understanding & Visual Puzzles highlighted

Kimi K3

Open-weight flagship · Frontier Vision Agentic · 2.8T parameters · native multimodal · 1M context

Vision Singlestep & Agentic RL

  • Long-run RL for vision singlestep and agentic, for multiple rounds of iterative optimization (K3 Report Chap4.2.3)
  • Optimization on multiple harnesses generalization
  • Scaling Mid-train data and crafting RL Promptset
  • At launch: Top-2 on Vision Agentic, e.g.: CharXiv RQ w/ python, only behind Claude Fable 5. Generalized vision agentic capability on multiple harnesses.
2026.01 / 04 K2.5 vision RL training curves on vision benchmarks

Kimi K2.5 / K2.6

Kimi's First Unified Vision–Text Model -> First Opensourced Frontier-level VTIR

Vision Post-training & VTIR

  • RL on Zero-Vision SFT: Text-only SFT model, followed by a large and long-run RL, activates close-to-SOTA vision capabilities. And then producing lots of on-policy vision data. (K2.5 Report Chap2.2, ZeroVision ColdStart -> Vision-Centric RL)
  • Coldstart & RL on K2.6 VTIR: enhancing visual perception, understanding and reasoning capabilities by general ipython integration, achieving Top-3 VTIR performance (Top-1 in Opensourced).
  • Large-scale Chart-to-Code
🔥 News
  • 2026.07: We are releasing Kimi K3, Open Frontier Intelligence, Frontier Vision Agentic! Vision Agentic Showcase
  • 2026.04: We release Kimi-K2.6: Frontier-Level Vision TIR! Visualize: K2.6 VTIR Trajectories
  • 2026.01: We release Kimi-K2.5!
  • 2025.11: We release IE-Critic-R1, a Pointwise, Generative Reward Model optimized by RLVR, specialized in assessing the quality of text-driven image editing results.
  • 2025.07: We release Kimi-K2!
  • 2025.02: ChartMoE is selected as ICLR 2025 Oral · Top 1.8%. It is a MLLM with MoE connector, for advanced chart 1️⃣understanding, 2️⃣replot, 3️⃣editing, 4️⃣highlighting and 5️⃣transformation.
  • 2024.10: We release Aria, a native LMM that excels on text, code, image, video, PDF and more!
  • 2023.12: Releasing personal project MPP-LLaVA series, clean repo on LLaVA-like MLLM training, conducted on RTX3090 GPUs by Pipeline Parallel.
Selected Publication

* Equal Contribution (i.e.: Co-First Author). 📧 Corresponding Author.

(Moonshot AI, 2026). Kimi K3: Open Frontier Intelligence. Technical Report.
(Moonshot AI, 2026). PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models. arXiv preprint.
(Moonshot AI, 2026). Kimi-K2.5: Visual Agentic Intelligence. Technical Report.
(Moonshot AI, 2025). Kimi K2: Open Agentic Intelligence. Technical Report.
(Moonshot AI, 2025). Kimi-VL Technical Report. arXiv preprint.
(2025). ChartMoE: Mixture of Diversely Aligned Expert Connector for Chart Understanding. ICLR 2025 Oral · Top 1.8%
(01.AI & Rhymes.AI, 2024). Aria: An Open Multimodal Native Mixture-of-Experts Model. Technical Report.