Skip to content

Xiaomi Unveils MiMo-V2.6, Claiming Major Reinforcement-Learning Gains Across Coding, Design and Cybersecurity

Sep 22, 2026
alphaXiv
Article image for Xiaomi Unveils MiMo-V2.6, Claiming Major Reinforcement-Learning Gains Across Coding, Design and Cybersecurity

Summary

Xiaomi unveils MiMo-V2.6, an omni-modal AI model family it says significantly improves coding, design and cybersecurity through groupwise reinforcement learning, with its Pro version reaching 72.6 average@3 on DeepSWE v1.1 after the team freezes an unstable Mixture-of-Experts router.

Key Points

  • Xiaomi's LLM-Core introduces MiMo-V2.6, an omni-modal model family that scales reinforcement learning with groupwise quality grading across code, workflow, visual-design and cybersecurity tasks.
  • On the 113-task DeepSWE v1.1 benchmark, MiMo-V2.6-Pro rises from 58.4 to 72.6 average@3 during RL, while Flash climbs from 48.7 to 65.7.
  • The team freezes MiMo-V2.6's Mixture-of-Experts router after trainable routing drives peak expert load from 6 times to 16 times the average and leaves 22% of experts cold.

Tags

Read Original Article