Researcher @ Xiaomi MiMo

Qingkai Fang 房庆凯

I am a Researcher at Xiaomi MiMo. I received my Ph.D. degree in June 2026 from Institute of Computing Technology, Chinese Academy of Sciences (ICT/CAS), advised by Prof. Yang Feng. Before that, I received my B.E. degree in Computer Science and Technology from Beijing University of Posts and Telecommunications (BUPT) in Jun. 2021.

Contact
Phone / WeChat
+86 13120328199
Qingkai Fang

My research focuses on large language models and related areas. Currently, I am involved in building:

03

Omni Multimodal Models

Multimodal understanding models supporting image, video, and audio inputs.

I successfully defended my Ph.D. thesis and joined Xiaomi MiMo as a full-time researcher.

One paper is accepted to ACL 2026.

One paper is accepted to NeurIPS 2025.

One paper is accepted to ACL 2025.

Two papers are accepted to ICLR 2025.

Four papers are accepted to ACL 2024.

One paper is accepted to EMNLP 2023.

One paper is accepted to NeurIPS 2023.

Three papers are accepted to ACL 2023.

One paper is accepted to EMNLP 2022.

Two papers are accepted to ACL 2022.

2026

BayLing-Duplex: Native Full-Duplex Speech Dialogue with a Single Autoregressive LLM

Qingkai Fang, Shoutao Guo, Yang Feng.

Preprint

Efficient Training for Cross-lingual Speech Language Models

Yan Zhou, Qingkai Fang, Yun Hong, Yang Feng.

Findings of ACL 2026

MiMo-V2-Flash Technical Report

Technical ReportContributor

2025

LLaMA-Omni 2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis

Qingkai Fang, Yan Zhou, Shoutao Guo, Shaolei Zhang, Yang Feng.

Proceedings of ACL 2025 CCF-A

LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Qingkai Fang, Shoutao Guo, Yan Zhou, Zhengrui Ma, Shaolei Zhang, Yang Feng.

ICLR 2025 CCF-A

LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Shaolei Zhang, Qingkai Fang, Zhe Yang, Yang Feng.

ICLR 2025 CCF-A

FastLongSpeech: Enhancing Large Speech-Language Models for Efficient Long-Speech Processing

Shoutao Guo, Shaolei Zhang, Qingkai Fang, Zhengrui Ma, Min Zhang, Yang Feng.

NeurIPS 2025 CCF-A

MiMo-Audio: Audio Language Models are Few-Shot Learners

Technical ReportCore Contributor

MiMo-VL Technical Report

Technical ReportContributor

MiMo: Unlocking the Reasoning Potential of Language Model--From Pretraining to Posttraining

Technical ReportContributor

2024

Can We Achieve High-quality Direct Speech-to-Speech Translation Without Parallel Speech Data?

Qingkai Fang, Shaolei Zhang, Zhengrui Ma, Min Zhang, Yang Feng.

Proceedings of ACL 2024 CCF-A

CTC-based Non-autoregressive Textless Speech-to-Speech Translation

Qingkai Fang, Zhengrui Ma, Yan Zhou, Min Zhang, Yang Feng.

Findings of ACL 2024

StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task Learning

Shaolei Zhang, Qingkai Fang, Shoutao Guo, Zhengrui Ma, Min Zhang, Yang Feng.

Proceedings of ACL 2024 CCF-A

A Non-autoregressive Generation Framework for Simultaneous Speech-to-x Translation

Zhengrui Ma, Qingkai Fang, Shaolei Zhang, Shoutao Guo, Yang Feng, Min Zhang.

Proceedings of ACL 2024 CCF-A

BayLing 2: A Multilingual Large Language Model with Efficient Language Alignment

Shaolei Zhang, Kehao Zhang, Qingkai Fang, Shoutao Guo, Yan Zhou, Xiaodong Liu, Yang Feng.

Preprint

2023

DASpeech: Directed Acyclic Transformer for Fast and High-quality Speech-to-Speech Translation

Qingkai Fang, Yan Zhou, Yang Feng.

NeurIPS 2023 CCF-A

Understanding and Bridging the Modality Gap for Speech Translation

Qingkai Fang, Yang Feng.

Proceedings of ACL 2023 CCF-A

Back Translation for Speech-to-text Translation Without Transcripts

Qingkai Fang, Yang Feng.

Proceedings of ACL 2023 CCF-A

CMOT: Cross-modal Mixup via Optimal Transport for Speech Translation

Yan Zhou, Qingkai Fang, Yang Feng.

Proceedings of ACL 2023 CCF-A

Bridging the Gap between Synthetic and Authentic Images for Multimodal Machine Translation

Wenyu Guo, Qingkai Fang, Dong Yu, Yang Feng.

Proceedings of EMNLP 2023 CCF-B

BayLing: Bridging Cross-lingual Alignment and Instruction Following through Interactive Translation for Large Language Models

Shaolei Zhang, Qingkai Fang, Zhuocheng Zhang, Zhengrui Ma, Yan Zhou, Langlin Huang, Mengyu Bu, Shangtong Gui, Yunji Chen, Xilin Chen, Yang Feng.

Preprint

2022

STEMM: Self-learning with Speech-text Manifold Mixup for Speech Translation

Qingkai Fang, Rong Ye, Lei Li, Yang Feng, Mingxuan Wang.

Proceedings of ACL 2022 CCF-A

Neural Machine Translation with Phrase-Level Universal Visual Representations

Qingkai Fang, Yang Feng.

Proceedings of ACL 2022 CCF-A

Low-resource Neural Machine Translation with Cross-modal Alignment

Zhe Yang, Qingkai Fang, Yang Feng.

Proceedings of EMNLP 2022 CCF-B

  • President Scholarship of CAS, at ICT/CAS, May 2026
  • National Scholarship, at ICT/CAS, Nov. 2024
  • ICT's Special Scholarship (Highest award in ICT/CAS), at ICT/CAS, Jan. 2024
  • First Academic Scholarship, at ICT/CAS, Sep. 2023/2024
  • Merit Student, at ICT/CAS, May. 2023/2024
  • Outstanding Graduates in Beijing, at BUPT, Jun. 2021
  • CCF Elite Collegiate Award, Aug. 2020
  • National Scholarship (Top 1%), at BUPT, Dec. 2019
  • Silver Medal, ACM-ICPC Asia Regional Contest, Shenyang Site, Oct. 2018
  • Silver Medal, China Collegiate Programming Contest (CCPC), Guilin Site, Oct. 2018
  • Bronze Medal, National Olympiad in Informatics (NOI), Jul. 2016