Yuanbao, Qian Qiao's cat

Qian Qiao

Independent Researcher
Shanghai, China

Researching video generation, multimodal models, and efficient learning.

👋 I am Qian Qiao (乔谦), an independent researcher. I received my M.Sc. and B.Eng. from Soochow University, advised by Professor Fanzhang Li. My previous research focused on multimodal understanding, image generation, text spotting, and few-shot learning.

My current interests are real-time video generation, joint video–audio generation, efficient fine-tuning, and multimodal models. I also have four years of experience in blockchain technology and investment. I am actively seeking collaborators in these fields.

Experience

Soochow University logo M.Sc. & B.Eng. Soochow University
Soul AI Lab logo Researcher Soul AILab Real-time Interactive Video Generation
SAC
DeFi Researcher SAC 2017 – 2021
🎓Google Scholar|Selected papers and citations

My research develops practical generative and multimodal systems, with an emphasis on high-quality real-time generation, efficient model adaptation, and visual-language understanding.

Agent Models, Long-Context RL & Generative UI
Real-Time Video & Audio Generation
Multimodal Understanding & Generation
Efficient Learning & Model Compression

Selected publications

LongStraw and Macaron-V1 project visual
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU BudgetarXiv, 2026 · #1 Hugging Face Paper of the Day[Paper] [GitHub] [NVIDIA NeMo]
UI4A benchmark visual
UI4A: A Component-Native Harness for Generative UIMind Lab, 2026 · Core Contributor[Article] [GitHub]
Macaron-V1 project visual
Introducing Macaron-V1Mind Lab, 2026 · Research Release[Article]
Uni-ViGU paper visual
Uni-ViGU: Towards Unified Video Generation and Understanding via a Diffusion-Based Video GeneratorTechnical Report, 2026[Project] [arXiv]
SoulX-FlashHead paper visual
SoulX-FlashHead: Oracle-Guided Real-Time Streaming Talking HeadSoul AILab, 2026[arXiv] [GitHub]
SoulX-FlashTalk paper visual
SoulX-FlashTalk: Audio-Driven Real-Time Streaming Digital HumanSoul AILab, 2026[arXiv] [GitHub]
TextFlux paper visual
TextFlux: An OCR-Free DiT Model for High-Fidelity Multilingual Scene Text SynthesisEurographics, 2026[GitHub]
Large language model compression paper visual
Large Language Model Compression with Global Rank and Sparsity OptimizationICLR, 2026
TransDiff paper visual
Marrying Autoregressive Transformer and Diffusion with Multi-Reference AutoregressionTechnical Report, 2025
AIM paper visual
AIM: Let Any Multimodal Large Language Models Embrace Efficient In-Context LearningAAAI, 2025
TALDS-Net paper visual
TALDS-Net: Task-Aware Adaptive Local Descriptors Selection for Few-Shot Image ClassificationICASSP, 2024
Lie group Laplacian support vector machine paper visual
A Lie Group Laplacian Support Vector Machine for Semi-Supervised LearningNeurocomputing, 2025
DeepTTS paper visual
DeepTTS: Enhanced Transformer-Based Text Spotter via Deep Interaction Between Detection and Recognition TasksPRICAI, 2024
QPruner paper visual
QPruner: Probabilistic Decision Quantization for Structured Pruning in Large Language ModelsNAACL, 2025[arXiv]
DNTextSpotter paper visual
DNTextSpotter: Arbitrary-Shaped Scene Text Spotting via Improved Denoising TrainingACM MM, 2024[arXiv] [Project]

Other Publications

  • ACL 2026 — Deputy: Accelerating Large Language Model Inference with Dynamic Low-Rank Substitution
  • ACL 2026 — CoRE: A Fine-Grained Code Reasoning Benchmark Beyond Output Prediction
  • ACL 2026 — Balancing Fidelity and Plasticity: Aligning Mixed-Precision Fine-Tuning with Linguistic Hierarchies
  • BVRCC: Bootstrapping Video Retrieval via Cross-Matching Correction
  • Temporal-Conditional Referring Video Object Segmentation with Noise-Free Text-to-Video Diffusion Model
  • STRA: A Simple Token Replacement Strategy Alleviating Exposure Bias in Text Generation
  • Efficient Fine-Tuning of Quantized Models via Adaptive Rank and Bitwidth

Academic Services

  • Conference Area Chair: ICME 2026, PRCV 2025
  • Program Committee Member: ICLR, ICML, NeurIPS, ACL ARR, AAAI, IJCAI, ACM MM, CVPR, ECCV, ICME, ICASSP, and more.
  • Journal Reviewer: IEEE TIP, IEEE TMM, IEEE TCSVT, Pattern Recognition, Neurocomputing.

Collaboration

I welcome discussion and collaboration on real-time video generation, audio-driven avatars, multimodal models, and efficient fine-tuning.

✉️ Contact me by email.

Honors & Awards

  • 2025 — Outstanding Graduate Award (Top 5%)
  • 2024 — National Scholarship, M.Sc. student
  • 2023 — Special Grade Scholarship, M.Sc. student

Academic Competitions

ICADR 2024 — Text Map Challenge (Click to view results ↗)

With my collaborator Yu Xie, we won three first places and one second place, and were invited to present a technical report.