M.Sc. & B.Eng.
Soochow University
👋 I am Qian Qiao (乔谦), an independent researcher. I received my M.Sc. and B.Eng. from Soochow University, advised by Professor Fanzhang Li. My previous research focused on multimodal understanding, image generation, text spotting, and few-shot learning.
My current interests are real-time video generation, joint video–audio generation, efficient fine-tuning, and multimodal models. I also have four years of experience in blockchain technology and investment. I am actively seeking collaborators in these fields.
Experience
M.Sc. & B.Eng.
Soochow University
Researcher
Soul AILab
Real-time Interactive Video Generation
Featured Open-Source Projects
Soul AILab / SoulX-FlashTalk
Audio-driven real-time streaming digital humans, released as an open project.
Soul AILab / SoulX-FlashHead
Oracle-guided real-time streaming talking heads, trained on the VividHead dataset.
yyyyyxie / TextFlux
OCR-free diffusion-transformer synthesis for high-fidelity multilingual scene text.
MindLab-Research / LongStraw
Long-context reinforcement-learning infrastructure beyond 2M tokens under a fixed GPU budget, using resident-prefix response-only GRPO.
Uni-ViGU Project Page
A unified video generation and understanding framework based on a diffusion video generator.
MindLab Research / UI4A
A component-native Generative UI harness, with Macaron Artifacts providing local runtimes and streamed previews for Claude Code and Codex.
My research develops practical generative and multimodal systems, with an emphasis on high-quality real-time generation, efficient model adaptation, and visual-language understanding.
Agent Models, Long-Context RL & Generative UI
Real-Time Video & Audio Generation
Multimodal Understanding & Generation
Efficient Learning & Model Compression
Selected publications















Other Publications
- ACL 2026 — Deputy: Accelerating Large Language Model Inference with Dynamic Low-Rank Substitution
- ACL 2026 — CoRE: A Fine-Grained Code Reasoning Benchmark Beyond Output Prediction
- ACL 2026 — Balancing Fidelity and Plasticity: Aligning Mixed-Precision Fine-Tuning with Linguistic Hierarchies
- BVRCC: Bootstrapping Video Retrieval via Cross-Matching Correction
- Temporal-Conditional Referring Video Object Segmentation with Noise-Free Text-to-Video Diffusion Model
- STRA: A Simple Token Replacement Strategy Alleviating Exposure Bias in Text Generation
- Efficient Fine-Tuning of Quantized Models via Adaptive Rank and Bitwidth
Academic Services
- Conference Area Chair: ICME 2026, PRCV 2025
- Program Committee Member: ICLR, ICML, NeurIPS, ACL ARR, AAAI, IJCAI, ACM MM, CVPR, ECCV, ICME, ICASSP, and more.
- Journal Reviewer: IEEE TIP, IEEE TMM, IEEE TCSVT, Pattern Recognition, Neurocomputing.
Collaboration
I welcome discussion and collaboration on real-time video generation, audio-driven avatars, multimodal models, and efficient fine-tuning.
Honors & Awards
- 2025 — Outstanding Graduate Award (Top 5%)
- 2024 — National Scholarship, M.Sc. student
- 2023 — Special Grade Scholarship, M.Sc. student
Academic Competitions
ICADR 2024 — Text Map Challenge (Click to view results ↗)
With my collaborator Yu Xie, we won three first places and one second place, and were invited to present a technical report.