Fangxin Shang
Senior Algorithm Expert, Qifu Technology
My ideal is to build multimodal AGI: AI systems that can perceive, understand, and reason over visual, audio, textual, and structured signals in complex real-world environments.
I currently work on multimodal intelligence for credit and financial scenarios, including multimodal benchmarks and real-time audio-video analysis for financial due diligence. Previously, I worked on medical multimodal intelligence at Baidu, covering medical image analysis, biomedical vision-language models, synthetic medical data, and medical AI platforms.
[ Google Scholar ] [ GitHub ] [ Email ]
Google Scholar snapshot, September 15, 2026.
Short BIO
Fangxin Shang is a senior algorithm expert at Qifu Technology. His current work focuses on multimodal intelligence for credit and financial scenarios, especially systems that integrate visual, audio, textual, and structured evidence for reliable domain understanding and decision support.
Before joining Qifu Technology, he worked at Baidu from 2019 to 2025, where he led and contributed to multiple medical AI systems, including healthcare multimodal large models, medical image analysis platforms, lung CT analysis, fundus disease screening, and large-scale medical simulation systems. His long-term research interest is to build useful and trustworthy multimodal AGI for complex real-world environments.
Research Interests
- Multimodal AGI: perception, grounding, reasoning, and action across modalities.
- Financial multimodal intelligence: credit scenarios, multimodal evidence understanding, real-time audio-video analysis, and trustworthy decision support.
- Medical multimodal intelligence: biomedical VLMs, medical image analysis, segmentation, synthetic data, and medical evaluation benchmarks.
- Real-world AI systems: robustness, reliability, domain intelligence, deployment, and benchmark-to-application gaps.
Publication
Synced with Google Scholar after the latest profile update. Citation counts are not shown for every paper to keep the page easy to maintain.
2026
- Hypothesis-Driven Skill Optimization for LLM Agents. arXiv preprint arXiv:2606.22330, 2026. [ arXiv ]
- FCMBench-Video: Benchmarking Document Video Intelligence. arXiv preprint arXiv:2604.25186, 2026. [ arXiv ] [ GitHub ] [ HuggingFace ]
- From Image to Pixels: Towards Fine-Grained Medical Vision-Language Models. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026. [ scholar ]
- FCMBench: A Comprehensive Financial Credit Multimodal Benchmark for Real-world Applications. arXiv preprint arXiv:2601.00150, 2026. [ arXiv ] [ GitHub ] [ HuggingFace ]
2025
- Towards a Multimodal Large Language Model with Pixel-Level Insight for Biomedicine. Proceedings of the AAAI Conference on Artificial Intelligence, 39(4), 3779-3787, 2025. [ arXiv ] [ GitHub ] [ HuggingFace ] [ scholar ]
- MedRepBench: A Comprehensive Benchmark for Medical Report Interpretation. European Conference on Computer Vision (ECCV), 2026. arXiv:2508.16674. [ arXiv ] [ HuggingFace ]
2024
- A Refer-and-Ground Multimodal Large Language Model for Biomedicine. International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI), 2024. [ arXiv ] [ GitHub ] [ scholar ]
- SegICL: A Universal In-context Learning Framework for Enhanced Segmentation in Medical Imaging. arXiv preprint arXiv:2403.16578, 2024. [ arXiv ]
- Changes in masseter muscle morphology after surgical-orthodontic treatment in patients with skeletal Class III malocclusion with mandibular asymmetry: The automatic masseter muscle segmentation model. American Journal of Orthodontics and Dentofacial Orthopedics, 165(6), 638-651, 2024.
- Measurement Plane of the Cross-sectional Area of the Masseter Muscle in Patients with Skeletal Class III Malocclusion: An Artificial Intelligence Model. American Journal of Orthodontics and Dentofacial Orthopedics, 166(2), 112-124, 2024.
2023
- SynFundus-1M: A High-quality Million-scale Synthetic Fundus Images Dataset with Fifteen Types of Annotation. arXiv preprint arXiv:2312.00377, 2023. [ arXiv ] [ GitHub ]
- HC-Net: A Hybrid Convolutional Network for Non-human Primate Brain Extraction. Frontiers in Computational Neuroscience, 17, 1113381, 2023.
2022
- SeATrans: Learning Segmentation-Assisted Diagnosis Model via Transformer. International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI), 2022. [ arXiv ] [ scholar ]
- Learning Self-calibrated Optic Disc and Cup Segmentation from Multi-rater Annotations. International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI), 2022. [ arXiv ]
- Automatic Masseter Muscle Accurate Segmentation from CBCT Using Deep Learning-Based Model. Journal of Clinical Medicine, 12(1), 55, 2022.
- An Effective Transformer-based Solution for RSNA Intracranial Hemorrhage Detection Competition. arXiv preprint arXiv:2205.07556, 2022. [ arXiv ]
- Opinions Vary? Diagnosis First! International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI), 2022. [ arXiv ]
- One Hyper-Initializer for All Network Architectures in Medical Image Analysis. arXiv preprint arXiv:2206.03661, 2022. [ arXiv ]
2021
- Robust Collaborative Learning of Patch-level and Image-level Annotations for Diabetic Retinopathy Grading from Fundus Image. IEEE Transactions on Cybernetics, 52(11), 11407-11417, 2021. (*equal contribution) [ arXiv ] [ GitHub ] [ scholar ]
- Multi-modality Images Analysis: A Baseline for Glaucoma Grading via Deep Learning. International Workshop on Ophthalmic Medical Image Analysis, 139-147, 2021.
Experience
-
Senior Algorithm Expert, Qifu Technology, 2025.9 - present.
Working on multimodal intelligence for credit and financial scenarios, including FCMBench and real-time audio-video analysis systems for financial due diligence, interview summarization, and fintech projects. -
Senior Engineer, Baidu, 2019.7 - 2025.9.
Worked on medical AI and healthcare multimodal systems. Led and built 2D/3D medical image analysis systems, contributed to large-scale medical AI infrastructure and simulation projects, developed multimodal dialogue and analysis systems for healthcare scenarios, and led small cross-functional teams from research prototyping to product delivery.
Patents & Community
30 granted invention patents, including 21 as first inventor.
Selected Granted Patents
- Image Desensitization Method, Model Training Method, Apparatus, Device, and Storage Medium. Chinese Invention Patent CN116228896B, granted 2025. [ patent ]
- Medical Image Generation Method, Model Training Method, Apparatus, Device, and Medium. Chinese Invention Patent CN116402913B, granted 2024. [ patent ]
- Image Registration Method, Image Registration Model Training Method, and Apparatus. Chinese Invention Patent CN115908515B, granted 2024. [ patent ]
- Method and Apparatus for Generating Convolutional Neural Networks, and Image Recognition Method and Apparatus. Chinese Invention Patent CN113361693B, granted 2022. [ patent ]
- Semantic Segmentation and Model Training Methods, Apparatuses, Device, and Storage Medium. Chinese Invention Patent CN113920314B, granted 2022. [ patent ]
Community
- PaddlePaddle Senior Developer Technical Expert (PPDE) and Paddle Framework Contributor Club (PFCC) member.
- Founding developer member of PaddlePaddle MedicalSeg; contributed to PaddlePaddle, PaddleClas, PaddleDetection, PaddleSeg, and AIStudio open-source projects.