Research
Publications
* co-first, † corresponding author.
Mind the Rarities: Can Rare Skin Diseases Be Reliably Diagnosed via Diagnostic
Reasoning?
arXiv Paper
A 6,354-case clinical benchmark with 26,030 multimodal image-text pairs. SFT improves
final-diagnosis accuracy by 52.8%–151.6% relative across InternVL2.5-4B, MedGemma-4B, and MedGemma-27B;
DPO adds only marginal gains.
Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in
Space?
arXiv Paper
SpatialMed — a CT-based benchmark with 31,253 QA pairs across 2,375 CT scans, 117
anatomical structures, and 7 tumor types. Benchmarks 24 SoTA MLLMs and reveals major limitations in
numerical estimation and reasoning faithfulness on volumetric data.