Ranking-PE: Prompt Optimization for Multimodal Clinical Diagnosis under Extreme Class Imbalance
臨床診斷 MLLM 提示詞優化:Ranking-PE 解決醫療資料極端不平衡問題
In clinical diagnostics, extreme class imbalance makes traditional accuracy-based prompt optimization ineffective, as models can score high by simply predicting negative. To solve this, the authors propose Ranking-PE, which reformulates prompt search using pairwise ordering. By optimizing how well candidate prompts rank positive instances over negative ones, it directly targets empirical AUROC. Tested on the MIMIC dataset, Ranking-PE boosts AUROC by +5.8 pp on Qwen3-VL-8B and +16.2 pp on MedGemma-4B without any additional model-call overhead.
Key points
AUROC over Accuracy
Addresses the issue where accuracy looks high due to predicting the majority negative class. Instead, it optimizes the model's ability to rank positives above negatives.
Pairwise Pareto Evolution
Applies pairwise ordering across all three layers of prompt search (Pareto dominance, LLM reflection, and final selection) to directly maximize empirical AUROC.
Zero Extra Overhead
Directly transforms the evaluation correctness matrix into a pairwise matrix without requiring surrogate losses or extra inference costs.
Vision Backbone is Critical
Ablation studies show prompt search cannot substitute for a weak vision encoder; a high-quality medical visual backbone remains a prerequisite.
How it works
| 傳統提示詞優化 (例如 GEPA) | Ranking-PE (本論文提出) | |
|---|---|---|
| Core Objective | 準確度 (Accuracy) | 排序能力 (AUROC) |
| Evaluation Base | 單一實例正確性 (1=對, 0=錯) | 陽性-陰性實例對排序正確性 |
| Under Class Imbalance | 容易退化、盲目預測多數類別 | 保持穩健、不受類別分佈影響 |
| Extra Compute Cost | 無 | 無 (免除代理損失函數) |
Why it matters
This research addresses a major bottleneck in deploying clinical AI: extreme data imbalance. Traditional prompt-tuning tools optimize for accuracy, yielding models that fail at screening rare diseases. Ranking-PE proves that optimizing prompt ranking ability, without costly parameter fine-tuning, dramatically improves the clinical utility of multimodal LLMs. It offers a practical, zero-overhead paradigm for medical AI adaptation.
Who it affects
- AI Researcher
- AI Developer
- Product Manager
- Enterprise Leader
How to use it
- 1Clinical Decision Support: Enhancing rare disease screening sensitivity and specificity on highly imbalanced medical images and patient notes.
- 2Automated Prompt Tuning for Medical MLLMs: Automatically generating clinical-grade, high-AUROC prompts for specific medical departments or diseases.
Limitations & caveats
- Relies heavily on a capable base vision encoder; prompt optimization cannot compensate for fundamentally poor medical-image understanding.
- Evaluated only on three diseases within the MIMIC dataset; generalization to broader clinical departments and live workflows requires further validation.
Related

Overcoming Generative Recommender Latency: Deploying HSTU Models with NVIDIA Dynamo-Triton and PyTorch AOTI
突破生成式推薦延遲瓶頸:NVIDIA Dynamo-Triton 與 PyTorch AOTI 部署 HSTU 模型實戰
Learn how to deploy HSTU generative recommenders using NVIDIA Dynamo-Triton, PyTorch AOTI, and FlexKV caching to achieve up to a 5.93x speedup on Blackwell GPUs.
Fixing the "Timing Shortcut": A Breakthrough in Non-Invasive Brain-to-Text Decoding
排除「時間捷徑」漏洞:非侵入式腦機介面解碼技術的新突破
Researchers revealed that recent breakthroughs in non-invasive brain-to-text decoding relied on a "timing shortcut" of word durations rather than actual brain signals. Their SimpleB2T method eliminates this shortcut, slashing the word error rate to 36.6%.
VideoMSN: Turning Image Classifiers into Efficient Video Learners via Super-Images
免 3D 結構與重構解碼器!VideoMSN 利用超大型圖片將影像分類器轉化為高效視訊學習器
VideoMSN is a self-supervised Masked Siamese Network that represents videos as grid-based 'super images', enabling standard image ViTs to learn powerful spatio-temporal video representations with up to 160x fewer pretraining epochs.