Aivora
arXivAI ResearchIntermediate

Ranking-PE: Prompt Optimization for Multimodal Clinical Diagnosis under Extreme Class Imbalance

臨床診斷 MLLM 提示詞優化:Ranking-PE 解決醫療資料極端不平衡問題

2 min read
Ranking-PE: Prompt Optimization for Multimodal Clinical Diagnosis under Extreme Class Imbalance
The 30-second version

In clinical diagnostics, extreme class imbalance makes traditional accuracy-based prompt optimization ineffective, as models can score high by simply predicting negative. To solve this, the authors propose Ranking-PE, which reformulates prompt search using pairwise ordering. By optimizing how well candidate prompts rank positive instances over negative ones, it directly targets empirical AUROC. Tested on the MIMIC dataset, Ranking-PE boosts AUROC by +5.8 pp on Qwen3-VL-8B and +16.2 pp on MedGemma-4B without any additional model-call overhead.

Key points

01

AUROC over Accuracy

Addresses the issue where accuracy looks high due to predicting the majority negative class. Instead, it optimizes the model's ability to rank positives above negatives.

02

Pairwise Pareto Evolution

Applies pairwise ordering across all three layers of prompt search (Pareto dominance, LLM reflection, and final selection) to directly maximize empirical AUROC.

03

Zero Extra Overhead

Directly transforms the evaluation correctness matrix into a pairwise matrix without requiring surrogate losses or extra inference costs.

04

Vision Backbone is Critical

Ablation studies show prompt search cannot substitute for a weak vision encoder; a high-quality medical visual backbone remains a prerequisite.

How it works

Traditional Optimization vs. Ranking-PE
傳統提示詞優化 (例如 GEPA)Ranking-PE (本論文提出)
Core Objective準確度 (Accuracy)排序能力 (AUROC)
Evaluation Base單一實例正確性 (1=對, 0=錯)陽性-陰性實例對排序正確性
Under Class Imbalance容易退化、盲目預測多數類別保持穩健、不受類別分佈影響
Extra Compute Cost無無 (免除代理損失函數)

Why it matters

This research addresses a major bottleneck in deploying clinical AI: extreme data imbalance. Traditional prompt-tuning tools optimize for accuracy, yielding models that fail at screening rare diseases. Ranking-PE proves that optimizing prompt ranking ability, without costly parameter fine-tuning, dramatically improves the clinical utility of multimodal LLMs. It offers a practical, zero-overhead paradigm for medical AI adaptation.

Who it affects

  • AI Researcher
  • AI Developer
  • Product Manager
  • Enterprise Leader

How to use it

  1. 1Clinical Decision Support: Enhancing rare disease screening sensitivity and specificity on highly imbalanced medical images and patient notes.
  2. 2Automated Prompt Tuning for Medical MLLMs: Automatically generating clinical-grade, high-AUROC prompts for specific medical departments or diseases.

Limitations & caveats

  • Relies heavily on a capable base vision encoder; prompt optimization cannot compensate for fundamentally poor medical-image understanding.
  • Evaluated only on three diseases within the MIMIC dataset; generalization to broader clinical departments and live workflows requires further validation.

Related

Fixing the "Timing Shortcut": A Breakthrough in Non-Invasive Brain-to-Text Decoding
arXivAI Research

Fixing the "Timing Shortcut": A Breakthrough in Non-Invasive Brain-to-Text Decoding

排除「時間捷徑」漏洞:非侵入式腦機介面解碼技術的新突破

Researchers revealed that recent breakthroughs in non-invasive brain-to-text decoding relied on a "timing shortcut" of word durations rather than actual brain signals. Their SimpleB2T method eliminates this shortcut, slashing the word error rate to 36.6%.

2 min read
VideoMSN: Turning Image Classifiers into Efficient Video Learners via Super-Images
arXivAI Research

VideoMSN: Turning Image Classifiers into Efficient Video Learners via Super-Images

免 3D 結構與重構解碼器!VideoMSN 利用超大型圖片將影像分類器轉化為高效視訊學習器

VideoMSN is a self-supervised Masked Siamese Network that represents videos as grid-based 'super images', enabling standard image ViTs to learn powerful spatio-temporal video representations with up to 160x fewer pretraining epochs.

2 min read