**Human–AI Collaboration in Dermatology Diagnostics: Enhancing Accuracy and Understanding AI Deference**
The integration of artificial intelligence (AI) into medical diagnostics has opened new avenues for improving accuracy, especially in fields like dermatology where visual pattern recognition plays a critical role. A recent study explored how human–AI collaboration impacts diagnostic performance among both the general public and medical professionals. By employing two large-scale experiments, the research evaluated how different AI assistance methods and decision-making paradigms influence diagnostic accuracy, biases, and trust in AI explanations.
—
### Study Design and Objectives
The study was designed as two complementary experiments. The first involved **623 members of the general public** tasked with a binary classification challenge—distinguishing melanoma from nevus. The second focused on **153 primary care physicians (PCPs)** and **320 medical students**, requiring them to tackle a more complex differential diagnosis task involving four skin conditions: atopic dermatitis, pityriasis rosea, Lyme disease, and cutaneous T cell lymphoma (CTCL).
Participants were randomly assigned to one of four AI assistance methods:
– **Basic**: AI provides predictions and confidence scores.
– **GradCAM**: AI offers heatmap visualizations highlighting relevant image regions.
– **CBIR**: AI shows visual similarities to reference cases.
– **Multimodal LLM**: AI provides natural language explanations.
Each participant evaluated 12 clinical images, with decisions either made **Human-First** (before reviewing AI suggestions) or **AI-First** (after reviewing AI inputs).
—
### Key Findings
#### 1. **AI Improves Diagnostic Performance**
Collaboration with AI significantly improved diagnostic accuracy for the general public. Average accuracy in melanoma detection rose from **69.7% to 75.8%**, with the largest gains observed when using **multimodal LLM explanations**.
#### 2. **AI Reduces Diagnostic Disparities Across Skin Tones**
Both the AI models and human-AI collaboration led to more balanced diagnostic performance across different skin tones. The fairness-constrained deep learning models reduced disparities by up to **76.9%** compared to traditional approaches.
#### 3. **AI Deference and the Double-Edged Sword of LLM Explanations**
While LLM explanations boosted performance when AI was correct, they also increased the risk of misplaced trust. When AI predictions were incorrect, participants relying on LLM explanations were most likely to follow misleading guidance, highlighting a need for critical evaluation skills.
#### 4. **Medical Professionals Show Greater Resilience**
Unlike the general public, **PCPs maintained stable performance** even when AI provided incorrect suggestions. Their ability to resist overreliance on AI was more pronounced, particularly with LLM-based explanations. Medical students, however, showed higher deference to AI, underscoring the role of expertise.
#### 5. **Decision-Making Paradigms Matter**
Placing AI suggestions **before human input (AI-First)** increased overall reliance on AI across all groups. While this did not significantly impact final accuracy, it amplified deferential behaviors, especially with LLM explanations.
—
### FAQ Section
**Q1: What types of AI explanations were tested in the study?**
The study tested four types of AI explanations:
– Basic (predictions and confidence scores)
– GradCAM (heatmap visualizations)
– CBIR (visual similarity to reference cases)
– Multimodal LLM (natural language explanations)
**Q2: How did the general public’s performance compare to that of PCPs?**
The general public showed improvements in accuracy when aided by AI, particularly with LLM explanations. PCPs, however, were more resilient to incorrect AI guidance and relied less on AI explanations, demonstrating the importance of medical expertise.
**Q3: Did the order of decision-making (Human-First vs. AI-First) affect outcomes?**
While AI-First paradigms led to higher initial accuracy and increased deference to AI, the final diagnostic performance remained largely unaffected by the order. However, AI-First increased the likelihood of participants trusting and following AI suggestions.
**Q4: What role did fairness-constrained models play?**
Fairness-constrained deep learning models significantly reduced diagnostic disparities across skin tones, improving equity in AI-assisted diagnostics.
**Q5: How do AI and humans complement each other in dermatology?**
AI excels in cases with subtle disease presentations, while humans perform better with atypical symptoms or unexpected features. This suggests that human–AI collaboration can leverage the strengths of both parties.
—
### Conclusion
The study highlights the transformative potential of human–AI collaboration in dermatology, particularly in improving diagnostic accuracy and reducing biases related to skin tone. While AI assistance proves invaluable, its implementation requires careful consideration of trust, explanation quality, and user expertise. For the general public, LLM explanations can amplify overreliance on AI, even when incorrect, whereas medical professionals demonstrate greater resilience. Ultimately, effective human–AI collaboration lies in combining AI’s computational strengths with human expertise and critical thinking, paving the way for more equitable and accurate diagnostics in clinical practice.



