2ⁿᵈ Edition of the Cancer R&D World Conference 2026

Speakers - CRDWC 2026

Kirti Harsha, Cancer R&D World Conference, Miami, Florida, USA

Kirti Harsha

Kirti Harsha

  • Designation: Indian Institute of Technology Roorkee
  • Country: India
  • Title: From Thousands of Labels to Dozens: Label-Efficient Breast Cancer Histopathology Using Pathology Foundation Models and Vision Language Zero Shot Diagnosis

Abstract

The principal obstacle to deploying artificial intelligence in diagnostic pathology is not algorithmic performance but the cost of expert annotation. Conventional deep learning models require thousands of pathologist labelled specimens for each diagnostic task, and this requirement recurs whenever a model encounters a new institution, a different staining protocol, a new scanner, or a tumour subtype that was rare in the original training data. Pathologist time is the scarcest resource in the diagnostic pipeline, and any approach that reduces the annotation demand has a more immediate route to clinical impact than an incremental gain in accuracy on an already saturated benchmark.

This presentation examines how far that annotation requirement can be reduced. Over the past two years, pathology foundation models have emerged as a distinct class of tools: large vision encoders trained by self-supervision on hundreds of thousands of whole-slide images, producing general-purpose representations of tissue morphology without task-specific labels. In parallel, vision–language models trained on paired histopathology images and diagnostic text have made it possible, in principle, to classify a specimen from a natural-language description of what the pathologist is looking for, using no labelled training examples at all. Neither capability has been systematically characterised for breast cancer histopathology, and it is not yet known how many labelled specimens a clinically useful system genuinely needs. We address this using BreaKHis, a public collection of 7,909 hematoxylin and eosin breast tissue images acquired from 82 patients at four magnifications.

Our evaluation is conducted under a strict patient-disjoint protocol, in which no patient contributes images to more than one partition. This point deserves emphasis: because several dozen images originate from each patient, partitions drawn at the level of individual images allow a model to encounter the same patient's tissue during both training and testing, which inflates apparent performance and is a recurring weakness in the published literature. Patient-level separation is applied throughout. The study compares three regimes. The first is zero-shot classification, in which a pathology vision–language model is prompted with diagnostic text and receives no training labels whatsoever. The second is linear probing on frozen foundation-model embeddings across a graded series of label budgets, from one percent of the available training data up to the full set, with subsampling performed at the patient rather than the image level so that reduced budgets reflect fewer patients rather than fewer views of the same patient. The third is a conventionally trained convolutional network using complete supervision, representing current practice.

Performance is assessed by accuracy, sensitivity, specificity and area under the ROC curve, at both image and patient level, with particular attention to false negatives, since a malignant specimen reported as benign is the clinically consequential error. We additionally examine whether label efficiency differs across magnification, which bears on scanning protocol selection. The question the presentation seeks to answer is a practical one: what is the smallest number of annotated specimens at which a foundation-model approach becomes indistinguishable from full supervision, and can a vision–language model provide a useful diagnostic signal with no labels at all? If the answer favours these methods, the barrier to deploying diagnostic support in new laboratories, for rare subtypes, and in resource-limited settings is considerably lower than currently assumed. Results will be presented in full, together with an assessment of the sensitivity levels required before any decision-support role can responsibly be considered.