PEAL shortcut analysis
Upload an ONNX image classifier and a zip with one folder per class. DiDAE distils the classifier into a sparse concept space, renders counterfactuals with a pretrained representation autoencoder, and asks you which edits are spurious. (Handing back a corrected classifier is switched off in this demo for now.)
1. Files
single input [batch, 3, H, W]; up to MB
<class_a>/*.jpg, <class_b>/*.jpg, …; up to MB
2. What the model expects (all fields required)
Job
log
Your verdicts are needed
Each card is one direction: a concept of the sparse dictionary that DiDAE switched on or off, which flipped the classifier. It shows how often the edit flipped the classifier, how accurate your classifier is on images with and without the concept, and every verified pair (original, edited counterfactual, where the image changed). Judge the direction as a whole: did the edit change the true class feature (), or something spurious the classifier should not rely on? Mark it not plausible if the edited images are broken.
Results
| # | direction | latent flips | verified flips | your verdict | examples |
|---|