Neural classifier inversion: decoding latent visual semantics
This paper presents semantic inversion of pretrained image classifiers as an experimental probe of the view that late neural representations behave as class-dependent error-correcting codewords. The method reconstructs ImageNet validation images from late classifier activations by optimizing pixels to match one tail-proximal feature distribution under a symmetric Kullback–Leibler (Jeffreys) objective. The protocol is not a generative-prior attack and not a data-free synthesis method: it uses only the frozen classifier, a smoothed optimization path, a feature-mean constraint, a total-variation prior, and a class-logit KL constraint. The experiment evaluates 21 Torchvision models spanning RegNet, ResNet, EfficientNet, ConvNeXt, and ViT families on 48 class-capped high-margin validation samples per model. Reconstruction quality is measured objectively with OpenCLIP ViT-L/14 image-image retrieval metrics, feature-KL reduction in bits, and top-1 classification accuracy of the reconstructed images. The results show strong semantic recovery for ViT-B/16 and EfficientNet models, moderate recovery for RegNet models and ConvNeXt-Tiny, weak retrieval for the larger ResNet and ConvNeXt variants, and adversarial nonsemantic failures for ResNet18 and ResNet34. Qualitative reconstructions further illustrate that successful recoveries preserve class-level form and color without copying exact pixel-level appearance.