A ViT feature extractor computer vision model to classify images of food into 101 different classes.
Pretrained ViT model finetuned on Food101 dataset