LLaVA 13B

LLaVA 13B is a Vision-language model which allows both image and text as inputs.

Sign in to see your generation history for this model.