GPT4V-level open-source multi-modal model based on Llama3-8B
- Updated
Mar 3, 2025 - Python
GPT4V-level open-source multi-modal model based on Llama3-8B
Tag manager and captioner for image datasets
Famous Vision Language Models and Their Architectures
Python scripts to use for captioning images with VLMs
Tiny-scale experiment showing that CLIP models trained using detailed captions generated by multimodal models (CogVLM and LLaVA 1.5) outperform models trained using the original alt-texts on a range of classification and retrieval tasks.
A comparitive study between the two of the best performing open source Vision Language Models - Google Gemini Vision and CogVLM